mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-05 15:09:42 +02:00
0e8b1981af5969825a4883eff72dc1355dc13f43
141
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0e8b1981af |
fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker: 1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same config directory and warns about it (confirmed cosmetic - the API key wins for actual requests, verified via a real session's own API Usage Billing line). A custom-model claude session now gets an isolated CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty, no files written into it) so there is nothing to conflict with. projects is symlinked (junction on Windows) back into the real config dir so the response viewer, subagent windows and Read My Mind keep working for that session, best-effort. 2. Context-window overflow. Claude Code assumes a large default context window for a model id it doesn't recognise and never compacts, so a custom endpoint's real, much smaller context (verified live: a 400 exceeding a 16384-token llama-swap model with a stock ~33.7K-token system prompt) silently overflows. Discovery now also learns each model's real context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx), but ONLY for a model llama-swap's own /v1/models response already marks status.value === 'loaded' - never an unloaded one, since llama-swap treats ?model= as a routing hint and probing an unloaded model risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery. A server with no status field at all gets no enrichment rather than a guess; a model not probed this round keeps its previously-learned value until it disappears from the list entirely. Stored per model (CustomModelHost.modelContextLengths) and applied via a new contextLengthVar registry field, set to CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude. Both new fields live on the existing env-kind customModelInjection capability shape, declared only on claude's registry entry - every other CLI's injection is unaffected (pinned by test). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
25fae9ad10 |
feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment: "generate those entries from the saved profiles rather than a fixed duplicate per harness, and put it in a follow-up PR so this one stays the backend... The Run-menu picker is yours if you want it." Adds the frontend surface the backend has been waiting on: - Run menu: a "Custom Endpoints" section lists one entry per (harness that supports customModelInjection, saved endpoint) pair, e.g. "Claude Code (llama.cpp)". The harness list comes from window.__codemanCustomModelClis, injected at page render straight off the CLI registry's own capabilities (never a hardcoded id list in the frontend), so a CLI whose injection recipe lands later appears with no frontend change. Picking an entry runs that harness's own existing run*() function unmodified (case creation, env overrides, everything, forced to a single instance) and then applies the endpoint's default model to the session it creates via the existing POST /api/sessions/:id/custom-model route. Entries are hidden for a remote/docker active case, since that route already refuses both. - Settings: App Settings -> Models gets a "Custom model endpoints" group wiring up the customModelEndpointsEnabled toggle (declared since #393, read by nothing until now) plus CRUD against the existing /api/model-endpoints routes: list, add/edit (inline form), delete, discover models. - Backend: CustomModelHost gains an optional defaultModelId, the model the picker applies with no further choice per endpoint (one generated menu entry per CLI+endpoint pair, not per CLI+endpoint+model). The route refuses a value that isn't one of the endpoint's own discovered models, and a fresh discovery drops a default that no longer appears rather than carrying an invalid one forward. Docs: docs/custom-model-endpoints.md describes the new picker and settings panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the "backend-only" status note and documents the picker's generation mechanism. Tests: four new route tests cover defaultModelId validation, acceptance, and the drop/keep behaviour across a re-discovery; a new render-index-html test pins the __codemanCustomModelClis injection (present, agent CLIs supporting the capability, antigravity and shell excluded) and its solo-window skip. No browser test was added for the Run-menu picker itself or the settings CRUD panel (this box has no tmux, so the live server used by test:browser/test:mobile could not be exercised here) -- worth a Playwright pass before merge, same as any other frontend PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
5b920cb43d |
feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream and named tabs wrong five ways (every prompt, every write path, a bare Esc eating the next prompt's first character, pasted newlines as Enter, any CSI clearing the draft) and replaced the whole name, which dropped the case from the tab and reset the w<n> counter. This lands the feature with each of those closed: - First prompt means the first: applyAutoName() flips a placeholder to `auto` whether or not the string changed. nameSource is now the tri-state placeholder | auto | manual; the name setter is the only manual path. - Only user-originated input counts: write()/writeViaMux() take SessionWriteOptions.fromUser, set by the browser WS path and POST /input only, so Ralph, respawn, cron, approvals and the trust-dialog keys can never name a tab. A startMode 'shell' CLI never feeds the tracker (a capability, not an id check); the send-key route feeds trackUserInput() because its line feed bypasses the session. - Prefix form `w3-case: title`: parseSessionPrefix() already renders it as the title with the prefix in the tooltip and the next-session counter still matches it. Composed within MAX_SESSION_NAME_LENGTH. - Tracker rules per key: bare Esc resolves at chunk end; mouse/focus reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R taint the draft so Enter submits nothing rather than a fragment; bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space; the draft keeps its head past 8192 code points; an escape past 64 bytes is abandoned. - Title: slash commands by shape (a path is a prompt), `!` escapes refused, first sentence only past 8 code points ("e.g." is not a title), 72 code points on a word boundary. - Synced `autoNameSessions` setting, default OFF (the prompt reaches mux-sessions.json, session:updated and /api/search), App Settings -> Appearance -> Tabs, read fresh per prompt after the eligibility check. Tests: test/session-auto-name.test.ts (tracker, title, composition, ownership, emit gating), the wiring test (once, prefix, setting off, manual protected), test/routes/session-name-routes.test.ts (PUT /name flips to manual and persists). Verified live on an isolated instance: API and browser-typed prompts name the tab, a second prompt does not, shells and renamed tabs are untouched, nameSource survives a restart. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
942bf37e48 |
fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a custom OpenAI-compatible endpoint by injecting env vars or a config file and restarting the CLI in place. Review of the apply path found four things, two of them destructive. This lands all four plus the smaller items from the same review. 1. Clearing a selection did not clear it. The injected vars reach the CLI via `tmux setenv`, which persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so deleting the keys from the session's envOverrides relaunched the CLI still pointed at the old endpoint, and for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just been deleted. `Session.setCustomModel()` now reports the removed keys, queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys` carries them into `applyEnvOverrides()`, which `setenv -u`s them before re-applying the live overrides, on the same path that already unsets the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket that `setenv -u HOME` hands the next respawn the global HOME back. 2. Applying a model to a local claude session killed the pane. The relaunch was `claude --session-id <id>` and Claude refuses an id that already has a transcript, and unlike the dead-pane respawn this one kills a working pane first. `restartCli()` now pins the live conversation id as the resume id for that respawn when the CLI's launch declares a `fallback` chain, which renders the same `--resume <id> || --session-id <id>` shape the docker and remote pane commands use. Gated on the registry shape, not the CLI id: an entry whose resume id is minted by the CLI itself never declares that chain. 3. pi, omp and grok wrote their config file and then launched without the `--model` that selects it, so the file was ignored. The registry entry now declares `customModelInjection.launchModel` (`custom/{modelId}` for pi and omp, grok's `[model.codeman-custom]` block name), the builder renders it, and `_withCustomModelLaunchModel()` applies it onto the respawn options through `legacyConfigField`, leaving the stored <Mode>Config untouched so a clear falls back to the user's own model. A model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. 4. Remote (SSH) and Docker sessions reported `restarted: true` and changed nothing: their `restartCli()` reattaches the durable tmux rather than relaunching the agent, and the env lands on the local pane. Both are refused with a 400 until those paths are plumbed. Smaller items from the same review: - The selection survives a Codeman restart as the disk-only `__customModel` bookkeeping (endpoint, model, injected key NAMES, config dir, launch model; never the values, which carry the API key). Recovery re-derives the values from the endpoint store through the same apply path the route uses and keeps the bookkeeping even when the endpoint is gone, so a later clear still has keys to unset. - Discovery goes through `webviewFetch()`, so the RESOLVED address is judged by the same egress guard the web-tab proxy uses, and `baseUrl` reuses `webviewUrlSchema` (http(s) only, no embedded credentials, link-local and cloud-metadata addresses refused). undici's `fetch failed` wrapper is unwrapped so the user sees the ECONNREFUSED underneath. - `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session config dir 0700/0600 (pi and omp embed the key literally), and that dir is removed with the session. - `PR.md` is gone from the repo root and the design doc moved to `docs/custom-model-endpoints-plan.md` with the LAN address and the personal name scrubbed; every reference follows. The guide's `authStyle` text matches the shipped schema (`bearer | api-key`, default `bearer`) and says that `customModelEndpointsEnabled` is read by nothing until the picker lands. - `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts` (four real type errors fixed). It is not yet wired into `npm run typecheck` because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json` there is the one-line follow-up. Tests: `test/session-custom-model-restart.test.ts` drives a real Session and fails on the unfixed code for items 1 to 3; the route suite covers item 4 and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets run before the overrides and that a shell-metachar key never reaches tmux. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1e42cb4e2d |
Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses) |
||
|
|
707ea345eb |
fix(statusline): GET /api/settings never writes, and a save sends the collection switch only on a flip
Two follow-ups to #361's sticky telemetry switch. GET /api/settings reconciled an absent showPlanUsageLimits by persisting true, but readJsonConfig() answers {} for ANY read failure (a parse error, EACCES, EMFILE, a read landing inside PUT's non-atomic write), not only ENOENT, and every page load calls this route, so one unlucky read replaced the whole settings file with a one-key file. The route is a plain read again and the default moved into the reader: readPlanUsageTelemetryEnabled() treats an absent key as ON, the same way readWorkspaceHooksEnabled() does, which is what the desktop chip already shows for an install that never touched the setting. saveAppSettings() sent showPlanUsageLimits on every save. The chip defaults OFF on handhelds, so a phone saving its font size persisted false and switched collection off for every desktop, whose chip then went stale with no error anywhere. The key is now stripped like the other per-device display keys and re-added only when the save FLIPS the chip relative to what the device had (planUsageCollectionFlip), so an explicit toggle on any device still writes it in either direction. Tests pin both: the GET route with a mocked filesystem (absent, missing, EACCES, garbage, explicit), the reader default, and the flip helper plus its wiring in saveAppSettings. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
b2b2c767ea |
Merge pull request #361 from timkjr/fix/statusline-injection-opt-out
fix(statusline): inject plan-usage telemetry via ephemeral CLI flag, never disk |
||
|
|
41416566aa |
feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. Covers local hardware (llama.cpp, Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter). Off by default (customModelEndpointsEnabled, synced, default OFF). - Registry: capabilities.customModelInjection per CLI entry (env / configContentEnv / configDir / unsupported kinds) - Pure injection builder (custom-model-injection.ts) turning an endpoint + model id into the real env vars / config content per CLI - Endpoint store + CRUD routes (custom-model-hosts.ts, custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded - Session integration: Session.setCustomModel()/restartCli() (POST /api/sessions/:id/custom-model), reusing the existing respawn-pane -k primitive to restart the CLI process with new env - Multi-user hardening: every new redirect-capable env var added to its CLI's privilegedEnvKeys, closing a pre-existing gap where several were already reachable via the generic envOverrides field's prefix allowlist - Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries against a real endpoint outside the web UI, independent of tmux/sessions - Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying every CLI's injected values through a real HTTP shape Real end-to-end validation against a live llama-swap server (inside a codeman/agent:llm-test Docker image with all 9 CLI binaries) found and fixed three real bugs before they shipped: - Codex's config.toml schema was wrong ([model].default table instead of a top-level model string + [model_providers.custom]); fixing it then surfaced a genuine, documented protocol incompatibility (Codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement) - Claude Code's async session-title-generation call validates ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and hangs the whole -p invocation on an unrecognized name; documented for chunk 6, worked around in the standalone script only (--bare is NOT safe for a real interactive session, which needs hooks) - The discovery route's authStyle: 'both' option (send both Authorization and api-key headers) reliably hung a real server; removed the option entirely rather than just changing the default Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see PR.md and deployment_plan.md for the full chunk breakdown and confidence table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3 |
||
|
|
8b23f3e260 |
feat(rail): sort the vertical tab rail by activity, and give its rows the home screen's card
The vertical rail lists exactly the sessions both home screens list, so it now answers their question the same way instead of showing the raw tab order. Order: new per-device `tabRailSort` (App Settings -> Appearance -> Tabs -> Vertical Rail Order, default "By activity"). It runs `CodemanSessionOrder` over rows classified by `_mobileOverviewState`, i.e. literally the home screens' comparator, `lastSubmitAt`-anchored running group included. It is applied as the flex `order` property, never by reordering the DOM. `#sessionTabs` stays in `sessionOrder`, which is what keeps the Alt+N badge honest (it names a shortcut, not a row position, so it deliberately does NOT run 1,2,3 down a sorted rail), and keeps drag-and-drop, the arrow-key walk, the sidebar filter and `_scrollActiveTabIntoView()` all reading the list they always read. A session changing state then moves one inline style instead of forcing the full rebuild that would restart every card's animation on every SSE tick. The incremental render path re-applies it, since a state flip adds no tab and never reaches the full rebuild, and an empty string is what clears it when sorting stops. Web tabs are pinned past the cards by a CSS `order: 9999`, since `renderWebviewTabs()` emits the same markup for every layout and the flex default of 0 would interleave them. Drag is switched off while sorting (the drop rewrites `sessionOrder` correctly and the sort puts the card straight back, so the affordance would be a lie); 'manual' is the way back. Cards: detailed rail rows become bordered cards on `--bg-card`, with the stamps line on its own full-width row and the pill at its right end. The state dot goes 6px to 9px, keeps its orbiting ring while working and gains the green halo; idle mutes toward `--text-muted` as the home rail does. Needs/error/ waiting reuse `home-sessions-blink-red`/`-yellow` rather than a second copy. These card rules are RAIL-SCOPED and deliberately absent from the comma-grouped selectors that carry both vertical surfaces: the rail is an occasional, resizable list you scan, while the sidebar is a permanently-docked nav column where 20 stacked cards read as a wall. Every state-dot rule also excludes `.tab-alert-action`/`.tab-alert-idle` by hand, because those alert rules are only (0,3,0) and these are (0,5,1)+. Lines: the lineage bracket already drew in the rail, but its track sat 6px from the left edge, so half of its 11px outer glow was clipped by the window frame and it read as a thread pinned to the frame. It now runs at 10px, mid-channel in the gutter the rail already reserves. Tests: test/tab-rail-order.test.ts (17) drives the real `isTabRailSorted()` and `_tabRailSortOrder()` out of app.js, covering the row model (a WORKING row ranked by `lastSubmitAt`, which would otherwise rank every running turn as freshly started and fail no rendering test), the Alt+N badge, and the opt-out. Verified in Chromium across sorted/manual/simple/header-strip/sidebar with the setting flipped at runtime: no page errors, and the header strip and sidebar render byte-identically to before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d5b75af628 |
fix(statusline): sticky telemetry collection, footer print-through, EOF fix
Responds to Ark0N's review round on the ephemeral-CLI-flag statusline injection rework: - Rebase-detail fixes: registry-gated telemetry eligibility via getCli(mode)?.capabilities.statusLineTelemetry instead of a hardcoded mode === 'claude' check, using the capability flag master's CLI-registry refactor already declares for exactly this purpose. - Design question settled: sticky (a). Rather than persisting the toggle as a new field and threading it through every session-creation path (cron, Ralph Loop API, quick-start), eliminated the per-session field entirely. readPlanUsageTelemetryEnabled() (hooks-config.ts) reads the existing showPlanUsageLimits setting fresh from settings.json at every claude create/respawn (TmuxManager.createSession/respawnPane) - no per-session state to survive a restart, and it applies uniformly to every creation path for free, since they all flow through the same TmuxManager methods. This required fixing a real bug found along the way: showPlanUsageLimits was not actually round-tripping through settings.json on save - settings-ui.js explicitly excluded it from the PUT body as a pure per-device display key. It now flows through normally (both true and false); the load-side per-device merge behavior is unchanged. Removed entirely as a result: the statusLineTelemetry field from CreateSessionSchema/SettingsUpdateSchema, CreateSessionOptions/ RespawnPaneOptions, Session._statusLineTelemetry (this is what makes the restart-persistence bug moot rather than patched), and the frontend send sites. - Footer print-through restored: the no-user-statusline branch of the exporter script now runs the telemetry POST in the foreground so its own stdout becomes the in-terminal footer, falling back to a plain "codeman" marker only on curl failure. - Background-subshell EOF fix: the wrap-a-real-statusline branch closes stdin too, not just stdout/stderr (`>/dev/null 2>&1 </dev/null &`) - the un-redirected subshell process itself, not curl, was what held a reader-to-EOF's pipe open for however long curl took to finish. Added curl --max-time 5 so a hung (not just refused) Codeman cannot wedge the render. Tests: real-shell-execution tests for the footer/EOF fixes (fake curl stand-in on PATH, real sh subprocess spawns, real elapsed-time measurements - verified non-vacuous against a hand-reconstructed old-style script), unit tests for readPlanUsageTelemetryEnabled. Adapted two existing tests whose payloads referenced the removed field. Fixed during independent code review: a stray indentation break and a test exercising the wrong (legacy) exporter code path. Docs synced: CLAUDE.md, docs/usage-limits-display-plan.md (old disk-based section marked superseded, kept for history), docs/architecture-invariants.md. Full suite green: 352 files, 6780 passed, 12 skipped, 0 failed. tsc/lint/format:check/frontend-syntax all clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8ee7926e27 |
feat(agent-cases): tag agent-spawned case dirs and sweep their leftovers
A long orchestration creates one case directory per worker and deleting the sessions never removed them, so ~/codeman-cases accumulated scratch folders that were indistinguishable from real projects. They are now labelled and have a cleanup path. - src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven spawn gets a .codeman-agent-case.json marker (when, by whom, parent session, mode). Only the create branch writes it, so a linked case, a cloned repo or any pre-existing path is never labelled; reading is total, so a malformed marker means "not agent-created" rather than a half-trusted entry. - The signal is the new X-Codeman-Agent-Origin header the skill preamble sets on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body field, falling back to a resolved parentSessionId so a worker spawned by a stale skill copy is still labelled. - GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage badges each case and offers a review-then-delete sweep that names every directory in its confirm and skips any case a live session is working in. Removal stays on the existing DELETE /api/cases/:name. - Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was written per claude session and never removed (236 leftovers measured on a working machine). Now deleted with the session and swept at boot, guarded by a live-session keep set plus a 7-day age floor. Verified end to end on an isolated instance: marker written for header, body and lineage-only spawns, absent with no agent signal and for a pre-existing directory; inUse flipping on session end; badge, sticky bar, confirm and sweep driven in a browser; preamble seeded on create, removed on delete, boot sweep taking only the aged orphans. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bca1b764cc |
Merge pull request #367 from shenlvkang-collab/pr/claude-conversation-first-hand
fix(session): learn the live Claude conversation from the CLI's own hook |
||
|
|
3d8ffcb9a2 |
Merge pull request #357 from dignfei/feat/docker-adopt-existing-container
feat(docker): attach a case to an already-running container Conflicts came from work that landed after the PR was opened, and each is resolved onto the newer abstraction rather than by keeping the older code: - `defaultDockerCommandForMode` is registry-driven since #347, so the PR's `runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the refusal is visible only inside the container, so an adopted root container otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and `test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch. - The probe's mode list and its mode -> binary table both duplicated the registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is also what fixes the merge's silent regression: the hand-written list predates `omp`, and the run menu gates every docker case on this probe, so owned containers would have lost that mode. `shell` needs no arm — it declares no binary, so it is dropped from the lookup and reported available regardless. - The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one `missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto it. Its test now pins the single gate instead of counting seven arms. - The create arm keeps #349's swap-limit warning filter, which the adopted arm never reaches; the run-mode list gains `omp` from #353. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1 |
||
|
|
550e08a791 |
fix(webview): refuse link-local and cloud-metadata targets on the resolved address
The web-tab proxy, its Test probe and its WebSocket relay accepted any http(s) host. A live PoC relayed an IMDSv2-shaped PUT with custom headers to a loopback echo server through a capability and no cookie, and 169.254.169.254 (decimal, hex, IPv6-mapped, or via a DNS name) was as valid a dashboard as any other. Loopback and RFC1918 stay allowed on purpose: a localhost Grafana is the feature. Only link-local and the fixed cloud-metadata addresses are refused (169.254.0.0/16, fe80::/10, fd00:ec2::254, 168.63.129.16, 100.100.100.200, metadata.google.internal), at three stages that are each load-bearing: - the Zod schema, so a save gets a clear refusal; - a synchronous hostname check at every connect site, because net.connect skips DNS for an IP literal and a lookup hook never sees one; - a `lookup` hook on an undici Agent (webviewFetch) and on the ws client, which judges the RESOLVED addresses of a name and refuses when any is blocked. This is what closes DNS rebinding, which a hostname-string check cannot. Adds undici@^6 so the proxy runs the package's own fetch with the package's own Agent; a package Agent handed to Node's bundled fetch can mismatch protocols. Verified live on an isolated beta: 169.254.169.254.nip.io (a real name resolving to the metadata address) is refused by probe, proxy (403) and WS relay (4003), while 127.0.0.1.nip.io still passes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE |
||
|
|
4830e662f9 |
refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. Code that used to ask "which CLI is this?" reads the entry instead. Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every spawn command as a literal string, captured from the hand-written builders before they were deleted, and `test/location-overlay-commands.test.ts` does the same for all 20 remote and in-container pane commands. Config can never contain shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only in this release. OMP is included as a registry entry rather than a tenth hand-written builder, so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of `buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen and doctor ladders all drop out. Guard rails: - `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id branching reappears outside `stock.ts`, in any of its four shapes (`===`, `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the negated forms, which is how 36 of them survived an earlier pass. Every allowlisted branch carries its reason. - `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities; deriving one from another shipped the `until=stop`-hangs-on-shell bug. - `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config` wire field is separate, bridged only by `legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. - Registry data resolves AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks). A module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. - Six fields are annotated DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/ `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured. A test pins the list so it cannot quietly grow. Three user-visible changes, all deliberate and named: - `probeDockerCliVersion()` derives the in-container binary from the registry rather than assuming it equals the mode name (`antigravity` runs `agy`). - The remote CLI version probe now covers grok and deepseek, which the hardcoded map it replaces omitted while its own comment said the rule was "every mode except shell". - `codeman doctor`'s CLI rows are generated from the entries, so Claude's install hint is the install command rather than a docs URL, five CLIs gain hints they never had, and the row order follows the catalog. Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars (matching the `cliId` pattern) before its failure message quotes the value back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading the hand-editable `clis.json`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
ccfda623fe |
fix(session): learn the live Claude conversation from the CLI's own hook
Which conversation a pane is on was re-derived by correlating ~/.claude/history.jsonl against Session.lastSubmitAt — and lastSubmitAt is bumped only by input that flows through Codeman's own write path (Session.write / writeViaMux). A user who attaches to the pane's tmux session directly never set it, so resolveActiveClaudeSessionIdFromHistory() returned at its first line for that pane's whole life and the response viewer stayed pinned to the launch conversation, showing a pre-/clear transcript indefinitely. A UserPromptSubmit hook reports the live conversation id from inside the CLI process, delivered under the pane's own $CODEMAN_SESSION_ID. That binding is a fact rather than a correlation: it never consults workingDir, so it cannot be claimed by a sibling pane on the same folder, a closed tab, or a bare `claude` in the user's terminal. A pane holding such an id skips the correlation entirely, so the number of prompts eligible for cwd-based guessing goes DOWN, never up — the naive alternative (relax the guard, or synthesize an anchor from PTY activity) is the reverted bug the resolver's own comment describes. The hook also stamps lastSubmitAt, so it finally means "a prompt was submitted" rather than "typed into Codeman's web terminal". Conversations vouched for first-hand — and only those — extend a persisted claudeSessionChain, whose tail re-pins the conversation when a surviving tmux session is re-attached after a restart. ⚠️ start() resets the id at THREE points and the last one runs unconditionally after the mux branch, so the tail is applied there too; patching only the mux branch looks right and silently does nothing. ⚠️ The hook's stdout is discarded with curl's own -o /dev/null. Claude Code injects a UserPromptSubmit hook's stdout into the model's context ("Exit code 0 - stdout shown to Claude"), and a trailing >/dev/null does NOT work: curlCmd already ends `... 2>/dev/null || true`, and in `pipeline || true >/dev/null` the shell binds the redirection to `true`, which never runs on the success path. The discard is opt-in so the five SSE-fed events keep byte-identical command text and no workspace's settings file is rewritten for them. The staleness marker is quote-free for the matching reason: hooksJson is JSON.stringify'd, so a quoted needle never matches and the gate would rewrite every workspace on every spawn. Existing workspaces heal on their next Claude spawn through the staleness sweep. |
||
|
|
da91b4353b |
Merge pull request #353 from timkjr/omp-mode
feat: add OMP (Oh My Pi) as a new session backend |
||
|
|
5452ad5c5a |
feat(docker): add a folder picker to both path fields
Both paths in the adoption form had to be typed. Each gets a Browse button using the same path-input-group markup Link Existing uses, so the two look and behave alike. What they can browse differs, and that is the point. The host workspace path reuses the existing host picker. The container workdir cannot: an adopted container has nothing mounted at a matching host path, so a host listing would be a different filesystem — and getting this field wrong is the source of the opaque OCI chdir error at launch, which makes it the field that most needs to be clickable. Adds a read-only POST /api/docker-cases/browse: one `ls` through docker exec, no writes, no lifecycle, path shell-escaped like every other value. `ls -Ap` marks directories with a trailing slash and keeps names with spaces intact. PathPicker takes an optional fetchListing source rather than being forked: the container variant only swaps where the rows come from, and reuses the rendering, navigation, Up and Choose/Select unchanged. |
||
|
|
c98a59d709 |
fix(docker): verify the container workdir and end the probe with exit 0
Two defects that only a real container exposes. The probe chained `command -v X && echo X` with semicolons, and a script's exit status is its last command's. A container without the last probed CLI made the whole `sh -lc` exit 1, so a perfectly healthy container with tmux and claude was reported as "could not exec into the container". A missing CLI is data here, not failure, so the script now ends with `exit 0`. containerWorkdir defaulted to hostWorkspacePath. That default holds for an owned container only because the create-time bind mount puts the host directory at that exact path; attaching mounts nothing, so the two are independent facts. A host path absent inside the container makes `docker exec --workdir` fail with an OCI chdir error that surfaces in the pane as a bare "execvp failed". The preflight now proves the directory exists inside the container and refuses at link time. |
||
|
|
15eebde832 |
feat(docker): attach a case to an already-running container
Docker cases could only run in a container Codeman created itself. Attaching to one the user already built and runs means Codeman must leave that container's lifecycle completely alone, which the launch chain could not do: it was `image inspect` -> `inspect || create` -> `start` -> `exec`. Adds `DockerCase.owned`, mirroring the `owned:false` contract remote-SSH already uses for attached sessions. Absent (every existing case) means owned, so current behaviour is byte-identical. `false` means the container belongs to the user and Codeman may only exec into it. The launch chain for an attached container only looks, then execs: no image gate (the image is theirs), no create, and no `start` — starting a container we do not own is the very mutation attaching promises not to perform. A missing or stopped container fails closed with an actionable message instead. Credential seeding is skipped too: those copies read from create-time read-only mounts that do not exist here, and writing host credentials into someone's container is not ours to do, so its CLIs must already be authenticated inside it. Four fail-closed guards. buildDockerStopCommand and buildDockerRemoveCommand throw during pure string construction, so no caller bug can turn into a `docker stop`/`rm` on a container we do not own; removeDockerContainer refuses again at the lowest layer; drift reports "none" for an attached container, which carries no `codeman.confighash` label and would otherwise always look drifted and 409 the launch gate forever; and the orphan reaper skips attached containers through a check deliberately independent of the two conditions already covering them. `owned` is applied AFTER the config hash is computed. dockerConfigHash takes an explicit field list, so ownership can never shift an existing case's hash — if it did, every pre-existing case would trip the drift gate at once, and the remedy the UI offers is "recreate the container". Adds POST /api/cases/docker-adopt and a read-only POST /api/docker-cases/adopt-preflight. The preflight refuses at LINK time rather than at session launch, where the only ways out would be a dead pane or starting a container we do not own. Tests assert the negative guarantee directly — that create, start, stop, rm, restart and kill are absent from the generated commands while `docker exec -it` and `new-session -A` remain — since it cannot be observed by using the feature. |
||
|
|
3e1a0e679f |
fix(omp): resume by mode, not silently as claude, and support --continue
resumeHistorySession() never sent mode when recreating a session from a history/session-manager row, so the server default silently opened a plain Claude session for every non-claude row -- reproduced live: OMP rows spawned Claude sessions on click. Thread the row's mode through every call site (welcome list, session manager, mobile overview) and only send the Claude-specific resumeSessionId for claude rows. Codeman has no live PTY-reattach outside server boot, and it's moot for OMP anyway (exiting it kills the pane's only process), so route the non-claude relaunch through each CLI's own continue-most-recent flag instead of a context-free fresh start. OMP never got one: buildOmpCommand only implemented --model/--resume despite omp --help documenting -c/--continue. Added continueSession to OmpConfig end-to-end (type, schema, builder) mirroring the existing opencode/pi/grok/deepseek fields, and wired resumeHistorySession to use it. Verified live: told a real omp session a secret, exited it, closed the tab without killing tmux, relaunched with --continue in the same directory, and had it recall the secret. |
||
|
|
4f5678fac4 | feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) | ||
|
|
b00ab3ceea | feat(web): show Codex plan usage in header | ||
|
|
ca5fe1ab3e |
Merge pull request #338 from Ark0N/feat/vertical-rail-detailed-rows
Vertical tab rail: detailed rows (created / working / status), plus a rename-cancel fix |
||
|
|
c30dfaf0e7 |
fix(deepseek): run the web UI server in the background, not in a shell tab
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a shell tab running the server next to it. The shell was deliberate - the server lived in an ordinary session so it was visible, scrollable, killable and died with its tab, and nothing new had to supervise a long-lived HTTP server. That reasoning was sound and the result was still wrong in use: opening a dashboard should open one tab, and after the first launch the terminal is pure noise. The server moves to a background child process owned by a new `src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session gave away for free is now explicit, which is most of the module: - Exactly one server. A second click reuses the running one instead of racing it for a port; the session flow could not do this at all, because two clicks were simply two sessions. - Restarted when the requested authority changes. `--trusted-host` fences dsh's own /api against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two. Reusing a server fenced for the other origin renders a page whose every call 403s, which reads as a broken dashboard rather than a misconfigured one, so a mismatch restarts instead. - Killed on shutdown. The child is detached so its whole plugin tree can be signalled at once, which also means it would outlive Codeman and hold its port against the next start - the exact EADDRINUSE this feature already got wrong once. - Boot output captured and returned. With no shell tab there is nowhere else for a stack trace to land, so a failed spawn reports its own tail. The endpoint is fenced at the same bar as the profile installer and for the same reason: booting a dsh profile executes the plugin code in it, so this is a privileged action even though it reads as "open a page". `authority` comes from the client (`location.host`) because only the browser knows which origin is in play, and it is regex-confined at the schema boundary - defence in depth behind the argv-array spawn, admitting host:port in the shapes a browser authority can take and nothing readable as a second argument. `GET /api/deepseek/web-port` is gone; port selection moved into the supervisor, which is the thing that knows whether a server is already running. The two client-side probe helpers went with it, since the server now owns the wait. Verified over the tailnet authority end to end: no session is created (session count unchanged, one tab), the server runs on 3081 beside the user's own dsh web on 3080, status reports the tailnet authority, and the proxied dashboard renders with zero 4xx. Full gate green (6148 passed, +6). |
||
|
|
15ae5f5d81 |
fix(deepseek): make the web-UI shortcut pick a free port, verify it, and trust its frame
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.
1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
precisely the port a DeepSeek user is most likely to be serving on already,
so the launch died with EADDRINUSE against the user's own server. The port
now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
free loopback port by BINDING it (a connect probe cannot tell "free" from
"listening but not answering yet").
2. It opened the tab unconditionally. The crashed server left a saved dashboard
pointing at nothing, with the failure only visible in a shell tab nobody had
a reason to look at. The launch now polls the existing webview probe until
the URL answers, and on timeout reports the error naming the shell tab
instead of persisting a dead dashboard.
3. The saved tab was untrusted, so the frame was sandboxed without
`allow-same-origin` and the dashboard was broken twice over: the dsh
client-runtime reads `localStorage` while loading its plugins and died there
("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
every `/api` call no matter which authority `--trusted-host` named. Passing
`location.host` only means anything once the frame actually carries that
origin, so `--trusted-host` had never once done its job. The managed tab is
now created `trusted: true`.
That trade is real and deliberate: a trusted proxied frame is same-origin
with Codeman and can reach Codeman's API. It is defensible only because this
dashboard is an agent harness Codeman just started itself, on loopback, which
can already run code as the user. It is not a precedent for trusting
third-party dashboards, which is why it is set at this one call site rather
than defaulted.
Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.
`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.
Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
|
||
|
|
b330f1d9e8 |
feat(tabs): give the vertical rail the home screen's per-session detail
The vertical tab rail (tabOrientation 'vertical') listed names and nothing
else, while the rich sidebar and both home screens already answered the
question a docked column exists to answer: which of these sessions wants me
next, and how long has it been like that. The rail is a docked column too, so
it now draws the same row.
- New per-device setting tabRailDetail ('rich' | 'simple', default rich),
App Settings -> Appearance -> Tabs, in SettingsUpdateSchema + displayKeys and
stamped as data-tab-rail-detail by the pre-paint script, so a detailed rail
does not flash through simple rows on every load.
- ONE gate for both vertical surfaces: isRichTabRows() =
isSessionSidebarRich() || isTabRailRich(). The row model, the markup and the
20s in-place clock are the existing rich-sidebar ones, classified by
_mobileOverviewState/_mobileOverviewSince, so the rail, the sidebar, the
desktop home rail and the phone overview cannot disagree about what
"working" means or which stamp measures it.
- Detail rides on its OWN attribute, exactly as the sidebar's does, so every
existing [data-tab-orientation='vertical'] rule keeps matching both variants
untouched. A flip of detail ALONE still forces a full render (the stamps line
is emitted by the row template, not toggled by CSS) and re-runs
applyTabWrapSettings(), which owns the folder line and is now rail-aware.
- CSS: every rich paint rule gains a rail twin as a COMMA-GROUPED selector,
never :is() - an :is() list takes its most specific argument, which would
lift the sidebar arm from (0,3,1) to the rail's (0,5,1) and let these rules
outrank things they never used to.
- Width is why there are thresholds. At 256px the stamps line ellipsizes
mid-word, the same reason the rich sidebar is 300px, so a rail that has never
been sized defaults to 320 (RICH_DEFAULT_WIDTH, the existing Wide preset,
which also keeps the settings select on a named choice). A width the user has
chosen is never overridden: below 288px the created stamp is dropped rather
than truncated (tab-rail-tight, CSS only) and below 240px the rows go back to
simple (tab-rail-compact, which re-renders).
- The rich clock is armed and disarmed by applyTabOrientation() as well as
applySessionListLayout(); a leaked interval would rewrite stamps in a list
that no longer has any.
Also fixes a data-loss bug in the inline tab rename that predates the rail and
reproduces in every layout, header strip included: Escape set the input to ''
and blurred it, and the blur handler commits - so cancelling a rename PUT an
empty name, and the tab fell back to its folder label (measured against a live
server: ["rail-alpha","","rail-gamma"]). Escape now calls cancelRename(), which
invalidates the edit so the blur that follows the input's removal is a no-op.
Tests: rail-detail gate, the three ways it turns back off (simple, compact,
horizontal), sidebar-wins, render-on-detail-flip and the plumbing/CSS guards in
test/session-list-layout.test.ts; the rename cancel in test/inline-rename.test.ts
(browser suite), pinned by running it against the old code first. Verified live
against a real server on an isolated instance: detailed/simple/compact/header/
sidebar variants, click-select, the ... menu, inline rename, Alt+N, the in-place
stamp tick and a full settings-picker round-trip including reload.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
4cda150493 |
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/ antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI as a Codeman web tab. DeepSeek is wired unlike its siblings in three ways, each of which is the reason for a design decision rather than an accident: 1. The agent is a PROFILE, not the binary. `dsh` is a launcher over $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and `base` -- the interactive terminal front door is always a third-party plugin. So availability is two questions: `isDeepSeekAvailable()` (binary) and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run button gates on the latter, because reporting only the binary would spawn a pane that dies on arrival. When the binary is present but no profile is, the run menu offers to install one (POST /api/deepseek/install-profile). 2. The permission switch is an env var, not a flag. The harness has no command-line permission option; its sandbox/approval rows read DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access). Exported via `tmux setenv`, never on the spawn line. Absent = the harness's own workspace-write, which still asks, so the multi-user clamp is the only-if-sent branch and clamps to workspace-write, never read-only. 3. It is the only non-claude mode that passes hooksAvailableForMode(), and it earned that. The terminal front door reports idle/working/blocked to a supervising process over a generic env-gated contract; a generated shim (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each report to /api/hook-event as stop / agent_working / permission_prompt. So a dsh session gets definitive respawn triggers, real wait-endpoint signals and real Approvals Inbox items instead of output-stabilization guesswork. `agent_working` is new (157th SSE constant) and joins APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its alert at once. The resolver needs the strictest identity probe of the family: `dsh` is not merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell), so `dsh --help` must print the harness's own banner before a candidate is handed a spawn line. Model is deliberately not a session field -- it is a composition entry in the profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider keys named by a settings-file `apiKeyEnv` stay out, which is pi's 34-provider-key problem in a new shape. Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the status endpoint's two-part answer, the no-profile refusal, the profile bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the permission mode injected via setenv, and the full status bridge -- a send-and-wait returned signal "stop" from a real turn, and blocked/working created and cleared an Approvals Inbox item. Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md (decisions + honest gaps). Tests: test/deepseek-mode.test.ts, test/deepseek-cli-resolver.test.ts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c173ae0264 |
Merge remote-tracking branch 'origin/master' into worktree-grok-mode
# Conflicts: # src/web/public/app.js |
||
|
|
e3a2fb767f | feat(tabs): COD-358 add resizable vertical session rail | ||
|
|
3f8c8e99d1 |
feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
98e37bf895 |
feat(sidebar): add a rich session sidebar that carries the home screen's row detail
Session List Layout gains a third option. The old "Left sidebar" becomes "Left sidebar simple" and is unchanged down to the byte; the new "Left sidebar" puts on each row what the desktop home rail and the phone overview already show: when the session was first created, how long it has been in the state it is in, and a status pill naming that state. A docked column is not a tab strip. It has width to spare and a row per session either way, and "name + folder" is the whole story a TAB can tell, not the whole story there is. This is the information that was missing, and it already existed one surface over. Both sidebar values are the same layout, and both set data-session-list="sidebar"; the row detail rides on a separate data-sidebar-detail attribute. That split is the load-bearing decision here: every one of the ~25 isSessionSidebarActive() call sites and every html[data-session-list="sidebar"] rule in styles.css and mobile.css keeps matching both variants without being touched. A third data-session-list value would have meant auditing and editing all of them. - Stored values: 'header', 'sidebar' (simple), 'sidebar-rich'. Anyone already on 'sidebar' keeps exactly the layout they picked — the rename is label-only. - State classification and the "how long has it been like this" anchor come from _mobileOverviewState() / _mobileOverviewSince(), not re-derived, so the three surfaces cannot disagree about what "working" means. A working pane repaints ~1/s, so its duration is measured from the turn's last Enter: a running turn reads "working 12m", not "0m". - Stamps refresh in place on a 20s clock rather than by re-rendering — a rebuild would restart every load spinner and alert animation in the list, twice a minute. The clock runs only while rich rows are on screen, and is stopped from both render paths and from applySessionListLayout(). - The incremental render path updates the pill, the accent class and the since anchor; a tick alone cannot see a state change, and a new turn re-stamps lastSubmitAt without changing state. - applySessionListLayout() now re-renders on a DETAIL change too. simple <-> rich leaves data-session-list on 'sidebar' both times, and the meta line is emitted by the row template rather than toggled by CSS, so the old layout-only test would have flipped the setting and repainted nothing. - Width: 300px for the extra line. The collapsed 44px rail and the handheld drawer are both explicitly held back from it — the desktop rule is (0,3,1) and would otherwise out-specify mobile.css's (0,2,1) drawer base and pin a 320px phone's drawer to 300px. - Missing/stale mobile-overview.js degrades to a row with no meta line rather than throwing and taking the whole tab strip down. 15 new tests cover the attribute split, the solo-window override, the detail-change re-render, the row model, both render paths, the clock lifecycle and the mobile width guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
05c94f5ac0 |
Merge pull request #304 from Ark0N/fix/workspace-hooks-install
fix: install Codeman hooks into every claude workspace, not just cases Codeman created |
||
|
|
210da991d5 |
Merge christianhaberl's session-sidebar branch, ported to current master
Brings in https://github.com/christianhaberl/Codeman/pull/4 (three commits, authorship preserved) and adapts it across the 211 commits master gained since the branch was cut: - App Settings control re-authored for the set-* surface (PR #278): a set-row in Layout -> Tabs, replacing the old settings-item markup the branch targeted. i18n description synced. - Lineage arcs (PR #291, post-branch) are SKIPPED in sidebar layout: computeLineagePath()'s U-bridge geometry hangs from the horizontal strip's bottom edge and has no meaning against a vertical list. The lineage strip-scroll listener now also redraws subagent/ultracode connectors while the sidebar scrolls vertically. - The desktop home tab rail (post-branch) defers to the sidebar: both dock the session list flush left, and the rail would render z-ordered under it. - Active-row reveal unified into _scrollActiveTabIntoView() (#257 landed on master after the branch): sidebar mode branches to scrollIntoView block:'nearest', and _fullRenderSessionTabs() restores scrollTop alongside the #257 scrollLeft restore so ambient rebuilds cannot yank a mid-scroll sidebar back to the top. - Mobile active-tab hoisting the branch guarded against no longer exists on master (removed by #257); kept master's order-stable render. Verified: typecheck, lint, format:check, check:frontend-syntax, check:public-assets, PostCSS parse of both merged stylesheets, the 26 new jsdom tests, the structural guard suites, and the headless-Chromium harness (scripts/verify-session-sidebar.mts) green across all seven layout states at 1600/1000/393px against current master. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f485085174 |
feat: workspaceHooksEnabled setting as the opt-out for workspace hook installs
Installing hooks into any workspace a Claude session runs in is the right default, but it takes a decision away from a user who deliberately removed them: nothing on disk distinguishes "removed on purpose" from "never had any", so they would come back on the next session create. Adds the synced workspaceHooksEnabled setting (App Settings -> Agents & CLIs -> Claude), default ON. OFF restores the older behavior exactly: a Codeman hooks block that is already present is still refreshed when stale (COD-91), but one is never added. Every create path routes through one applyWorkspaceHooks() helper so the gate cannot apply to some paths only, and the boot-time recovery sweep honours it too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c5b59633d8 |
feat(pi): add Pi (pi.dev) as a sixth CLI run mode (#206)
SessionMode gains 'pi', a first-class backend alongside Claude Code, OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose tab identity, welcome button, run-mode entry, cron agentType, Docker and remote-SSH command defaults, and clone-repo Brain option. Pi is a different shape of CLI from the other four, and three decisions follow from that: - It has NO permission prompts and no sandbox, so there is no --dangerously-skip-permissions analog and none was invented. The privilege-shaped knob is the tri-state approveProjectTrust, which makes pi load and EXECUTE repo-local .pi/extensions TypeScript and install missing project packages. clampExternalCliBypassForOwner() therefore puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets --no-approve even when no config was sent, because pi's own default is a prompt the session user could answer themselves. That helper had zero test coverage; it now has coverage for all four CLIs. - Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode context, so admitting them would widen the allowlist for every mode at once. Auth goes through pi's /login or the server's own environment. --api-key is deliberately never wired: it would put a provider secret on the spawn command line. - pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the main screen with terminal-owned scrollback, and its 0.84.0 fullscreen mode is runtime-switchable via /settings; that flip was measured to put the pane into the alt screen, which the strip would have corrupted. pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires semver-shaped output, because `pi` is a short generic name a stray binary can shadow; GET /api/pi/status surfaces path and version so a misresolution is diagnosable rather than presenting as a broken mode. Docker installs pi in its own --ignore-scripts step so that flag cannot affect the other four CLIs, and seeds its credentials per-file rather than whole-dir (~/.pi/agent also holds sessions, extensions and package trees). Verified end to end against pi 0.84.1 on an isolated instance: resolver search-dir fallback, flag construction, piConfig persistence across a full server restart, the trust prompt and its --no-approve suppression, the rose Run button on the default daylight-blue skin (the nested skin block eats per-mode gradients unless the rule lives inside it), and the buffer local-echo policy, which pi tolerates where codex did not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f39beb3326 |
chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4b51ba306e |
feat(voice): dictate through the server's Claude Code login, no API key
The mic button previously needed a Deepgram API key, or fell back to the browser's Web Speech engine. It can now transcribe through the same speech-to-text service Claude Code's own /voice mode uses, so anyone signed in to Claude Code on the server gets dictation with no third-party account. Claude Code's voice mode cannot be driven directly: it opens the HOST's microphone (sox/arecord), and the CLI runs in a headless tmux pane while the human is in a browser somewhere else. So capture stays in the browser and only the transcription backend is borrowed. Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since MediaRecorder cannot emit raw PCM) and receives text. - GET /api/voice/status reports readiness and never the token - GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade guard as the terminal socket, plus caps on concurrency, stream length and frame size - credentials are read-only: Codeman never refreshes them, since a refresh rotates the refresh token and could sign the user out of their own CLI - claudeVoiceEnabled (synced, default OFF) gates the whole server side - voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers Claude, then a configured Deepgram key, then the browser Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5671c20076 |
Merge remote-tracking branch 'origin/master' into feat/readmymind-phase2
# Conflicts: # CLAUDE.md |
||
|
|
1692238531 |
Merge remote-tracking branch 'origin/master' into pr251-review-fixes
# Conflicts: # CLAUDE.md |
||
|
|
94abcf29dc |
feat: Read My Mind phase 2, the predictor and the brain button
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.
Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
sources: pending approval dialog, user goals, last assistant turn tail,
recent prompts, tool activity, git workspace signals, away context,
sibling sessions, rethink state; 30 KB budget, whole-section drop from
the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
keeps only a 500-char snippet) and git signal collection (execFile,
2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
session, opus by default (readMyMindModel setting), strict JSON
contract with 1-3 suggestions (continue / verify / redirect), newline
stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
prediction in flight per session (409 CONFLICT), rethink body
{ steer, rejected }; ownership via findSessionOrFail
Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
until readMyMindEnabled is ON, desktop only (phone key is phase 3);
modal with editable suggestion + rationale and Send / Insert /
Rethink / Dismiss; suggestion text rendered via value/textContent only
and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
strings
Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
23d91a6ee1 |
feat(sessions): allow per-session CLAUDE_CONFIG_DIR env override (#255)
Adds an exact-key tier (ALLOWED_ENV_KEYS) beside ALLOWED_ENV_PREFIXES in schemas.ts, admitting CLAUDE_CONFIG_DIR so a case can run on a separate Claude subscription (client-billed accounts). Exact match only: other CLAUDE_* keys and near-misses like CLAUDE_CONFIG_DIR_EXTRA stay rejected, blocked keys stay blocked. The key also survives getEnvOverridesForPersist() (a path, not a secret; dropping it would silently switch a rebuilt session back to the default account after a reboot). Docs cover the transcript caveat: a relocated config dir writes transcripts outside ~/.claude/projects, so response viewer / subagent windows / ultracode / Read My Mind go blind for that session unless projects is symlinked back into the shared tree. Design and spec contributed by @jordan8037310 in #255. Closes #255. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
161f1da2eb |
feat: Read My Mind phase 1, per-case intent profiles (capture + API + skill)
Per-case profiles of user intent (docs/readmymind-plan.md): user-stated goals plus the user's recently submitted prompts, captured from the Claude session transcript behind the new synced readMyMindEnabled setting (default OFF). - intent-store.ts: keyed by owner + realpath(workingDir), FIFO/size caps, consecutive-dupe collapse, atomic 0600 writes to ~/.codeman/intents.json - transcript-watcher.ts: new transcript:user_prompt event for typed user turns (tool_result-only entries stay silent); capture wiring in server.ts is claude-only and gated on the setting per event - readmymind-routes.ts: GET/PUT/DELETE /api/sessions/:id/intent, ownership via findSessionOrFail, strict Zod schema - agent skill: SKILL.md recipe + endpoints.md rows so agents can read and record intent (PUT replaces: read + merge; never delete unprompted) - groundwork for the phase-2 predictor button; nothing is ever auto-sent Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e087198056 |
Merge pull request #250 from Ark0N/feat/path-picker-show-hidden
feat(path-picker): show hidden files and folders, and harden the secret blocklist |
||
|
|
6cc7b4328b |
feat(cases): clone a Git repository as a new case (#236)
Adds an Add Case -> "Clone Repo" tab plus two endpoints, implementing @DodgyBadger's proposal in #236: clone a public repository straight into codeman-cases/<name> and register it as a normal local case. POST /api/cases/clone is synchronous by design (request held open, bounded by GIT_CLONE_TIMEOUT_MS): no job store, no polling, no cancellation surface. Success broadcasts the usual case:created event, so the case still appears when a proxy idle-timeout kills the request mid-clone. POST /api/cases/clone-preflight runs `git ls-remote --symref` so the UI can say, while the user is still typing, whether the URL is cloneable without credentials, what its default branch is, and which branches/tags exist. Core lives in src/git-clone.ts, split into a pure half (URL parse, argv/env, ls-remote parse, stderr classification) and a thin IO half, so every security decision is unit-testable without spawning anything: - `<name>::<payload>` transports are refused as a family, not by name: ext:: is the famous one, but any of them dispatches to git-remote-<name> and turns a clone into arbitrary command execution. - A leading `-` is refused AND every spawn puts `--` before the operands. Either alone is one edit away from being a hole. - argv arrays, never a shell. URLs carrying user:password@ are refused. - gitNonInteractiveEnv() closes all four ways git can block on a prompt with no terminal attached (terminal prompt, askpass/GUI, ssh, GCM). HOME/PATH stay inherited, so a user's own credential helper or ssh agent keeps working; Codeman itself collects and stores nothing. - The timeout signals the process GROUP, since clone fans out into git-remote-https/index-pack children that outlive a signal to the parent. - Bounded output (redacted stderr tail, capped ls-remote stdout, 500 refs each) and a global 2-op pool, so N large clones cannot exhaust the host. Repository contents beat scaffolding: an existing CLAUDE.md is kept, hooks are merged into whatever .claude/settings.local.json the repo shipped, and a repo that ships its own Claude settings is reported back as a warning (those hooks run locally as soon as a session starts there). A failed clone removes only the directory the attempt created, and refuses a pre-existing destination outright, so it can never squat on a case name. Not admin-gated in multi-user mode, unlike /api/cases/link: it writes only inside the caller's own case space. Local-path/file:// sources are the exception and stay admin-only there. UI: live verdict under the URL field, case name filled from the parsed repo until the user types their own, branch/tag as a datalist of the remote's real refs, optional shallow clone, and a Brain picker (installed CLIs only) that points the Run button at the chosen agent. Starting a session stays opt-in. The tab hides itself when the server reports no git. Tests: the pure half exhaustively (every refusal has a case), plus real git against a real local bare repo for clone/ref/timeout/cleanup, and a route-level suite with unmocked fs that clones through the endpoint. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ce22c2a608 |
feat(path-picker): show hidden files and folders, and harden the secret blocklist
The picker behind Link Existing's "Browse" and the mobile keyboard's Path key refused every path with a dot-prefixed segment, so `.github/workflows/ci.yml` could not be selected and a hidden folder could not be opened at all. It gains the same `.*` toggle as the File Viewer: default OFF, per-device, and applied to both the listing and the preview endpoint, which re-resolves the path independently. That dotfile filter was quietly doing security work. The picker's roots include Home, so with every hidden path unreachable the shared blocklist never had to name the credentials that live in dot-directories. Lifting the filter removes that accident, so `isSensitivePath` now covers them explicitly: SSH keys at any depth rather than only under $HOME, GPG keyrings, AWS/GCloud/Azure/Docker/ Kubernetes credentials, npm, Yarn, git, gh, netrc, PyPI, RubyGems, Cargo and Terraform tokens, .pgpass and .my.cnf, and the Claude and Codeman agent credentials. `~/.codeman/` and `~/.claude/` stay attachable as trees, since the publish skill and the review-card loop read from them; only their secret-bearing members are named. Everything else still applies with the toggle on: blocked trees, sensitive files, root confinement, ownership scoping and symlink-escape checks. A hidden entry whose realpath is a secret is dropped from the listing, and opening it is refused. Follows #221 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
338f0e460d |
feat(approvals): gate push Approve/Deny buttons on the opt-in setting too
One switch now governs the whole feature: with approvalsInboxEnabled off (the default), sendPushNotifications strips the actions and approvalId from permission push payloads, so the buttons no longer render at all (pre-inbox they rendered and did nothing). The page-side action relay is gated the same way for stale notifications sent before the toggle flipped. Only the store and answer endpoints keep running, so enabling the toggle surfaces anything already pending immediately. sendPushNotifications is async now (cached settings read); all call sites were already fire-and-forget. Covered by three new payload tests alongside the existing hostTitle suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6c744f8677 |
feat(approvals): make the inbox opt-in (default OFF) and drop em-dashes
Owner decision: every Approvals Inbox UI surface (header bell, drawer, phone overview answer strips, reload seeding) now requires enabling approvalsInboxEnabled in App Settings -> Panels; only an explicit true turns it on. The store, endpoints, and push Approve/Deny actions keep running regardless (the push buttons are already opt-in per subscription). Also replaces em-dashes with plain punctuation across the newly authored comments, docs, and strings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ff10a50bc0 |
feat: Approvals Inbox, one cross-session queue for prompts waiting on a human
Permission dialogs, AskUserQuestion questions and idle prompts from every session now land in a server-side inbox (web/approval-inbox.ts, one item per session, claude-mode only) and are answerable in place: a header bell + drawer on desktop, inline answer strips on the phone overview's NEEDS YOU rows, and working push Approve/Deny buttons (previously dead ends, now answered straight from sw.js with no tab open). Pending alerts survive reloads because the frontend seeds from GET /api/approvals on init. Answering sends the digit / Esc / prompt text through the existing tmux input path; option digits are accepted only when they match options parsed from the captured pane frame, and the answer path re-captures the pane first so a dialog that already left the screen refuses with 409 instead of typing into the composer. New elicitation_complete / elicitation_response hook matchers resolve question items the moment they are answered in the terminal; refreshStaleCodemanHooks heals existing cases. Verified end-to-end against a live claude session: a real AskUserQuestion dialog parsed into 5 option buttons and was answered from the drawer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e88b971bb7 |
feat(skill): add the agent-skill install layer and harden the packaged skill
Ship `skills/codeman` as an installable Claude Code skill rather than a repo-only reference, and fix six defects found while verifying it live. Install layer: - `codeman skill install [--case <name>]` / `codeman skill uninstall`. Case names resolve through linked-cases.json first, mirroring the server's resolveCasePath(), so a case linked in from outside ~/codeman-cases no longer fails with "Case not found". - applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in hooks-config.ts. Copies are marker-owned, so an unmarked user-authored skill is never touched, and a symlinked skill dir is refused (this repo's own .claude/skills/codeman is a symlink to the source). - Synced `agentSkillEnabled` setting, default OFF: schemas.ts, ports/config-port.ts, server.ts, session-routes.ts (add-only injection on Claude session create and quick-start), plus the App Settings toggle. Skill content fixes, each reproduced before and after: - Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`. Shell state does not survive between agent tool calls, and an undefined is_self exited 127, firing the `||` branch and deleting the caller's own session with the one guard bypassed. The request now lives inside the guard, so a lost preamble deletes nothing. - clientId is a fixed literal instead of `agent-$$`. The pid changes per tool call, so the documented resend-identical-request loop stopped being a duplicate and retyped the prompt, submitting the turn twice. - `last-response` is now the documented read path for claude and codex workers. It returns clean transcript text; the terminal scrape it replaces returns a wall of TUI repaint noise. Its transcript flush lags the stop signal, so the recipes poll it rather than reading once. - quick-start examples branch on `.success`. Previously a failed spawn yielded the literal session id "null" and burned the whole readiness budget before reporting jq noise instead of the cause. - Documented that turning `agentSkillEnabled` off sweeps nothing, and corrected the hooks-config comment that claimed a toggle-off sweep exists. Per-case cleanup is `codeman skill uninstall --case <name>`. - Documented that SESSION_BUSY means the 50-session cap on quick-start, and that caseName resolves linked cases, so a generic name can land a worker in a real repo. Tests: test/agent-skill.test.ts covers install, refresh, idempotence, marker ownership and symlink refusal against the real packaged source; test/quick-start.test.ts covers injection behind the setting. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |