mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-03 14:09:42 +02:00
492f8d8ddf493b2b9141bb37dd03682667541fd2
353
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bb8ada7e5f |
fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does. 1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred it from the name: a session the user renamed by hand to something shaped like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the next prompt overwrote their name. The route persists right after, so the loss went to disk. `restoreMuxSessions()` already passes it. 2. The already-live sets were snapshotted once before a loop that awaits a real `startInteractive()` per entry, so by the tenth entry the snapshot was tens of seconds old and a conversation resumed by hand from the Resume list in that window was invisible to it: two panes on one transcript, the exact thing the check exists to prevent. Both sets are now read per iteration, and the late case is spent rather than re-offered for the same reason the batch case is. 3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path. The stamp predates the reboot and the pane is new, so honouring it meant one click had every restored session type `continue` into itself about a minute later, unattended, against the route header's own promise that a restored session comes back idle and disarmed. The setting stays ENABLED, so it re-arms on the next real limit message. A Codeman restart still re-arms from the stamp, because the limit footer will not reprint on its own; the new option exists only to tell the two paths apart. 4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()` and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession` performs that it was missing. Cosmetic, but a run left open reads as still going in the away digest. 5. A restored claude session gets `seedAgentSessionPreamble()` like both create paths, so the agent skill's bootstrap stays a two-line loader. 6. The heuristic's container comment was wrong in one direction and quiet about the real gap: after a genuine host reboot a containerized Codeman sees the host's short uptime and the banner does appear. What it cannot see is a container-only restart, which is where this would help most. 7. The banner is hidden in a solo window, which shows one session and has no tab strip to put restored ones in. Also reverts 17 of the 18 hunks in docs/api-reference.md, which were Prettier reformatting of prose the PR does not otherwise touch (docs/ is outside the format glob), keeping only the Reboot restore section and repairing the two continuation lines that reformat de-indented; renumbers reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which loads after it; and gives the feature its CLAUDE.md entry plus a route test for the multi-user workspace-forbidden branch, the only new rule that had nothing behind it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5f55f9cb65 |
fix(sessions): never let the reboot-restore plan fail recovery
The plan build runs inside the try that decides whether restoreMuxSessions() succeeded, so a throw would be caught there, report restoration as failed, and block the stale cleanup and layout reconciliation that follow. An optional convenience would then break the recovery it exists to help. It is guarded on its own now: the correct way for this to fail is an offer nobody gets. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
39976041e0 |
fix(sessions): let a dismiss reach the entries a restore is holding
Fourth review of the reboot-restore branch, and the third to find a defect in the previous round's fix. This one is the same shape as its predecessor: a counter keyed on one thing, compared against a set keyed on another. The generation counter was indexed by the entry's owner, while the in-flight set holds the caller doing the restoring. Those are the same person exactly when a user restores their own sessions, which is every case the tests covered. The route deliberately supports the other case: an admin may spend another user's entries. So when an admin restored Bob's sessions and Bob dismissed the banner, nothing matched, the entries came back, and a plan Bob had explicitly dismissed was re-armed for another twenty-four hours. Rather than reconcile the two key spaces, the counter is gone. `take()` now parks the entries it hands out, remembering which caller is spending them, and they stay parked until that restore ends. A dismiss filters the parked entries by `canAccess(entry.owner)` — the same predicate it already applies to the plan — so it reaches them wherever they are. `releaseFlight()` puts back only what is still parked. Expiry and a fresh boot plan unpark everything, for the same reason. There is one key space now, the entry's owner, and the spender is only ever used to tell two concurrent flights apart. That removes `generations`, `snapshotGenerations()`, `bump()`, `bumpAll()` and the argument threaded through the route. The discard grew the teardown it still lacked. A rebuild can fail after startInteractive() resolved, and a restored workspace still carries Codeman's hooks, so the CLI can post a hook event within milliseconds; the transcript watcher that starts from it, the attachment registry, the wait registry and the approvals inbox all outlive the listeners and would meet the retry, which reuses the session id by design. Its steps also run in reverse order now, so no live listener can reach a tracker that has already stopped, and the mux kill has its own guard, because stop() kills the pane in its last block after destroying four trackers. Tests. The run-summary test named an interval and asserted a map entry, so dropping stop() left it green; it now spies on stop(). Nothing pinned that before-spawn must precede setupSessionListeners, which reads the flag that phase restores, so swapping the two lines was silent; the ordering test now includes the listener setup. The retry assertion was a tautology and now asserts a different refs object. Both strengthened tests were verified by reverting their fix. Two new tests cover the admin-restores-another-owner cases this round was about. The server in the discard test is built once and stopped, since its constructor registers handlers on module-level watchers, and the workspace is removed through safeRmHomeTree. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
71ed7b127c |
fix(sessions): make the discard a real inverse of the construction
Third review of the reboot-restore branch. The narrow discard the previous commit introduced avoided everything cleanupSession() did wrongly, and in dropping so much of it also dropped four things it had to keep. The worst broke the retry the whole design rests on. setupSessionListeners() returns early while sessionListenerRefs still holds the session id, and the discard never cleared that entry. So the advertised flow — a rebuild fails because the agent binary is missing, the user fixes their PATH and clicks again — reused the same id, wired no listeners at all, and produced a tab that never showed output, never updated its status and never persisted. That is worse than the leak the discard was added to prevent. Three more registrations leaked with it: a RunSummaryTracker and its interval, an image watcher on the workspace, and the Ralph fix-plan watcher. The discard now undoes each registration setupSessionListeners() makes, in its order, and the per-session custom-model config directory, which holds the endpoint's API key literally and which nothing else would ever remove. The image-watcher flag was restored after the code that reads it, so a session came back reporting the feature as on with nothing watching. It moves to the before-spawn phase, and that phase now runs before the listeners rather than after them. The generation counter that lets a mid-restore dismiss win was global while clear() is ownership-scoped, so one user's dismiss discarded another user's unspent entries, permanently, because nothing rebuilds an in-memory plan. It is now per owner. Bumping only the owners of entries the dismiss removed was not enough either: take() has already emptied the plan by then, so a dismiss landing mid-restore saw nothing of that owner's to remove and invalidated nothing. The owners that matter are those with a restore in flight, filtered by what the dismissing user may access, and that is what clear() now bumps. Plan expiry bumps too, so a restore straddling the 24-hour boundary cannot hand entries back and give an expired plan another full day. Tests. discardPartiallyBuiltSession had no test at all: the only implementation any test ran was the mock's one-line stub, which is why every defect above was invisible. test/discard-partially-built-session.ts drives the real WebServer, and the retry assertion fails if the listener refs are left behind — verified by reverting the fix. The dismiss-race test drove the registry by hand, so deleting the route's generation argument left it green; it now goes through the route, and two further tests cover the multi-user cases. The mock context has now gone stale twice, because route tests pass it as `ctx as never` and tsconfig.json includes only src, so nothing ever compares it to the ports. A type-level guard is therefore inert — I wrote one and confirmed it never fires. test/mocks/mock-route-context-completeness.ts compares the mock's keys against WebServer.createRouteContext() at runtime instead, and names what is missing. Also: the API reference now says workspace-forbidden is judged against the owner's grant, the banner's module header no longer claims Restore always dismisses it, and the detail span gets the same min-width: 0 the phone rule already needed. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fa52753e8b |
fix(sessions): undo a failed rebuild without deleting the user's data
A second review of the previous commit found that its own repair for the session leak introduced three defects, all from reaching for cleanupSession() to undo a half-built session. That function is the user-initiated delete, not an undo. It banked the session's historical token and cost totals into the lifetime figures, and a reboot never runs cleanup, so those totals had never been counted before; every failed rebuild added them again. It saw the pin that had just been restored and demoted the record to `stopped`, which this pass reads as the durable marker of a deliberate kill, so a pinned session whose rebuild failed became permanently unrestorable. And it recursively removed `.claude-images` from the working directory, which belongs to the workspace rather than to the session, so a failed rebuild destroyed the pasted images of any other live session in that repo. discardPartiallyBuiltSession() now undoes only what the construction did: the map entry, the tab-layout slot, the listeners and any pane the launch created before throwing. The persisted record, the lifetime totals, the Ralph state and the workspace's files are left alone. Re-applying the persisted state also splits in two, which removes the first two defects at the root rather than only at the call site. The half that shapes the pane, the custom-model environment and the nice priority, still runs before the spawn. The half that is the session's own history now runs after it, so a session whose pane never started carries no totals and no pin for anything downstream to misread. The rest of that review. The multi-user workspace confinement re-check read the requesting user's grant, and returns true for an admin, so the case its own comment described was the one it missed; it now resolves the entry owner's grant through isWorkingDirAllowedForUsername, the way cron does. A forbidden workspace goes back on offer, matching both the registry's stated contract and the API reference. The client re-reads the plan after a restore instead of blanking the banner, so entries the server put back stay reachable, and a 409 now says a restore is already running rather than reporting a failure. A dismiss arriving mid-restore wins, through a generation counter the route carries across its take. The re-application also restores the tab colour, the image-watcher flag and the original pinnedAt, via a new Session.restorePin that does not re-stamp the pin time. The phone breakpoint gains min-width: 0, without which a nowrap flex item never shrinks and the buttons still overflow, and it folds into the existing phone block. Ralph's loop configuration still does not survive a restore, because toState() reads it off a live tracker and there is no way to keep it without arming the loop. The method now says so rather than leaving it implied. Tests. The capacity test could not fail on the property it existed for: it filled the board past the cap before the loop, so a single pre-loop check would have passed it. It now leaves one seat, so only a per-iteration check restores exactly one entry. New tests cover the ordering around the spawn, a throw before the loop returning the whole plan and releasing the flight, the dismiss-during-restore race, and that the failure path calls the narrow discard rather than the delete. The shared mock context gains the port method it was missing, which is what made the first run of these tests fail for the wrong reason. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fbede5cd2a |
fix(sessions): act on the dual review of the reboot-restore route
Fifteen findings from two independent reviews of #442, three of them blocking. Every one is addressed here. The three blockers all sat in the restore route. A rebuild that threw after addSession left a registered session with no pane behind it, visible on the board, holding a layout slot and written to state.json, with its plan entry already spent; the catch now cleans the session up and puts the entry back. The loop checked neither the global nor the per-user session cap, so one click could take a board past a documented limit; capacity is now re-checked per iteration, because the loop is itself creating the sessions it counts. Worst of the three, a rebuilt session carried none of the state its constructor has no parameter for and then persisted itself over the record that held it, zeroing token and cost totals and dropping the pin. The pin matters most: pruning keeps a record only while it is pinned, so discarding it handed the record to the next stale sweep. A new reapplyPersistedSessionState() on the session port restores the pin, the token totals, auto-compact, auto-clear, auto-resume, nice priority, the flicker filter and the custom-model selection, and it runs before both startInteractive and the first persist. The rest, in the order they bite a user. Every rebuild failure was reported as workspace-missing, so the banner told users their repo was gone when the agent had simply failed to start; there are now distinct reasons, and the toast names each one. The client read restored and skipped off the outer response object rather than through the uniform envelope, so every count came back zero and neither toast ever fired. A board left open across the reboot never learned an offer existed, because the banner was seeded only on the page-load path; it now re-reads on every SSE init. The workspace check was existence-only, skipping the multi-user confinement that the create route applies, so a withdrawn grant would not be noticed. The banner had no phone breakpoint while its text was nowrap and its buttons could not shrink. Smaller: a missing workspace is now re-offered rather than dropped, while an already-open conversation is dropped rather than re-offered forever; a throw anywhere in the route returns the unspent entries instead of discarding the plan; the single flight is keyed by owner, since take() already stops two callers receiving one entry; the env clamp's header no longer claims a protection it cannot provide on this path today, and names the check that does bite; the three endpoints are documented in docs/api-reference.md; and the module header now says that os.uptime() reads the host's clock, so the feature is effectively off inside a container. The review also explained why the tests missed all of this: they proved the construction claim through their own copy of the construction rather than through the route, and the route tests used workspaces that did not exist, so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts mocks the Session module to drive the route's real path, and covers the cleanup, the reason reported, the re-application ordering, the broadcast and the caps. The mock route context gains the port method and the mux call the route needs. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
da933d70be |
feat(sessions): offer to rebuild the sessions a host reboot destroyed
A host reboot takes the tmux server down with it, so every pane dies, reconciliation finds nothing to attach to, and the board comes up empty. Picking yesterday's work back up meant finding each conversation in history and resuming it by hand, one at a time. The boot pass now works out what the reboot killed and leaves it on offer. It runs inside restoreMuxSessions(), in the window where reconciliation has reported the dead sessions and cleanupStaleSessions() has not pruned their records yet, which is the only place the records can still be read. The board shows a banner, and nothing is created until the user clicks it. A click rather than an automatic restore is what makes the reboot heuristic acceptable. The heuristic cannot tell a reboot from a crash that took tmux down inside the same window, so it decides whether to ASK, never whether to act: a wrong yes costs a line of text the user dismisses instead of N CLI processes nobody asked for. Four things are re-checked when the click arrives rather than trusted from boot, because hours can pass and the board moves on. The owner's privilege grant re-resolves through the env clamp. The workspace must still be on disk. A conversation the user already resumed by hand from the Resume list is skipped, since two panes running --resume on one conversation would fight over the same transcript. Entries leave the plan synchronously before the first await, and the route is single-flighted, so a double-click or two devices cannot both reach the same entry. A restored session comes back attached, idle and disarmed. Respawn controllers and Ralph loops are deliberately not re-armed: a machine that just came up is the worst moment to turn an autonomous run loose. Its workspace hooks are installed by the restore route itself, because the boot-time sweep sits behind a gate that is false after a reboot and has finished long before the click; without them a session goes silently blind, with no stop or idle events for respawn, no Approvals Inbox item and no red tab on a blocking dialog. Stats collection starts the same way. The pane is new, so the conversation continues and the terminal scrollback does not. The banner says so rather than letting an empty pane read as a broken restore. The plan lives in memory only. A server restart drops it, which costs the convenience this adds and never the conversation: the conversation is the transcript under ~/.claude/projects, which the Welcome screen's Resume list and the Session Manager already read, so a dropped plan returns the user to resuming by hand. clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the question it answers is about session privilege rather than about HTTP and it now has a caller outside the route layer. Its test hook stays re-exported from session-routes.ts. Claude sessions only for this pass. The other CLIs name their thread in their own config object, which this does not thread through yet. Remote and docker sessions are skipped on purpose, because both need another host or a container to be up and a freshly booted machine cannot promise either. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5b920cb43d |
feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream and named tabs wrong five ways (every prompt, every write path, a bare Esc eating the next prompt's first character, pasted newlines as Enter, any CSI clearing the draft) and replaced the whole name, which dropped the case from the tab and reset the w<n> counter. This lands the feature with each of those closed: - First prompt means the first: applyAutoName() flips a placeholder to `auto` whether or not the string changed. nameSource is now the tri-state placeholder | auto | manual; the name setter is the only manual path. - Only user-originated input counts: write()/writeViaMux() take SessionWriteOptions.fromUser, set by the browser WS path and POST /input only, so Ralph, respawn, cron, approvals and the trust-dialog keys can never name a tab. A startMode 'shell' CLI never feeds the tracker (a capability, not an id check); the send-key route feeds trackUserInput() because its line feed bypasses the session. - Prefix form `w3-case: title`: parseSessionPrefix() already renders it as the title with the prefix in the tooltip and the next-session counter still matches it. Composed within MAX_SESSION_NAME_LENGTH. - Tracker rules per key: bare Esc resolves at chunk end; mouse/focus reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R taint the draft so Enter submits nothing rather than a fragment; bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space; the draft keeps its head past 8192 code points; an escape past 64 bytes is abandoned. - Title: slash commands by shape (a path is a prompt), `!` escapes refused, first sentence only past 8 code points ("e.g." is not a title), 72 code points on a word boundary. - Synced `autoNameSessions` setting, default OFF (the prompt reaches mux-sessions.json, session:updated and /api/search), App Settings -> Appearance -> Tabs, read fresh per prompt after the eligibility check. Tests: test/session-auto-name.test.ts (tracker, title, composition, ownership, emit gating), the wiring test (once, prefix, setting off, manual protected), test/routes/session-name-routes.test.ts (PUT /name flips to manual and persists). Verified live on an isolated instance: API and browser-typed prompts name the tab, a second prompt does not, shells and renamed tabs are untouched, nameSource survives a restart. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
c4322513d9 | Merge pull request #376 from shenlvkang-collab/feat/auto-session-names-upstream | ||
|
|
942bf37e48 |
fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a custom OpenAI-compatible endpoint by injecting env vars or a config file and restarting the CLI in place. Review of the apply path found four things, two of them destructive. This lands all four plus the smaller items from the same review. 1. Clearing a selection did not clear it. The injected vars reach the CLI via `tmux setenv`, which persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so deleting the keys from the session's envOverrides relaunched the CLI still pointed at the old endpoint, and for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just been deleted. `Session.setCustomModel()` now reports the removed keys, queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys` carries them into `applyEnvOverrides()`, which `setenv -u`s them before re-applying the live overrides, on the same path that already unsets the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket that `setenv -u HOME` hands the next respawn the global HOME back. 2. Applying a model to a local claude session killed the pane. The relaunch was `claude --session-id <id>` and Claude refuses an id that already has a transcript, and unlike the dead-pane respawn this one kills a working pane first. `restartCli()` now pins the live conversation id as the resume id for that respawn when the CLI's launch declares a `fallback` chain, which renders the same `--resume <id> || --session-id <id>` shape the docker and remote pane commands use. Gated on the registry shape, not the CLI id: an entry whose resume id is minted by the CLI itself never declares that chain. 3. pi, omp and grok wrote their config file and then launched without the `--model` that selects it, so the file was ignored. The registry entry now declares `customModelInjection.launchModel` (`custom/{modelId}` for pi and omp, grok's `[model.codeman-custom]` block name), the builder renders it, and `_withCustomModelLaunchModel()` applies it onto the respawn options through `legacyConfigField`, leaving the stored <Mode>Config untouched so a clear falls back to the user's own model. A model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. 4. Remote (SSH) and Docker sessions reported `restarted: true` and changed nothing: their `restartCli()` reattaches the durable tmux rather than relaunching the agent, and the env lands on the local pane. Both are refused with a 400 until those paths are plumbed. Smaller items from the same review: - The selection survives a Codeman restart as the disk-only `__customModel` bookkeeping (endpoint, model, injected key NAMES, config dir, launch model; never the values, which carry the API key). Recovery re-derives the values from the endpoint store through the same apply path the route uses and keeps the bookkeeping even when the endpoint is gone, so a later clear still has keys to unset. - Discovery goes through `webviewFetch()`, so the RESOLVED address is judged by the same egress guard the web-tab proxy uses, and `baseUrl` reuses `webviewUrlSchema` (http(s) only, no embedded credentials, link-local and cloud-metadata addresses refused). undici's `fetch failed` wrapper is unwrapped so the user sees the ECONNREFUSED underneath. - `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session config dir 0700/0600 (pi and omp embed the key literally), and that dir is removed with the session. - `PR.md` is gone from the repo root and the design doc moved to `docs/custom-model-endpoints-plan.md` with the LAN address and the personal name scrubbed; every reference follows. The guide's `authStyle` text matches the shipped schema (`bearer | api-key`, default `bearer`) and says that `customModelEndpointsEnabled` is read by nothing until the picker lands. - `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts` (four real type errors fixed). It is not yet wired into `npm run typecheck` because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json` there is the one-line follow-up. Tests: `test/session-custom-model-restart.test.ts` drives a real Session and fails on the unfixed code for items 1 to 3; the route suite covers item 4 and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets run before the overrides and that a shell-metachar key never reaches tmux. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1e42cb4e2d |
Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses) |
||
|
|
792a251e35 |
Merge pull request #421 from Randalix/fix/remote-file-access
fix(files): read remote-case previews, downloads and attachments over ssh |
||
|
|
1306f731cf |
fix(webview): recover a proxied dashboard that reloads on its landing page
The runtime shim masks `/webview/<cap>/` off a proxied page's URL so its router boots on the path it expects, and the landing page masks to exactly `/`. A `location.reload()` there (a Vite dev server on a config change or a failed HMR update, the likeliest case in the feature's own motivating scenario) therefore asks for Codeman's root as an iframe navigation. `serveLostWebviewFrame()` returned early for `/`, so on a passwordless install the frame received Codeman's own app shell and rendered it inside the web tab, and with a password it got a 401 in the frame. Either way no `codeman:webview-lost` message was posted, and because the document loaded fine the load handler cleared the failed-frame panel, so the Reload / Open in new tab affordances never appeared. Before masking the frame's URL was the prefixed one, so a reload worked; this was a regression. `/` is the one lost-frame path a registered route also serves, so the route table cannot tell that reload from a real navigation. Credentials can: nothing in Codeman frames its own root, and a sandboxed frame is opaque-origin with no cookie and no Authorization header. `carriesAuthCredentials()` (pure, in webview-proxy.ts) makes that test, and `/` is now admitted by the auth hook only when it fails; a framed `/` that does carry credentials still gets the shell. Without a password no auth hook runs at all, so the index route applies the same test itself (`isLostWebviewRootFrame`) before rendering the shell, and the three places that emitted the recovery page share `sendLostWebviewFramePage()`. Tests: the password form in webview-auth-exemption (recovery page for a credential-free framed `/`, shell with valid Basic auth, 401 with a stale cookie or a top-level navigation), the passwordless form against a real WebServer in webview-lost-root-frame (port 3198), and the credential predicate in webview-proxy. All three fail without the fix. Verified against a live isolated instance as well: a framed `GET /` with no credentials answers the 470-byte recovery page, a top-level `GET /` and a framed one carrying a cookie answer the shell. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
b0dddc9c57 |
Merge pull request #402 from shenlvkang-collab/pr/webview-route-masking
fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads |
||
|
|
63aafdf274 |
fix(files): serve remote-case attachments, the path a click takes outside the case
A clicked path that points OUTSIDE the case directory goes through the attachment routes (the frontend's `_isExternalPreviewPath` sends every absolute path not under `workingDir` to `POST /attachments`), and those had the same local-`fs` assumption as file-raw: `realpathSync`/`fs.stat` on a path that only exists on the remote host, so the file never opened — the case the #415 report was actually about. - `registerExternalAttachment()` accepts `remote` and resolves through `remoteProbePaths` (canonical path, size/mtime, kind, plus the workspace root for the confinement check). Everything around it — blocklist, extension allowlist, workspace confinement, registry/dedupe — is now shared by both branches, so the remote path cannot drift from the local one. - The by-id routes (`raw`, `preview`, `thumbnail`), the metadata poll and the attachment history list resolve over ssh too. `raw` streams with the same Range contract as file-raw; `preview` (office) and `thumbnail` answer 400 for a remote record; an unreachable host answers 502, a vanished file 404. - Which host a record is read from follows the SESSION, never the path string: the same absolute path is a different file on each host, and a remote session never falls back to a local file with that name. - Codex generated artifacts keep force-workspace confinement for a remote case: the well-known artifact directories are anchored at THIS host's home, so only a file inside the remote workspace is trusted. Still local-only by design: writes, office conversion, thumbnails, the file tree/picker and tail-file. |
||
|
|
41416566aa |
feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. Covers local hardware (llama.cpp, Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter). Off by default (customModelEndpointsEnabled, synced, default OFF). - Registry: capabilities.customModelInjection per CLI entry (env / configContentEnv / configDir / unsupported kinds) - Pure injection builder (custom-model-injection.ts) turning an endpoint + model id into the real env vars / config content per CLI - Endpoint store + CRUD routes (custom-model-hosts.ts, custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded - Session integration: Session.setCustomModel()/restartCli() (POST /api/sessions/:id/custom-model), reusing the existing respawn-pane -k primitive to restart the CLI process with new env - Multi-user hardening: every new redirect-capable env var added to its CLI's privilegedEnvKeys, closing a pre-existing gap where several were already reachable via the generic envOverrides field's prefix allowlist - Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries against a real endpoint outside the web UI, independent of tmux/sessions - Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying every CLI's injected values through a real HTTP shape Real end-to-end validation against a live llama-swap server (inside a codeman/agent:llm-test Docker image with all 9 CLI binaries) found and fixed three real bugs before they shipped: - Codex's config.toml schema was wrong ([model].default table instead of a top-level model string + [model_providers.custom]); fixing it then surfaced a genuine, documented protocol incompatibility (Codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement) - Claude Code's async session-title-generation call validates ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and hangs the whole -p invocation on an unrecognized name; documented for chunk 6, worked around in the standalone script only (--bare is NOT safe for a real interactive session, which needs hooks) - The discovery route's authStyle: 'both' option (send both Authorization and api-key headers) reliably hung a real server; removed the option entirely rather than just changing the default Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see PR.md and deployment_plan.md for the full chunk breakdown and confidence table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3 |
||
|
|
349a89ec3b |
fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads
A dashboard served through a web tab saw `/webview/<cap>/` as its
`location.pathname`, and no app has a route for that: a React Router, Vue
Router or Vite dev-server page painted its HTML and CSS and then replaced
them with its own "page not found" the moment its script ran (reproduced
with a minimal history-routed page).
The proxy's runtime shim now rewrites the history entry to the path the
page would see on its own origin, before any page script runs. The base
element still resolves relative URLs inside the prefix and every root-
absolute sink is rewritten back into it, so only what the page READS
changes. With the document URL masked the Referer-keyed 404 rescue can no
longer help a request the shim misses, so the remaining URL-taking entry
points (`Worker`, `SharedWorker`, `navigator.sendBeacon`, `window.open`)
are covered by the shim as well.
A navigation the page starts itself afterwards — `location.reload()`
(a dev server's full-reload HMR), a root-absolute `location.href` — lands
on Codeman's root with no capability anywhere: no prefix in the path, no
cookie in an opaque-origin frame, a Referer naming the masked page. It is
recognised by shape (a top-level iframe navigation asking for HTML, for a
path Codeman does not serve) and answered with a static page whose only
script posts `{type:'codeman:webview-lost', path}` to the parent; the tab
that owns the frame (matched by `event.source`, never by the payload)
remounts it inside the prefix at that path, bounded per frame. The
unauthenticated form is answered in the auth middleware before the
credential checks, so a dev server that reloads on every save cannot
rate-limit its own user out of Codeman; the authenticated form (Basic
auth, trusted mode) is answered by the 404 handler.
Verified end to end against a history-routed page: boots on `/`, its
API call succeeds, a reload inside the frame comes back routed on the
path it had pushed, `location.href = '/about'` comes back on `/about`,
and a deep link opens on its path.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
d5b75af628 |
fix(statusline): sticky telemetry collection, footer print-through, EOF fix
Responds to Ark0N's review round on the ephemeral-CLI-flag statusline injection rework: - Rebase-detail fixes: registry-gated telemetry eligibility via getCli(mode)?.capabilities.statusLineTelemetry instead of a hardcoded mode === 'claude' check, using the capability flag master's CLI-registry refactor already declares for exactly this purpose. - Design question settled: sticky (a). Rather than persisting the toggle as a new field and threading it through every session-creation path (cron, Ralph Loop API, quick-start), eliminated the per-session field entirely. readPlanUsageTelemetryEnabled() (hooks-config.ts) reads the existing showPlanUsageLimits setting fresh from settings.json at every claude create/respawn (TmuxManager.createSession/respawnPane) - no per-session state to survive a restart, and it applies uniformly to every creation path for free, since they all flow through the same TmuxManager methods. This required fixing a real bug found along the way: showPlanUsageLimits was not actually round-tripping through settings.json on save - settings-ui.js explicitly excluded it from the PUT body as a pure per-device display key. It now flows through normally (both true and false); the load-side per-device merge behavior is unchanged. Removed entirely as a result: the statusLineTelemetry field from CreateSessionSchema/SettingsUpdateSchema, CreateSessionOptions/ RespawnPaneOptions, Session._statusLineTelemetry (this is what makes the restart-persistence bug moot rather than patched), and the frontend send sites. - Footer print-through restored: the no-user-statusline branch of the exporter script now runs the telemetry POST in the foreground so its own stdout becomes the in-terminal footer, falling back to a plain "codeman" marker only on curl failure. - Background-subshell EOF fix: the wrap-a-real-statusline branch closes stdin too, not just stdout/stderr (`>/dev/null 2>&1 </dev/null &`) - the un-redirected subshell process itself, not curl, was what held a reader-to-EOF's pipe open for however long curl took to finish. Added curl --max-time 5 so a hung (not just refused) Codeman cannot wedge the render. Tests: real-shell-execution tests for the footer/EOF fixes (fake curl stand-in on PATH, real sh subprocess spawns, real elapsed-time measurements - verified non-vacuous against a hand-reconstructed old-style script), unit tests for readPlanUsageTelemetryEnabled. Adapted two existing tests whose payloads referenced the removed field. Fixed during independent code review: a stray indentation break and a test exercising the wrong (legacy) exporter code path. Docs synced: CLAUDE.md, docs/usage-limits-display-plan.md (old disk-based section marked superseded, kept for history), docs/architecture-invariants.md. Full suite green: 352 files, 6780 passed, 12 skipped, 0 failed. tsc/lint/format:check/frontend-syntax all clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8ee7926e27 |
feat(agent-cases): tag agent-spawned case dirs and sweep their leftovers
A long orchestration creates one case directory per worker and deleting the sessions never removed them, so ~/codeman-cases accumulated scratch folders that were indistinguishable from real projects. They are now labelled and have a cleanup path. - src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven spawn gets a .codeman-agent-case.json marker (when, by whom, parent session, mode). Only the create branch writes it, so a linked case, a cloned repo or any pre-existing path is never labelled; reading is total, so a malformed marker means "not agent-created" rather than a half-trusted entry. - The signal is the new X-Codeman-Agent-Origin header the skill preamble sets on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body field, falling back to a resolved parentSessionId so a worker spawned by a stale skill copy is still labelled. - GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage badges each case and offers a review-then-delete sweep that names every directory in its confirm and skips any case a live session is working in. Removal stays on the existing DELETE /api/cases/:name. - Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was written per claude session and never removed (236 leftovers measured on a working machine). Now deleted with the session and swept at boot, guarded by a live-session keep set plus a 7-day age floor. Verified end to end on an isolated instance: marker written for header, body and lineage-only spawns, absent with no agent signal and for a pre-existing directory; inUse flipping on session end; badge, sticky bar, confirm and sweep driven in a browser; preamble seeded on create, removed on delete, boot sweep taking only the aged orphans. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cfc8fe7e41 |
Merge pull request #381 from mtiller/feat/reverse-proxy-base-url
feat(web): support a reverse-proxy base URL |
||
|
|
7e4914d991 |
feat(web): support a reverse-proxy base URL (--base-url / CODEMAN_BASE_URL)
Codeman can now be mounted under a sub-path behind a reverse proxy that forwards the prefix unchanged (e.g. https://host/codeman/). Default is `/` (root), which is byte-identical to the historical behavior. Design — few choke points, mirrored ingress/egress: - src/config/base-path.ts: pure single-source normalize/validate/join/strip. - Server ingress: stripBasePath() inside Fastify rewriteUrl, so routes stay declared prefix-agnostic; un-prefixed requests (hooks, health, docker bridge hitting the raw port) pass through unchanged. - Server egress: one onSend hook prepends the base to root-absolute Location headers (covers all redirects). - HTML: renderIndexHtml points <base href> at the mount and injects window.__CODEMAN_BASE__ — ONLY when a base is set (inert at root). - Frontend runtime URLs: CodemanBase.url() route builder in constants.js, applied transparently by a fetch wrapper and explicitly at the EventSource/WebSocket/window.open/<img|iframe|a>-src sites. - sw.js derives its base from self.location; manifest uses relative start_url/scope. - Web-tab proxy: proxyPrefixFor(cap, basePath) is the single base-aware root that cascades to the injected <base>, HTML/attr rewrites, runtimeUrlShim, Set-Cookie Path and Location; capabilityFromReferer strips the base off the browser Referer, while the ingress parsers stay base-agnostic (rewriteUrl already stripped it). --base-url rides the daemon relaunch (buildWebArgs) and the service unit (resolveServicePlan). constants.js is guarded against a missing `window` for isolated unit-test contexts. Tests: test/base-path.test.ts (pure helpers), base-path coverage in webview-proxy/render-index-html/daemon-control; CodemanBase stubbed in the vm-isolated panels-ui test contexts. Docs: Remote-Access.md (sub-path section + nginx example), security-architecture.md env table, CLAUDE.md pattern. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XUkPBxbumnct6qSrx4JDju |
||
|
|
268e4819ff | feat: auto-name sessions from first prompt | ||
|
|
ccfda623fe |
fix(session): learn the live Claude conversation from the CLI's own hook
Which conversation a pane is on was re-derived by correlating ~/.claude/history.jsonl against Session.lastSubmitAt — and lastSubmitAt is bumped only by input that flows through Codeman's own write path (Session.write / writeViaMux). A user who attaches to the pane's tmux session directly never set it, so resolveActiveClaudeSessionIdFromHistory() returned at its first line for that pane's whole life and the response viewer stayed pinned to the launch conversation, showing a pre-/clear transcript indefinitely. A UserPromptSubmit hook reports the live conversation id from inside the CLI process, delivered under the pane's own $CODEMAN_SESSION_ID. That binding is a fact rather than a correlation: it never consults workingDir, so it cannot be claimed by a sibling pane on the same folder, a closed tab, or a bare `claude` in the user's terminal. A pane holding such an id skips the correlation entirely, so the number of prompts eligible for cwd-based guessing goes DOWN, never up — the naive alternative (relax the guard, or synthesize an anchor from PTY activity) is the reverted bug the resolver's own comment describes. The hook also stamps lastSubmitAt, so it finally means "a prompt was submitted" rather than "typed into Codeman's web terminal". Conversations vouched for first-hand — and only those — extend a persisted claudeSessionChain, whose tail re-pins the conversation when a surviving tmux session is re-attached after a restart. ⚠️ start() resets the id at THREE points and the last one runs unconditionally after the mux branch, so the tail is applied there too; patching only the mux branch looks right and silently does nothing. ⚠️ The hook's stdout is discarded with curl's own -o /dev/null. Claude Code injects a UserPromptSubmit hook's stdout into the model's context ("Exit code 0 - stdout shown to Claude"), and a trailing >/dev/null does NOT work: curlCmd already ends `... 2>/dev/null || true`, and in `pipeline || true >/dev/null` the shell binds the redirection to `true`, which never runs on the success path. The discard is opt-in so the five SSE-fed events keep byte-identical command text and no workspace's settings file is rewritten for them. The staleness marker is quote-free for the matching reason: hooksJson is JSON.stringify'd, so a quoted needle never matches and the gate would rewrite every workspace on every spawn. Existing workspaces heal on their next Claude spawn through the staleness sweep. |
||
|
|
da91b4353b |
Merge pull request #353 from timkjr/omp-mode
feat: add OMP (Oh My Pi) as a new session backend |
||
|
|
54a930c80e |
feat(omp): survive a full session kill by reading omp's own transcripts
Claude conversations survive "Kill Tmux & Claude" because Codeman reads them back independently from ~/.claude/projects, not from its own session bookkeeping. omp conversations had no equivalent: kill the Codeman session and the conversation vanished from Past Sessions entirely, even though omp itself never forgot it on disk. Adds omp-transcript.ts, a scanner over omp's own ~/.omp/agent/sessions/<mangled-cwd>/<uuid>.jsonl files (the same shape as Claude Code's own transcript scanner, but simpler -- these files are small enough to read whole instead of doing head/tail windows). Each file's own "session" header line carries the real cwd and session id directly, so unlike Claude's mangled-directory-name decoding this never has to guess. Wired into gatherUnifiedInputs() as a second history source alongside the Claude scan, and HistoryInput/ mergeUnifiedSessions() now carry an optional `mode` so a non-claude history-only row still gets a real mode badge. Also fixes the ambiguity behind the "continue picks the wrong conversation" report from this session's testing: omp mints its OWN session uuid, unrelated to Codeman's, so a live/persisted row and its own history-scan row would otherwise show up as two separate entries for the same conversation the moment the id gets resolved. Reuses the existing claudeSessionId alias field (mergeUnifiedSessions' fold-into- owner mechanism) to point at the resolved omp id, threading it through every place `_claudeSessionId` gets (re)computed -- the constructor, _resolvedOmpRespawnConfig, and a new _maybeCaptureOmpSessionId() that opportunistically resolves it the first time a brand-new omp session (one that has never gone through a respawn) goes idle. Also closes a THIRD instance of the "ompConfig never got wired in here" gap this session kept finding: restoreMuxSessions() in server.ts restores every sibling CLI's config from persisted state on boot except omp's, so a boot-recovered omp session always lost its resolved resume id and fell back to guessing again. Verified live end-to-end: told a session a secret, killed it fully (Kill Tmux equivalent, killMux=true -- the Codeman session AND its tmux pane both gone), and the conversation still showed up in the unified list as a history-sourced row with the real first prompt as its title and an omp mode badge, keyed by omp's own session id. Known remaining gap, not fixed here: the claudeSessionId alias doesn't yet resolve reliably on every boot-recovery path for a session that was never respawned while alive (e.g. a plain re-attach to a pane that was never dead) -- worth a follow-up, but doesn't affect the two things that matter most: the conversation surviving a kill, and continuation correctness once an id has been resolved (which happens on the very next respawn either way). |
||
|
|
4f5678fac4 | feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) | ||
|
|
b00ab3ceea | feat(web): show Codex plan usage in header | ||
|
|
c30dfaf0e7 |
fix(deepseek): run the web UI server in the background, not in a shell tab
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a shell tab running the server next to it. The shell was deliberate - the server lived in an ordinary session so it was visible, scrollable, killable and died with its tab, and nothing new had to supervise a long-lived HTTP server. That reasoning was sound and the result was still wrong in use: opening a dashboard should open one tab, and after the first launch the terminal is pure noise. The server moves to a background child process owned by a new `src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session gave away for free is now explicit, which is most of the module: - Exactly one server. A second click reuses the running one instead of racing it for a port; the session flow could not do this at all, because two clicks were simply two sessions. - Restarted when the requested authority changes. `--trusted-host` fences dsh's own /api against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two. Reusing a server fenced for the other origin renders a page whose every call 403s, which reads as a broken dashboard rather than a misconfigured one, so a mismatch restarts instead. - Killed on shutdown. The child is detached so its whole plugin tree can be signalled at once, which also means it would outlive Codeman and hold its port against the next start - the exact EADDRINUSE this feature already got wrong once. - Boot output captured and returned. With no shell tab there is nowhere else for a stack trace to land, so a failed spawn reports its own tail. The endpoint is fenced at the same bar as the profile installer and for the same reason: booting a dsh profile executes the plugin code in it, so this is a privileged action even though it reads as "open a page". `authority` comes from the client (`location.host`) because only the browser knows which origin is in play, and it is regex-confined at the schema boundary - defence in depth behind the argv-array spawn, admitting host:port in the shapes a browser authority can take and nothing readable as a second argument. `GET /api/deepseek/web-port` is gone; port selection moved into the supervisor, which is the thing that knows whether a server is already running. The two client-side probe helpers went with it, since the server now owns the wait. Verified over the tailnet authority end to end: no session is created (session count unchanged, one tab), the server runs on 3081 beside the user's own dsh web on 3080, status reports the tailnet authority, and the proxied dashboard renders with zero 4xx. Full gate green (6148 passed, +6). |
||
|
|
2034719d61 |
fix(deepseek): close the env-var clamp hole, bound the profile install, make the hook gate per-session
Three review findings on the DeepSeek Harness mode, plus one the third exposed. 1. The multi-user clamp was bypassable by a sibling field on the same request. clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER _configureDeepSeek(), so a non-granted owner sending envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated instance: a session created with permissionMode "read-only" and that override ran with DSH_PERMISSION_MODE=danger-full-access in its pane. Every other CLI's bypass is a command-line flag reachable only through the per-CLI config, which is why the config clamp alone is the whole gate for them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on that list because it aims the launcher at a profile tree whose plugin code runs at boot, before any approval row can apply. Verified end to end in real multi-user mode: a non-granted user sending both now gets workspace-write and no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched. 2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout` signals only the direct child, and a plugin install fans out into package-manager children that keep the inherited stdio pipes open, so `close` never fires and the held-open request leaks with no route-level deadline. Reproduced: with a 1.5s built-in timeout the promise was still unsettled after 6s and both fan-out children were alive. Now detached: true plus negative-pid SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a last-resort reap for a grandchild that escaped the group. Same probe after the change: close fires, direct child and both grandchildren dead. 3. hooksAvailableForMode() promised more than a dsh session can deliver. deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that triple is the only reason a dsh session posts hook events, so `until=stop` was accepted and then blocked for the caller's whole timeout: the exact infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now takes HookCapabilityOptions and every call site passes sessionHookOptions(), with the deepseek arm reading `!== false` so a forgotten one degrades to the old behaviour. The refusal names the setting rather than saying "no Claude Code hooks", which would send the caller hunting a bug that is really a setting they chose. Profile conformance stays unknowable at request time and is documented as such. The stale "True for `claude` and nothing else" docblock is corrected. 4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent capture read Claude's own transcript, and adding deepseek silently widened both to a mode that has none. They compare mode === 'claude' directly now, and a static check pins them there. Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 -- bridge off plus explicit until=stop is a 400 naming the setting, bridge off with no `until` still 200s on idle/exit, bridge on accepts stop. |
||
|
|
4cda150493 |
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/ antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI as a Codeman web tab. DeepSeek is wired unlike its siblings in three ways, each of which is the reason for a design decision rather than an accident: 1. The agent is a PROFILE, not the binary. `dsh` is a launcher over $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and `base` -- the interactive terminal front door is always a third-party plugin. So availability is two questions: `isDeepSeekAvailable()` (binary) and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run button gates on the latter, because reporting only the binary would spawn a pane that dies on arrival. When the binary is present but no profile is, the run menu offers to install one (POST /api/deepseek/install-profile). 2. The permission switch is an env var, not a flag. The harness has no command-line permission option; its sandbox/approval rows read DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access). Exported via `tmux setenv`, never on the spawn line. Absent = the harness's own workspace-write, which still asks, so the multi-user clamp is the only-if-sent branch and clamps to workspace-write, never read-only. 3. It is the only non-claude mode that passes hooksAvailableForMode(), and it earned that. The terminal front door reports idle/working/blocked to a supervising process over a generic env-gated contract; a generated shim (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each report to /api/hook-event as stop / agent_working / permission_prompt. So a dsh session gets definitive respawn triggers, real wait-endpoint signals and real Approvals Inbox items instead of output-stabilization guesswork. `agent_working` is new (157th SSE constant) and joins APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its alert at once. The resolver needs the strictest identity probe of the family: `dsh` is not merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell), so `dsh --help` must print the harness's own banner before a candidate is handed a spawn line. Model is deliberately not a session field -- it is a composition entry in the profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider keys named by a settings-file `apiKeyEnv` stay out, which is pi's 34-provider-key problem in a new shape. Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the status endpoint's two-part answer, the no-profile refusal, the profile bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the permission mode injected via setenv, and the full status bridge -- a send-and-wait returned signal "stop" from a real turn, and blocked/working created and cleared an Approvals Inbox item. Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md (decisions + honest gaps). Tests: test/deepseek-mode.test.ts, test/deepseek-cli-resolver.test.ts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7a340fe7bc |
fix(tabs): review-driven hardening for the tab-layout foundation and vertical rail
Post-merge follow-ups from the deep review of #334 and #335, so they ship in the same release as the features. Tab-layout foundation (#335): - PUT /api/session-order drops unknown/foreign ids again instead of 400ing the whole write, in both the owner and the admin path (single-user requests are the synthetic admin, so that path is the one the browser hits). The frontend debounces its reorder push and swallows errors, so a session deleted inside the debounce window silently cost the user the entire reorder - and the endpoint sits on the stable /api/v1 surface, where the pre-layout server merged leniently. - A failed mux restore no longer locks explicit deletions into 500s for the process lifetime: runSessionDeletion and webviewDeleted degrade to best-effort without layout coordination, while the automated stale sweep (runStaleSessionCleanup) stays fail-closed. - sse-events doc comment: no 'suppressed' hook event exists; hooks stay 8. - registerSessionWithLayout resolves its owner through ownerLayoutKey() instead of a hardcoded '@single'. Vertical rail (#334) - all rail-awareness gaps in sidebar-only predicates, unified behind the new _isVerticalTabList() (sidebar OR rail): - Drag-reorder read the insertion side from clientX in the rail, so before/after was effectively arbitrary on vertical rows; the drag-over indicators now draw as top/bottom edges there like the sidebar's. - The active tab is scrolled into view in the rail (Alt+N/palette selection used to leave the row below the fold). - Floating subagent/ultracode windows anchor to the RIGHT of rail tabs, and the connector redraw gates (render tail + strip scroll) cover the rail. - Server-seeded tabOrientation is applied when the async settings load resolves, not only at boot, so a fresh device shows the rail immediately. - The pre-paint script stamps data-tab-orientation and --tab-rail-width (sidebar-wins and solo carve-outs included), removing the flash of the header strip on every vertical-mode load. - The session name font defaults to 12px, the sidebar's historical 0.75rem size, so installs that never touch the new slider are not restyled. Also documents the rail in CLAUDE.md (second #sessionTabs host, mover ordering, the axis-predicate rule) and gives tab-rail-resize.js its @dependency/@loadorder header. Full gate green (6093 tests); the excluded browser suite was run by hand - only the known environmental failures (opencode/codex binaries) remain, identical to pristine master. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c173ae0264 |
Merge remote-tracking branch 'origin/master' into worktree-grok-mode
# Conflicts: # src/web/public/app.js |
||
|
|
74194e4fc0 | feat(tabs): COD-359 add owner-scoped tab layouts | ||
|
|
3f8c8e99d1 |
feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
14a911b4f6 |
fix: color the server startup line and its security warning
The startup banner is now the only one (the CLI printed a duplicate) and is painted like the rest of the CLI. The non-loopback-without-password warning was plain console.warn while the CLI's copy of the same warning was yellow; chalk degrades off a TTY, so journald and web.log stay free of escape codes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8a54b331e3 |
fix(deps): clear production npm advisories, fix sw.js caching regression
Resolves the four advisories that reach the production dependency tree. The other 16 npm audit reports are devDependencies-only (Remotion, Puppeteer, postcss, the eslint/tsx toolchain) and never ship to users. - @fastify/static 9.1.3 -> 10.1.3 GHSA-8pvw-jcv7-9cmj (authz bypass via non-canonical URL paths). Covers <=10.1.1, so all of 9.x is affected and the fix exists only on the 10.x line. - find-my-way 9.6.0 -> 9.8.0 GHSA-c96f-x56v-gq3h (HTTP/2 DDoS) - fast-uri 3.1.2 -> 3.1.5 GHSA-v2hh-gcrm-f6hx (host confusion) - brace-expansion -> 5.0.9/1.1.18 GHSA-3jxr-9vmj-r5cp (expansion DoS) The last three are transitive and needed only a lockfile re-resolve, so no overrides were introduced. The @fastify/static major changes setHeaders' first argument from a Node ServerResponse to a FastifyReply. Two consequences: 1. res.setHeader() -> reply.header(). The v9 body throws TypeError from inside the plugin on every static request. 2. Precedence flips, silently. The callback used to write to the raw response and lose to the route's staged reply headers; it now writes to the reply and wins. That gave /sw.js a year of immutable in place of the no-cache, no-store its route sets, pinning a service worker on every client with no server-side recovery. A route that already set Cache-Control now keeps it. Verified against v9 to confirm the sw.js behaviour is a regression and not a pre-existing bug. ws appears in npm audit but production is on 8.21.0, outside the vulnerable range; the only affected copy is bundled under @remotion/renderer (dev-only, and remotion is pinned at 4.0.473 because the compositor refuses to start on a version mismatch). Adds test/static-cache-headers.test.ts, which drives a real server and covers a caching contract that had no test at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b0e493d462 |
Merge pull request #311 from Ark0N/fix/workspace-hooks-followups
Workspace hooks follow-ups: one decision core for every claude create path, docker shell gate, boot-sweep and statusLine guards |
||
|
|
cb9149879d |
restore the activity stamp across restarts: the quiet ordering no longer flattens on deploy
Root cause of the reviewer's mass-bump measurement (17 of 17 sessions with an identical lastActivityAt): every restart restamps all sessions in the constructor loop, and the boot auto-attach's repaint re-bumps the rest within the same second. A 12-minute steady-state sample shows NO ambient mass bump, so restarts are the whole story, and Codeman restarts on every deploy. The stamp now has a display twin: recovery threads the previous run's lastActivityAt from state.json into the wire-visible stamp (getter + toState), and a 15s settle window keeps the attach repaint from overwriting it. Real actions (input, task assignment, respawn) always write through. The private stamp keeps its boot-anchored semantics untouched, because the idle confirmation reads it as how long the pane has been quiet. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f44d597450 |
review fixes: every claude create path routes through the workspace-hooks decision
Post-#304 follow-ups. The install-vs-refresh decision (workspaceHooksEnabled, default ON) moved from a session-routes-local helper into hooks-config.ts as applyWorkspaceHooks(workspace, install?), and the claude session-create sites that bypassed it now go through it: cron job fires (cron-service), legacy scheduled-run iterations (runScheduledLoop), and the plan-orchestrator research and planner one-shots. A cron or scheduled run firing in a linked case that never had an interactive session ran hook-blind (no stop for completion detection, no tab alert on a blocking dialog). The shared core also carries the two guards every caller needs: a workspace that no longer exists is skipped (ensureCodemanHooks mkdir -p's, so the boot recovery sweep used to resurrect a deleted repo as an empty tree holding only .claude/settings.local.json), and all errors are swallowed since a create must never fail on hooks. Route handlers keep resolving the setting through their ConfigPort and pass it in; non-route callers omit it and the core reads settings.json itself (absent key or unreadable file = ON). Two adjacent gates tightened in session-routes: - the docker quick-start hooks branch excluded the five external CLIs but let `shell` through, contradicting its own rule that only claude reads .claude hooks; it is now gated on mode === 'claude' - the statusLine exporter call in POST /api/sessions got the same !remote && body.workingDir guard the hooks call got in |
||
|
|
f485085174 |
feat: workspaceHooksEnabled setting as the opt-out for workspace hook installs
Installing hooks into any workspace a Claude session runs in is the right default, but it takes a decision away from a user who deliberately removed them: nothing on disk distinguishes "removed on purpose" from "never had any", so they would come back on the next session create. Adds the synced workspaceHooksEnabled setting (App Settings -> Agents & CLIs -> Claude), default ON. OFF restores the older behavior exactly: a Codeman hooks block that is already present is still refreshed when stale (COD-91), but one is never added. Every create path routes through one applyWorkspaceHooks() helper so the gate cannot apply to some paths only, and the boot-time recovery sweep honours it too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
98fa8c00d1 |
fix: install Codeman hooks into every claude workspace, not just cases Codeman created
A session in a linked case (or any pre-existing repo) ran with no hooks block at all: writeHooksConfig only fires when Codeman CREATES the case directory, and refreshStaleCodemanHooks deliberately never adds one. Every hook-driven surface was therefore dead in exactly the place most sessions run: no tab alert or phone-overview NEEDS YOU row when a dialog blocks the pane, no Approvals Inbox item, no push, no definitive stop/idle_prompt for respawn, and no stop/blocked for the agent wait endpoints. Both session-create paths and restoreMuxSessions() now call ensureCodemanHooks(), an add-only merge that keeps a user's own handlers and leaves a malformed settings file untouched. Claude Code re-reads settings.local.json, so a session already running in the workspace starts firing hooks without a restart. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c5b59633d8 |
feat(pi): add Pi (pi.dev) as a sixth CLI run mode (#206)
SessionMode gains 'pi', a first-class backend alongside Claude Code, OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose tab identity, welcome button, run-mode entry, cron agentType, Docker and remote-SSH command defaults, and clone-repo Brain option. Pi is a different shape of CLI from the other four, and three decisions follow from that: - It has NO permission prompts and no sandbox, so there is no --dangerously-skip-permissions analog and none was invented. The privilege-shaped knob is the tri-state approveProjectTrust, which makes pi load and EXECUTE repo-local .pi/extensions TypeScript and install missing project packages. clampExternalCliBypassForOwner() therefore puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets --no-approve even when no config was sent, because pi's own default is a prompt the session user could answer themselves. That helper had zero test coverage; it now has coverage for all four CLIs. - Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode context, so admitting them would widen the allowlist for every mode at once. Auth goes through pi's /login or the server's own environment. --api-key is deliberately never wired: it would put a provider secret on the spawn command line. - pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the main screen with terminal-owned scrollback, and its 0.84.0 fullscreen mode is runtime-switchable via /settings; that flip was measured to put the pane into the alt screen, which the strip would have corrupted. pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires semver-shaped output, because `pi` is a short generic name a stray binary can shadow; GET /api/pi/status surfaces path and version so a misresolution is diagnosable rather than presenting as a broken mode. Docker installs pi in its own --ignore-scripts step so that flag cannot affect the other four CLIs, and seeds its credentials per-file rather than whole-dir (~/.pi/agent also holds sessions, extensions and package trees). Verified end to end against pi 0.84.1 on an isolated instance: resolver search-dir fallback, flag construction, piConfig persistence across a full server restart, the trust prompt and its --no-approve suppression, the rose Run button on the default daylight-blue skin (the nested skin block eats per-mode gradients unless the rule lives inside it), and the buffer local-echo policy, which pi tolerates where codex did not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f39beb3326 |
chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cf3183abf7 |
chore: version packages
Release 1.16.6: phone overview started/idle stamps, plus fixes for the selection-dialog keyboard lockout, the accessory bar arrows bypassing the local-echo overlay, and recovered sessions being restamped as newly created on every server restart. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4b51ba306e |
feat(voice): dictate through the server's Claude Code login, no API key
The mic button previously needed a Deepgram API key, or fell back to the browser's Web Speech engine. It can now transcribe through the same speech-to-text service Claude Code's own /voice mode uses, so anyone signed in to Claude Code on the server gets dictation with no third-party account. Claude Code's voice mode cannot be driven directly: it opens the HOST's microphone (sox/arecord), and the CLI runs in a headless tmux pane while the human is in a browser somewhere else. So capture stays in the browser and only the transcription backend is borrowed. Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since MediaRecorder cannot emit raw PCM) and receives text. - GET /api/voice/status reports readiness and never the token - GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade guard as the terminal socket, plus caps on concurrency, stream length and frame size - credentials are read-only: Codeman never refreshes them, since a refresh rotates the refresh token and could sign the user out of their own CLI - claudeVoiceEnabled (synced, default OFF) gates the whole server side - voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers Claude, then a configured Deepgram key, then the browser Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5671c20076 |
Merge remote-tracking branch 'origin/master' into feat/readmymind-phase2
# Conflicts: # CLAUDE.md |
||
|
|
1692238531 |
Merge remote-tracking branch 'origin/master' into pr251-review-fixes
# Conflicts: # CLAUDE.md |
||
|
|
94abcf29dc |
feat: Read My Mind phase 2, the predictor and the brain button
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.
Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
sources: pending approval dialog, user goals, last assistant turn tail,
recent prompts, tool activity, git workspace signals, away context,
sibling sessions, rethink state; 30 KB budget, whole-section drop from
the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
keeps only a 500-char snippet) and git signal collection (execFile,
2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
session, opus by default (readMyMindModel setting), strict JSON
contract with 1-3 suggestions (continue / verify / redirect), newline
stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
prediction in flight per session (409 CONFLICT), rethink body
{ steer, rejected }; ownership via findSessionOrFail
Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
until readMyMindEnabled is ON, desktop only (phone key is phase 3);
modal with editable suggestion + rationale and Send / Insert /
Rethink / Dismiss; suggestion text rendered via value/textContent only
and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
strings
Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
161f1da2eb |
feat: Read My Mind phase 1, per-case intent profiles (capture + API + skill)
Per-case profiles of user intent (docs/readmymind-plan.md): user-stated goals plus the user's recently submitted prompts, captured from the Claude session transcript behind the new synced readMyMindEnabled setting (default OFF). - intent-store.ts: keyed by owner + realpath(workingDir), FIFO/size caps, consecutive-dupe collapse, atomic 0600 writes to ~/.codeman/intents.json - transcript-watcher.ts: new transcript:user_prompt event for typed user turns (tool_result-only entries stay silent); capture wiring in server.ts is claude-only and gated on the setting per event - readmymind-routes.ts: GET/PUT/DELETE /api/sessions/:id/intent, ownership via findSessionOrFail, strict Zod schema - agent skill: SKILL.md recipe + endpoints.md rows so agents can read and record intent (PUT replaces: read + merge; never delete unprompted) - groundwork for the phase-2 predictor button; nothing is ever auto-sent Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6cc7b4328b |
feat(cases): clone a Git repository as a new case (#236)
Adds an Add Case -> "Clone Repo" tab plus two endpoints, implementing @DodgyBadger's proposal in #236: clone a public repository straight into codeman-cases/<name> and register it as a normal local case. POST /api/cases/clone is synchronous by design (request held open, bounded by GIT_CLONE_TIMEOUT_MS): no job store, no polling, no cancellation surface. Success broadcasts the usual case:created event, so the case still appears when a proxy idle-timeout kills the request mid-clone. POST /api/cases/clone-preflight runs `git ls-remote --symref` so the UI can say, while the user is still typing, whether the URL is cloneable without credentials, what its default branch is, and which branches/tags exist. Core lives in src/git-clone.ts, split into a pure half (URL parse, argv/env, ls-remote parse, stderr classification) and a thin IO half, so every security decision is unit-testable without spawning anything: - `<name>::<payload>` transports are refused as a family, not by name: ext:: is the famous one, but any of them dispatches to git-remote-<name> and turns a clone into arbitrary command execution. - A leading `-` is refused AND every spawn puts `--` before the operands. Either alone is one edit away from being a hole. - argv arrays, never a shell. URLs carrying user:password@ are refused. - gitNonInteractiveEnv() closes all four ways git can block on a prompt with no terminal attached (terminal prompt, askpass/GUI, ssh, GCM). HOME/PATH stay inherited, so a user's own credential helper or ssh agent keeps working; Codeman itself collects and stores nothing. - The timeout signals the process GROUP, since clone fans out into git-remote-https/index-pack children that outlive a signal to the parent. - Bounded output (redacted stderr tail, capped ls-remote stdout, 500 refs each) and a global 2-op pool, so N large clones cannot exhaust the host. Repository contents beat scaffolding: an existing CLAUDE.md is kept, hooks are merged into whatever .claude/settings.local.json the repo shipped, and a repo that ships its own Claude settings is reported back as a warning (those hooks run locally as soon as a session starts there). A failed clone removes only the directory the attempt created, and refuses a pre-existing destination outright, so it can never squat on a case name. Not admin-gated in multi-user mode, unlike /api/cases/link: it writes only inside the caller's own case space. Local-path/file:// sources are the exception and stay admin-only there. UI: live verdict under the URL field, case name filled from the parsed repo until the user types their own, branch/tag as a datalist of the remote's real refs, optional shallow clone, and a Brain picker (installed CLIs only) that points the Run button at the chosen agent. Starting a session stays opt-in. The tab hides itself when the server reports no git. Tests: the pure half exhaustively (every refusal has a case), plus real git against a real local bare repo for clone/ref/timeout/cleanup, and a route-level suite with unmocked fs that clones through the endpoint. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |