Commit Graph
155 Commits
Author SHA1 Message Date
Codeman maintainer 01f403dc1b feat(claude): advisor tool support, per session and as an App Settings default
Claude Code's advisor tool (code.claude.com/docs/en/advisor) lets the session's
main model consult a second, stronger model at decision points: before
committing to an approach, on a recurring error, and before declaring a task
done. Codeman can now start claude sessions with one.

- `advisorModel` field on POST /api/sessions, /api/quick-start and
  /api/ralph-loop/start (fable, opus, sonnet or a full model id in those
  families; haiku cannot advise and is refused). Stored on the session and
  persisted, so respawn, boot restore and reboot restore keep it. Remote and
  docker quick-starts refuse it, as they refuse effort.
- App Settings, Models, "Advisor" segment (Default / Sonnet / Opus / Fable),
  synced as `claudeAdvisorModel`. Run, resume and the Ralph wizard send it.
  Default sends nothing, leaving the CLI's own /advisor choice in charge.
- Carried as the `advisorModel` key in the launch's single --settings JSON,
  merged with ultracode and the statusLine exporter, never the --advisor
  flag: `claude --advisor haiku` exits 1 at launch, which would leave a dead
  pane on every respawn, while the settings key degrades to no advisor. A
  launch without an advisor is byte-identical to before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-04 19:05:43 +02:00
Codeman maintainer d67da5c9d0 fix(input): an oversized paste no longer poisons the durable input queue (#484)
A single input over MAX_INPUT_LENGTH (64 KiB) was queued for reliable
delivery, refused by both transports (the WebSocket silently, POST with a
400), and never dropped: the client treated the 400 as transient, so the
frame was re-sent every 2 s forever, blocked every later input for that
session, and came back from localStorage on every reload.

- Client: a paste over the frame limit is split into in-limit frames
  (never cutting a surrogate pair) delivered in seq order; over 1 MiB, or
  an oversized mux write, it is refused with a toast and never queued.
- Client: the POST drain drops a frame answered 400/413; a WS error ACK
  drops it too; frames over the limit persisted by an older build are
  pruned on load.
- Server: the WebSocket answers an oversized sequenced frame with
  {t:'ia',seq,err:'too_large',max} instead of silence (an older client
  reads that as a plain ACK and drops it); the POST schema uses
  MAX_INPUT_LENGTH instead of a second 100000 limit.

Verified end to end on an isolated instance: a 110 KB paste reached the
PTY byte-identical over both the WebSocket and the POST path, a poisoned
120 KB persisted frame was pruned on load, and a 2 MB paste showed the
refusal toast with nothing queued.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 18:17:06 +02:00
DevvynandClaude Opus 5.5 0a52a99ca9 feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis

Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI

Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).

Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.

Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).

Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.

Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.

Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.

27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else

window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.

Fixed in two places:

- server.ts: after building `available`, intersect the nine real
  SessionMode ids against `enabledClis()`. git/cloudflared (utility
  binaries, not CLI registry entries) and deepseekBinary (a secondary
  installed-only flag for the "add a profile" affordance) are deliberately
  left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
  `window.__codemanCliAvailable` in place and refreshes the welcome screen,
  the mobile overview and an already-open Run menu, mirroring the existing
  `installDeepSeekProfile()` pattern for the same "injected once, needs an
  explicit patch" reason — without this half, the server-side fix alone
  still left every surface stale until the next reload.

New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.

Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap

Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:

- A new "Status as of 2026-09-22" section up top: what's actually
  implemented (verified by grepping the routes/schema/UI, not just trusting
  the commit messages), the availability-flag staleness bug found and fixed
  in this session (commit 0c77dd0a) with its devbox verification record, and
  one real outstanding gap.

- The outstanding gap: a custom CLI created via Phase 5's write API has no
  way to actually be launched. The Run menu is static per-mode markup with
  no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
  custom entry, confirm it can be launched" verify step was never actually
  exercised against this. Documented with two candidate fixes, neither
  started.

- Each phase's checklist flipped to [x] where confirmed present in the tree,
  Status lines updated from TODO to DONE, and the two originally-open
  questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
  marked resolved against what actually shipped.

No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* docs: add the CLI-registry deployment plan and the parked Copilot plan

Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:

- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
  series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
  is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
  name in docs/cli-enable-disable-plan.md's own header as a sibling plan
  tracked separately — kept for continuity, not active on this branch.

The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* fix(cli-registry): render enabled CLIs in launch surfaces

* test(cli-registry): update frontend branch guard

* fix(test): isolate suite from deployment environment

* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent

shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.

Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): hide shell's toggle entirely instead of greying it out

A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical

renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* style: prettier fixes from the master merge

* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels

Four gaps found verifying #476 against the #343 review trail:

- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
  (the nine per-CLI resolvers and the generic registry one) caches in its own
  closure, with a negative-cache backoff of up to 5 minutes, and nothing
  cleared them. invalidateCliExecutableResolvers(binaries) now drops those
  caches per binary; install (success or failure), create, edit and delete
  call it plus invalidateCliResolverCache(id). Before this, a CLI installed
  from Settings could fail to launch for minutes, and an edited custom entry
  kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
  ignoring the entry's searchDirs and the login-shell lookup that spawn and
  the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
  sit behind an explicit confirm; the confirm now names the exact command,
  which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
  of the word ("Claude", "Codex"). It uses the registry label again, which is
  identical to the old static table for every stock CLI (now pinned).

14 new tests; 9 of them fail against the previous head and pass here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs

Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
  does not parse or has group/world permission bits instead of overwriting it
  (isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
  existence/duplicate checks inside the serialized step, plus a unique tmp
  name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
  section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md

Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
  body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
  map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
  label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
  "Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* test(cli-registry): pin the #343/#347 findings #476 makes reachable

A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 01:48:26 +02:00
Codeman maintainer fe3bd0074c fix(custom-model): split the two confirmation questions, and seed the API key the way claude reads it
Two findings from the review of 5fc391a4, both fixed here rather than sent back.

**The API-key trust seed never matched a real key.** `seedApiKeyTrustFile()` wrote the
key verbatim into `customApiKeyResponses.approved`, but Claude Code stores and compares
only the last 20 characters (`key.trim().slice(-20)`, applied on both the write and the
lookup). For any real key the seed missed, so claude stopped at the interactive
"Detected a custom API key in your environment" prompt, whose default is
"No (recommended)": the launch hangs, or silently refuses the key this feature just
injected and falls through to an OAuth login the isolated config dir does not have. It
survived review because a keyless llama.cpp/llama-swap endpoint uses DEFAULT_API_KEY
('local-dummy-key', 15 chars), where slice(-20) returns the whole string and the seed
matches by accident, and every test used a key shorter than that. Now truncated through
`truncateApiKeyForTrustFile()`, with a test using a 57-character key that also asserts
the full credential never reaches that second file.

**One `confirmed` flag answered two different questions.** The context-floor warning
("this model's window is below what this CLI needs") and the swap-conflict warning
("loading this unloads the model another session is using") shared it, and the context
check runs first, so a user clicking "launch anyway" past the context warning silently
consented to evicting someone else's model. They are about different people, so an
answer to one is not consent to the other. Both routes now read `confirmedContext` and
`confirmedSwap` independently; the legacy `confirmed` still means both, because it
shipped in this feature's HTTP-API-only cut and an existing caller must keep working.
The frontend answers each question with its own flag and accumulates them, on the
one-shot path, the restart path and the batch carry-forward alike.

Also from the same review: the swap-confirm dialog no longer renders " are currently
using ..." when multi-user scoping leaves the affected-session list empty (the swap is
blocked regardless of ownership; only the NAMES are scoped), and the per-endpoint
llama-swap log tails are closed in `WebServer.stop()` instead of only by the idle sweep
whose interval that same teardown disposes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:32:50 +02:00
Codeman maintainer 1a99b5836c Merge pull request #430 from opticon454/custom-model-run-menu 2026-09-19 12:25:11 +02:00
Codeman maintainer 035bfbc2fe fix(remote): merge-time fixes for Wake-on-LAN
The MAC-count limit lived in two places that disagreed. RemoteHostSchema.wakeMac's
128-character cap admits seven comma-separated MACs while parseMacList takes at most
four, all-or-nothing, so a five-MAC value validated, was written to remote-hosts.json,
and then resolved to NO wake target: POST /api/sessions/:id/wake answered
"No wake-on-LAN target configured for this host" and the banner offered "Configure WoL"
for a host the user had just configured. MAX_WAKE_MACS now lives in
src/config/remote-wake-limits.ts and both sides refine against it. Its own module
because src/remote-wake.ts is import-fenced to session-routes.ts and server.ts (the
wiring guard that stops a watcher waking a host), and because schemas.ts must not drag
dgram/net/child_process into every request-validating module.

The documented 40 s request budget also omitted the wake's own cost. A `command` target
is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS and runs BEFORE the readiness poll, so a
slow one pushed a wakeCommand host's worst case to ~68 s, past the 60 s
proxy_read_timeout the budget exists to stay under. _wakeAndWait now subtracts the
wake's measured elapsed time from the readiness budget, floored at one poll interval so
a wake that ate the whole budget still gets one probe. A magic packet is effectively
instant and is unaffected, which is why live testing never saw it.

Also: the two new endpoints are documented in docs/api-reference.md with the import
fence stated as the rule it is, CLAUDE.md's frontend module count moves to 34, and the
release changesets carry the Thanks section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:40 +02:00
RandalixandClaude Opus 5 e271a65e79 Merge origin/master into feat/remote-host-wake
Resolves CLAUDE.md count tables (route counts recounted on the merged
tree: 235 handlers, sessions 37) and keeps both the host-wake and the
reboot-restore banner in index.html.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:20:41 +02:00
Devvyn 56209e7829 Merge remote-tracking branch 'upstream/master' into feature/run-menu-custom-model-picker 2026-09-19 03:25:26 +08:00
DevvynandClaude Sonnet 5 55dae31530 feat(custom-model): estimate model load time from its discovered size
Discovery now also parses a GB figure out of an auto-discovered model's own
description (llama-swap writes "Auto-discovered 16.35 GB - parameters
auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike
context length this needs no /props probe (the figure is right there in
/v1/models) so it is populated for every model regardless of loaded state.
A hand-configured profile's own description has no such figure and
correctly gets no entry.

The loading banner (_watchLlamaSwapLoading) now looks this up and, when
known, shows it plus a rough estimate from a small size->time matrix
(_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) -
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on
llama-swap... this can take a while" - and uses that same estimate's own
bracket to scale the banner's default give-up timeout for a very large
model, instead of a flat 5 minutes for everything. Explicitly labelled as
an UNMEASURED, typical-hardware estimate in every relevant comment - this
is not benchmarked against any real endpoint's actual storage/GPU, just a
reasonable expectation-setter. A model with no discoverable size (a
hand-configured profile) gets no size/estimate shown at all, matching the
"never a guess" convention modelContextLengths already established.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 19:58:46 +08:00
DevvynandClaude Sonnet 5 f865f74a0f feat(custom-model): launch directly on the endpoint, no restart, for 7 of 8 CLIs
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.

POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.

Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.

Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:04:10 +08:00
DevvynandClaude Sonnet 5 bcebc81fcd feat(custom-model): detect llama-swap model conflicts before switching
Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.

1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
   server) via its own GET /running, which plain llama.cpp has no concept of
   at all. New GET /api/model-endpoints/:id/running-status route exposes this
   read-only, for the frontend's polling loop below.

2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
   what llama-swap currently has loaded. If it differs from the requested
   model AND another live session's own customModel selection is actively
   using that loaded model, the apply is refused with a
   {requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
   instead of silently switching. A "confirmed: true" field on the retry
   skips the check. Switching with nothing else affected proceeds
   immediately, no confirmation asked, only ever when there is something to
   warn about.

3. The frontend (runCustomModelEntry) shows a native confirm() naming the
   affected session(s) and the model they'd lose, matching this codebase's
   existing convention for this class of decision (delete case, kill
   session, etc.) rather than a new modal. On a successful apply the response
   also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
   poll shows a sticky "Loading <model>..." toast via the new running-status
   route until llama-swap reports the target model ready (bounded at 2
   minutes), so a prompt sent mid-swap reads as "loading", never as silence
   or an answer from whatever was loaded a moment before.

Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 14:35:07 +08:00
Michael GrundbergandClaude Opus 5 da933d70be feat(sessions): offer to rebuild the sessions a host reboot destroyed
A host reboot takes the tmux server down with it, so every pane dies,
reconciliation finds nothing to attach to, and the board comes up empty.
Picking yesterday's work back up meant finding each conversation in history
and resuming it by hand, one at a time.

The boot pass now works out what the reboot killed and leaves it on offer.
It runs inside restoreMuxSessions(), in the window where reconciliation has
reported the dead sessions and cleanupStaleSessions() has not pruned their
records yet, which is the only place the records can still be read. The
board shows a banner, and nothing is created until the user clicks it.

A click rather than an automatic restore is what makes the reboot heuristic
acceptable. The heuristic cannot tell a reboot from a crash that took tmux
down inside the same window, so it decides whether to ASK, never whether to
act: a wrong yes costs a line of text the user dismisses instead of N CLI
processes nobody asked for.

Four things are re-checked when the click arrives rather than trusted from
boot, because hours can pass and the board moves on. The owner's privilege
grant re-resolves through the env clamp. The workspace must still be on
disk. A conversation the user already resumed by hand from the Resume list
is skipped, since two panes running --resume on one conversation would
fight over the same transcript. Entries leave the plan synchronously before
the first await, and the route is single-flighted, so a double-click or two
devices cannot both reach the same entry.

A restored session comes back attached, idle and disarmed. Respawn
controllers and Ralph loops are deliberately not re-armed: a machine that
just came up is the worst moment to turn an autonomous run loose. Its
workspace hooks are installed by the restore route itself, because the
boot-time sweep sits behind a gate that is false after a reboot and has
finished long before the click; without them a session goes silently blind,
with no stop or idle events for respawn, no Approvals Inbox item and no red
tab on a blocking dialog. Stats collection starts the same way.

The pane is new, so the conversation continues and the terminal scrollback
does not. The banner says so rather than letting an empty pane read as a
broken restore.

The plan lives in memory only. A server restart drops it, which costs the
convenience this adds and never the conversation: the conversation is the
transcript under ~/.claude/projects, which the Welcome screen's Resume list
and the Session Manager already read, so a dropped plan returns the user to
resuming by hand.

clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the
question it answers is about session privilege rather than about HTTP and
it now has a caller outside the route layer. Its test hook stays re-exported
from session-routes.ts.

Claude sessions only for this pass. The other CLIs name their thread in
their own config object, which this does not thread through yet. Remote and
docker sessions are skipped on purpose, because both need another host or a
container to be up and a freshly booted machine cannot promise either.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 08:05:55 +02:00
DevvynandClaude Sonnet 5 0e8b1981af fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker:

1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still
   coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same
   config directory and warns about it (confirmed cosmetic - the API key
   wins for actual requests, verified via a real session's own API Usage
   Billing line). A custom-model claude session now gets an isolated
   CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty,
   no files written into it) so there is nothing to conflict with. projects
   is symlinked (junction on Windows) back into the real config dir so the
   response viewer, subagent windows and Read My Mind keep working for that
   session, best-effort.

2. Context-window overflow. Claude Code assumes a large default context
   window for a model id it doesn't recognise and never compacts, so a
   custom endpoint's real, much smaller context (verified live: a 400
   exceeding a 16384-token llama-swap model with a stock ~33.7K-token system
   prompt) silently overflows. Discovery now also learns each model's real
   context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx),
   but ONLY for a model llama-swap's own /v1/models response already marks
   status.value === 'loaded' - never an unloaded one, since llama-swap
   treats ?model= as a routing hint and probing an unloaded model risks
   triggering an actual, slow, GPU-swapping load as a side effect of
   read-only discovery. A server with no status field at all gets no
   enrichment rather than a guess; a model not probed this round keeps its
   previously-learned value until it disappears from the list entirely.
   Stored per model (CustomModelHost.modelContextLengths) and applied via a
   new contextLengthVar registry field, set to
   CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude.

Both new fields live on the existing env-kind customModelInjection
capability shape, declared only on claude's registry entry - every other
CLI's injection is unaffected (pinned by test).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 12:08:26 +08:00
DevvynandClaude Sonnet 5 25fae9ad10 feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment:
"generate those entries from the saved profiles rather than a fixed
duplicate per harness, and put it in a follow-up PR so this one stays the
backend... The Run-menu picker is yours if you want it."

Adds the frontend surface the backend has been waiting on:

- Run menu: a "Custom Endpoints" section lists one entry per (harness that
  supports customModelInjection, saved endpoint) pair, e.g.
  "Claude Code (llama.cpp)". The harness list comes from
  window.__codemanCustomModelClis, injected at page render straight off the
  CLI registry's own capabilities (never a hardcoded id list in the
  frontend), so a CLI whose injection recipe lands later appears with no
  frontend change. Picking an entry runs that harness's own existing run*()
  function unmodified (case creation, env overrides, everything, forced to
  a single instance) and then applies the endpoint's default model to the
  session it creates via the existing POST /api/sessions/:id/custom-model
  route. Entries are hidden for a remote/docker active case, since that
  route already refuses both.
- Settings: App Settings -> Models gets a "Custom model endpoints" group
  wiring up the customModelEndpointsEnabled toggle (declared since #393,
  read by nothing until now) plus CRUD against the existing
  /api/model-endpoints routes: list, add/edit (inline form), delete,
  discover models.
- Backend: CustomModelHost gains an optional defaultModelId, the model the
  picker applies with no further choice per endpoint (one generated menu
  entry per CLI+endpoint pair, not per CLI+endpoint+model). The route
  refuses a value that isn't one of the endpoint's own discovered models,
  and a fresh discovery drops a default that no longer appears rather than
  carrying an invalid one forward.

Docs: docs/custom-model-endpoints.md describes the new picker and settings
panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the
"backend-only" status note and documents the picker's generation mechanism.

Tests: four new route tests cover defaultModelId validation, acceptance,
and the drop/keep behaviour across a re-discovery; a new render-index-html
test pins the __codemanCustomModelClis injection (present, agent CLIs
supporting the capability, antigravity and shell excluded) and its
solo-window skip. No browser test was added for the Run-menu picker itself
or the settings CRUD panel (this box has no tmux, so the live server used
by test:browser/test:mobile could not be exercised here) -- worth a
Playwright pass before merge, same as any other frontend PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
Codeman maintainer 5b920cb43d feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream
and named tabs wrong five ways (every prompt, every write path, a bare Esc
eating the next prompt's first character, pasted newlines as Enter, any CSI
clearing the draft) and replaced the whole name, which dropped the case from
the tab and reset the w<n> counter. This lands the feature with each of those
closed:

- First prompt means the first: applyAutoName() flips a placeholder to
  `auto` whether or not the string changed. nameSource is now the tri-state
  placeholder | auto | manual; the name setter is the only manual path.
- Only user-originated input counts: write()/writeViaMux() take
  SessionWriteOptions.fromUser, set by the browser WS path and POST /input
  only, so Ralph, respawn, cron, approvals and the trust-dialog keys can
  never name a tab. A startMode 'shell' CLI never feeds the tracker (a
  capability, not an id check); the send-key route feeds trackUserInput()
  because its line feed bypasses the session.
- Prefix form `w3-case: title`: parseSessionPrefix() already renders it as
  the title with the prefix in the tooltip and the next-session counter
  still matches it. Composed within MAX_SESSION_NAME_LENGTH.
- Tracker rules per key: bare Esc resolves at chunk end; mouse/focus
  reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R
  taint the draft so Enter submits nothing rather than a fragment;
  bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space;
  the draft keeps its head past 8192 code points; an escape past 64 bytes
  is abandoned.
- Title: slash commands by shape (a path is a prompt), `!` escapes
  refused, first sentence only past 8 code points ("e.g." is not a title),
  72 code points on a word boundary.
- Synced `autoNameSessions` setting, default OFF (the prompt reaches
  mux-sessions.json, session:updated and /api/search), App Settings ->
  Appearance -> Tabs, read fresh per prompt after the eligibility check.

Tests: test/session-auto-name.test.ts (tracker, title, composition,
ownership, emit gating), the wiring test (once, prefix, setting off,
manual protected), test/routes/session-name-routes.test.ts (PUT /name
flips to manual and persists). Verified live on an isolated instance: API
and browser-typed prompts name the tab, a second prompt does not, shells
and renamed tabs are untouched, nameSource survives a restart.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 17:59:16 +02:00
Randalix 8b5a13435a feat(remote): host-unreachable banner, manual wake, and native MAC wake-on-LAN
The reactive wake (typing into a session whose host slept) left the state invisible:
nothing told the user the machine was asleep, and with no wake target configured
there was nothing to do about it. Adds:

- RemoteHost.wakeMac (comma-separated) - Codeman builds and broadcasts the magic
  packet itself (UDP port 9), so the common case needs no external script. The
  existing wakeCommand stays as the explicit override.
- GET /api/sessions/:id/reachability - probes (throttled, cached, and it never
  wakes) and reports HOW the host can be woken, or that nothing is configured.
- POST /api/sessions/:id/wake - wakes, waits, reattaches the pane and flushes
  buffered input; 400 with a routable message when no target is configured.
- The amber host-unreachable banner + its 'Wake' / 'Configure WoL' action, and a
  small config dialog that saves via PUT /api/remote-hosts/:id.
- RemoteWakeDeps.resolveRemote: host config is re-resolved for LIVE sessions
  (throttled + cached), so saving the dialog takes effect without a restart.
2026-09-15 14:20:31 +02:00
Randalix a81f430e41 feat(remote): wake a sleeping host from user input (Wake-on-LAN)
A durable remote session survives SSH drops (COD-104/108), but nothing brought
the HOST back: after the remote machine suspended, the local tmux pane's ssh
child stalled silently and `send-keys` SUCCEEDS against it, so typed input
vanished with no error anywhere.

Add an optional per-host `wakeCommand` (Wake-on-LAN wrapper, e.g. whuff) that
the input route runs when a wake-enabled host is unreachable: input is buffered,
the host is woken, the pane is reattached, and the buffer is flushed in order.
Detection is a throttled bare TCP probe on wake-enabled hosts only, and only
REAL user input may wake a host - the auto-reconnect watcher and boot recovery
deliberately cannot, or the host would be re-woken seconds after every suspend
and could never stay asleep.
2026-09-15 10:25:52 +02:00
Codeman maintainer 942bf37e48 fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a
custom OpenAI-compatible endpoint by injecting env vars or a config file
and restarting the CLI in place. Review of the apply path found four
things, two of them destructive. This lands all four plus the smaller
items from the same review.

1. Clearing a selection did not clear it. The injected vars reach the CLI
   via `tmux setenv`, which persists at the tmux-session level and is
   inherited by `respawn-pane` (measured: `setenv FOO bar` survived two
   successive `respawn-pane -k`), so deleting the keys from the session's
   envOverrides relaunched the CLI still pointed at the old endpoint, and
   for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just
   been deleted. `Session.setCustomModel()` now reports the removed keys,
   queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys`
   carries them into `applyEnvOverrides()`, which `setenv -u`s them before
   re-applying the live overrides, on the same path that already unsets
   the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket
   that `setenv -u HOME` hands the next respawn the global HOME back.

2. Applying a model to a local claude session killed the pane. The
   relaunch was `claude --session-id <id>` and Claude refuses an id that
   already has a transcript, and unlike the dead-pane respawn this one
   kills a working pane first. `restartCli()` now pins the live
   conversation id as the resume id for that respawn when the CLI's launch
   declares a `fallback` chain, which renders the same
   `--resume <id> || --session-id <id>` shape the docker and remote pane
   commands use. Gated on the registry shape, not the CLI id: an entry
   whose resume id is minted by the CLI itself never declares that chain.

3. pi, omp and grok wrote their config file and then launched without the
   `--model` that selects it, so the file was ignored. The registry entry
   now declares `customModelInjection.launchModel` (`custom/{modelId}` for
   pi and omp, grok's `[model.codeman-custom]` block name), the builder
   renders it, and `_withCustomModelLaunchModel()` applies it onto the
   respawn options through `legacyConfigField`, leaving the stored
   <Mode>Config untouched so a clear falls back to the user's own model.
   A model id the CLI's `model` token pattern cannot carry is refused
   with a 400 rather than silently dropped by the argv engine.

4. Remote (SSH) and Docker sessions reported `restarted: true` and changed
   nothing: their `restartCli()` reattaches the durable tmux rather than
   relaunching the agent, and the env lands on the local pane. Both are
   refused with a 400 until those paths are plumbed.

Smaller items from the same review:

- The selection survives a Codeman restart as the disk-only `__customModel`
  bookkeeping (endpoint, model, injected key NAMES, config dir, launch
  model; never the values, which carry the API key). Recovery re-derives
  the values from the endpoint store through the same apply path the route
  uses and keeps the bookkeeping even when the endpoint is gone, so a
  later clear still has keys to unset.
- Discovery goes through `webviewFetch()`, so the RESOLVED address is
  judged by the same egress guard the web-tab proxy uses, and `baseUrl`
  reuses `webviewUrlSchema` (http(s) only, no embedded credentials,
  link-local and cloud-metadata addresses refused). undici's `fetch failed`
  wrapper is unwrapped so the user sees the ECONNREFUSED underneath.
- `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session
  config dir 0700/0600 (pi and omp embed the key literally), and that dir
  is removed with the session.
- `PR.md` is gone from the repo root and the design doc moved to
  `docs/custom-model-endpoints-plan.md` with the LAN address and the
  personal name scrubbed; every reference follows. The guide's `authStyle`
  text matches the shipped schema (`bearer | api-key`, default `bearer`)
  and says that `customModelEndpointsEnabled` is read by nothing until
  the picker lands.
- `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts`
  (four real type errors fixed). It is not yet wired into `npm run typecheck`
  because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json`
  there is the one-line follow-up.

Tests: `test/session-custom-model-restart.test.ts` drives a real Session and
fails on the unfixed code for items 1 to 3; the route suite covers item 4
and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets
run before the overrides and that a shell-metachar key never reaches tmux.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:46:28 +02:00
Codeman maintainer 1e42cb4e2d Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses)
2026-09-14 23:46:27 +02:00
Codeman maintainer 707ea345eb fix(statusline): GET /api/settings never writes, and a save sends the collection switch only on a flip
Two follow-ups to #361's sticky telemetry switch.

GET /api/settings reconciled an absent showPlanUsageLimits by persisting
true, but readJsonConfig() answers {} for ANY read failure (a parse
error, EACCES, EMFILE, a read landing inside PUT's non-atomic write), not
only ENOENT, and every page load calls this route, so one unlucky read
replaced the whole settings file with a one-key file. The route is a
plain read again and the default moved into the reader:
readPlanUsageTelemetryEnabled() treats an absent key as ON, the same way
readWorkspaceHooksEnabled() does, which is what the desktop chip already
shows for an install that never touched the setting.

saveAppSettings() sent showPlanUsageLimits on every save. The chip
defaults OFF on handhelds, so a phone saving its font size persisted
false and switched collection off for every desktop, whose chip then
went stale with no error anywhere. The key is now stripped like the
other per-device display keys and re-added only when the save FLIPS the
chip relative to what the device had (planUsageCollectionFlip), so an
explicit toggle on any device still writes it in either direction.

Tests pin both: the GET route with a mocked filesystem (absent, missing,
EACCES, garbage, explicit), the reader default, and the flip helper plus
its wiring in saveAppSettings.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:59:00 +02:00
Ark0N b2b2c767ea Merge pull request #361 from timkjr/fix/statusline-injection-opt-out
fix(statusline): inject plan-usage telemetry via ephemeral CLI flag, never disk
2026-09-14 15:58:48 +02:00
DevvynandClaude Sonnet 5 41416566aa feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).

- Registry: capabilities.customModelInjection per CLI entry (env /
  configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
  model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
  custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
  (POST /api/sessions/:id/custom-model), reusing the existing
  respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
  CLI's privilegedEnvKeys, closing a pre-existing gap where several were
  already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
  against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
  every CLI's injected values through a real HTTP shape

Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
  a top-level model string + [model_providers.custom]); fixing it then
  surfaced a genuine, documented protocol incompatibility (Codex only
  speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
  don't implement)
- Claude Code's async session-title-generation call validates
  ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
  hangs the whole -p invocation on an unrecognized name; documented for
  chunk 6, worked around in the standalone script only (--bare is NOT
  safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
  and api-key headers) reliably hung a real server; removed the option
  entirely rather than just changing the default

Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00
Codeman maintainer 8b23f3e260 feat(rail): sort the vertical tab rail by activity, and give its rows the home screen's card
The vertical rail lists exactly the sessions both home screens list, so it now
answers their question the same way instead of showing the raw tab order.

Order: new per-device `tabRailSort` (App Settings -> Appearance -> Tabs ->
Vertical Rail Order, default "By activity"). It runs `CodemanSessionOrder` over
rows classified by `_mobileOverviewState`, i.e. literally the home screens'
comparator, `lastSubmitAt`-anchored running group included.

It is applied as the flex `order` property, never by reordering the DOM.
`#sessionTabs` stays in `sessionOrder`, which is what keeps the Alt+N badge
honest (it names a shortcut, not a row position, so it deliberately does NOT
run 1,2,3 down a sorted rail), and keeps drag-and-drop, the arrow-key walk, the
sidebar filter and `_scrollActiveTabIntoView()` all reading the list they
always read. A session changing state then moves one inline style instead of
forcing the full rebuild that would restart every card's animation on every SSE
tick. The incremental render path re-applies it, since a state flip adds no tab
and never reaches the full rebuild, and an empty string is what clears it when
sorting stops. Web tabs are pinned past the cards by a CSS `order: 9999`, since
`renderWebviewTabs()` emits the same markup for every layout and the flex
default of 0 would interleave them. Drag is switched off while sorting (the
drop rewrites `sessionOrder` correctly and the sort puts the card straight
back, so the affordance would be a lie); 'manual' is the way back.

Cards: detailed rail rows become bordered cards on `--bg-card`, with the stamps
line on its own full-width row and the pill at its right end. The state dot
goes 6px to 9px, keeps its orbiting ring while working and gains the green
halo; idle mutes toward `--text-muted` as the home rail does. Needs/error/
waiting reuse `home-sessions-blink-red`/`-yellow` rather than a second copy.

These card rules are RAIL-SCOPED and deliberately absent from the comma-grouped
selectors that carry both vertical surfaces: the rail is an occasional,
resizable list you scan, while the sidebar is a permanently-docked nav column
where 20 stacked cards read as a wall. Every state-dot rule also excludes
`.tab-alert-action`/`.tab-alert-idle` by hand, because those alert rules are
only (0,3,0) and these are (0,5,1)+.

Lines: the lineage bracket already drew in the rail, but its track sat 6px from
the left edge, so half of its 11px outer glow was clipped by the window frame
and it read as a thread pinned to the frame. It now runs at 10px, mid-channel
in the gutter the rail already reserves.

Tests: test/tab-rail-order.test.ts (17) drives the real `isTabRailSorted()` and
`_tabRailSortOrder()` out of app.js, covering the row model (a WORKING row
ranked by `lastSubmitAt`, which would otherwise rank every running turn as
freshly started and fail no rendering test), the Alt+N badge, and the opt-out.

Verified in Chromium across sorted/manual/simple/header-strip/sidebar with the
setting flipped at runtime: no page errors, and the header strip and sidebar
render byte-identically to before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 06:09:28 +02:00
timkjrandClaude Sonnet 5 d5b75af628 fix(statusline): sticky telemetry collection, footer print-through, EOF fix
Responds to Ark0N's review round on the ephemeral-CLI-flag statusline
injection rework:

- Rebase-detail fixes: registry-gated telemetry eligibility via
  getCli(mode)?.capabilities.statusLineTelemetry instead of a hardcoded
  mode === 'claude' check, using the capability flag master's CLI-registry
  refactor already declares for exactly this purpose.

- Design question settled: sticky (a). Rather than persisting the toggle
  as a new field and threading it through every session-creation path
  (cron, Ralph Loop API, quick-start), eliminated the per-session field
  entirely. readPlanUsageTelemetryEnabled() (hooks-config.ts) reads the
  existing showPlanUsageLimits setting fresh from settings.json at every
  claude create/respawn (TmuxManager.createSession/respawnPane) - no
  per-session state to survive a restart, and it applies uniformly to
  every creation path for free, since they all flow through the same
  TmuxManager methods.

  This required fixing a real bug found along the way: showPlanUsageLimits
  was not actually round-tripping through settings.json on save -
  settings-ui.js explicitly excluded it from the PUT body as a pure
  per-device display key. It now flows through normally (both true and
  false); the load-side per-device merge behavior is unchanged.

  Removed entirely as a result: the statusLineTelemetry field from
  CreateSessionSchema/SettingsUpdateSchema, CreateSessionOptions/
  RespawnPaneOptions, Session._statusLineTelemetry (this is what makes
  the restart-persistence bug moot rather than patched), and the
  frontend send sites.

- Footer print-through restored: the no-user-statusline branch of the
  exporter script now runs the telemetry POST in the foreground so its
  own stdout becomes the in-terminal footer, falling back to a plain
  "codeman" marker only on curl failure.

- Background-subshell EOF fix: the wrap-a-real-statusline branch closes
  stdin too, not just stdout/stderr (`>/dev/null 2>&1 </dev/null &`) -
  the un-redirected subshell process itself, not curl, was what held a
  reader-to-EOF's pipe open for however long curl took to finish. Added
  curl --max-time 5 so a hung (not just refused) Codeman cannot wedge
  the render.

Tests: real-shell-execution tests for the footer/EOF fixes (fake curl
stand-in on PATH, real sh subprocess spawns, real elapsed-time
measurements - verified non-vacuous against a hand-reconstructed
old-style script), unit tests for readPlanUsageTelemetryEnabled.
Adapted two existing tests whose payloads referenced the removed field.
Fixed during independent code review: a stray indentation break and a
test exercising the wrong (legacy) exporter code path.

Docs synced: CLAUDE.md, docs/usage-limits-display-plan.md (old
disk-based section marked superseded, kept for history),
docs/architecture-invariants.md.

Full suite green: 352 files, 6780 passed, 12 skipped, 0 failed.
tsc/lint/format:check/frontend-syntax all clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 20:37:29 -05:00
Codeman maintainer 8ee7926e27 feat(agent-cases): tag agent-spawned case dirs and sweep their leftovers
A long orchestration creates one case directory per worker and deleting the
sessions never removed them, so ~/codeman-cases accumulated scratch folders
that were indistinguishable from real projects. They are now labelled and
have a cleanup path.

- src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven
  spawn gets a .codeman-agent-case.json marker (when, by whom, parent session,
  mode). Only the create branch writes it, so a linked case, a cloned repo or
  any pre-existing path is never labelled; reading is total, so a malformed
  marker means "not agent-created" rather than a half-trusted entry.
- The signal is the new X-Codeman-Agent-Origin header the skill preamble sets
  on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body
  field, falling back to a resolved parentSessionId so a worker spawned by a
  stale skill copy is still labelled.
- GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is
  a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage
  badges each case and offers a review-then-delete sweep that names every
  directory in its confirm and skips any case a live session is working in.
  Removal stays on the existing DELETE /api/cases/:name.
- Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was
  written per claude session and never removed (236 leftovers measured on a
  working machine). Now deleted with the session and swept at boot, guarded by
  a live-session keep set plus a 7-day age floor.

Verified end to end on an isolated instance: marker written for header, body
and lineage-only spawns, absent with no agent signal and for a pre-existing
directory; inUse flipping on session end; badge, sticky bar, confirm and sweep
driven in a browser; preamble seeded on create, removed on delete, boot sweep
taking only the aged orphans.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 19:09:24 +02:00
Ark0N bca1b764cc Merge pull request #367 from shenlvkang-collab/pr/claude-conversation-first-hand
fix(session): learn the live Claude conversation from the CLI's own hook
2026-09-06 23:02:33 +02:00
Codeman maintainer 3d8ffcb9a2 Merge pull request #357 from dignfei/feat/docker-adopt-existing-container
feat(docker): attach a case to an already-running container

Conflicts came from work that landed after the PR was opened, and each is
resolved onto the newer abstraction rather than by keeping the older code:

- `defaultDockerCommandForMode` is registry-driven since #347, so the PR's
  `runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude
  Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the
  refusal is visible only inside the container, so an adopted root container
  otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and
  `test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch.

- The probe's mode list and its mode -> binary table both duplicated the
  registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is
  also what fixes the merge's silent regression: the hand-written list predates
  `omp`, and the run menu gates every docker case on this probe, so owned
  containers would have lost that mode. `shell` needs no arm — it declares no
  binary, so it is dropped from the lookup and reported available regardless.

- The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one
  `missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto
  it. Its test now pins the single gate instead of counting seven arms.

- The create arm keeps #349's swap-limit warning filter, which the adopted arm
  never reaches; the run-mode list gains `omp` from #353.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
2026-09-05 16:21:51 +02:00
Codeman maintainer 550e08a791 fix(webview): refuse link-local and cloud-metadata targets on the resolved address
The web-tab proxy, its Test probe and its WebSocket relay accepted any http(s)
host. A live PoC relayed an IMDSv2-shaped PUT with custom headers to a loopback
echo server through a capability and no cookie, and 169.254.169.254 (decimal,
hex, IPv6-mapped, or via a DNS name) was as valid a dashboard as any other.

Loopback and RFC1918 stay allowed on purpose: a localhost Grafana is the feature.
Only link-local and the fixed cloud-metadata addresses are refused
(169.254.0.0/16, fe80::/10, fd00:ec2::254, 168.63.129.16, 100.100.100.200,
metadata.google.internal), at three stages that are each load-bearing:

- the Zod schema, so a save gets a clear refusal;
- a synchronous hostname check at every connect site, because net.connect skips
  DNS for an IP literal and a lookup hook never sees one;
- a `lookup` hook on an undici Agent (webviewFetch) and on the ws client, which
  judges the RESOLVED addresses of a name and refuses when any is blocked. This
  is what closes DNS rebinding, which a hostname-string check cannot.

Adds undici@^6 so the proxy runs the package's own fetch with the package's own
Agent; a package Agent handed to Node's bundled fetch can mismatch protocols.

Verified live on an isolated beta: 169.254.169.254.nip.io (a real name resolving
to the metadata address) is refused by probe, proxy (403) and WS relay (4003),
while 127.0.0.1.nip.io still passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:12 +02:00
DevvynandClaude Opus 5 4830e662f9 refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.

Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.

Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.

OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.

Guard rails:

- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
  branching reappears outside `stock.ts`, in any of its four shapes (`===`,
  `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
  negated forms, which is how 36 of them survived an earlier pass. Every
  allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
  deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
  `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
  wire field is separate, bridged only by `legacyConfigAliases`. Getting
  `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
  clamp's only handle on a CLI's privilege switch, and a wrong name clamps
  nothing with no error and no failing test — so `schema.ts` rejects an entry
  naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
  `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
  thunks). A module-level const freezes at first import, so a CLI enabled while
  the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
  (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
  `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
  rather than measured. A test pins the list so it cannot quietly grow.

Three user-visible changes, all deliberate and named:

- `probeDockerCliVersion()` derives the in-container binary from the registry
  rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
  hardcoded map it replaces omitted while its own comment said the rule was
  "every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
  install hint is the install command rather than a docs URL, five CLIs gain
  hints they never had, and the row order follows the catalog.

Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 08:26:45 +08:00
shenlvkang-collab ccfda623fe fix(session): learn the live Claude conversation from the CLI's own hook
Which conversation a pane is on was re-derived by correlating
~/.claude/history.jsonl against Session.lastSubmitAt — and lastSubmitAt is
bumped only by input that flows through Codeman's own write path
(Session.write / writeViaMux). A user who attaches to the pane's tmux session
directly never set it, so resolveActiveClaudeSessionIdFromHistory() returned at
its first line for that pane's whole life and the response viewer stayed pinned
to the launch conversation, showing a pre-/clear transcript indefinitely.

A UserPromptSubmit hook reports the live conversation id from inside the CLI
process, delivered under the pane's own $CODEMAN_SESSION_ID. That binding is a
fact rather than a correlation: it never consults workingDir, so it cannot be
claimed by a sibling pane on the same folder, a closed tab, or a bare `claude`
in the user's terminal. A pane holding such an id skips the correlation
entirely, so the number of prompts eligible for cwd-based guessing goes DOWN,
never up — the naive alternative (relax the guard, or synthesize an anchor from
PTY activity) is the reverted bug the resolver's own comment describes.

The hook also stamps lastSubmitAt, so it finally means "a prompt was submitted"
rather than "typed into Codeman's web terminal". Conversations vouched for
first-hand — and only those — extend a persisted claudeSessionChain, whose tail
re-pins the conversation when a surviving tmux session is re-attached after a
restart. ⚠️ start() resets the id at THREE points and the last one runs
unconditionally after the mux branch, so the tail is applied there too; patching
only the mux branch looks right and silently does nothing.

⚠️ The hook's stdout is discarded with curl's own -o /dev/null. Claude Code
injects a UserPromptSubmit hook's stdout into the model's context ("Exit code 0
- stdout shown to Claude"), and a trailing >/dev/null does NOT work: curlCmd
already ends `... 2>/dev/null || true`, and in `pipeline || true >/dev/null` the
shell binds the redirection to `true`, which never runs on the success path. The
discard is opt-in so the five SSE-fed events keep byte-identical command text
and no workspace's settings file is rewritten for them. The staleness marker is
quote-free for the matching reason: hooksJson is JSON.stringify'd, so a quoted
needle never matches and the gate would rewrite every workspace on every spawn.

Existing workspaces heal on their next Claude spawn through the staleness sweep.
2026-09-01 12:33:24 +08:00
Codeman maintainer da91b4353b Merge pull request #353 from timkjr/omp-mode
feat: add OMP (Oh My Pi) as a new session backend
2026-08-30 16:16:26 +02:00
d fei 5452ad5c5a feat(docker): add a folder picker to both path fields
Both paths in the adoption form had to be typed. Each gets a Browse button
using the same path-input-group markup Link Existing uses, so the two look and
behave alike.

What they can browse differs, and that is the point. The host workspace path
reuses the existing host picker. The container workdir cannot: an adopted
container has nothing mounted at a matching host path, so a host listing would
be a different filesystem — and getting this field wrong is the source of the
opaque OCI chdir error at launch, which makes it the field that most needs to
be clickable.

Adds a read-only POST /api/docker-cases/browse: one `ls` through docker exec, no
writes, no lifecycle, path shell-escaped like every other value. `ls -Ap` marks
directories with a trailing slash and keeps names with spaces intact.

PathPicker takes an optional fetchListing source rather than being forked: the
container variant only swaps where the rows come from, and reuses the rendering,
navigation, Up and Choose/Select unchanged.
2026-08-29 23:31:28 -07:00
d fei c98a59d709 fix(docker): verify the container workdir and end the probe with exit 0
Two defects that only a real container exposes.

The probe chained `command -v X && echo X` with semicolons, and a script's exit
status is its last command's. A container without the last probed CLI made the
whole `sh -lc` exit 1, so a perfectly healthy container with tmux and claude was
reported as "could not exec into the container". A missing CLI is data here, not
failure, so the script now ends with `exit 0`.

containerWorkdir defaulted to hostWorkspacePath. That default holds for an owned
container only because the create-time bind mount puts the host directory at that
exact path; attaching mounts nothing, so the two are independent facts. A host
path absent inside the container makes `docker exec --workdir` fail with an OCI
chdir error that surfaces in the pane as a bare "execvp failed". The preflight now
proves the directory exists inside the container and refuses at link time.
2026-08-29 21:02:19 -07:00
d fei 15eebde832 feat(docker): attach a case to an already-running container
Docker cases could only run in a container Codeman created itself. Attaching to
one the user already built and runs means Codeman must leave that container's
lifecycle completely alone, which the launch chain could not do: it was
`image inspect` -> `inspect || create` -> `start` -> `exec`.

Adds `DockerCase.owned`, mirroring the `owned:false` contract remote-SSH already
uses for attached sessions. Absent (every existing case) means owned, so current
behaviour is byte-identical. `false` means the container belongs to the user and
Codeman may only exec into it.

The launch chain for an attached container only looks, then execs: no image gate
(the image is theirs), no create, and no `start` — starting a container we do not
own is the very mutation attaching promises not to perform. A missing or stopped
container fails closed with an actionable message instead. Credential seeding is
skipped too: those copies read from create-time read-only mounts that do not
exist here, and writing host credentials into someone's container is not ours to
do, so its CLIs must already be authenticated inside it.

Four fail-closed guards. buildDockerStopCommand and buildDockerRemoveCommand
throw during pure string construction, so no caller bug can turn into a
`docker stop`/`rm` on a container we do not own; removeDockerContainer refuses
again at the lowest layer; drift reports "none" for an attached container, which
carries no `codeman.confighash` label and would otherwise always look drifted and
409 the launch gate forever; and the orphan reaper skips attached containers
through a check deliberately independent of the two conditions already covering
them.

`owned` is applied AFTER the config hash is computed. dockerConfigHash takes an
explicit field list, so ownership can never shift an existing case's hash — if it
did, every pre-existing case would trip the drift gate at once, and the remedy
the UI offers is "recreate the container".

Adds POST /api/cases/docker-adopt and a read-only
POST /api/docker-cases/adopt-preflight. The preflight refuses at LINK time rather
than at session launch, where the only ways out would be a dead pane or starting
a container we do not own.

Tests assert the negative guarantee directly — that create, start, stop, rm,
restart and kill are absent from the generated commands while `docker exec -it`
and `new-session -A` remain — since it cannot be observed by using the feature.
2026-08-29 21:02:19 -07:00
timkjr 3e1a0e679f fix(omp): resume by mode, not silently as claude, and support --continue
resumeHistorySession() never sent mode when recreating a session from a
history/session-manager row, so the server default silently opened a
plain Claude session for every non-claude row -- reproduced live: OMP
rows spawned Claude sessions on click. Thread the row's mode through
every call site (welcome list, session manager, mobile overview) and
only send the Claude-specific resumeSessionId for claude rows.

Codeman has no live PTY-reattach outside server boot, and it's moot for
OMP anyway (exiting it kills the pane's only process), so route the
non-claude relaunch through each CLI's own continue-most-recent flag
instead of a context-free fresh start. OMP never got one: buildOmpCommand
only implemented --model/--resume despite omp --help documenting
-c/--continue. Added continueSession to OmpConfig end-to-end (type,
schema, builder) mirroring the existing opencode/pi/grok/deepseek
fields, and wired resumeHistorySession to use it.

Verified live: told a real omp session a secret, exited it, closed the
tab without killing tmux, relaunched with --continue in the same
directory, and had it recall the secret.
2026-08-28 11:32:30 -05:00
timkjr 4f5678fac4 feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) 2026-08-26 20:05:48 -05:00
Jack Stuart b00ab3ceea feat(web): show Codex plan usage in header 2026-08-26 18:43:52 +08:00
Ark0N ca5fe1ab3e Merge pull request #338 from Ark0N/feat/vertical-rail-detailed-rows
Vertical tab rail: detailed rows (created / working / status), plus a rename-cancel fix
2026-08-25 19:12:57 +02:00
Codeman maintainer c30dfaf0e7 fix(deepseek): run the web UI server in the background, not in a shell tab
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a
shell tab running the server next to it. The shell was deliberate - the server
lived in an ordinary session so it was visible, scrollable, killable and died
with its tab, and nothing new had to supervise a long-lived HTTP server. That
reasoning was sound and the result was still wrong in use: opening a dashboard
should open one tab, and after the first launch the terminal is pure noise.

The server moves to a background child process owned by a new
`src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session
gave away for free is now explicit, which is most of the module:

- Exactly one server. A second click reuses the running one instead of racing
  it for a port; the session flow could not do this at all, because two clicks
  were simply two sessions.
- Restarted when the requested authority changes. `--trusted-host` fences dsh's
  own /api against the browser authority, and a Codeman reachable at both
  loopback and a tailnet name has two. Reusing a server fenced for the other
  origin renders a page whose every call 403s, which reads as a broken
  dashboard rather than a misconfigured one, so a mismatch restarts instead.
- Killed on shutdown. The child is detached so its whole plugin tree can be
  signalled at once, which also means it would outlive Codeman and hold its
  port against the next start - the exact EADDRINUSE this feature already got
  wrong once.
- Boot output captured and returned. With no shell tab there is nowhere else
  for a stack trace to land, so a failed spawn reports its own tail.

The endpoint is fenced at the same bar as the profile installer and for the
same reason: booting a dsh profile executes the plugin code in it, so this is a
privileged action even though it reads as "open a page". `authority` comes from
the client (`location.host`) because only the browser knows which origin is in
play, and it is regex-confined at the schema boundary - defence in depth behind
the argv-array spawn, admitting host:port in the shapes a browser authority can
take and nothing readable as a second argument.

`GET /api/deepseek/web-port` is gone; port selection moved into the supervisor,
which is the thing that knows whether a server is already running. The two
client-side probe helpers went with it, since the server now owns the wait.

Verified over the tailnet authority end to end: no session is created (session
count unchanged, one tab), the server runs on 3081 beside the user's own dsh
web on 3080, status reports the tailnet authority, and the proxied dashboard
renders with zero 4xx. Full gate green (6148 passed, +6).
2026-08-25 03:08:15 +02:00
Codeman maintainer 15ae5f5d81 fix(deepseek): make the web-UI shortcut pick a free port, verify it, and trust its frame
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.

1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
   precisely the port a DeepSeek user is most likely to be serving on already,
   so the launch died with EADDRINUSE against the user's own server. The port
   now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
   free loopback port by BINDING it (a connect probe cannot tell "free" from
   "listening but not answering yet").

2. It opened the tab unconditionally. The crashed server left a saved dashboard
   pointing at nothing, with the failure only visible in a shell tab nobody had
   a reason to look at. The launch now polls the existing webview probe until
   the URL answers, and on timeout reports the error naming the shell tab
   instead of persisting a dead dashboard.

3. The saved tab was untrusted, so the frame was sandboxed without
   `allow-same-origin` and the dashboard was broken twice over: the dsh
   client-runtime reads `localStorage` while loading its plugins and died there
   ("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
   opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
   every `/api` call no matter which authority `--trusted-host` named. Passing
   `location.host` only means anything once the frame actually carries that
   origin, so `--trusted-host` had never once done its job. The managed tab is
   now created `trusted: true`.

   That trade is real and deliberate: a trusted proxied frame is same-origin
   with Codeman and can reach Codeman's API. It is defensible only because this
   dashboard is an agent harness Codeman just started itself, on loopback, which
   can already run code as the user. It is not a precedent for trusting
   third-party dashboards, which is why it is set at this one call site rather
   than defaulted.

Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.

`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.

Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
2026-08-25 02:39:57 +02:00
Codeman maintainer b330f1d9e8 feat(tabs): give the vertical rail the home screen's per-session detail
The vertical tab rail (tabOrientation 'vertical') listed names and nothing
else, while the rich sidebar and both home screens already answered the
question a docked column exists to answer: which of these sessions wants me
next, and how long has it been like that. The rail is a docked column too, so
it now draws the same row.

- New per-device setting tabRailDetail ('rich' | 'simple', default rich),
  App Settings -> Appearance -> Tabs, in SettingsUpdateSchema + displayKeys and
  stamped as data-tab-rail-detail by the pre-paint script, so a detailed rail
  does not flash through simple rows on every load.
- ONE gate for both vertical surfaces: isRichTabRows() =
  isSessionSidebarRich() || isTabRailRich(). The row model, the markup and the
  20s in-place clock are the existing rich-sidebar ones, classified by
  _mobileOverviewState/_mobileOverviewSince, so the rail, the sidebar, the
  desktop home rail and the phone overview cannot disagree about what
  "working" means or which stamp measures it.
- Detail rides on its OWN attribute, exactly as the sidebar's does, so every
  existing [data-tab-orientation='vertical'] rule keeps matching both variants
  untouched. A flip of detail ALONE still forces a full render (the stamps line
  is emitted by the row template, not toggled by CSS) and re-runs
  applyTabWrapSettings(), which owns the folder line and is now rail-aware.
- CSS: every rich paint rule gains a rail twin as a COMMA-GROUPED selector,
  never :is() - an :is() list takes its most specific argument, which would
  lift the sidebar arm from (0,3,1) to the rail's (0,5,1) and let these rules
  outrank things they never used to.
- Width is why there are thresholds. At 256px the stamps line ellipsizes
  mid-word, the same reason the rich sidebar is 300px, so a rail that has never
  been sized defaults to 320 (RICH_DEFAULT_WIDTH, the existing Wide preset,
  which also keeps the settings select on a named choice). A width the user has
  chosen is never overridden: below 288px the created stamp is dropped rather
  than truncated (tab-rail-tight, CSS only) and below 240px the rows go back to
  simple (tab-rail-compact, which re-renders).
- The rich clock is armed and disarmed by applyTabOrientation() as well as
  applySessionListLayout(); a leaked interval would rewrite stamps in a list
  that no longer has any.

Also fixes a data-loss bug in the inline tab rename that predates the rail and
reproduces in every layout, header strip included: Escape set the input to ''
and blurred it, and the blur handler commits - so cancelling a rename PUT an
empty name, and the tab fell back to its folder label (measured against a live
server: ["rail-alpha","","rail-gamma"]). Escape now calls cancelRename(), which
invalidates the edit so the blur that follows the input's removal is a no-op.

Tests: rail-detail gate, the three ways it turns back off (simple, compact,
horizontal), sidebar-wins, render-on-detail-flip and the plumbing/CSS guards in
test/session-list-layout.test.ts; the rename cancel in test/inline-rename.test.ts
(browser suite), pinned by running it against the old code first. Verified live
against a real server on an isolated instance: detailed/simple/compact/header/
sidebar variants, click-select, the ... menu, inline rename, Alt+N, the in-place
stamp tick and a full settings-picker round-trip including reload.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 04:45:02 +02:00
Codeman maintainer 4cda150493 feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.

DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:

1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
   $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
   `base` -- the interactive terminal front door is always a third-party
   plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
   and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
   button gates on the latter, because reporting only the binary would spawn a
   pane that dies on arrival. When the binary is present but no profile is,
   the run menu offers to install one (POST /api/deepseek/install-profile).

2. The permission switch is an env var, not a flag. The harness has no
   command-line permission option; its sandbox/approval rows read
   DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
   Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
   own workspace-write, which still asks, so the multi-user clamp is the
   only-if-sent branch and clamps to workspace-write, never read-only.

3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
   earned that. The terminal front door reports idle/working/blocked to a
   supervising process over a generic env-gated contract; a generated shim
   (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
   report to /api/hook-event as stop / agent_working / permission_prompt. So a
   dsh session gets definitive respawn triggers, real wait-endpoint signals and
   real Approvals Inbox items instead of output-stabilization guesswork.
   `agent_working` is new (157th SSE constant) and joins
   APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
   alert at once.

The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.

Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.

Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.

Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 03:37:56 +02:00
Codeman maintainer c173ae0264 Merge remote-tracking branch 'origin/master' into worktree-grok-mode
# Conflicts:
#	src/web/public/app.js
2026-08-24 00:32:25 +02:00
Aamer Akhter e3a2fb767f feat(tabs): COD-358 add resizable vertical session rail 2026-08-23 12:21:02 -04:00
Codeman maintainer 3f8c8e99d1 feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.

Grok mixes two existing shapes and the wiring follows from that:

- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
  (--always-approve, grok's bypassPermissions mode; config-level deny rules
  still apply on top). The Run button sends it true, like runAntigravity(),
  and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
  a bare grok spawn is grok's own ask-mode default, which is already safe,
  so only a sent config needs the flag forced off. Cron needs nothing for
  the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
  with mouse support, so it stays OUT of isAltScreenStripMode() and lands
  on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
  (unmeasured against an authenticated composer; documented fallback is the
  'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
  also installs a grok bin), so grok-cli-resolver.ts version-probes every
  candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
  GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
  shared with the dependency registry so doctor and run mode cannot drift.

Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.

Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.

Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).

Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:39:03 +02:00
Codeman maintainer 98e37bf895 feat(sidebar): add a rich session sidebar that carries the home screen's row detail
Session List Layout gains a third option. The old "Left sidebar" becomes
"Left sidebar simple" and is unchanged down to the byte; the new "Left sidebar"
puts on each row what the desktop home rail and the phone overview already show:
when the session was first created, how long it has been in the state it is in,
and a status pill naming that state.

A docked column is not a tab strip. It has width to spare and a row per session
either way, and "name + folder" is the whole story a TAB can tell, not the whole
story there is. This is the information that was missing, and it already existed
one surface over.

Both sidebar values are the same layout, and both set data-session-list="sidebar";
the row detail rides on a separate data-sidebar-detail attribute. That split is
the load-bearing decision here: every one of the ~25 isSessionSidebarActive()
call sites and every html[data-session-list="sidebar"] rule in styles.css and
mobile.css keeps matching both variants without being touched. A third
data-session-list value would have meant auditing and editing all of them.

- Stored values: 'header', 'sidebar' (simple), 'sidebar-rich'. Anyone already on
  'sidebar' keeps exactly the layout they picked — the rename is label-only.
- State classification and the "how long has it been like this" anchor come from
  _mobileOverviewState() / _mobileOverviewSince(), not re-derived, so the three
  surfaces cannot disagree about what "working" means. A working pane repaints
  ~1/s, so its duration is measured from the turn's last Enter: a running turn
  reads "working 12m", not "0m".
- Stamps refresh in place on a 20s clock rather than by re-rendering — a rebuild
  would restart every load spinner and alert animation in the list, twice a
  minute. The clock runs only while rich rows are on screen, and is stopped from
  both render paths and from applySessionListLayout().
- The incremental render path updates the pill, the accent class and the since
  anchor; a tick alone cannot see a state change, and a new turn re-stamps
  lastSubmitAt without changing state.
- applySessionListLayout() now re-renders on a DETAIL change too. simple <-> rich
  leaves data-session-list on 'sidebar' both times, and the meta line is emitted
  by the row template rather than toggled by CSS, so the old layout-only test
  would have flipped the setting and repainted nothing.
- Width: 300px for the extra line. The collapsed 44px rail and the handheld
  drawer are both explicitly held back from it — the desktop rule is (0,3,1) and
  would otherwise out-specify mobile.css's (0,2,1) drawer base and pin a 320px
  phone's drawer to 300px.
- Missing/stale mobile-overview.js degrades to a row with no meta line rather
  than throwing and taking the whole tab strip down.

15 new tests cover the attribute split, the solo-window override, the
detail-change re-render, the row model, both render paths, the clock lifecycle
and the mobile width guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:26:34 +02:00
Ark0N 05c94f5ac0 Merge pull request #304 from Ark0N/fix/workspace-hooks-install
fix: install Codeman hooks into every claude workspace, not just cases Codeman created
2026-08-16 19:31:06 +02:00
Codeman maintainer 210da991d5 Merge christianhaberl's session-sidebar branch, ported to current master
Brings in https://github.com/christianhaberl/Codeman/pull/4 (three commits,
authorship preserved) and adapts it across the 211 commits master gained
since the branch was cut:

- App Settings control re-authored for the set-* surface (PR #278): a
  set-row in Layout -> Tabs, replacing the old settings-item markup the
  branch targeted. i18n description synced.
- Lineage arcs (PR #291, post-branch) are SKIPPED in sidebar layout:
  computeLineagePath()'s U-bridge geometry hangs from the horizontal
  strip's bottom edge and has no meaning against a vertical list. The
  lineage strip-scroll listener now also redraws subagent/ultracode
  connectors while the sidebar scrolls vertically.
- The desktop home tab rail (post-branch) defers to the sidebar: both dock
  the session list flush left, and the rail would render z-ordered under it.
- Active-row reveal unified into _scrollActiveTabIntoView() (#257 landed on
  master after the branch): sidebar mode branches to scrollIntoView
  block:'nearest', and _fullRenderSessionTabs() restores scrollTop alongside
  the #257 scrollLeft restore so ambient rebuilds cannot yank a mid-scroll
  sidebar back to the top.
- Mobile active-tab hoisting the branch guarded against no longer exists on
  master (removed by #257); kept master's order-stable render.

Verified: typecheck, lint, format:check, check:frontend-syntax,
check:public-assets, PostCSS parse of both merged stylesheets, the 26 new
jsdom tests, the structural guard suites, and the headless-Chromium harness
(scripts/verify-session-sidebar.mts) green across all seven layout states
at 1600/1000/393px against current master.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 18:28:15 +02:00
Codeman maintainer f485085174 feat: workspaceHooksEnabled setting as the opt-out for workspace hook installs
Installing hooks into any workspace a Claude session runs in is the right
default, but it takes a decision away from a user who deliberately removed
them: nothing on disk distinguishes "removed on purpose" from "never had any",
so they would come back on the next session create.

Adds the synced workspaceHooksEnabled setting (App Settings -> Agents & CLIs ->
Claude), default ON. OFF restores the older behavior exactly: a Codeman hooks
block that is already present is still refreshed when stale (COD-91), but one
is never added.

Every create path routes through one applyWorkspaceHooks() helper so the gate
cannot apply to some paths only, and the boot-time recovery sweep honours it too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 07:05:10 +02:00
Codeman maintainer c5b59633d8 feat(pi): add Pi (pi.dev) as a sixth CLI run mode (#206)
SessionMode gains 'pi', a first-class backend alongside Claude Code,
OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose
tab identity, welcome button, run-mode entry, cron agentType, Docker and
remote-SSH command defaults, and clone-repo Brain option.

Pi is a different shape of CLI from the other four, and three decisions
follow from that:

- It has NO permission prompts and no sandbox, so there is no
  --dangerously-skip-permissions analog and none was invented. The
  privilege-shaped knob is the tri-state approveProjectTrust, which makes
  pi load and EXECUTE repo-local .pi/extensions TypeScript and install
  missing project packages. clampExternalCliBypassForOwner() therefore
  puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets
  --no-approve even when no config was sent, because pi's own default is
  a prompt the session user could answer themselves. That helper had zero
  test coverage; it now has coverage for all four CLIs.
- Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars
  share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode
  context, so admitting them would widen the allowlist for every mode at
  once. Auth goes through pi's /login or the server's own environment.
  --api-key is deliberately never wired: it would put a provider secret on
  the spawn command line.
- pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the
  main screen with terminal-owned scrollback, and its 0.84.0 fullscreen
  mode is runtime-switchable via /settings; that flip was measured to put
  the pane into the alt screen, which the strip would have corrupted.

pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires
semver-shaped output, because `pi` is a short generic name a stray binary
can shadow; GET /api/pi/status surfaces path and version so a
misresolution is diagnosable rather than presenting as a broken mode.

Docker installs pi in its own --ignore-scripts step so that flag cannot
affect the other four CLIs, and seeds its credentials per-file rather than
whole-dir (~/.pi/agent also holds sessions, extensions and package trees).

Verified end to end against pi 0.84.1 on an isolated instance: resolver
search-dir fallback, flag construction, piConfig persistence across a full
server restart, the trust prompt and its --no-approve suppression, the
rose Run button on the default daylight-blue skin (the nested skin block
eats per-mode gradients unless the rule lives inside it), and the buffer
local-echo policy, which pi tolerates where codex did not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:54:47 +02:00