Commit Graph
229 Commits
Author SHA1 Message Date
Codeman maintainer 06c4c7da16 fix(sessions): scope the launch model to claude and pin it with the advisor (#514, #515, #530 landing)
Maintainer merge-time fixes for the three PRs that landed together on the
session create / launch / persistence path.

#514 findings (bot verdict merge-with-fixes):
- minor, fixed: SessionState.model was published and persisted for every
  mode, so a codex/opencode cron session reported the app-wide Claude
  default it never ran on. toState() now emits it only where the new
  cliTakesSessionModel() holds (registry capability model.source ===
  'claude-settings-file', no CLI id branch). POST /api/sessions uses the
  same helper for its non-claude refusal, so refusal and publication cannot
  drift. Recovery then hands back undefined for other modes on its own.
- nit, fixed: the `model` schema admitted a leading dash (and '.', '[').
  The first character must now be a letter or digit; still a subset of the
  registry's model-claude pattern, so nothing accepted is refused at launch.
- nit, fixed (reject, the consistent choice): `model` with
  attachRemoteSession was silently dropped. Now a 400 INVALID_INPUT, as
  #514 does for non-claude CLIs and quick-start does for remote cases.
  advisorModel (#530) gets the same refusal there. effort and envOverrides
  keep their older silent ignore on that branch so no existing caller breaks.

#515 finding (bot verdict merge, one nit):
- nit, fixed: the types/session.ts @fileoverview described CodexConfig as
  (model, resumeSessionId); it now lists reasoningEffort, bypass,
  animations and renderMode too.

Audit of the merged combination (not reviewed before):
- The conflict resolutions in session.ts (toState), types/session.ts,
  reboot-restore-routes.ts, server.ts (restoreMuxSessions), CLAUDE.md and
  skills/codeman/reference/endpoints.md (+ plugin mirror) keep both sides
  correctly; nothing was lost or doubled.
- A claude session with both `model` and `advisorModel` launches with
  `--model <id>` and ONE merged `--settings` JSON (ultracode + advisorModel,
  or advisorModel beside `--effort <level>`), on the tmux template
  (including the resume || new variant and with the statusLine exporter)
  and on the direct-PTY fallback. Both values (and effort) survive
  restoreMuxSessions onto a dead pane, a reboot restore into a fresh pane,
  and restartCli/dead-pane respawn via _buildRespawnPaneOptions.
- quick-start and ralph-loop take no per-session `model` (matching #514's
  scope, POST /api/sessions only) and launch on the app-wide default, which
  toState now persists for claude, so recovery stays consistent.
- No defect found in the combination beyond the findings above. Noted, not
  changed: advisorModel is still published for any mode a caller sends it
  with (launch-inert there; the UI and skill send it for claude only).

Tests: test/session-model-recovery.test.ts pins the pair through both
recovery shapes for effort ultracode/high/none, the recovery constructors'
fields, the tmux-manager builder hop, and the codex/opencode/shell
non-publication; test/advisor-model.test.ts pins the launch lines and a
real direct-PTY Session's pty.spawn argv; the route test covers flag-shaped
models, attach refusals and the published fields. Docs: SessionState.model
docstring, the reboot-restore-registry header, the golden test comment and
the CLAUDE.md model/advisor bullets.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-04 23:52:41 +02:00
Codeman maintainer b451b3851e fix(mcp): no file text in sync errors, follow relocated config dirs, docs and Settings polish (#521 review)
Maintainer merge-time fixes for the MCP server sync (opt-in mcpSyncEnabled, synced, default OFF).

M1, parse errors echoed config text (secrets included) into the HTTP response and Settings:
smol-toml's TomlError carries a code frame of the offending lines and V8's JSON "Unexpected
token" errors quote source. Both catch sites now go through describeMcpSyncError(): a parse
failure is reported by line/column only ("not valid TOML (line 3, column 21)", "not valid
JSON"), an errno failure by Node's own message (code, syscall, path), the module's own
messages via a McpConfigError class, anything else as "unexpected error". Tests put a secret
on the broken line (TOML, both JSON message shapes, and a write refused at the re-parse that
would have quoted a copied server's env) and assert it is absent from the result and from the
route's response body; they fail against the old code.

M2, CODEX_HOME / CLAUDE_CONFIG_DIR / XDG_CONFIG_HOME were ignored, so a sync could create a
file the CLI never reads and report success: new optional registry field
capabilities.mcpConfig.relocation { envVar, path } (registry data, no id branch; schema
reuses the env-name and no-traversal path rules). Declared for claude (CLAUDE_CONFIG_DIR,
checked in the 2.1.289 binary), codex (CODEX_HOME), opencode (XDG_CONFIG_HOME) and gemini
(GEMINI_CLI_HOME, gemini-cli paths.ts); antigravity follows $HOME only (agy 1.1.12 has no
relocation var). Resolved from the server process env at call time: absolute moves the file,
empty means unset, anything else reports the target with the new status "skipped" plus the
reason and writes nothing. Dedupe is now by resolved file. When a caller overrides `home`
without passing `env`, process.env is not consulted, and the route tests clear those vars so
a CI runner's XDG_CONFIG_HOME can never aim a write outside the temp HOME.

M3, feature undocumented: CLAUDE.md Key Patterns paragraph (opt-in, admin-only, additive
only, backups, re-parse validation, 0600 for copied secrets, names-only responses with
position-only parse errors, capabilities.mcpConfig and relocation), a Settings-Reference row
in the wiki, and docs/cli-registry.md + docs/api-reference.md updated for relocation, the
"skipped" status and the error policy.

Nits:
- N1 Preview/Sync before Save: the UI remembers the saved value on open and says "Save
  settings to turn MCP sync on first" instead of calling the routes; the 403 message also
  says to turn it on and save.
- N2 non-admins in multi-user mode: _applyMcpSyncAdminGate() hides the whole MCP group, called
  from applyMcpSyncVisibility() and the codeman:me event like the CLI-management gate.
- N3 scope chip says "synced".
- N4 "(1 servers)" pluralised; the unsupported list only names installed CLIs (route test
  pins it with a per-test installed set).
- N5 McpSyncResult / McpSyncTargetResult moved to src/types/mcp-sync.ts (barrel export); only
  the route imported them, so no churn.

Verified with an isolated instance (throwaway HOME, own instance and tmux socket) and
Playwright: chip, save-first message, preview rendering and the admin gate.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-04 23:52:41 +02:00
Codeman maintainer ccd6df893f Merge pull request #514 from irisitymichaelgrundberg/feat/claude-session-model
feat(sessions): accept a per-session Claude model on POST /api/sessions
2026-10-04 23:24:43 +02:00
Codeman maintainer e082d8e438 Merge pull request #515 from irisitymichaelgrundberg/feat/codex-reasoning-effort
feat(codex): start a codex session at a chosen reasoning effort
2026-10-04 23:24:39 +02:00
Codeman maintainer 5724e0c8b4 Merge pull request #523 from opticon454/feat/webhook-notifications
feat(notifications): ntfy/Slack/Discord/generic webhook for the push events

# Conflicts:
#	config/test-suites.ts
#	docs/api-reference.md
#	src/web/public/settings-ui.js
#	src/web/routes/index.ts
#	src/web/server.ts
2026-10-04 23:24:39 +02:00
Codeman maintainer 9649b5019b Merge pull request #521 from opticon454/feat/mcp-sync
feat(mcp): sync MCP servers across enabled CLIs

# Conflicts:
#	src/config/cli-registry/schema.ts
#	src/config/cli-registry/types.ts
#	src/web/public/settings-ui.js
2026-10-04 23:24:26 +02:00
DevvynandClaude Sonnet 5.5 2cf37529e9 fix(terminal): address review: Key tester isolates shortcuts, Codex stays on line feed
- app.js: the shortcut dispatcher returns early for events aimed at a data-raw-keys
  field, so Ctrl+W / Ctrl+L / Escape / Alt+1 / Ctrl+K pressed in the Key tester no
  longer kill the session, clear the terminal or close Settings
- stock.ts: drop Codex's esc-enter (a line feed works); no stock CLI declares a chord.
  The esc-enter path is tested through a clis.json override
- tests: unused port (3194), Ctrl+Enter asserts no keypress, shortcut-isolation test
  (verified to fail without the guard)
- docs/comments point at capabilities.newline; set-input class, trailing whitespace

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-03 11:19:42 +08:00
DevvynandClaude Sonnet 5.5 0ff57ce304 feat(notifications): ntfy/Slack/Discord/generic webhook for the push events
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 22:06:29 +08:00
DevvynandClaude Sonnet 5.5 f39c66e4e8 feat(terminal): newline chord as registry data, plus a Key tester in Settings
capabilities.newline replaces choosing the Shift+Enter bytes in the send-key
route. Key tester shows the keydown/keypress/keyup a browser reports.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 21:53:31 +08:00
DevvynandClaude Sonnet 5.5 4398dbfad0 feat(mcp): make sync opt-in and address review
Opt-in (mcpSyncEnabled, default OFF; routes 403 until on). Review fixes:
- codex TOML read/validated with smol-toml: CRLF, inline tables and
  command-less tables no longer yield a duplicate [mcp_servers.x]; the new
  text is re-parsed before writing
- null-prototype tables and own-key checks; unsafe names ignored at every level
- servers switched off in their own CLI (codex/opencode/antigravity) are not copied
- only CLIs that are installed or already have a config file take part
- files receiving env/headers are left 0600; symlinked configs are written through
- one apply at a time (409), unique tmp files cleaned on failure, failed status
- routes set real HTTP status codes; api-reference section; format type single-sourced

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 21:15:31 +08:00
DevvynandClaude Sonnet 5.5 af032fc81a fix(mcp): block __proto__ server names, fix lint; add route and registry tests
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 18:27:07 +08:00
Michael GrundbergandClaude Opus 5.5 3df113fc54 fix(sessions): keep a session's model through recovery, refuse it off claude
SessionState now carries the model a session launched with, and both
recovery constructors (mux recovery and reboot restore) pass it back, so a
recovered session relaunches on the same --model rather than the account
default. A top-level `model` sent with any other CLI is refused, since
those take their model in their own config object, and an empty string
means no per-session model, as it does for modelOverride.

CLAUDE.md now describes both routes for a Claude model. The tests pin
which of `model` and `modelOverride` reaches the launch and which the
case file, and that a model opening with a dash renders as --model's value.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 17:24:12 +02:00
Michael GrundbergandClaude Opus 5.5 0ae39cdd94 test(codex): pin reasoning effort through quick-start and the multi-user clamp
Both create schemas now refuse an unknown level, and a non-granted owner's
codexConfig keeps its reasoningEffort when the clamp forces bypass off.
docs/architecture-invariants.md lists the two --config values codex now
takes from codexConfig.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 17:16:18 +02:00
Michael GrundbergandClaude Opus 5.5 4123d229f4 feat(sessions): accept a per-session Claude model on POST /api/sessions
POST /api/sessions takes an optional `model`, and a Claude session
launches with `claude --model <id>`. It wins over the app-wide default
model and writes nothing to disk, unlike `modelOverride`, which stays as
it is and still writes the case's .claude/settings.local.json.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 16:59:45 +02:00
Codeman maintainer 5f5de5827e Merge pull request #503 from JDProfresh/feat/file-viewer-markdown
feat(files): render markdown in the File Viewer, with Lines/Wrap toggles
2026-10-01 11:09:26 +02:00
Codeman maintainer d5ffc22f4a Merge the input-delivery fix from master
fix(input): deliver API prompts through tmux so their Enter is not lost
2026-09-28 16:45:27 +02:00
Codeman maintainer c2dfc775a3 fix(input): deliver API prompts through tmux so their Enter is not lost
A prompt posted to /api/sessions/:id/input without `useMux` was written
into the pane in one piece. Claude Code (measured on 2.1.283) takes a
`<text>\r` burst of about a hundred characters or more as a paste, so the
trailing `\r` landed as a newline in the composer and the prompt sat there
unsent while the route answered 200. A later raw `\r` did not recover it;
a tmux `send-keys Enter` did. Short prompts submitted, which is why it
looked random. The same stranding was seen with Codex and OpenCode.

A plain prompt (printable text plus exactly one trailing `\r`, detected by
`isPlainPromptInput()`) now goes through `writeViaMux` even without
`useMux`: the text is typed, Enter is pressed as its own key, and the
SubmitVerifier re-presses it while the prompt is still on the composer.
The write is awaited, since the browser's POST fallback sends frames one
at a time and a following keystroke must not overtake the Enter. Raw
frames (escape sequences, bracketed paste, a line feed, a bare `\r`) and
an explicit `useMux: false` keep the direct write.

Verified on an isolated instance: the 239- and 104-character prompts that
stranded (at +1 s, at +50 s on ultracode, and on a warm session) all
submitted on the first Enter with no `useMux`. The phone's local-echo
flow (a burst, then its `\r` as a separate write) was measured unaffected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:44:59 +02:00
Codeman maintainer 714050fe8a fix(terminal): merge-time fixes for #494
- Skip and latch a bounded Shell window once the browser is at xterm's
  scrollback cap (scrollback + rows): a 1 MiB window of short lines can carry
  more rows than the browser can ever hold, so it replayed and re-captured on
  every scroll-to-top with no 60 s back-off.
- Label a replayed bounded window 'tail' even when the capture was byte-capped,
  so the banner keeps offering Load full history instead of calling the rest
  unrecoverable.
- Pin GET /terminal?full=1&tail=<n> in the route tests: full-history source,
  truncationReason 'tail', and the closing relative cursor move survive the cut.
- Log the bounded skip via _logScrollRouting('repull-skipped-bounded').

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:30:40 +02:00
JD 5e27043bf7 feat(files): render markdown in the File Viewer, with Lines/Wrap toggles
Clicking a .md in the Files panel showed wrapped source with an Edit
pencil and no way to see it rendered, although marked + DOMPurify were
already on the page for the Response Viewer. The viewer now renders
.md/.markdown through that same pipeline (one parser, one click
delegate) with an MD pill back to source, and the plain-text view gains
Lines (CSS-counter gutter) and Wrap toggles. All three persist per device
in their own localStorage keys.

- Relative images are rebased onto the workspace-confined file-raw route
  under the document's directory, built inside a <template> so no fetch
  fires before the rewrite; a failed load degrades to alt text. Relative
  links become a.rv-path so the existing delegate opens them in the
  viewer; fragment and http(s) links are untouched.
- The rendered container carries data-i18n-skip so the translator does
  not rewrite the document's prose.
- Markdown fetches the route's 10000-line ceiling; other text keeps 500.
- avif renders inline (file-content image set, file-raw MIME map), and
  avif/ico printed paths open the viewer instead of tailing bytes. .md
  deliberately stays with the tail viewer for printed paths.
2026-09-28 01:49:37 -04:00
DevvynandClaude Sonnet 5 8cef31086b fix(docker): persist CLIs installed from Settings across container updates
The image sets NPM_CONFIG_PREFIX=/opt/codeman-cli, which is image content, so
Update-Codeman.sh discarded every npm-installed CLI (dsh, pi). In the Compose
container, POST /api/clis/:id/install now installs into ~/.local on the
persistent home mount, and ~/.local/bin is appended to the image PATH.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-25 08:56:20 +08:00
Codeman maintainer d67da5c9d0 fix(input): an oversized paste no longer poisons the durable input queue (#484)
A single input over MAX_INPUT_LENGTH (64 KiB) was queued for reliable
delivery, refused by both transports (the WebSocket silently, POST with a
400), and never dropped: the client treated the 400 as transient, so the
frame was re-sent every 2 s forever, blocked every later input for that
session, and came back from localStorage on every reload.

- Client: a paste over the frame limit is split into in-limit frames
  (never cutting a surrogate pair) delivered in seq order; over 1 MiB, or
  an oversized mux write, it is refused with a toast and never queued.
- Client: the POST drain drops a frame answered 400/413; a WS error ACK
  drops it too; frames over the limit persisted by an older build are
  pruned on load.
- Server: the WebSocket answers an oversized sequenced frame with
  {t:'ia',seq,err:'too_large',max} instead of silence (an older client
  reads that as a plain ACK and drops it); the POST schema uses
  MAX_INPUT_LENGTH instead of a second 100000 limit.

Verified end to end on an isolated instance: a 110 KB paste reached the
PTY byte-identical over both the WebSocket and the POST path, a poisoned
120 KB persisted frame was pruned on load, and a 2 MB paste showed the
refusal toast with nothing queued.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 18:17:06 +02:00
69a71287e6 fix(sessions): stop pinning the w1-myapp placeholder as Claude's /resume title (#457)
* fix(sessions): stop pinning the w1-myapp placeholder as Claude's /resume title

Local claude spawns passed the tab name as `--name`. That flag is not only the
cross-session peer name: it is also the prompt-box label, the `/resume` picker
entry and the terminal title, and a pinned title stops Claude generating its own
(`customTitle ?? aiTitle`). So every conversation of a case was listed in
`/resume` as the same `w1-myapp`, and none of them got a generated title. On one
workspace, 34 of 34 conversations spawned with `--name` had no ai-title, while
every conversation spawned without it had one.

Only a name the user chose is pinned now: `Session.cliPinnedName` is the name
when `nameSource === 'manual'`, carried to the builders as a separate `cliName`
so the tab/mux name is untouched. Placeholder and auto names let Claude title
the conversation again.

A rename in Codeman also reaches `/resume`: the new name is appended to the
conversation's transcript as the `custom-title` row `/rename` writes (never
creating the file, never writing an empty title). For a pane spawned without
`--name` this holds immediately; a pane spawned with one re-appends its own
title each turn, so there the new name holds from the next spawn.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(sessions): skip no-op renames and docker sessions when syncing the /resume title

A same-name PUT (the Session Options field saves on blur and recomposes the
unchanged placeholder) no longer flips nameSource to manual or appends a
custom-title row, and docker sessions skip the host transcript scan since their
transcript lives in the container. The skill pages no longer use a w<N>- name
as the peer-name example, and the changeset notes the re-append caveat.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: record that nameSource decides --name and renames reach /resume

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: codeman-local <codeman@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-24 01:48:33 +02:00
DevvynandClaude Opus 5.5 0a52a99ca9 feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis

Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI

Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).

Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.

Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).

Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.

Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.

Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.

27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else

window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.

Fixed in two places:

- server.ts: after building `available`, intersect the nine real
  SessionMode ids against `enabledClis()`. git/cloudflared (utility
  binaries, not CLI registry entries) and deepseekBinary (a secondary
  installed-only flag for the "add a profile" affordance) are deliberately
  left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
  `window.__codemanCliAvailable` in place and refreshes the welcome screen,
  the mobile overview and an already-open Run menu, mirroring the existing
  `installDeepSeekProfile()` pattern for the same "injected once, needs an
  explicit patch" reason — without this half, the server-side fix alone
  still left every surface stale until the next reload.

New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.

Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap

Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:

- A new "Status as of 2026-09-22" section up top: what's actually
  implemented (verified by grepping the routes/schema/UI, not just trusting
  the commit messages), the availability-flag staleness bug found and fixed
  in this session (commit 0c77dd0a) with its devbox verification record, and
  one real outstanding gap.

- The outstanding gap: a custom CLI created via Phase 5's write API has no
  way to actually be launched. The Run menu is static per-mode markup with
  no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
  custom entry, confirm it can be launched" verify step was never actually
  exercised against this. Documented with two candidate fixes, neither
  started.

- Each phase's checklist flipped to [x] where confirmed present in the tree,
  Status lines updated from TODO to DONE, and the two originally-open
  questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
  marked resolved against what actually shipped.

No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* docs: add the CLI-registry deployment plan and the parked Copilot plan

Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:

- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
  series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
  is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
  name in docs/cli-enable-disable-plan.md's own header as a sibling plan
  tracked separately — kept for continuity, not active on this branch.

The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* fix(cli-registry): render enabled CLIs in launch surfaces

* test(cli-registry): update frontend branch guard

* fix(test): isolate suite from deployment environment

* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent

shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.

Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): hide shell's toggle entirely instead of greying it out

A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical

renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* style: prettier fixes from the master merge

* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels

Four gaps found verifying #476 against the #343 review trail:

- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
  (the nine per-CLI resolvers and the generic registry one) caches in its own
  closure, with a negative-cache backoff of up to 5 minutes, and nothing
  cleared them. invalidateCliExecutableResolvers(binaries) now drops those
  caches per binary; install (success or failure), create, edit and delete
  call it plus invalidateCliResolverCache(id). Before this, a CLI installed
  from Settings could fail to launch for minutes, and an edited custom entry
  kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
  ignoring the entry's searchDirs and the login-shell lookup that spawn and
  the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
  sit behind an explicit confirm; the confirm now names the exact command,
  which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
  of the word ("Claude", "Codex"). It uses the registry label again, which is
  identical to the old static table for every stock CLI (now pinned).

14 new tests; 9 of them fail against the previous head and pass here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs

Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
  does not parse or has group/world permission bits instead of overwriting it
  (isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
  existence/duplicate checks inside the serialized step, plus a unique tmp
  name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
  section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md

Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
  body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
  map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
  label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
  "Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* test(cli-registry): pin the #343/#347 findings #476 makes reachable

A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 01:48:26 +02:00
Codeman maintainer 8536aaef7b Merge pull request #473 from irisitymichaelgrundberg/feat/session-watching-badge
feat(approvals): let a session watching its own background work keep quiet (#468)

# Conflicts:
#	src/config/cli-registry/stock.ts
2026-09-23 11:32:11 +02:00
Codeman maintainer 6a01412af9 Merge pull request #466 from irisitymichaelgrundberg/feat/pane-exit-reporting
feat(tmux): report that a pane's agent has exited (#446, part 1)
2026-09-23 11:32:02 +02:00
DevvynandClaude Opus 5.5 02e40f506b fix(docker): gate gh/az seeding on its switch; no shared git sign-in for non-admin clones
Addresses the review on #472.

- CRED_STORES: `.config/gh` and `.azure` now carry `enabledByEnv`
  (CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ), and resolveDockerCredentialArtifacts
  skips a store unless that variable is exactly `1`, read at container
  create. A host that merely has ~/.config/gh/hosts.yml or a plaintext MSAL
  cache no longer copies them into every case container. Tests: the default
  environment seeds neither even with the files present, and each store
  follows only its own switch.
- Multi-user mode: a non-admin's Clone Repo clone and preflight run with
  `git -c credential.helper=` (GIT_NO_CREDENTIAL_HELPERS, placed before the
  subcommand), so the server account's helpers are never lent to them.
  Verified against a real private repo that it also clears the URL-scoped
  credential.<url>.helper entries, and that public clones still work.
  Tests: the argv in test/git-clone.test.ts, and the route decision
  (non-admin cleared; admin and single-user kept) in
  test/routes/case-clone-credential-helpers.test.ts.
- Docs: recreate the case container to pick up seeds (docker/README.md,
  Docker-Cases wiki, docker-cases.md); the multi-user behaviour in
  docker/README.md and security-architecture.md; "functionally unchanged"
  instead of "unchanged" for an image built with both switches off
  (server.Dockerfile comment, README, changeset).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw
2026-09-23 14:44:26 +08:00
Michael GrundbergandClaude Opus 5 9c286eeddf fix(session): persist an exit retraction, and let tests reach the watcher
Ten findings from a two-model review of this branch. Both reviewers cleared the
detection logic itself; everything here is a gap around it.

A route that starts a command in a pane now PERSISTS as well as broadcasts.
`/interactive` and `/shell` did neither before, and the pane-exit watcher cannot
cover for them: its next tick finds `paneExit` already cleared in memory,
reports no change and writes nothing, so `state.json` kept saying the agent had
exited for as long as the session stayed quiet. Nothing reads that record for a
decision yet, which is exactly why it had to be fixed now — part 2 is designed
to read it. The `clearPaneExitForNewPane()` docstring claimed its callers
already persisted; that claim was false for these two, and now says what the
caller owes instead.

The watcher's four guards were unreachable by any test. `refreshPaneExits()`
opened with `if (IS_TEST_MODE) return;`, so the read gate, the in-flight
suppression, the generation counter and the empty-read rule could each be
deleted with the whole suite green. The tmux call moves into `readPaneRows()`,
which a test subclass overrides — the shape `runRemoteReconnectTick` already
uses in this file for the same reason — and the test-mode gate moves with it, so
what a test cannot do is spawn a process rather than exercise the bookkeeping.
Each of the four guards now has a test that fails when it is deleted.

The muted status dot turned out to be a specificity fight on three surfaces, not
two. `.tab-status.error` was not excluded, so a session whose agent exited and
whose PTY-exit breaker then tripped lost its red dot to the mute — the state the
browser answers with a "restart it?" confirm, and a needs-you colour by the same
argument that protects the two alert classes. And mobile.css gives a `busy` dot
a 9px size and a green glow with `!important`, while `status` stays `busy` for a
pane whose agent died mid-turn, so a phone rendered a grey dot still wearing the
green halo beside a badge reading "exited". Both measured against the real
stylesheets, both now excluded, and the CSS test reads mobile.css too instead of
being structurally blind to half the problem.

Six comments said things that were not true. Two named the stats collector as
what replaces a restored reading, which is the opposite of the design. The
interval constant argued that 2000 ms keeps a read inside a tick, when the
5000 ms exec timeout means it cannot — which is why the in-flight guard exists.
`MuxSession.discovered` did not say the flag is permanent, though `saveSessions()`
serializes it. The empty-read docstring claimed a distinction that `|| true`
makes impossible. The invariants doc promised more than its drift test delivers.
And CLAUDE.md had no pointer at all, leaving its two hardest prohibitions
("never set `status: 'error'`", "never null the pid") only in the file it is
meant to route people to.

Refs Ark0N/Codeman#446.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 15:24:09 +02:00
Michael GrundbergandClaude Opus 5 74884a20eb feat(approvals): let a watching session keep quiet, and fix the tui gate
The badge alone left the row in NEEDS YOU, which is the thing the issue was
about. The fix is the alert that does not fire.

An idle prompt from a session that is watching its own background work now
opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to
`notePrompt()`, which sets `acknowledgedAt` and records why in a new
`acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has
always meant "the alert this prompt armed is spent", and the prompt itself
stays pending, answerable and available as Read My Mind context. A wrong label
therefore costs a card that does not blink, never an alert that was never
created.

Every surface follows from that. The broadcast carries the reason, so a live
page declines to arm the tab alert and raises no desktop notification. The push
is skipped, since a false alarm is hardest to ignore on a phone. A reloading
page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And
`classifySession()` now reads it too, which is a pre-existing bug fixed here:
acknowledging on one device cleared the alert everywhere except `codeman tui`.
It re-arms for free, because the next idle prompt supersedes the item and is
built fresh. Only `idle` is eligible, so a dialog that blocks the agent still
goes red whatever else it started.

The label is pane-derived and therefore prompt-injectable, so it is now read
from the last two rows of the screen only, with Claude's pattern anchored on
the `·` its footer joins items with, ANSI-stripped and length-capped at the
source. An agent that prints `· 1 monitor ·` into its own output finds no
match.

Verified on an isolated beta: a session that armed a monitor took its idle
prompt acknowledged with no alert on any surface, wore the badge, and showed
"quiet, watching 1 monitor" on its still-answerable card; the same session with
the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts`
pins both directions across all four surfaces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 08:10:05 +02:00
Codeman maintainer 4205f6930f fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine.

**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.

**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.

**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.

**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.

The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:58:33 +02:00
Codeman maintainer fe3bd0074c fix(custom-model): split the two confirmation questions, and seed the API key the way claude reads it
Two findings from the review of 5fc391a4, both fixed here rather than sent back.

**The API-key trust seed never matched a real key.** `seedApiKeyTrustFile()` wrote the
key verbatim into `customApiKeyResponses.approved`, but Claude Code stores and compares
only the last 20 characters (`key.trim().slice(-20)`, applied on both the write and the
lookup). For any real key the seed missed, so claude stopped at the interactive
"Detected a custom API key in your environment" prompt, whose default is
"No (recommended)": the launch hangs, or silently refuses the key this feature just
injected and falls through to an OAuth login the isolated config dir does not have. It
survived review because a keyless llama.cpp/llama-swap endpoint uses DEFAULT_API_KEY
('local-dummy-key', 15 chars), where slice(-20) returns the whole string and the seed
matches by accident, and every test used a key shorter than that. Now truncated through
`truncateApiKeyForTrustFile()`, with a test using a 57-character key that also asserts
the full credential never reaches that second file.

**One `confirmed` flag answered two different questions.** The context-floor warning
("this model's window is below what this CLI needs") and the swap-conflict warning
("loading this unloads the model another session is using") shared it, and the context
check runs first, so a user clicking "launch anyway" past the context warning silently
consented to evicting someone else's model. They are about different people, so an
answer to one is not consent to the other. Both routes now read `confirmedContext` and
`confirmedSwap` independently; the legacy `confirmed` still means both, because it
shipped in this feature's HTTP-API-only cut and an existing caller must keep working.
The frontend answers each question with its own flag and accumulates them, on the
one-shot path, the restart path and the batch carry-forward alike.

Also from the same review: the swap-confirm dialog no longer renders " are currently
using ..." when multi-user scoping leaves the affected-session list empty (the swap is
blocked regardless of ownership; only the NAMES are scoped), and the per-endpoint
llama-swap log tails are closed in `WebServer.stop()` instead of only by the idle sweep
whose interval that same teardown disposes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:32:50 +02:00
Codeman maintainer 1a99b5836c Merge pull request #430 from opticon454/custom-model-run-menu 2026-09-19 12:25:11 +02:00
Ark0N 2c3ccdf030 Merge pull request #439
feat(remote): wake a sleeping host (Wake-on-LAN) from input, banner and native magic packet
2026-09-19 12:18:18 +02:00
DevvynandClaude Sonnet 5 5fc391a47c fix(custom-model): address fourth pre-merge review + merge upstream master (Ark0N)
Merged upstream/master (22 commits: reboot-restore recovery feature,
terminal keycode229 recovery work, install.sh/CLI-catalog generator
changes, CHANGELOG/version bump to 1.30.0) into this branch. No
conflicts; git auto-merged every overlapping file (CLAUDE.md,
docs/api-reference.md, app.js, index.html, styles.css, routes/index.ts,
session-routes.ts, schemas.ts, server.ts).

Two required fixes from the latest review:

1. privilegedEnvKeys widening (stock.ts) changes behaviour outside this
   feature. The reviewer decided to keep both CLAUDE_CODE_MAX_CONTEXT_TOKENS
   and CLAUDE_CONFIG_DIR listed (types.ts's rule that every traffic-
   redirecting var this feature introduces must appear there stays
   literally true), and asked for the real consequences documented
   instead of hidden:
   - Corrected session-env-clamp.ts's fileoverview, which stated the
     opposite of what the code now does (reboot-restore's clamp call
     used to be able to strip nothing for claude; it now strips a
     persisted CLAUDE_CONFIG_DIR for a non-granted owner).
   - Corrected the rationale comments in stock.ts: privilegedEnvKeys
     has exactly one consumer (ownerClampedEnvKeys, feeding the
     generic envOverrides clamp on create/quick-start/reboot-restore),
     not the custom-model routes.
   - Added a CLAUDE.md line to the CLAUDE_CONFIG_DIR gotcha covering
     the admin-only-in-multi-user-mode and reboot-restore-strips-it
     consequences.
   - Added a "Claude multi-user clamp" test next to the existing
     DeepSeek/OMP ones, pinning the new stripping behaviour.

2. GET .../running-status (custom-model-routes.ts) no longer passes
   the raw llama-swap `cmd` field (the literal launch line, which can
   carry model paths and --api-key) to the browser -- the frontend
   only ever reads model/state, cmd exists solely for server-side
   parseCtxFromCmd() during discovery. Added a test asserting the
   response never contains cmd or a planted secret.

Also regenerated config/clis.stock.json and install.sh's catalogue
block (npm run generate:cli-catalog) to clear drift introduced by the
upstream merge, since it was failing the sync check.

Left to the reviewer, as they said they'd take at merge: the two
"comments pointing at removed code" cleanups, the two stale CLAUDE.md
counts, and the small items list (mode==='claude' frontend branch,
isCliAvailable() unknown-id gap, shared confirmed flag ordering,
one-shot cancel toast severity, pumpLlamaSwapLogTail buffer cap).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-19 18:05:51 +08:00
RandalixandClaude Opus 5 5bb489addb fix(remote): authorize the attach wake first; tell the caller what happened to its bytes
Review round 3 on #439.

- The attachRemoteSession branch of POST /api/sessions ran `ensureHostAwake`
  before the multi-user gates, so a non-admin could have any configured
  host's `wakeCommand` spawned (or a packet broadcast) and the request held
  for the wake budget, then be refused for the workingDir. The admin gate
  now comes first, before the host is even looked up; remote hosts are
  admin-only infrastructure everywhere else. Route test: wake spy empty,
  403.
- The non-wait input route answers `{buffered:true}` when the registry took
  the chunk and `{buffered:true, dropped:true}` when it was over the cap
  and is gone (`RemoteInputOutcome` gains 'dropped'); additive to the bare
  `{}`.
- The send-and-wait path answers OPERATION_FAILED when the host never comes
  back, like create and attach, instead of writing into the stalled pane
  and reporting delivered:true plus a timeout.
- The flush writes with `fromUser: true`, so a first prompt buffered
  through a wake can still name the tab.

Docs: api-reference (input route), remote-sessions.md (two invariants),
CLAUDE.md key pattern.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-19 11:39:50 +02:00
Michael GrundbergandClaude Opus 5 e0d4477edc fix(terminal): keep the geometry replay to the pass that can converge
Three follow-ups to the source gate, each one measured rather than reasoned.

A pane already drawing at the size the client just requested is left alone. The
replay runs at `dimsAfterLoad`, so it can only change what is on screen if the
pane was drawing at some other size; when the reported geometry already IS that
size, the second pass captures the identical frame and pays a full reload to do
it, including a visible re-flash, a dropped and reopened WebSocket and a deleted
xterm snapshot. That equality is the signature of a clamp rather than a race:
`getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` does not, so a
terminal narrower than 40 columns or shorter than 10 rows reports a pane
permanently bigger than itself and replayed on every tab switch without ever
converging. A race never produces the equality, since its premise is that the
pane was still at the size it was asked to leave. The declined-resize case does
not produce it either, so that one still costs the single capped attempt and
needs the pane-ownership question this does not touch.

The full-history re-arm is unreachable and now says so. A pass that consumed the
flag sent `full=1`, and the route answers `full=1` with `mux-full-history` or
`history`, never `mux-visible`, so the source gate already rules out every such
pass. The line stays for the invariant, but its comment no longer reads as if a
page load retries, and the suite pins that it does not.

The response no longer reports geometry for a body that carries no capture. The
full-history path writes `capturedGeometry` from the cursor query and then
returns '' for a pane holding nothing visible, which drops the source to
`history` with the geometry already recorded: a `full=1` request whose capture
reported 100x50 and returned nothing answered `source: "history"` with both
fields set. Nothing acted on it, because the client ignores geometry on any
other source, but the field said a frame had been drawn at a size when none had.

The browser stub now derives `source` from the request the way the route does,
rather than answering `full=1` with `mux-visible`, which the route cannot
produce. Each case reaches a visible-frame response the way production does, by
not being the first select of the page. Three cases pin the new behaviour and
each fails without its guard: the clamp case sees two fetches instead of one,
the scope case and the full-history case both see a replay the gate forbids, and
the width case sees one fetch instead of two.

The changeset now describes the change from 1.29.x rather than the difference
between the two commits on this branch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
Michael GrundbergandClaude Opus 5 5cfb98fb8b fix(terminal): compare capture geometry only on a visible-frame response
Only a visible-frame capture positions its rows absolutely, so only that frame
can be damaged by a terminal of the wrong size. A `full=1` body is linear
scrollback closed by a relative cursor move, which is relative precisely so the
browser's row count need not match the pane's, and a `history` body is the byte
stream, which carries no row alignment to protect. The geometry comparison ran
on all three, so it fired most often on the one response it cannot help:
`_fullHistoryLoaded` is empty on the first select of every non-shell session per
page, and a session whose pane a desktop tab holds too tall to ever fit then
paid a second whole-scrollback capture, reset and replay on every page load and
every first tab switch.

`framePositionsRowsAbsolutely` gates both the captured-geometry comparison and
`sizeMovedUnderLoad`. A size that moved under a byte-stream or scrollback replay
is healed by xterm's own reflow plus the SIGWINCH the trailing `sendResize`
already sends.

A pane WIDER than the terminal damages the same frame a second way, so
`captureCols` is now compared rather than only logged. `formatPaneSnapshot`
paints each row out to the pane's own width, so a narrower browser wraps every
painted row, and the wrap on the last one scrolls the whole frame up by a row.

The terminal response no longer falls back to `session.ptyCols`/`ptyRows` when
the capture reported no geometry. The cursor query is what produces the absolute
addressing in the first place, so a capture that lost it returned a raw frame
that was never positioned, and a byte-history response was never positioned
either. Naming the session's own PTY size there described a frame that does not
exist and invited a repair for damage that is not present. `_ptyCols` is also
written only by `resize()` while the PTY is spawned at the size queried from
tmux, so it can be wrong on its own terms. Both fields are now absent instead,
and the `Session` getters added for that fallback go with it.

Two browser cases cover the new behaviour and each fails without its fix: a
`mux-full-history` response with both dimensions mismatched asserts one fetch
(two without the gate), and a `mux-visible` response wider than the terminal
but short enough to fit asserts two (one without the width comparison).

Corrects a claim in the comment above `capturedGeometry` in tmux-manager.ts.
Both replay paths do not address rows absolutely; the full-history one ends in a
relative move, which is the whole reason the gate is right.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
Michael GrundbergandClaude Opus 5 3edf9aae2f fix(terminal): replay a pane capture at the geometry it was taken at
A visible-frame capture repaints each row at an absolute position, counting up
to the pane's height. A terminal shorter than that clamps every address past its
own height onto its last line. The overflow rows then overwrite one another, and
the rows underneath are lost. Replaying a real 50-row capture into a 30-row
terminal rendered 28 lines of a 45-line command and drew the frame twice.

Nothing in the response said what height the frame was built for, so the client
could not detect this. A capture now reports the geometry it was really taken at
through `capturedGeometry` on `PaneCaptureOptions`, and the terminal response
carries it as `captureCols` and `captureRows`. When the captured pane is taller
than the terminal, or the size that produced the capture did not survive the
load, `selectSession` replays once at the size that stuck. `resizeRetry` caps
that at one attempt, so two competing fits cannot trade replays forever.

The retry re-arms the full-history flag only when the pass that ran had consumed
it. A tab switch takes the bounded tail, so its retry takes the tail too:
clearing the flag unconditionally would upgrade that switch into a fresh
scrollback capture the user never asked for, which the route's own comments put
at tens of megabytes.

What this repairs is a capture that won a race against the resize meant to
precede it. It does not repair a capture whose pane was too tall because
`Session.resize` declined the resize outright, which it does for a small
viewport while a desktop viewport's size claim is live. The retry re-sends the
same declined resize and captures the same pane, and `resizeRetry` then stops
it. Repairing that means changing who owns the pane size, which is a policy
question this does not touch. The reported geometry still helps there, because
the client can see the mismatch at all rather than being blind to it.

Follows #395, #396 and #397, which fixed the other ways the replayed frame and
the terminal could disagree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
RandalixandClaude Opus 5 1040f6c489 fix(remote): a proxied host is reachability-unknown; scope remote: SSE per session
Review round 2 on #439.

1. The bare TCP probe connects to host:port, which a host behind a jump host
   or SOCKS proxy does not answer even while ssh works. Acting on that
   verdict drew a permanent banner over a healthy session, replaced a real
   "needs tmux" error with "not reachable" in quick-start, and - with a wake
   target - buffered every HTTP input for the life of the session, since the
   readiness poll could never succeed. `WakeableRemote` now carries
   `jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such
   a host into reachability-UNKNOWN: input is delivered, `checkReachable` /
   `checkHostReachable` answer `null` (never `false`), `ensureHostAwake`
   returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate
   fires on `=== false` only, and `GET …/reachability` reports
   `reachable: null, probeable: false` so the banner has nothing to key on.
   A wake target can still be fired for it, blind: no readiness poll, no
   reattach, no toast - the response says only whether the packet went out.

2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake
   has no session yet, so the registry names the requesting user
   (`ensureHostAwake({ requestedBy })` -> `username` in the payload) and
   `deriveSseHint` routes on it; with neither it fails closed to admins.
   Single-user mode is unaffected.

Smaller, from the same review:

- A flush write that fails now drops the remaining buffer (logged) instead
  of retaining it: the wake still resolved and marked the host reachable, so
  the retained chunk waited for the NEXT wake and was replayed hours later,
  after everything typed since. Same policy as the oversized paste.
- The banner polls on tab activation (a user action) and on its 30 s timer
  only for a host with a wake target; a timer connecting to a host Codeman
  cannot wake is the traffic invariant #2 rejects keepalives for. A proxied
  host is never polled.
- `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP
  socket refuse under VITEST, as remote-files.ts does. The guard caught a
  leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the
  probe but still polled readiness with the real one, so the shutdown test
  had been connecting to a production address. The poll now uses the
  injected probe.
- docs/remote-sessions.md is additions only again (the reformatting is
  gone); the architecture-invariants overlap resolved itself in the merge.

Live, against a throwaway instance with a non-routable ghost host: proxied
-> no probe, no wake, the genuine ssh error after 10 s; direct (control) ->
probe, magic packet, "did not come back" after the 40 s budget.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:46:11 +02:00
RandalixandClaude Opus 5 e271a65e79 Merge origin/master into feat/remote-host-wake
Resolves CLAUDE.md count tables (route counts recounted on the merged
tree: 235 handlers, sessions 37) and keeps both the host-wake and the
reboot-restore banner in index.html.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:20:41 +02:00
Devvyn 56209e7829 Merge remote-tracking branch 'upstream/master' into feature/run-menu-custom-model-picker 2026-09-19 03:25:26 +08:00
DevvynandClaude Sonnet 5 afb6754453 fix(custom-model): address third pre-merge review (Ark0N)
Blocker 1: the loading banner hides itself ~200ms after it reopens.

- _showCenterStatus reuses one shared DOM node; dismiss() scheduled
  el.hidden = true 200ms later with nothing to cancel it. On the
  Claude path, switchingToast.dismiss() is followed by one same-
  origin request (5-30ms locally) before _watchLlamaSwapLoading opens
  the new banner -- well inside that window -- so the stale timer
  fired against the shared node and hid the fresh banner, leaving the
  whole model-load wait with no progress text, no log line and no
  reachable Cancel button.
- Fixed by parking the pending timeout on the element and clearing it
  at the top of _showCenterStatus. Added a regression test that
  reproduces the exact repro (open, dismiss, reopen 20ms later,
  advance past 200ms) alongside the existing Cancel-button DOM tests;
  confirmed it fails without the fix and passes with it.

Blocker 2: the swap-conflict warning named other users' sessions.

- Both affectedSessions scans (POST .../custom-model and quick-start)
  walked the whole session map with no ownership filter, so in multi-
  user mode a non-admin pointing their own session at a shared
  endpoint learned another user's session name and id -- which with
  autoNameSessions on is that user's own prompt.
- The swap is still blocked pending confirmation regardless of
  ownership (a foreign session is just as real a disruption); only
  which ones get NAMED back to the caller is scoped, via the
  already-imported canAccessOwned. Added a two-owner test to
  test/routes/session-custom-model.test.ts covering both the
  foreign-owner (blocked, not named) and same-owner (named) cases.

Smaller ride-along fixes:

- server.ts boot recovery now passes contextLength into
  applyCustomModelInjection, so CLAUDE_CODE_MAX_CONTEXT_TOKENS is
  correctly rebuilt into _envOverrides after a restart instead of
  surviving only because tmux retains the old setenv.
- pumpLlamaSwapLogTail's finally now deletes by IDENTITY, not just by
  key, so an aborted pump finishing after a newer entry was created
  for the same endpoint can no longer delete that newer entry and
  orphan its connection.
- docs/custom-model-endpoints.md now notes that clearing a custom
  model removes injected keys by name, including CLAUDE_CONFIG_DIR --
  so a session that also had CLAUDE_CONFIG_DIR set via envOverrides
  (the per-client-account case) silently falls back to the default
  account on clear.

Left for later, as flagged in the review itself: the quick-start
case-scaffolding/cancel ordering (real behavioural reordering across
a large handler, too risky to make without a live re-test), and
retiring runCustomModelEntry's mode === 'claude' branch behind a
launchStrategy registry field (explicitly deferred by the reviewer to
"the next one").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-19 03:09:36 +08:00
Codeman maintainer bb8ada7e5f fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does.

1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred
   it from the name: a session the user renamed by hand to something shaped
   like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the
   next prompt overwrote their name. The route persists right after, so the
   loss went to disk. `restoreMuxSessions()` already passes it.

2. The already-live sets were snapshotted once before a loop that awaits a
   real `startInteractive()` per entry, so by the tenth entry the snapshot
   was tens of seconds old and a conversation resumed by hand from the
   Resume list in that window was invisible to it: two panes on one
   transcript, the exact thing the check exists to prevent. Both sets are
   now read per iteration, and the late case is spent rather than re-offered
   for the same reason the batch case is.

3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path.
   The stamp predates the reboot and the pane is new, so honouring it meant
   one click had every restored session type `continue` into itself about a
   minute later, unattended, against the route header's own promise that a
   restored session comes back idle and disarmed. The setting stays ENABLED,
   so it re-arms on the next real limit message. A Codeman restart still
   re-arms from the stamp, because the limit footer will not reprint on its
   own; the new option exists only to tell the two paths apart.

4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()`
   and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession`
   performs that it was missing. Cosmetic, but a run left open reads as
   still going in the away digest.

5. A restored claude session gets `seedAgentSessionPreamble()` like both
   create paths, so the agent skill's bootstrap stays a two-line loader.

6. The heuristic's container comment was wrong in one direction and quiet
   about the real gap: after a genuine host reboot a containerized Codeman
   sees the host's short uptime and the banner does appear. What it cannot
   see is a container-only restart, which is where this would help most.

7. The banner is hidden in a solo window, which shows one session and has
   no tab strip to put restored ones in.

Also reverts 17 of the 18 hunks in docs/api-reference.md, which were
Prettier reformatting of prose the PR does not otherwise touch (docs/ is
outside the format glob), keeping only the Reboot restore section and
repairing the two continuation lines that reformat de-indented; renumbers
reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which
loads after it; and gives the feature its CLAUDE.md entry plus a route
test for the multi-user workspace-forbidden branch, the only new rule that
had nothing behind it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:46:04 +02:00
Michael GrundbergandClaude Opus 5 18ab2ab595 docs(sessions): correct what a failed rebuild is actually likely to be
Ran the feature against a real server for the first time, on an isolated
instance, and two claims in the code turned out to be wrong.

A rebuild that fails after the session is registered was documented as
commonly caused by a CLI binary missing from a freshly booted machine's
PATH. It is not: the resolver finds its binary by absolute path, so PATH
never enters into it, and a server started without claude on PATH restored
every session normally. Nor does an un-enterable workspace fail — tmux falls
back to another directory and the pane comes up there. Neither obvious cause
throws, so the discard path is defended rather than expected, and the
comments now say that instead of naming a cause that cannot happen.

The four review rounds that shaped this path all reasoned about a trigger
none of them could test. The path itself is still worth having, since a mux
failure would reach it, but its comments should not claim a likelihood the
machine disagrees with.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 08:26:29 +02:00
DevvynandClaude Sonnet 5 b45a96358e feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.

- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
  registry entry declares contextLengthVar (currently only claude) and
  the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
  (40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
  quick-start customModel path) check this before the swap-conflict
  check and before launching/restarting anything, returning
  {requiresContextWarning, modelId, contextLength, minSafeContextTokens}
  — skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
  _resolveContextWarningConfirm (session-ui.js), wired into both
  _quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
  (the path Claude actually uses) ahead of the swap-confirmation check.
  Explains the fix in-modal: give the model an explicit larger -c/
  --ctx-size in llama-swap instead of relying on --fit-ctx, which
  optimizes for the biggest model that fits rather than the biggest
  context.

Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 07:36:57 +08:00
Randalix 7b947fa3f1 fix(remote): close the wake-state leaks and the dishonest wake budget
Review follow-up on the wake-on-LAN PR (five findings, all of them about the
state the feature keeps and the budgets it inherits):

- Wake state is dropped by `WebServer.cleanupSession` instead of the two delete
  routes, so it now goes with the session on EVERY cleanup path (cron, admin,
  scheduled-run teardown, error paths) instead of surviving with up to 4 KB of
  the user's buffered keystrokes. `registerSessionRoutes` returns the registry
  so the server can own its lifetime without the wake-capable code living in
  `server.ts`; the wiring guard is updated to allow that and gains a second
  assertion that `server.ts` calls nothing but `drop`/`stop` on it.
- `_effectiveRemote` returns before `_state`, so a LOCAL session no longer gets
  a wake-state entry — the input gate runs on every keystroke, so that entry
  used to be allocated for every session the user types in.
- An input chunk larger than the 4 KB cap is dropped OUTRIGHT instead of being
  head-trimmed and then written as a fragment: one paste is one `input` value
  and was never typed character by character, so its tail is a partial command
  the user never sent. The drop is logged.
- The manual wake button passes `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s)
  like the create/attach paths, instead of inheriting the 90 s session default
  that the dashboard's reverse proxy cuts off at 60 s.
- `RemoteWakeRegistry.stop()` aborts in-flight readiness polls (abortable
  sleep) and refuses new wakes, and `WebServer.stop()` calls it, so a restart
  during a wake no longer waits the poll out.
- The banner/toast wording keys off a new `queuedInput` flag on the two SSE
  events, which is true only when the server actually holds bytes: browser
  keystrokes travel over the WebSocket, which never passes through the
  registry, so the wake BUTTON must not promise queued input. The failed-wake
  path also stops pattern-matching the error message (it re-asks the
  reachability route) and the WoL dialog says "admin-only" instead of "host not
  found" for a non-admin in multi-user mode.
2026-09-16 20:44:39 +02:00
Michael GrundbergandClaude Opus 5 39976041e0 fix(sessions): let a dismiss reach the entries a restore is holding
Fourth review of the reboot-restore branch, and the third to find a defect
in the previous round's fix. This one is the same shape as its predecessor:
a counter keyed on one thing, compared against a set keyed on another.

The generation counter was indexed by the entry's owner, while the in-flight
set holds the caller doing the restoring. Those are the same person exactly
when a user restores their own sessions, which is every case the tests
covered. The route deliberately supports the other case: an admin may spend
another user's entries. So when an admin restored Bob's sessions and Bob
dismissed the banner, nothing matched, the entries came back, and a plan Bob
had explicitly dismissed was re-armed for another twenty-four hours.

Rather than reconcile the two key spaces, the counter is gone. `take()` now
parks the entries it hands out, remembering which caller is spending them,
and they stay parked until that restore ends. A dismiss filters the parked
entries by `canAccess(entry.owner)` — the same predicate it already applies
to the plan — so it reaches them wherever they are. `releaseFlight()` puts
back only what is still parked. Expiry and a fresh boot plan unpark
everything, for the same reason. There is one key space now, the entry's
owner, and the spender is only ever used to tell two concurrent flights
apart. That removes `generations`, `snapshotGenerations()`, `bump()`,
`bumpAll()` and the argument threaded through the route.

The discard grew the teardown it still lacked. A rebuild can fail after
startInteractive() resolved, and a restored workspace still carries
Codeman's hooks, so the CLI can post a hook event within milliseconds; the
transcript watcher that starts from it, the attachment registry, the wait
registry and the approvals inbox all outlive the listeners and would meet
the retry, which reuses the session id by design. Its steps also run in
reverse order now, so no live listener can reach a tracker that has already
stopped, and the mux kill has its own guard, because stop() kills the pane
in its last block after destroying four trackers.

Tests. The run-summary test named an interval and asserted a map entry, so
dropping stop() left it green; it now spies on stop(). Nothing pinned that
before-spawn must precede setupSessionListeners, which reads the flag that
phase restores, so swapping the two lines was silent; the ordering test now
includes the listener setup. The retry assertion was a tautology and now
asserts a different refs object. Both strengthened tests were verified by
reverting their fix. Two new tests cover the admin-restores-another-owner
cases this round was about. The server in the discard test is built once and
stopped, since its constructor registers handlers on module-level watchers,
and the workspace is removed through safeRmHomeTree.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 16:51:00 +02:00
Michael GrundbergandClaude Opus 5 71ed7b127c fix(sessions): make the discard a real inverse of the construction
Third review of the reboot-restore branch. The narrow discard the previous
commit introduced avoided everything cleanupSession() did wrongly, and in
dropping so much of it also dropped four things it had to keep.

The worst broke the retry the whole design rests on. setupSessionListeners()
returns early while sessionListenerRefs still holds the session id, and the
discard never cleared that entry. So the advertised flow — a rebuild fails
because the agent binary is missing, the user fixes their PATH and clicks
again — reused the same id, wired no listeners at all, and produced a tab
that never showed output, never updated its status and never persisted. That
is worse than the leak the discard was added to prevent. Three more
registrations leaked with it: a RunSummaryTracker and its interval, an image
watcher on the workspace, and the Ralph fix-plan watcher. The discard now
undoes each registration setupSessionListeners() makes, in its order, and
the per-session custom-model config directory, which holds the endpoint's
API key literally and which nothing else would ever remove.

The image-watcher flag was restored after the code that reads it, so a
session came back reporting the feature as on with nothing watching. It
moves to the before-spawn phase, and that phase now runs before the
listeners rather than after them.

The generation counter that lets a mid-restore dismiss win was global while
clear() is ownership-scoped, so one user's dismiss discarded another user's
unspent entries, permanently, because nothing rebuilds an in-memory plan. It
is now per owner. Bumping only the owners of entries the dismiss removed was
not enough either: take() has already emptied the plan by then, so a dismiss
landing mid-restore saw nothing of that owner's to remove and invalidated
nothing. The owners that matter are those with a restore in flight, filtered
by what the dismissing user may access, and that is what clear() now bumps.
Plan expiry bumps too, so a restore straddling the 24-hour boundary cannot
hand entries back and give an expired plan another full day.

Tests. discardPartiallyBuiltSession had no test at all: the only
implementation any test ran was the mock's one-line stub, which is why every
defect above was invisible. test/discard-partially-built-session.ts drives
the real WebServer, and the retry assertion fails if the listener refs are
left behind — verified by reverting the fix. The dismiss-race test drove the
registry by hand, so deleting the route's generation argument left it green;
it now goes through the route, and two further tests cover the multi-user
cases.

The mock context has now gone stale twice, because route tests pass it as
`ctx as never` and tsconfig.json includes only src, so nothing ever compares
it to the ports. A type-level guard is therefore inert — I wrote one and
confirmed it never fires. test/mocks/mock-route-context-completeness.ts
compares the mock's keys against WebServer.createRouteContext() at runtime
instead, and names what is missing.

Also: the API reference now says workspace-forbidden is judged against the
owner's grant, the banner's module header no longer claims Restore always
dismisses it, and the detail span gets the same min-width: 0 the phone rule
already needed.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:34:47 +02:00
Michael GrundbergandClaude Opus 5 fa52753e8b fix(sessions): undo a failed rebuild without deleting the user's data
A second review of the previous commit found that its own repair for the
session leak introduced three defects, all from reaching for
cleanupSession() to undo a half-built session. That function is the
user-initiated delete, not an undo.

It banked the session's historical token and cost totals into the lifetime
figures, and a reboot never runs cleanup, so those totals had never been
counted before; every failed rebuild added them again. It saw the pin that
had just been restored and demoted the record to `stopped`, which this pass
reads as the durable marker of a deliberate kill, so a pinned session whose
rebuild failed became permanently unrestorable. And it recursively removed
`.claude-images` from the working directory, which belongs to the workspace
rather than to the session, so a failed rebuild destroyed the pasted images
of any other live session in that repo.

discardPartiallyBuiltSession() now undoes only what the construction did:
the map entry, the tab-layout slot, the listeners and any pane the launch
created before throwing. The persisted record, the lifetime totals, the
Ralph state and the workspace's files are left alone.

Re-applying the persisted state also splits in two, which removes the first
two defects at the root rather than only at the call site. The half that
shapes the pane, the custom-model environment and the nice priority, still
runs before the spawn. The half that is the session's own history now runs
after it, so a session whose pane never started carries no totals and no pin
for anything downstream to misread.

The rest of that review. The multi-user workspace confinement re-check read
the requesting user's grant, and returns true for an admin, so the case its
own comment described was the one it missed; it now resolves the entry
owner's grant through isWorkingDirAllowedForUsername, the way cron does. A
forbidden workspace goes back on offer, matching both the registry's stated
contract and the API reference. The client re-reads the plan after a restore
instead of blanking the banner, so entries the server put back stay
reachable, and a 409 now says a restore is already running rather than
reporting a failure. A dismiss arriving mid-restore wins, through a
generation counter the route carries across its take. The re-application
also restores the tab colour, the image-watcher flag and the original
pinnedAt, via a new Session.restorePin that does not re-stamp the pin time.
The phone breakpoint gains min-width: 0, without which a nowrap flex item
never shrinks and the buttons still overflow, and it folds into the existing
phone block.

Ralph's loop configuration still does not survive a restore, because
toState() reads it off a live tracker and there is no way to keep it without
arming the loop. The method now says so rather than leaving it implied.

Tests. The capacity test could not fail on the property it existed for: it
filled the board past the cap before the loop, so a single pre-loop check
would have passed it. It now leaves one seat, so only a per-iteration check
restores exactly one entry. New tests cover the ordering around the spawn,
a throw before the loop returning the whole plan and releasing the flight,
the dismiss-during-restore race, and that the failure path calls the narrow
discard rather than the delete. The shared mock context gains the port
method it was missing, which is what made the first run of these tests fail
for the wrong reason.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:04:35 +02:00
Michael GrundbergandClaude Opus 5 fbede5cd2a fix(sessions): act on the dual review of the reboot-restore route
Fifteen findings from two independent reviews of #442, three of them
blocking. Every one is addressed here.

The three blockers all sat in the restore route. A rebuild that threw after
addSession left a registered session with no pane behind it, visible on the
board, holding a layout slot and written to state.json, with its plan entry
already spent; the catch now cleans the session up and puts the entry back.
The loop checked neither the global nor the per-user session cap, so one
click could take a board past a documented limit; capacity is now re-checked
per iteration, because the loop is itself creating the sessions it counts.
Worst of the three, a rebuilt session carried none of the state its
constructor has no parameter for and then persisted itself over the record
that held it, zeroing token and cost totals and dropping the pin. The pin
matters most: pruning keeps a record only while it is pinned, so discarding
it handed the record to the next stale sweep. A new
reapplyPersistedSessionState() on the session port restores the pin, the
token totals, auto-compact, auto-clear, auto-resume, nice priority, the
flicker filter and the custom-model selection, and it runs before both
startInteractive and the first persist.

The rest, in the order they bite a user. Every rebuild failure was reported
as workspace-missing, so the banner told users their repo was gone when the
agent had simply failed to start; there are now distinct reasons, and the
toast names each one. The client read restored and skipped off the outer
response object rather than through the uniform envelope, so every count
came back zero and neither toast ever fired. A board left open across the
reboot never learned an offer existed, because the banner was seeded only on
the page-load path; it now re-reads on every SSE init. The workspace check
was existence-only, skipping the multi-user confinement that the create
route applies, so a withdrawn grant would not be noticed. The banner had no
phone breakpoint while its text was nowrap and its buttons could not shrink.

Smaller: a missing workspace is now re-offered rather than dropped, while an
already-open conversation is dropped rather than re-offered forever; a throw
anywhere in the route returns the unspent entries instead of discarding the
plan; the single flight is keyed by owner, since take() already stops two
callers receiving one entry; the env clamp's header no longer claims a
protection it cannot provide on this path today, and names the check that
does bite; the three endpoints are documented in docs/api-reference.md; and
the module header now says that os.uptime() reads the host's clock, so the
feature is effectively off inside a container.

The review also explained why the tests missed all of this: they proved the
construction claim through their own copy of the construction rather than
through the route, and the route tests used workspaces that did not exist,
so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts
mocks the Session module to drive the route's real path, and covers the
cleanup, the reason reported, the re-application ordering, the broadcast and
the caps. The mock route context gains the port method and the mux call the
route needs.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 11:44:37 +02:00
DevvynandClaude Sonnet 5 0929694012 fix(custom-model): actually trigger the llama-swap load, not just watch for it
Root cause of "it doesn't look like llama-swap is actually switching the
model" (confirmed live: no load_model line in llama-swap's own logs after
applying a selection). llama-swap has no "switch model" admin endpoint - the
ONLY thing that starts a swap is a real inference request naming the model.
Every previous fix (the conflict check, the loading banner) assumed a swap
would start on its own; nothing ever actually asked llama-swap to load
anything until the launched CLI's first real prompt did, which could be
much later than "applying the selection" implied.

Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest
real request that will start a load - POST <baseUrl>/v1/chat/completions,
max_tokens: 1, one throwaway message - fire-and-forget (never awaited by
the caller; the frontend's own running-status polling is what actually
confirms readiness). Wired into both apply paths (the dedicated restart
route and the one-shot quick-start route), fired whenever the target model
isn't already the one loaded and ready - a broader condition than the
existing swapNeeded (which only gates the "this will evict another
session's model" confirmation ask and deliberately stays narrow to that).
modelSwapInProgress in both routes' responses now reflects this same
broader condition too, so the frontend's loading banner actually correlates
with a real in-flight load rather than only firing when something else
happened to be loaded already.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:53:34 +08:00