Commit Graph
26 Commits
Author SHA1 Message Date
Codeman maintainer 4205f6930f fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine.

**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.

**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.

**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.

**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.

The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:58:33 +02:00
Codeman maintainer 9af12afb57 docs(custom-model): make the docs match the code, and trim the changeset
More from the review of 5fc391a4, all documentation rather than behaviour.

The changeset was 1602 words of development log, written as the PR grew, with bullets
and loose paragraphs interleaved. That text becomes CHANGELOG.md and the GitHub release
body verbatim, so it is now one user-facing account of what the feature does and what
the real-server work bought, at roughly a fifth the length.

docs/api-reference.md promised a `cmd` field on running-status that the route
deliberately strips (it carries model paths and can carry --api-key).

Two places claimed the apply routes validate `modelId` against the endpoint's
discovered models. Neither does. Dropped the claim rather than adding the check:
discovery can be up to five minutes stale, so a 400 there would refuse a launch that
actually works, and a typo'd id already fails on the CLI's own first request. CLAUDE.md
now says so explicitly, since the absence is the surprising part.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:34:24 +02:00
Codeman maintainer 1a99b5836c Merge pull request #430 from opticon454/custom-model-run-menu 2026-09-19 12:25:11 +02:00
Codeman maintainer 035bfbc2fe fix(remote): merge-time fixes for Wake-on-LAN
The MAC-count limit lived in two places that disagreed. RemoteHostSchema.wakeMac's
128-character cap admits seven comma-separated MACs while parseMacList takes at most
four, all-or-nothing, so a five-MAC value validated, was written to remote-hosts.json,
and then resolved to NO wake target: POST /api/sessions/:id/wake answered
"No wake-on-LAN target configured for this host" and the banner offered "Configure WoL"
for a host the user had just configured. MAX_WAKE_MACS now lives in
src/config/remote-wake-limits.ts and both sides refine against it. Its own module
because src/remote-wake.ts is import-fenced to session-routes.ts and server.ts (the
wiring guard that stops a watcher waking a host), and because schemas.ts must not drag
dgram/net/child_process into every request-validating module.

The documented 40 s request budget also omitted the wake's own cost. A `command` target
is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS and runs BEFORE the readiness poll, so a
slow one pushed a wakeCommand host's worst case to ~68 s, past the 60 s
proxy_read_timeout the budget exists to stay under. _wakeAndWait now subtracts the
wake's measured elapsed time from the readiness budget, floored at one poll interval so
a wake that ate the whole budget still gets one probe. A magic packet is effectively
instant and is unaffected, which is why live testing never saw it.

Also: the two new endpoints are documented in docs/api-reference.md with the import
fence stated as the rule it is, CLAUDE.md's frontend module count moves to 34, and the
release changesets carry the Thanks section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:40 +02:00
RandalixandClaude Opus 5 5bb489addb fix(remote): authorize the attach wake first; tell the caller what happened to its bytes
Review round 3 on #439.

- The attachRemoteSession branch of POST /api/sessions ran `ensureHostAwake`
  before the multi-user gates, so a non-admin could have any configured
  host's `wakeCommand` spawned (or a packet broadcast) and the request held
  for the wake budget, then be refused for the workingDir. The admin gate
  now comes first, before the host is even looked up; remote hosts are
  admin-only infrastructure everywhere else. Route test: wake spy empty,
  403.
- The non-wait input route answers `{buffered:true}` when the registry took
  the chunk and `{buffered:true, dropped:true}` when it was over the cap
  and is gone (`RemoteInputOutcome` gains 'dropped'); additive to the bare
  `{}`.
- The send-and-wait path answers OPERATION_FAILED when the host never comes
  back, like create and attach, instead of writing into the stalled pane
  and reporting delivered:true plus a timeout.
- The flush writes with `fromUser: true`, so a first prompt buffered
  through a wake can still name the tab.

Docs: api-reference (input route), remote-sessions.md (two invariants),
CLAUDE.md key pattern.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-19 11:39:50 +02:00
Devvyn 56209e7829 Merge remote-tracking branch 'upstream/master' into feature/run-menu-custom-model-picker 2026-09-19 03:25:26 +08:00
DevvynandClaude Sonnet 5 9982a1325f fix(custom-model): address second pre-merge review (Ark0N)
Blocker: .center-status-banner never actually disappears.

- Add `.center-status-banner[hidden] { display: none; }`, same trap as
  `.home-sessions[hidden]`: the author-level `display: flex` beat the
  UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
  card stayed laid out at `opacity: 0` with its text/cancel/close
  children still `pointer-events: auto` -- an invisible 442x67 click
  blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
  banner (10001) and the swap-confirm/context-warning modals (10010)
  in CLAUDE.md's Z-index layers list.

Stale wording pointed at the reverted sticky-toast default:

- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
  `.toast-message` comment in styles.css all still said "toasts
  default to sticky" after 1f32128c put the flat 3s default back.
  Reworded all three to describe the actual behaviour: one call site
  passes an explicit `duration: 0`.

Smaller items from the same review:

- docs/api-reference.md said discovery failures answer
  `502 OPERATION_FAILED`; OPERATION_FAILED is 422 per src/types/api.ts
  and the error-code table earlier in the same file.
- The periodic re-discovery sweep (server.ts) never read
  customModelEndpointsEnabled, so turning the feature off left
  Codeman polling every saved endpoint forever. Added
  readCustomModelEndpointsEnabled() (custom-model-routes.ts, same
  shape as readPlanUsageTelemetryEnabled) and gated the interval
  callback on it.
- Reverted the formatting-only Prettier pass docs/api-reference.md
  picked up (table padding, *x* to _x_, JSON re-indent) by re-merging
  the new Custom Model Endpoints section onto the pre-PR file, so the
  diff is reviewable. No prose content was lost -- verified by diffing
  the result against the pre-revert file (formatting-only) and against
  the merge-base file (only the new section added).
- docs/custom-model-endpoints.md now states that a custom-model Claude
  session's isolated CLAUDE_CONFIG_DIR loses the user's global
  settings.json, user-level skills/agents/commands, and MCP servers
  from ~/.claude.json -- only `projects` is symlinked back.

Design question left open in the review (does `confirmed: true` need
to be two flags so "launch anyway" on the context warning doesn't also
skip the llama-swap displacement warning): keeping the single flag, as
offered. The 20s displacement sweep still catches a resulting swap
after the fact, so it's a surprise rather than a silent failure, and
splitting it is real behavioural surface I have no way to verify live
in this environment.

`npm run test:browser` could not be run in this environment (no tmux,
no downloaded Playwright browser binary) -- none of its suite's files
touch code this fix changes, but it still needs a real pass before
merge, same as any frontend change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-18 21:45:50 +08:00
Codeman maintainer bb8ada7e5f fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does.

1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred
   it from the name: a session the user renamed by hand to something shaped
   like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the
   next prompt overwrote their name. The route persists right after, so the
   loss went to disk. `restoreMuxSessions()` already passes it.

2. The already-live sets were snapshotted once before a loop that awaits a
   real `startInteractive()` per entry, so by the tenth entry the snapshot
   was tens of seconds old and a conversation resumed by hand from the
   Resume list in that window was invisible to it: two panes on one
   transcript, the exact thing the check exists to prevent. Both sets are
   now read per iteration, and the late case is spent rather than re-offered
   for the same reason the batch case is.

3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path.
   The stamp predates the reboot and the pane is new, so honouring it meant
   one click had every restored session type `continue` into itself about a
   minute later, unattended, against the route header's own promise that a
   restored session comes back idle and disarmed. The setting stays ENABLED,
   so it re-arms on the next real limit message. A Codeman restart still
   re-arms from the stamp, because the limit footer will not reprint on its
   own; the new option exists only to tell the two paths apart.

4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()`
   and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession`
   performs that it was missing. Cosmetic, but a run left open reads as
   still going in the away digest.

5. A restored claude session gets `seedAgentSessionPreamble()` like both
   create paths, so the agent skill's bootstrap stays a two-line loader.

6. The heuristic's container comment was wrong in one direction and quiet
   about the real gap: after a genuine host reboot a containerized Codeman
   sees the host's short uptime and the banner does appear. What it cannot
   see is a container-only restart, which is where this would help most.

7. The banner is hidden in a solo window, which shows one session and has
   no tab strip to put restored ones in.

Also reverts 17 of the 18 hunks in docs/api-reference.md, which were
Prettier reformatting of prose the PR does not otherwise touch (docs/ is
outside the format glob), keeping only the Reboot restore section and
repairing the two continuation lines that reformat de-indented; renumbers
reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which
loads after it; and gives the feature its CLAUDE.md entry plus a route
test for the multi-user workspace-forbidden branch, the only new rule that
had nothing behind it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:46:04 +02:00
DevvynandClaude Sonnet 5 8520925e76 docs(custom-model): bring CLAUDE.md and api-reference.md up to date
Full documentation review pass across the branch's 30 commits.
CLAUDE.md's Custom Model Endpoint Profiles entry hadn't been touched
since the initial backend+picker cut (3 early commits) despite 27
follow-up commits adding real behavior — it described restart-in-place
as universal (now claude-only; 7 other CLIs launch one-shot) and
claimed codex's Responses-API gap as a flat protocol break (now
re-verified as a more precise tool-calling gap). Corrected both and
added a new paragraph covering everything landed since: the llama-swap
conflict check, the after-the-fact swap-displacement sweep, the
/running-cmd-based context-length fix, the context-window floor
warning, skipFirstRunPrompts, the real-time /api/events-based log
status, and the countdown-to-Cancel-button change.

docs/api-reference.md's custom-model-endpoints section was missing the
running-status route, the requiresConfirmation/requiresContextWarning
response shapes, and POST /api/quick-start's customModel field
entirely (the primary launch path for 7 of 8 supported CLIs) — added
all three. Also fixed a real markdown bug in custom-model-endpoints.md:
an inline code span (`POST <baseUrl>/v1/chat/completions`) split across
a line break, which CommonMark renders with the line ending collapsed
to a space, so it displayed as ".../v1/chat/ completions" with a
spurious space inside the path.

Verified: origin/master and upstream/master are both already an
ancestor of this branch (identical at bd286bf5, no new commits since
this branch was cut) — nothing to merge, no conflicts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:18:04 +08:00
Michael GrundbergandClaude Opus 5 71ed7b127c fix(sessions): make the discard a real inverse of the construction
Third review of the reboot-restore branch. The narrow discard the previous
commit introduced avoided everything cleanupSession() did wrongly, and in
dropping so much of it also dropped four things it had to keep.

The worst broke the retry the whole design rests on. setupSessionListeners()
returns early while sessionListenerRefs still holds the session id, and the
discard never cleared that entry. So the advertised flow — a rebuild fails
because the agent binary is missing, the user fixes their PATH and clicks
again — reused the same id, wired no listeners at all, and produced a tab
that never showed output, never updated its status and never persisted. That
is worse than the leak the discard was added to prevent. Three more
registrations leaked with it: a RunSummaryTracker and its interval, an image
watcher on the workspace, and the Ralph fix-plan watcher. The discard now
undoes each registration setupSessionListeners() makes, in its order, and
the per-session custom-model config directory, which holds the endpoint's
API key literally and which nothing else would ever remove.

The image-watcher flag was restored after the code that reads it, so a
session came back reporting the feature as on with nothing watching. It
moves to the before-spawn phase, and that phase now runs before the
listeners rather than after them.

The generation counter that lets a mid-restore dismiss win was global while
clear() is ownership-scoped, so one user's dismiss discarded another user's
unspent entries, permanently, because nothing rebuilds an in-memory plan. It
is now per owner. Bumping only the owners of entries the dismiss removed was
not enough either: take() has already emptied the plan by then, so a dismiss
landing mid-restore saw nothing of that owner's to remove and invalidated
nothing. The owners that matter are those with a restore in flight, filtered
by what the dismissing user may access, and that is what clear() now bumps.
Plan expiry bumps too, so a restore straddling the 24-hour boundary cannot
hand entries back and give an expired plan another full day.

Tests. discardPartiallyBuiltSession had no test at all: the only
implementation any test ran was the mock's one-line stub, which is why every
defect above was invisible. test/discard-partially-built-session.ts drives
the real WebServer, and the retry assertion fails if the listener refs are
left behind — verified by reverting the fix. The dismiss-race test drove the
registry by hand, so deleting the route's generation argument left it green;
it now goes through the route, and two further tests cover the multi-user
cases.

The mock context has now gone stale twice, because route tests pass it as
`ctx as never` and tsconfig.json includes only src, so nothing ever compares
it to the ports. A type-level guard is therefore inert — I wrote one and
confirmed it never fires. test/mocks/mock-route-context-completeness.ts
compares the mock's keys against WebServer.createRouteContext() at runtime
instead, and names what is missing.

Also: the API reference now says workspace-forbidden is judged against the
owner's grant, the banner's module header no longer claims Restore always
dismisses it, and the detail span gets the same min-width: 0 the phone rule
already needed.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:34:47 +02:00
Michael GrundbergandClaude Opus 5 fbede5cd2a fix(sessions): act on the dual review of the reboot-restore route
Fifteen findings from two independent reviews of #442, three of them
blocking. Every one is addressed here.

The three blockers all sat in the restore route. A rebuild that threw after
addSession left a registered session with no pane behind it, visible on the
board, holding a layout slot and written to state.json, with its plan entry
already spent; the catch now cleans the session up and puts the entry back.
The loop checked neither the global nor the per-user session cap, so one
click could take a board past a documented limit; capacity is now re-checked
per iteration, because the loop is itself creating the sessions it counts.
Worst of the three, a rebuilt session carried none of the state its
constructor has no parameter for and then persisted itself over the record
that held it, zeroing token and cost totals and dropping the pin. The pin
matters most: pruning keeps a record only while it is pinned, so discarding
it handed the record to the next stale sweep. A new
reapplyPersistedSessionState() on the session port restores the pin, the
token totals, auto-compact, auto-clear, auto-resume, nice priority, the
flicker filter and the custom-model selection, and it runs before both
startInteractive and the first persist.

The rest, in the order they bite a user. Every rebuild failure was reported
as workspace-missing, so the banner told users their repo was gone when the
agent had simply failed to start; there are now distinct reasons, and the
toast names each one. The client read restored and skipped off the outer
response object rather than through the uniform envelope, so every count
came back zero and neither toast ever fired. A board left open across the
reboot never learned an offer existed, because the banner was seeded only on
the page-load path; it now re-reads on every SSE init. The workspace check
was existence-only, skipping the multi-user confinement that the create
route applies, so a withdrawn grant would not be noticed. The banner had no
phone breakpoint while its text was nowrap and its buttons could not shrink.

Smaller: a missing workspace is now re-offered rather than dropped, while an
already-open conversation is dropped rather than re-offered forever; a throw
anywhere in the route returns the unspent entries instead of discarding the
plan; the single flight is keyed by owner, since take() already stops two
callers receiving one entry; the env clamp's header no longer claims a
protection it cannot provide on this path today, and names the check that
does bite; the three endpoints are documented in docs/api-reference.md; and
the module header now says that os.uptime() reads the host's clock, so the
feature is effectively off inside a container.

The review also explained why the tests missed all of this: they proved the
construction claim through their own copy of the construction rather than
through the route, and the route tests used workspaces that did not exist,
so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts
mocks the Session module to drive the route's real path, and covers the
cleanup, the reason reported, the re-application ordering, the broadcast and
the caps. The mock route context gains the port method and the mux call the
route needs.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 11:44:37 +02:00
DevvynandClaude Sonnet 5 5a9ff07f57 feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:

1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
   apply the endpoint's defaultModelId (or the first discovered model)
   silently. Now, via the new selectCustomModelEntry() (session-ui.js):
   - exactly one discovered model launches straight away, same as before
   - two or more open a new #customModelPickModal listing every discovered
     model; defaultModelId (if set) is marked but never auto-chosen, since
     the point of asking is letting ONE launch deliberately differ from
     the saved default, not just confirming it
   The endpoint is re-fetched at click time rather than trusting anything
   cached from the dropdown's own render, since the model list can have
   changed (the sweep below, or a settings-panel edit) since it opened.
   runCustomModelEntry() itself — the actual launch, routed through run()
   for the in-flight lock, snapshot-guarded against applying to the wrong
   session — is unchanged; it now just always receives an explicit model
   id from one of these two paths instead of computing one itself.

2. Periodic re-discovery. Every saved endpoint's models now refresh
   automatically every 5 minutes in the background
   (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
   as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
   off under testMode), so a model the server starts or stops serving shows
   up without another manual "Discover" click. The manual POST
   .../discover-models route and the new refreshAllCustomModelHosts()
   sweep (custom-model-routes.ts) now share one pure merge step
   (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
   that no longer appears) rather than two copies that could drift. The
   sweep is best-effort per host — one endpoint being unreachable on a
   cycle never blocks the others — and re-reads the store before each
   host's write, keyed by id, so a concurrent edit or delete from the
   settings panel always wins over a sweep that started before it.

Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.

Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 08:50:18 +08:00
DevvynandClaude Sonnet 5 60e1bd52f7 fix(custom-model): act on the draft review — unparseable onclick, unwrapped envelope, wrong-session apply, missing lock, no tests
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.

Blockers:

1. Every generated inline onclick was unparseable. JSON.stringify's own
   double quotes terminated the double-quoted HTML attribute at the first
   one, leaving btn.onclick null on every picker entry and every Discover/
   Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
   argument, the same idiom deleteCase's onclick already uses four lines
   away in session-ui.js. This also closes the live-HTML-injection route
   through modelId (server-controlled, from the endpoint's own /v1/models
   reply): with quoting intact, a `>` inside it can no longer terminate the
   <button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
   like every other /api route (server.ts's preSerialization hook applies
   to arrays too), so Array.isArray(hosts) was always false in production
   and the picker/settings panel silently saw nothing. Both call sites now
   go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
   returns normally without ever changing activeSessionId, so the apply
   step used to silently re-point and restart whatever session the user was
   already looking at. runCustomModelEntry() now snapshots activeSessionId
   before the launch and requires it to have actually changed.

Majors:

4. Routes the launch through run() itself via a temporary _runMode swap
   (never persisted — setRunMode() would sync it to the server) instead of
   a parallel hardcoded dispatch table, so a custom-model launch now holds
   the same _runInFlight lock every other Run click gets. This also
   resolves the "hardcoded runners map contradicts the PR's own design"
   minor: dispatch is run()'s own, so a CLI whose customModelInjection
   recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
   against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
   parses markup this module generated itself) for exactly the DOM-level
   facts the review said needed no Playwright and no tmux: a generated
   button's onclick genuinely compiles and fires, a dangerous modelId never
   produces a live element, the envelope unwrap works, the session-changed
   guard holds, run() actually gets called (proving the in-flight lock
   engages), and _runMode is restored afterward. Confirmed against the
   pre-fix code first (reproduces btn.onclick === null exactly) so this
   isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
   render-index-html.test.ts for the other fixes below.

Minors:

- Generated entries now filter through isCliAvailable(), matching
  _refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
  (applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
  and to settings-modal open) instead of always rendering; the endpoint GET
  no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
  redactApiKey() replaces the field with a computed apiKeySet: boolean, and
  a PUT with no apiKey now keeps the stored one server-side
  (applyStoredApiKey()) instead of the client resending a value it was
  never given. New tests cover both directions (kept vs. replaced) by
  observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
  (_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
  event, since the real role can resolve after settings were first opened)
  — endpoint writes were already admin-only server-side, but the button
  used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
  toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
  route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
  that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
  (CLAUDE.md already records that exact literal turning the settings
  preview into a grey slab on light skins), .run-mode-custom-models gets
  the same gap: 2px .run-mode-menu's own flex gap only applies one level
  up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
  </script> (CliEntry.label is user-clis.json-settable, unlike
  __codemanCliAvailable's booleans-only payload) via a new exported
  escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
  docs/api-reference.md; left the "no zh-CN for the new Models-section
  group" minor unaddressed only insofar as the wider Models section (task
  routing, thinking effort, etc.) has never had zh-CN coverage either —
  everything this PR itself introduces (labels, hints, button text, the
  Run-menu's "Custom Endpoints" header) IS translated in i18n.js.

Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
Codeman maintainer aaa93d4252 fix(session): answer Claude Code 2.1.252's reversed folder-trust dialog
Every claude session in a directory claude had not seen before died about six
seconds after it started (`Pane is dead (status 1)`), before the agent drew a
composer. Reproduced on a fresh case and measured.

Claude Code 2.1.252 rewrote the dialog. It used to be

  ❯ 1. Yes, I trust this folder
    2. No, exit

and is now unnumbered, reversed, and highlights the option that quits:

  ❯ No, exit
    Yes, I trust this folder

Detection still worked (the confirm affordance carries the match once the
numbered option text is gone), so the failure was entirely in the answer: the
auto-accept pressed Enter on the highlighted default, which is now exit.

trustDialogNextKey() reads the ❯ marker off the rendered pane and returns ONE
keystroke at a time: an arrow while the cursor is on the wrong option, Enter
only once the screen shows it on the trust option, and null for a frame that
does not say. Both layouts are handled, and which way the trust option lies is
read from the frame rather than assumed, so a further reordering costs a
repaint instead of a session. The last marked option wins, because the
direct-PTY fallback reads an append-only buffer where an older frame must not
out-vote the freshest one.

Two things only a live pane showed:

- The scan ran solely from the PTY onData handler. The arrow that moves the
  cursor is the last output the pane produces, so the first fix parked every
  session with the cursor sitting on the right option and no Enter ever sent.
  It now schedules its own follow-up read (_trustDialogTimer, cleared in
  _clearAllTimers()), offset past the scan throttle so the chain cannot break
  on a boundary.
- The keystroke cap goes 3 -> 6, since answering is no longer one press.

The bundled codeman skill had the same blind \r as its bounded fallback, so
preamble 1.21.0 replaces it with _trust_key/_accept_trust: read
terminal?full=1, steer onto the trust option, re-read, then confirm. Those
keystrokes go out under their own clientId, because input sequence numbers are
monotonic per client and spending prompt numbers on dialog keys would make the
next send-and-wait look like a stale duplicate and vanish while reporting
success. The readiness recipes in docs/extending-codeman.md,
docs/api-reference.md and the skill's own reference carry the corrected answer,
plus a symptom-table entry for a worker whose pane is dead seconds after spawn.

Verified live on an isolated instance (own data dir and tmux socket): fresh
case -> arrow at 5 s -> Enter at 7 s -> composer, with hasTrustDialogAccepted
recorded. With the server-side auto-accept disabled in a throwaway copy, the
skill's fallback cleared a genuinely parked dialog in 1.1 s and spawn_worker
took a brand-new case to a live composer in 7.2 s; spawn_workers + sendwait +
last_text then ran end to end.
2026-09-01 02:31:26 +02:00
Codeman maintainer f60bf93c99 chore: version packages
Red tab alerts track the dialog, not the keyboard: typing no longer clears them, and a dialog answered in the terminal resolves itself on the next listing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:46:48 +02:00
Codeman maintainer f4ba4d2cb1 chore: version packages
Persist the 'I checked it' state of yellow idle tab alerts across reloads and devices.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:46:03 +02:00
Codeman maintainer 4bbe2b7ff6 docs: add pi to the mode lists the sixth-backend sweep missed
PR #282 added pi across the prominent surfaces but left the enumerations
that read as exhaustive: the env-prefix allowlist (missing PI_*), the
external-CLI list for stop/blocked, cron's agent types (also missing
antigravity), the narrow-strip mode list, and the claude-only caveats in the
cron and Read My Mind guides. Both READMEs and the four affected docs now agree
with the schema.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 19:40:54 +02:00
Codeman maintainer c790166564 feat(sse): heal a stalled SSE stream with a heartbeat + client watchdog
An EventSource that stops delivering does not always error. A proxy that
idle-closed the connection, a laptop resumed from sleep, a tailnet reconnect:
`onerror` never fires, the header dot stays green, and every SSE-driven surface
(tab status dots, sessions created on another device, renames) freezes until the
user reloads. Nothing on the client tracked stream liveness at all.

The server already wrote a keepalive every 15s, but as an SSE `:keepalive`
COMMENT, and comments are invisible to `EventSource` by spec, so there was
nothing a client could observe.

Server:
- `sse:heartbeat` under a new Transport category in the event registry
  (155 constants now, both counts updated).
- `cleanupDeadClients()` writes that named frame (`{"t":<epoch ms>}`) instead of
  the comment. Interval, tunnel padding and dead-socket eviction are unchanged.
  The write stays per-client rather than going through `broadcast()`: the frame
  carries no session data, so it needs no multi-user owner routing.

Client:
- `computeSseStale()` in constants.js, a pure policy beside
  `computeConnectionLossUi`. Stale only when the transport believes it is
  `connected`, the device is online, and no frame has arrived for 45s (three
  missed heartbeats). The `connected`-only guard is also the loop breaker: a
  forced reconnect leaves that state immediately, so the watchdog cannot re-fire
  while one is in flight.
- The liveness stamp is applied inside `addListener` itself, so the
  `_SSE_HANDLER_MAP` wrappers and the directly-registered listeners all feed it
  from one place instead of three that can drift. The heartbeat's own listener
  is a no-op that exists only to be registered, since `EventSource` drops named
  events nobody listens for.
- A 5s watchdog forces `connectSSE()` when the policy says stale, and is cleared
  at the top of `connectSSE()` and nowhere else (its only teardown path).
  Recovery needs no new sync path: the reconnect re-runs `handleInit`, which
  already rebuilds from the server. `visibilitychange` -> visible checks too,
  riding the existing listener, since a background tab's timers are throttled
  and a wake is exactly when a stream comes back zombie.
- The forced reconnect logs one diagnostic line: if a middlebox ever strips or
  delays heartbeats, the failure mode is "silently reconnects every 45s", which
  is undebuggable from a field report without it.

Tests: `test/sse-staleness.test.ts` (node VM over constants.js, threshold
boundaries and every not-stale guard) and `test/sse-heartbeat.test.ts` (drives
`cleanupDeadClients()` with fake replies: named frame not a comment, parseable
payload, padding only with a tunnel, dead clients still evicted).

Verified end to end on an isolated instance: with the stream closed client-side
(no `onerror`), a rename sticks, an out-of-band session stays invisible, then
the watchdog reconnects on its own and it appears without a reload.

Event names are part of the stable API contract, so this is a MINOR bump.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 17:28:35 +02:00
Codeman maintainer f39beb3326 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 02:30:08 +02:00
Codeman maintainer 4b51ba306e feat(voice): dictate through the server's Claude Code login, no API key
The mic button previously needed a Deepgram API key, or fell back to the
browser's Web Speech engine. It can now transcribe through the same
speech-to-text service Claude Code's own /voice mode uses, so anyone signed
in to Claude Code on the server gets dictation with no third-party account.

Claude Code's voice mode cannot be driven directly: it opens the HOST's
microphone (sox/arecord), and the CLI runs in a headless tmux pane while the
human is in a browser somewhere else. So capture stays in the browser and only
the transcription backend is borrowed.

Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the
page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since
MediaRecorder cannot emit raw PCM) and receives text.

- GET /api/voice/status reports readiness and never the token
- GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade
  guard as the terminal socket, plus caps on concurrency, stream length and
  frame size
- credentials are read-only: Codeman never refreshes them, since a refresh
  rotates the refresh token and could sign the user out of their own CLI
- claudeVoiceEnabled (synced, default OFF) gates the whole server side
- voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers
  Claude, then a configured Deepgram key, then the browser

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 11:19:51 +02:00
Codeman maintainer 94abcf29dc feat: Read My Mind phase 2, the predictor and the brain button
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.

Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
  sources: pending approval dialog, user goals, last assistant turn tail,
  recent prompts, tool activity, git workspace signals, away context,
  sibling sessions, rethink state; 30 KB budget, whole-section drop from
  the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
  keeps only a 500-char snippet) and git signal collection (execFile,
  2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
  session, opus by default (readMyMindModel setting), strict JSON
  contract with 1-3 suggestions (continue / verify / redirect), newline
  stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
  prediction in flight per session (409 CONFLICT), rethink body
  { steer, rejected }; ownership via findSessionOrFail

Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
  until readMyMindEnabled is ON, desktop only (phone key is phase 3);
  modal with editable suggestion + rationale and Send / Insert /
  Rethink / Dismiss; suggestion text rendered via value/textContent only
  and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
  strings

Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 23:05:11 +02:00
Codeman maintainer 6bb3d66004 docs: Read My Mind user guide (enable, capture rules, privacy, API, troubleshooting)
docs/readmymind.md covers phase 1 as a user guide: how to enable the synced
readMyMindEnabled setting via the API (no UI checkbox until phase 2), exactly
what is and is not captured, the hooks dependency (Docker bridge / remote-SSH
caveats), storage and wipe paths, curl examples for the three endpoints, the
agent-skill ground rules, and a troubleshooting table. Cross-linked from the
CLAUDE.md Key Patterns entry and the api-reference section.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 18:22:52 +02:00
Codeman maintainer 161f1da2eb feat: Read My Mind phase 1, per-case intent profiles (capture + API + skill)
Per-case profiles of user intent (docs/readmymind-plan.md): user-stated goals
plus the user's recently submitted prompts, captured from the Claude session
transcript behind the new synced readMyMindEnabled setting (default OFF).

- intent-store.ts: keyed by owner + realpath(workingDir), FIFO/size caps,
  consecutive-dupe collapse, atomic 0600 writes to ~/.codeman/intents.json
- transcript-watcher.ts: new transcript:user_prompt event for typed user turns
  (tool_result-only entries stay silent); capture wiring in server.ts is
  claude-only and gated on the setting per event
- readmymind-routes.ts: GET/PUT/DELETE /api/sessions/:id/intent, ownership
  via findSessionOrFail, strict Zod schema
- agent skill: SKILL.md recipe + endpoints.md rows so agents can read and
  record intent (PUT replaces: read + merge; never delete unprompted)
- groundwork for the phase-2 predictor button; nothing is ever auto-sent

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 18:03:13 +02:00
Codeman maintainer ff10a50bc0 feat: Approvals Inbox, one cross-session queue for prompts waiting on a human
Permission dialogs, AskUserQuestion questions and idle prompts from every
session now land in a server-side inbox (web/approval-inbox.ts, one item per
session, claude-mode only) and are answerable in place: a header bell + drawer
on desktop, inline answer strips on the phone overview's NEEDS YOU rows, and
working push Approve/Deny buttons (previously dead ends, now answered straight
from sw.js with no tab open). Pending alerts survive reloads because the
frontend seeds from GET /api/approvals on init.

Answering sends the digit / Esc / prompt text through the existing tmux input
path; option digits are accepted only when they match options parsed from the
captured pane frame, and the answer path re-captures the pane first so a
dialog that already left the screen refuses with 409 instead of typing into
the composer. New elicitation_complete / elicitation_response hook matchers
resolve question items the moment they are answered in the terminal;
refreshStaleCodemanHooks heals existing cases.

Verified end-to-end against a live claude session: a real AskUserQuestion
dialog parsed into 5 option buttons and was answered from the drawer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 14:04:34 +02:00
Codeman maintainer d26f26fe34 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:18:03 +02:00
arkonandClaude Opus 4.8 458fb81cbe feat(api): establish stable HTTP contract — uniform {success,data} envelope, status codes, /api/v1
Point 1 of the v1.0 lock-in: commit to a stable HTTP API (the cleanest, fullest form).

Core (centralized):
- Every JSON /api response now uses ONE envelope via a Fastify preSerialization hook (src/web/server.ts): success -> { success:true, data:<payload> }; error -> { success:false, error, errorCode } with a conventional HTTP status. Non-JSON routes (file-raw, tail-file SSE, download, screenshots, /q redirect, WS) are skipped.
- Error-code -> HTTP status is a single source of truth (httpStatusForErrorCode in src/types/api.ts): 400/401/404/409/422/429/500. Expanded ApiErrorCode (added UNAUTHORIZED, CONFLICT, RATE_LIMITED). Errors are no longer HTTP 200.
- Versioned alias: /api/v1/* rewrites to /api/* (rewriteApiV1Url), so external clients pin to a stable surface while the bundled UI keeps using /api/*.
- Handlers stripped of manual 'success:true' (50 across 14 route files) so they return bare payloads the hook wraps uniformly; fixed the mux DELETE {success:<bool>} envelope collision (-> {killed}).

Frontend (48 call sites across 10 files):
- _apiJson() auto-unwraps { success:true, data } -> data (null on error), so most bare-shape readers are transparent. Raw-fetch sites relocate payload reads under .data; success/res.ok/error checks unchanged.

Docs: new docs/api-reference.md (envelope, status table, error codes, /api/v1, SSE); versioning-policy.md flipped — the HTTP/SSE API is now part of the stable, SemVer-covered surface.

Verification: full unit/route suite green (2680 passed) incl. ~166 updated assertions across 24 test files; typecheck/lint/format/frontend-syntax clean; a headless-chromium smoke loaded the migrated UI and drove the panels with 0 console/page errors; /api/status and /api/v1/status confirmed returning the uniform envelope live.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 01:30:43 +02:00