Commit Graph
21 Commits
Author SHA1 Message Date
Codeman maintainer d3f2ec0220 fix(custom-model): merge-time fixes for the promoted-model picker (#459)
The maintainer's promised follow-ups to opticon454's picker promotion,
applied on the landing branch after the merge (ecb95b5d):

- session-ui.js: the promotion tag ("Currently loaded" / "Last used") and
  the "Default" pill are two separate spans, so a promoted row that is
  also the endpoint's defaultModelId shows both instead of silently
  losing its Default marking; two tests pin it (both fail on the old
  exclusive-slot rendering).
- styles.css: a dedicated #customModelPickModal .set-scope rule, since
  the pill was only styled inside the three settings modals and rendered
  as plain body text here; same skin tokens, modal layout untouched.
- docs/wiki/Custom-Model-Endpoints.md: describe the promotion (currently
  loaded, else last used per device), the separate Default pill, and
  that nothing is ever auto-chosen.
- CLAUDE.md + docs/custom-model-endpoints.md: credit the real "Last used"
  writers (_runCustomModelEntryViaRestart and
  _quickStartWithCustomModelConfirm; runCustomModelEntry only dispatches
  since 88e5b7b2) and drop the now-wrong "both defer to Default" sentence.
- Not done: moving the one-shot "last used" write into
  _runCustomModelEntryOneShot, because the existing one-shot tests assert
  that _quickStartWithCustomModelConfirm writes the key itself.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 0cb0f911adc14a852ba5c2951768a4aa87c25657)
2026-09-21 04:30:01 +02:00
DevvynandClaude Sonnet 5 88e5b7b200 fix(custom-model): address Ark0N's PR review — client-side probe timeout, defer "last used" past confirmation, docs, zh-CN
Four things from the maintainer's review on PR #459, all fixed:

1. Bound _getCustomModelCurrentlyLoaded's probe client-side (~800ms via
   Promise.race, on top of — never instead of — the route's own 5s
   server-side timeout). Without it, an asleep/firewalled endpoint behind
   a saved model list left the picker completely invisible for up to 5s
   after the Run menu had already closed, with no spinner or toast.
   `timeoutMs` is an optional param (default 800, real callers never pass
   it) so a test can drive it in milliseconds, same pattern as
   `_watchLlamaSwapLoading`'s own `pollIntervalMs` — this code runs in a
   JSDOM window's own realm, whose setTimeout vi.useFakeTimers() cannot
   patch.

2. "Last used" is now written only once a launch actually applies, never
   on the mere click. It moved out of runCustomModelEntry (unconditional)
   and into each path's own success point: _quickStartWithCustomModelConfirm
   after the final post succeeds, and _runCustomModelEntryViaRestart right
   after the apply's success check. A context-window-warning decline means
   this exact model cannot work with this CLI at all, so the old
   unconditional write would promote, next time the picker opened, the one
   model guaranteed to fail again.

3. Documented the promotion/tag precedence and the new
   codeman:customModelLastUsed:<mode>:<endpointId> localStorage key in both
   CLAUDE.md's Custom Model Endpoint Profiles section and
   docs/custom-model-endpoints.md's Run-menu picker section.

4. Added zh-CN entries for "Currently loaded" and "Last used" in i18n.js,
   next to this modal's existing "Choose a model"/"Custom Endpoints" pair.

New tests: the client-side timeout (endpoint that never answers, one that
answers within the bound, and a rejected-after-timeout probe settling
quietly), and "last used" recording on success vs. NOT recording on either
confirmation's decline, for both the restart and one-shot paths.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
2026-09-20 21:19:42 +08:00
DevvynandClaude Sonnet 5 73607663fd fix(custom-model): guard the picker's async open against a slower, superseded probe
Code review (high effort) on the previous commit found a real race: making
_openCustomModelPickModal async (it now awaits the currently-loaded-model
probe before rendering) meant a second, faster call for a different
endpoint could render first, only for the first call's slower probe to
resolve afterwards and overwrite the modal with the wrong endpoint's model
list — while _pendingCustomModelPick (set synchronously, before either
await) still named the second, correct endpoint. Picking a model in that
state would launch/apply the wrong model on the wrong endpoint.

Fixed with the same mutable-generation-counter guard
_watchLlamaSwapLoading already uses for an identical async-superseded-by-
newer-call shape: every DOM write, including _pendingCustomModelPick
itself, is deferred until after the awaited probe, and a call that finds
its generation already superseded bails out untouched instead of clobbering
whatever a newer call already rendered.

Added a regression test driving two overlapping opens with a controlled
promise so the earlier, slower probe resolves after the later, faster one
renders, asserting the late response is a no-op.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
2026-09-20 19:27:07 +08:00
DevvynandClaude Sonnet 5 458ca578e7 feat(custom-model): promote the currently-loaded/last-used model in the Run-menu picker
Custom Model Endpoint Profiles' "which model" picker (session-ui.js's
_openCustomModelPickModal) always listed models in their raw discovery
order, so on a host with several downloaded GGUFs the user had to
remember (or eyeball the "Default" tag) which one llama-swap actually
had hot before picking — the whole point of the picker being fast is
undone if it makes you think first.

The picker now promotes exactly one model to the top of the list:

- If llama-swap reports a model from this host's own list `ready`
  right now (via the existing GET /api/model-endpoints/:id/running-status
  route), that model is promoted and tagged "Currently loaded" — it's
  what a launch attaches to with zero wait.
- Otherwise, the last model actually launched on this exact
  (harness, endpoint) pair is promoted and tagged "Last used", read
  from a new per-device localStorage key
  (codeman:customModelLastUsed:<mode>:<endpointId>), written by
  runCustomModelEntry on every launch attempt regardless of outcome.
- A plain (non-llama-swap) OpenAI-compatible server, an unreachable
  endpoint, or a loaded-but-not-yet-ready model never promotes
  anything — the rest of the list keeps its discovery order.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
2026-09-20 19:20:45 +08:00
Codeman maintainer 4205f6930f fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine.

**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.

**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.

**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.

**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.

The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:58:33 +02:00
DevvynandClaude Sonnet 5 afb6754453 fix(custom-model): address third pre-merge review (Ark0N)
Blocker 1: the loading banner hides itself ~200ms after it reopens.

- _showCenterStatus reuses one shared DOM node; dismiss() scheduled
  el.hidden = true 200ms later with nothing to cancel it. On the
  Claude path, switchingToast.dismiss() is followed by one same-
  origin request (5-30ms locally) before _watchLlamaSwapLoading opens
  the new banner -- well inside that window -- so the stale timer
  fired against the shared node and hid the fresh banner, leaving the
  whole model-load wait with no progress text, no log line and no
  reachable Cancel button.
- Fixed by parking the pending timeout on the element and clearing it
  at the top of _showCenterStatus. Added a regression test that
  reproduces the exact repro (open, dismiss, reopen 20ms later,
  advance past 200ms) alongside the existing Cancel-button DOM tests;
  confirmed it fails without the fix and passes with it.

Blocker 2: the swap-conflict warning named other users' sessions.

- Both affectedSessions scans (POST .../custom-model and quick-start)
  walked the whole session map with no ownership filter, so in multi-
  user mode a non-admin pointing their own session at a shared
  endpoint learned another user's session name and id -- which with
  autoNameSessions on is that user's own prompt.
- The swap is still blocked pending confirmation regardless of
  ownership (a foreign session is just as real a disruption); only
  which ones get NAMED back to the caller is scoped, via the
  already-imported canAccessOwned. Added a two-owner test to
  test/routes/session-custom-model.test.ts covering both the
  foreign-owner (blocked, not named) and same-owner (named) cases.

Smaller ride-along fixes:

- server.ts boot recovery now passes contextLength into
  applyCustomModelInjection, so CLAUDE_CODE_MAX_CONTEXT_TOKENS is
  correctly rebuilt into _envOverrides after a restart instead of
  surviving only because tmux retains the old setenv.
- pumpLlamaSwapLogTail's finally now deletes by IDENTITY, not just by
  key, so an aborted pump finishing after a newer entry was created
  for the same endpoint can no longer delete that newer entry and
  orphan its connection.
- docs/custom-model-endpoints.md now notes that clearing a custom
  model removes injected keys by name, including CLAUDE_CONFIG_DIR --
  so a session that also had CLAUDE_CONFIG_DIR set via envOverrides
  (the per-client-account case) silently falls back to the default
  account on clear.

Left for later, as flagged in the review itself: the quick-start
case-scaffolding/cancel ordering (real behavioural reordering across
a large handler, too risky to make without a live re-test), and
retiring runCustomModelEntry's mode === 'claude' branch behind a
launchStrategy registry field (explicitly deferred by the reviewer to
"the next one").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-19 03:09:36 +08:00
DevvynandClaude Sonnet 5 9982a1325f fix(custom-model): address second pre-merge review (Ark0N)
Blocker: .center-status-banner never actually disappears.

- Add `.center-status-banner[hidden] { display: none; }`, same trap as
  `.home-sessions[hidden]`: the author-level `display: flex` beat the
  UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
  card stayed laid out at `opacity: 0` with its text/cancel/close
  children still `pointer-events: auto` -- an invisible 442x67 click
  blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
  banner (10001) and the swap-confirm/context-warning modals (10010)
  in CLAUDE.md's Z-index layers list.

Stale wording pointed at the reverted sticky-toast default:

- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
  `.toast-message` comment in styles.css all still said "toasts
  default to sticky" after 1f32128c put the flat 3s default back.
  Reworded all three to describe the actual behaviour: one call site
  passes an explicit `duration: 0`.

Smaller items from the same review:

- docs/api-reference.md said discovery failures answer
  `502 OPERATION_FAILED`; OPERATION_FAILED is 422 per src/types/api.ts
  and the error-code table earlier in the same file.
- The periodic re-discovery sweep (server.ts) never read
  customModelEndpointsEnabled, so turning the feature off left
  Codeman polling every saved endpoint forever. Added
  readCustomModelEndpointsEnabled() (custom-model-routes.ts, same
  shape as readPlanUsageTelemetryEnabled) and gated the interval
  callback on it.
- Reverted the formatting-only Prettier pass docs/api-reference.md
  picked up (table padding, *x* to _x_, JSON re-indent) by re-merging
  the new Custom Model Endpoints section onto the pre-PR file, so the
  diff is reviewable. No prose content was lost -- verified by diffing
  the result against the pre-revert file (formatting-only) and against
  the merge-base file (only the new section added).
- docs/custom-model-endpoints.md now states that a custom-model Claude
  session's isolated CLAUDE_CONFIG_DIR loses the user's global
  settings.json, user-level skills/agents/commands, and MCP servers
  from ~/.claude.json -- only `projects` is symlinked back.

Design question left open in the review (does `confirmed: true` need
to be two flags so "launch anyway" on the context warning doesn't also
skip the llama-swap displacement warning): keeping the single flag, as
offered. The 20s displacement sweep still catches a resulting swap
after the fact, so it's a surprise rather than a silent failure, and
splitting it is real behavioural surface I have no way to verify live
in this environment.

`npm run test:browser` could not be run in this environment (no tmux,
no downloaded Playwright browser binary) -- none of its suite's files
touch code this fix changes, but it still needs a real pass before
merge, same as any frontend change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-18 21:45:50 +08:00
DevvynandClaude Sonnet 5 db9729e1fc feat(custom-model): remove loading-banner countdown, add manual Cancel
Replaces the size-scaled expected-time estimate + matching auto-timeout
with a generic hardware/model-size disclaimer and a user-driven Cancel
button, per explicit request. Real load time depends on hardware this
feature has no way to know (VRAM, storage speed, GPU contention), so
the old estimate/timeout was a guess dressed up as a fact — worse, one
that could kill a genuinely slow load partway through on slower
hardware.

- _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline
  entirely — polls indefinitely until ready or cancelled, no automatic
  give-up. Message is now "Loading <model> (<size>) on <endpoint> —
  this can take a while depending on your hardware and the model
  size.", with the real llama.cpp log line still on its own second
  line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/
  _formatRemaining (dead code once the countdown is gone) —
  _lookupModelSizeGB is kept, the GB figure still shows.
- _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real
  "Cancel" button (distinct from the error-type "×" close button,
  since Cancel has a real consequence) that calls it on click. Caller
  owns what cancelling actually means, same split as the swap-confirm
  modal's promise-resolving buttons.
- Cancelling dismisses the banner, shows an info toast (not an error —
  this was deliberate), and closes the session, mirroring what the old
  timeout used to do automatically but now on the user's own call.
- New .center-status-cancel CSS (bordered pill button, distinct from
  the plain "×" close glyph).

Test changes: removed the now-invalid timeout-auto-close/estimate
tests, added cancel-flow tests (dismiss/toast-type/session-close,
never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests
for the new Cancel button (bootAppWithRealCenterStatus, evaluating
panels-ui.js instead of stubbing _showCenterStatus, since this button
is worth verifying for real rather than just through the stub every
other test in the file uses). Typecheck/lint/frontend-syntax clean;
full suite shows no new regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:00:50 +08:00
DevvynandClaude Sonnet 5 2d3fc65758 feat(custom-model): show real-time llama.cpp backend status in the loading banner
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.

- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
  (custom-model-routes.ts): one persistent GET /api/events (SSE)
  connection held open per endpoint, parsing logData frames and
  keeping the latest source:"upstream" (backend llama-server) line —
  filtering out llama-swap's own source:"proxy" request-access lines.
  Idle-closed after 30s of no polling, same 20s sweep as the existing
  swap-displacement check.
- running-status route now returns logLine alongside the existing
  isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
  ("llama.cpp: <line>", bootlog timestamp/level/component prefix
  stripped for display) that stays on the last real thing llama.cpp
  said rather than clearing to blank between polls.

⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).

12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 12:22:45 +08:00
DevvynandClaude Sonnet 5 b45a96358e feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.

- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
  registry entry declares contextLengthVar (currently only claude) and
  the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
  (40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
  quick-start customModel path) check this before the swap-conflict
  check and before launching/restarting anything, returning
  {requiresContextWarning, modelId, contextLength, minSafeContextTokens}
  — skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
  _resolveContextWarningConfirm (session-ui.js), wired into both
  _quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
  (the path Claude actually uses) ahead of the swap-confirmation check.
  Explains the fix in-modal: give the model an explicit larger -c/
  --ctx-size in llama-swap instead of relying on --fit-ctx, which
  optimizes for the biggest model that fits rather than the biggest
  context.

Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 07:36:57 +08:00
DevvynandClaude Sonnet 5 7bbe408e44 feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout
(updated every poll, so every second by default) instead of a static
"this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically
~1-3 min) on llama-swap - 47s remaining".

If the countdown reaches zero and the model still isn't ready, this is now
treated as a real failure rather than a "keep waiting" shrug:
- The banner turns into a sticky error (_showCenterStatus gains a `type`
  option - 'error' drops the spinner and adds a close button, since nothing
  is "in progress" anymore and a sticky message needs a way to dismiss it),
  naming the llama-swap server's own logs as where to look for detail.
- The session that load was for is closed automatically (closeSession) -
  requested explicitly: a console left open and pointed at a model that
  never finished loading is worse than no console at all. Both apply paths
  now thread the new session's id through to _watchLlamaSwapLoading for
  this (new required 3rd parameter, after endpointId/modelId).

_watchLlamaSwapGeneration's existing stale-call guard extends naturally to
this: a superseded call's own eventual timeout recognises it no longer owns
the banner and neither shows the error nor closes a session that may by
then belong to a different, newer launch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 20:22:02 +08:00
DevvynandClaude Sonnet 5 55dae31530 feat(custom-model): estimate model load time from its discovered size
Discovery now also parses a GB figure out of an auto-discovered model's own
description (llama-swap writes "Auto-discovered 16.35 GB - parameters
auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike
context length this needs no /props probe (the figure is right there in
/v1/models) so it is populated for every model regardless of loaded state.
A hand-configured profile's own description has no such figure and
correctly gets no entry.

The loading banner (_watchLlamaSwapLoading) now looks this up and, when
known, shows it plus a rough estimate from a small size->time matrix
(_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) -
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on
llama-swap... this can take a while" - and uses that same estimate's own
bracket to scale the banner's default give-up timeout for a very large
model, instead of a flat 5 minutes for everything. Explicitly labelled as
an UNMEASURED, typical-hardware estimate in every relevant comment - this
is not benchmarked against any real endpoint's actual storage/GPU, just a
reasonable expectation-setter. A model with no discoverable size (a
hand-configured profile) gets no size/estimate shown at all, matching the
"never a guess" convention modelContextLengths already established.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 19:58:46 +08:00
DevvynandClaude Sonnet 5 0af233c96c fix(custom-model): poll llama-swap readiness every 1s, check immediately, extend the cap
Reported: the "Loading..." banner stayed up past 2 minutes even though
llama-swap itself had already finished loading the model. Three fixes:

1. pollIntervalMs default 3000ms -> 1000ms (as asked).
2. The loop now checks readiness IMMEDIATELY on entry rather than sleeping
   a full interval first - a model that's already ready (a fast load, or a
   re-apply onto one already loaded) shouldn't sit on "Loading..." at all.
3. maxWaitMs default 120000ms (2 min) -> 300000ms (5 min): a large (20GB+)
   model reading from disk can genuinely take longer than 2 minutes, which
   would have looked identical to the reported symptom - "still stuck past
   the point it should have resolved" - except it would have actually
   flipped to a "still waiting" warning toast at the 2-minute mark rather
   than staying on "Loading" indefinitely, so this alone doesn't explain
   what was reported, but is a real, separate improvement worth making.

Also fixes a real, separate bug this surfaced while reasoning through the
report: _showCenterStatus's banner is ONE shared, reused DOM node. A second
call to _watchLlamaSwapLoading (e.g. switching models again before the
first switch's loop had finished) would take over that shared banner, but
the FIRST loop was still running and would eventually dismiss or overwrite
it once ITS OWN deadline or readiness check resolved - clobbering whatever
the second, current loop had put there. A generation counter
(_watchLlamaSwapGeneration) now lets each call recognise when it no longer
owns the banner and stop touching it silently, rather than only the last
call to actually start ever safely reading or writing it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 19:44:37 +08:00
DevvynandClaude Sonnet 5 01b32ee6cd fix(custom-model): move the switching/loading status to a centred banner
The "Claude started - switching to <endpoint>..." and "Loading <model> on
<endpoint>... this can take a while" messages lived in the top-right toast
corner along with everything else, easy to miss given they can each sit on
screen for well over a minute (a real llama-swap model load).

Adds _showCenterStatus() (panels-ui.js): a single, reused, screen-centred
banner with a spinner, non-blocking (no backdrop, pointer-events: none on
the wrapper) so it never gets in the way of using the app while it's up.
Both call sites (_runCustomModelEntryViaRestart's switching message,
_watchLlamaSwapLoading's loading message) now use it instead of showToast.
Every OTHER status in these two flows - the llama-swap conflict warning
already moved to its own modal, apply failures, cancellation, and
_watchLlamaSwapLoading's own final "ready"/"still waiting" outcome - stays
exactly where it was, in the corner.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:40:43 +08:00
DevvynandClaude Sonnet 5 2936ba6e3d fix(custom-model): replace the native confirm() popup with an in-app modal
The llama-swap "this will unload it for session X" warning used a native
browser confirm() popup, which looks out of place next to the rest of the
app's own modals.

Adds #customModelSwapConfirmModal (index.html) with Cancel/Switch-anyway
buttons, styled to match the app. _confirmModelSwap(message) shows it and
returns a promise that resolves true/false the same way confirm() would;
_resolveModelSwapConfirm(proceed) (wired to both buttons and the backdrop
click) settles it. Both llama-swap conflict call sites
(_quickStartWithCustomModelConfirm for the one-shot launch path,
_runCustomModelEntryViaRestart for Claude's restart path) now await this
instead of calling confirm() directly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:29:38 +08:00
DevvynandClaude Sonnet 5 83033b4299 fix(custom-model): show a status toast during Claude's native-boot-then-restart window
Claude stays on the launch-then-restart path (see runCustomModelEntry's own
comment for why), but with nothing on screen during that window, a native
boot that briefly talks to the cloud model read as "the endpoint didn't
apply" rather than "the switch hasn't happened yet".

A sticky "Claude started - switching to <endpoint>..." toast now covers the
whole window from the native launch through the apply call, updated in
place (never stacked) as the outcome resolves: dismissed on cancel or
failure (replaced by the existing cancellation/error toast), handed off to
_watchLlamaSwapLoading's own sticky toast when a model swap is in progress,
or updated to the existing "Pointed at ... - restarting" message and
auto-dismissed after 3s on a plain success.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:21:09 +08:00
DevvynandClaude Sonnet 5 bcebc81fcd feat(custom-model): detect llama-swap model conflicts before switching
Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.

1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
   server) via its own GET /running, which plain llama.cpp has no concept of
   at all. New GET /api/model-endpoints/:id/running-status route exposes this
   read-only, for the frontend's polling loop below.

2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
   what llama-swap currently has loaded. If it differs from the requested
   model AND another live session's own customModel selection is actively
   using that loaded model, the apply is refused with a
   {requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
   instead of silently switching. A "confirmed: true" field on the retry
   skips the check. Switching with nothing else affected proceeds
   immediately, no confirmation asked, only ever when there is something to
   warn about.

3. The frontend (runCustomModelEntry) shows a native confirm() naming the
   affected session(s) and the model they'd lose, matching this codebase's
   existing convention for this class of decision (delete case, kill
   session, etc.) rather than a new modal. On a successful apply the response
   also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
   poll shows a sticky "Loading <model>..." toast via the new running-status
   route until llama-swap reports the target model ready (bounded at 2
   minutes), so a prompt sent mid-swap reads as "loading", never as silence
   or an answer from whatever was loaded a moment before.

Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 14:35:07 +08:00
DevvynandClaude Sonnet 5 5c25a52f95 fix(custom-model): wait for a freshly launched session to go idle before applying
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.

Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.

Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 10:38:37 +08:00
DevvynandClaude Sonnet 5 409a6e65f9 fix(custom-model,toast): surface the real apply error, and make error toasts sticky with a close button
Two related fixes, both needed to actually diagnose 'Session started on
the native backend — could not apply the custom endpoint' reports from
live testing:

1. runCustomModelEntry()'s apply call went through _apiJson(), which
   unwraps a success body but SWALLOWS a failure response entirely and
   returns null — discarding the one thing (error, errorCode) that would
   tell 'endpoint unreachable' apart from 'not a discovered model',
   'remote/Docker session', or a dozen other real causes the apply route
   already reports distinctly. Switched to _api() so the actual response
   body is read on failure too, and the toast now includes the real
   message.
2. showToast() defaulted every toast, error or not, to a 3s auto-dismiss
   with no way to read it again — exactly what made the above generic
   message impossible to act on even before the fix above. Error toasts
   now default to sticky (duration: 0, no auto-dismiss) unless a caller
   opts into a duration, and every toast — sticky or not — gets an
   explicit close (x) button, since a sticky toast with no way to
   dismiss it would just accumulate across repeated failures.

Tests: custom-model-run-menu-ui.test.ts's two apply tests updated for the
_api() switch (their mocks previously stubbed _apiJson, which the apply
call no longer goes through), plus a new test pinning that the real
server error string reaches the toast on a failure.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 09:55:03 +08:00
DevvynandClaude Sonnet 5 5a9ff07f57 feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:

1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
   apply the endpoint's defaultModelId (or the first discovered model)
   silently. Now, via the new selectCustomModelEntry() (session-ui.js):
   - exactly one discovered model launches straight away, same as before
   - two or more open a new #customModelPickModal listing every discovered
     model; defaultModelId (if set) is marked but never auto-chosen, since
     the point of asking is letting ONE launch deliberately differ from
     the saved default, not just confirming it
   The endpoint is re-fetched at click time rather than trusting anything
   cached from the dropdown's own render, since the model list can have
   changed (the sweep below, or a settings-panel edit) since it opened.
   runCustomModelEntry() itself — the actual launch, routed through run()
   for the in-flight lock, snapshot-guarded against applying to the wrong
   session — is unchanged; it now just always receives an explicit model
   id from one of these two paths instead of computing one itself.

2. Periodic re-discovery. Every saved endpoint's models now refresh
   automatically every 5 minutes in the background
   (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
   as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
   off under testMode), so a model the server starts or stops serving shows
   up without another manual "Discover" click. The manual POST
   .../discover-models route and the new refreshAllCustomModelHosts()
   sweep (custom-model-routes.ts) now share one pure merge step
   (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
   that no longer appears) rather than two copies that could drift. The
   sweep is best-effort per host — one endpoint being unreachable on a
   cycle never blocks the others — and re-reads the store before each
   host's write, keyed by id, so a concurrent edit or delete from the
   settings panel always wins over a sweep that started before it.

Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.

Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 08:50:18 +08:00
DevvynandClaude Sonnet 5 60e1bd52f7 fix(custom-model): act on the draft review — unparseable onclick, unwrapped envelope, wrong-session apply, missing lock, no tests
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.

Blockers:

1. Every generated inline onclick was unparseable. JSON.stringify's own
   double quotes terminated the double-quoted HTML attribute at the first
   one, leaving btn.onclick null on every picker entry and every Discover/
   Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
   argument, the same idiom deleteCase's onclick already uses four lines
   away in session-ui.js. This also closes the live-HTML-injection route
   through modelId (server-controlled, from the endpoint's own /v1/models
   reply): with quoting intact, a `>` inside it can no longer terminate the
   <button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
   like every other /api route (server.ts's preSerialization hook applies
   to arrays too), so Array.isArray(hosts) was always false in production
   and the picker/settings panel silently saw nothing. Both call sites now
   go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
   returns normally without ever changing activeSessionId, so the apply
   step used to silently re-point and restart whatever session the user was
   already looking at. runCustomModelEntry() now snapshots activeSessionId
   before the launch and requires it to have actually changed.

Majors:

4. Routes the launch through run() itself via a temporary _runMode swap
   (never persisted — setRunMode() would sync it to the server) instead of
   a parallel hardcoded dispatch table, so a custom-model launch now holds
   the same _runInFlight lock every other Run click gets. This also
   resolves the "hardcoded runners map contradicts the PR's own design"
   minor: dispatch is run()'s own, so a CLI whose customModelInjection
   recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
   against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
   parses markup this module generated itself) for exactly the DOM-level
   facts the review said needed no Playwright and no tmux: a generated
   button's onclick genuinely compiles and fires, a dangerous modelId never
   produces a live element, the envelope unwrap works, the session-changed
   guard holds, run() actually gets called (proving the in-flight lock
   engages), and _runMode is restored afterward. Confirmed against the
   pre-fix code first (reproduces btn.onclick === null exactly) so this
   isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
   render-index-html.test.ts for the other fixes below.

Minors:

- Generated entries now filter through isCliAvailable(), matching
  _refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
  (applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
  and to settings-modal open) instead of always rendering; the endpoint GET
  no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
  redactApiKey() replaces the field with a computed apiKeySet: boolean, and
  a PUT with no apiKey now keeps the stored one server-side
  (applyStoredApiKey()) instead of the client resending a value it was
  never given. New tests cover both directions (kept vs. replaced) by
  observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
  (_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
  event, since the real role can resolve after settings were first opened)
  — endpoint writes were already admin-only server-side, but the button
  used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
  toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
  route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
  that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
  (CLAUDE.md already records that exact literal turning the settings
  preview into a grey slab on light skins), .run-mode-custom-models gets
  the same gap: 2px .run-mode-menu's own flex gap only applies one level
  up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
  </script> (CliEntry.label is user-clis.json-settable, unlike
  __codemanCliAvailable's booleans-only payload) via a new exported
  escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
  docs/api-reference.md; left the "no zh-CN for the new Models-section
  group" minor unaddressed only insofar as the wider Models section (task
  routing, thinking effort, etc.) has never had zh-CN coverage either —
  everything this PR itself introduces (labels, hints, button text, the
  Run-menu's "Custom Endpoints" header) IS translated in i18n.js.

Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00