fix(custom-model): actually trigger the llama-swap load, not just watch for it

Root cause of "it doesn't look like llama-swap is actually switching the
model" (confirmed live: no load_model line in llama-swap's own logs after
applying a selection). llama-swap has no "switch model" admin endpoint - the
ONLY thing that starts a swap is a real inference request naming the model.
Every previous fix (the conflict check, the loading banner) assumed a swap
would start on its own; nothing ever actually asked llama-swap to load
anything until the launched CLI's first real prompt did, which could be
much later than "applying the selection" implied.

Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest
real request that will start a load - POST <baseUrl>/v1/chat/completions,
max_tokens: 1, one throwaway message - fire-and-forget (never awaited by
the caller; the frontend's own running-status polling is what actually
confirms readiness). Wired into both apply paths (the dedicated restart
route and the one-shot quick-start route), fired whenever the target model
isn't already the one loaded and ready - a broader condition than the
existing swapNeeded (which only gates the "this will evict another
session's model" confirmation ask and deliberately stays narrow to that).
modelSwapInProgress in both routes' responses now reflects this same
broader condition too, so the frontend's loading banner actually correlates
with a real in-flight load rather than only firing when something else
happened to be loaded already.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 15:53:34 +08:00
co-authored by Claude Sonnet 5
parent 01b32ee6cd
commit 0929694012
6 changed files with 246 additions and 2 deletions
+29
View File
@@ -239,6 +239,35 @@ earlier launch in the same isolated directory (`userID`, `numStartups`,
earlier approved keys), and a missing or corrupt file is treated as empty
rather than failing the apply.
**llama-swap gets two more fixes on top of the context-length/config-dir
ones above, both from watching a real switch live.** llama.cpp only ever
runs one model at a time; llama-swap swaps the backing process on demand,
which can take anywhere from a few seconds to well over a minute:
- **The conflict check.** Both apply routes (the restart one here and the
one-shot `POST /api/quick-start` above) call llama-swap's own
`GET /running` first — feature-detected, so a plain llama.cpp/OpenAI-
compatible server (no such endpoint) is simply never checked. If a
*different* model is currently loaded and ready, and another **live
session's own selection** is using it, the apply returns
`{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}`
instead of silently switching — nothing is applied or created yet.
Retrying with `confirmed: true` skips the check. Switching with nothing
else affected proceeds immediately; this is a warning about disrupting
another session, never a gate on the switch itself.
- **Actually starting the load.** llama-swap has no "switch model" admin
call — the only thing that starts a swap is a real inference request
naming the model, and confirmed live: applying a selection alone never
reached llama-swap at all (nothing in its own server logs), since nothing
had actually asked it to load anything yet. Both apply routes now also
send the smallest real request that will — `POST <baseUrl>/v1/chat/
completions` with `max_tokens: 1` and one throwaway message — whenever the
target model isn't already the one loaded and ready, fire-and-forget (its
response is never read; `GET /api/model-endpoints/:id/running-status`,
polled client-side, is what actually confirms readiness). The response
also carries `modelSwapInProgress: true` in that case, which is what
drives the Run-menu picker's own "loading model" status banner.
Clear back to the harness's native cloud default with:
```bash