fix(custom-model): actually trigger the llama-swap load, not just watch for it

Root cause of "it doesn't look like llama-swap is actually switching the
model" (confirmed live: no load_model line in llama-swap's own logs after
applying a selection). llama-swap has no "switch model" admin endpoint - the
ONLY thing that starts a swap is a real inference request naming the model.
Every previous fix (the conflict check, the loading banner) assumed a swap
would start on its own; nothing ever actually asked llama-swap to load
anything until the launched CLI's first real prompt did, which could be
much later than "applying the selection" implied.

Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest
real request that will start a load - POST <baseUrl>/v1/chat/completions,
max_tokens: 1, one throwaway message - fire-and-forget (never awaited by
the caller; the frontend's own running-status polling is what actually
confirms readiness). Wired into both apply paths (the dedicated restart
route and the one-shot quick-start route), fired whenever the target model
isn't already the one loaded and ready - a broader condition than the
existing swapNeeded (which only gates the "this will evict another
session's model" confirmation ask and deliberately stays narrow to that).
modelSwapInProgress in both routes' responses now reflects this same
broader condition too, so the frontend's loading banner actually correlates
with a real in-flight load rather than only firing when something else
happened to be loaded already.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 15:53:34 +08:00
co-authored by Claude Sonnet 5
parent 01b32ee6cd
commit 0929694012
6 changed files with 246 additions and 2 deletions
+10
View File
@@ -73,6 +73,16 @@ redirecting those hasn't landed yet, see below. The picker also only appears in
**Run** dropdown; the phone home screen builds its own run picker separately and does not
currently offer these entries.
**Against llama-swap, applying a selection also starts the actual model load, rather than
waiting on your first prompt to do it.** llama-swap has no "switch model" button of its own
— the only thing that starts a swap is a real request naming the model, and confirmed live:
just applying a selection never reached llama-swap's own logs at all until something asked
it to load. Picking an entry now also sends the smallest real request that will trigger
that load, in the background, the moment the target model isn't already loaded and ready —
which is what the prominent **"Loading `<model>`… this can take a while"** banner
(centred on screen, not a corner toast — a real load can take well over a minute) is
actually watching for.
**Claude Code specifically gets two extra fixes applied automatically:**
- Its discovered context length (see above) is passed through as