mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-08 16:39:42 +02:00
fix(custom-model): actually trigger the llama-swap load, not just watch for it
Root cause of "it doesn't look like llama-swap is actually switching the model" (confirmed live: no load_model line in llama-swap's own logs after applying a selection). llama-swap has no "switch model" admin endpoint - the ONLY thing that starts a swap is a real inference request naming the model. Every previous fix (the conflict check, the loading banner) assumed a swap would start on its own; nothing ever actually asked llama-swap to load anything until the launched CLI's first real prompt did, which could be much later than "applying the selection" implied. Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest real request that will start a load - POST <baseUrl>/v1/chat/completions, max_tokens: 1, one throwaway message - fire-and-forget (never awaited by the caller; the frontend's own running-status polling is what actually confirms readiness). Wired into both apply paths (the dedicated restart route and the one-shot quick-start route), fired whenever the target model isn't already the one loaded and ready - a broader condition than the existing swapNeeded (which only gates the "this will evict another session's model" confirmation ask and deliberately stays narrow to that). modelSwapInProgress in both routes' responses now reflects this same broader condition too, so the frontend's loading banner actually correlates with a real in-flight load rather than only firing when something else happened to be loaded already. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
01b32ee6cd
commit
0929694012
@@ -73,6 +73,16 @@ redirecting those hasn't landed yet, see below. The picker also only appears in
|
||||
**Run** dropdown; the phone home screen builds its own run picker separately and does not
|
||||
currently offer these entries.
|
||||
|
||||
**Against llama-swap, applying a selection also starts the actual model load, rather than
|
||||
waiting on your first prompt to do it.** llama-swap has no "switch model" button of its own
|
||||
— the only thing that starts a swap is a real request naming the model, and confirmed live:
|
||||
just applying a selection never reached llama-swap's own logs at all until something asked
|
||||
it to load. Picking an entry now also sends the smallest real request that will trigger
|
||||
that load, in the background, the moment the target model isn't already loaded and ready —
|
||||
which is what the prominent **"Loading `<model>`… this can take a while"** banner
|
||||
(centred on screen, not a corner toast — a real load can take well over a minute) is
|
||||
actually watching for.
|
||||
|
||||
**Claude Code specifically gets two extra fixes applied automatically:**
|
||||
|
||||
- Its discovered context length (see above) is passed through as
|
||||
|
||||
Reference in New Issue
Block a user