feat(custom-model): estimate model load time from its discovered size

Discovery now also parses a GB figure out of an auto-discovered model's own
description (llama-swap writes "Auto-discovered 16.35 GB - parameters
auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike
context length this needs no /props probe (the figure is right there in
/v1/models) so it is populated for every model regardless of loaded state.
A hand-configured profile's own description has no such figure and
correctly gets no entry.

The loading banner (_watchLlamaSwapLoading) now looks this up and, when
known, shows it plus a rough estimate from a small size->time matrix
(_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) -
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on
llama-swap... this can take a while" - and uses that same estimate's own
bracket to scale the banner's default give-up timeout for a very large
model, instead of a flat 5 minutes for everything. Explicitly labelled as
an UNMEASURED, typical-hardware estimate in every relevant comment - this
is not benchmarked against any real endpoint's actual storage/GPU, just a
reasonable expectation-setter. A model with no discoverable size (a
hand-configured profile) gets no size/estimate shown at all, matching the
"never a guess" convention modelContextLengths already established.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 19:58:46 +08:00
co-authored by Claude Sonnet 5
parent 0af233c96c
commit 55dae31530
7 changed files with 329 additions and 41 deletions
+11
View File
@@ -61,6 +61,17 @@ export interface CustomModelHost {
* has no entry for simply gets no context-length env override applied — never a guess.
*/
modelContextLengths?: Record<string, number>;
/**
* Discovered file size (GB) per model id, keyed by the same strings as `models`.
* Populated during discovery by parsing llama-swap's own `description` field for an
* auto-discovered model ("Auto-discovered 16.35 GB - parameters auto-fitted by
* llama.cpp") — a hand-configured profile's own description has no such figure and
* correctly gets no entry, never a guess. Used only to label the Run-menu picker's
* "loading model" banner with a rough, unmeasured expected-time estimate
* (`estimateModelLoad()` in session-ui.js) — never a guarantee, and never anything a
* server-side check relies on.
*/
modelSizesGB?: Record<string, number>;
}
export function customModelHostsPath(configDir: string): string {