feat(custom-model): estimate model load time from its discovered size

Discovery now also parses a GB figure out of an auto-discovered model's own
description (llama-swap writes "Auto-discovered 16.35 GB - parameters
auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike
context length this needs no /props probe (the figure is right there in
/v1/models) so it is populated for every model regardless of loaded state.
A hand-configured profile's own description has no such figure and
correctly gets no entry.

The loading banner (_watchLlamaSwapLoading) now looks this up and, when
known, shows it plus a rough estimate from a small size->time matrix
(_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) -
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on
llama-swap... this can take a while" - and uses that same estimate's own
bracket to scale the banner's default give-up timeout for a very large
model, instead of a flat 5 minutes for everything. Explicitly labelled as
an UNMEASURED, typical-hardware estimate in every relevant comment - this
is not benchmarked against any real endpoint's actual storage/GPU, just a
reasonable expectation-setter. A model with no discoverable size (a
hand-configured profile) gets no size/estimate shown at all, matching the
"never a guess" convention modelContextLengths already established.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 19:58:46 +08:00
co-authored by Claude Sonnet 5
parent 0af233c96c
commit 55dae31530
7 changed files with 329 additions and 41 deletions
+14
View File
@@ -83,6 +83,20 @@ session" below) so a CLI that would otherwise assume a large default context
window for an unrecognized model id stops silently overflowing a much
smaller real one.
**File size is discovered too, when the server states one.** llama-swap
writes a GB figure into an auto-discovered model's own `description`
(`"Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"`), parsed
into `modelSizesGB` — unlike context length, this needs no `/props` probe
(the figure is right there in the `/v1/models` response) and so is populated
for every model regardless of loaded state. A hand-configured profile's own
description has no such figure and correctly gets no entry, never a guess.
Used only to label the Run-menu picker's "loading model" banner with a
rough, UNMEASURED expected-time estimate (`_estimateModelLoad()` in
session-ui.js, based on typical local NVMe/SSD throughput — not benchmarked
against any real endpoint's actual hardware/storage) and to scale that same
banner's own give-up timeout for a very large model; never anything a
server-side check relies on.
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
from the endpoint's own discovered `models`, and the route refuses a value