feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes

Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:

1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
   apply the endpoint's defaultModelId (or the first discovered model)
   silently. Now, via the new selectCustomModelEntry() (session-ui.js):
   - exactly one discovered model launches straight away, same as before
   - two or more open a new #customModelPickModal listing every discovered
     model; defaultModelId (if set) is marked but never auto-chosen, since
     the point of asking is letting ONE launch deliberately differ from
     the saved default, not just confirming it
   The endpoint is re-fetched at click time rather than trusting anything
   cached from the dropdown's own render, since the model list can have
   changed (the sweep below, or a settings-panel edit) since it opened.
   runCustomModelEntry() itself — the actual launch, routed through run()
   for the in-flight lock, snapshot-guarded against applying to the wrong
   session — is unchanged; it now just always receives an explicit model
   id from one of these two paths instead of computing one itself.

2. Periodic re-discovery. Every saved endpoint's models now refresh
   automatically every 5 minutes in the background
   (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
   as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
   off under testMode), so a model the server starts or stops serving shows
   up without another manual "Discover" click. The manual POST
   .../discover-models route and the new refreshAllCustomModelHosts()
   sweep (custom-model-routes.ts) now share one pure merge step
   (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
   that no longer appears) rather than two copies that could drift. The
   sweep is best-effort per host — one endpoint being unreachable on a
   cycle never blocks the others — and re-reads the store before each
   host's write, keyed by id, so a concurrent edit or delete from the
   settings panel always wins over a sweep that started before it.

Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.

Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 08:50:18 +08:00
co-authored by Claude Sonnet 5
parent 60e1bd52f7
commit 5a9ff07f57
12 changed files with 471 additions and 35 deletions
+5 -1
View File
@@ -551,7 +551,11 @@ user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md).
`lastDiscoveredAt`. A `defaultModelId` that no longer appears in the fresh
list is dropped rather than carried forward invalid. Failures answer
`502 OPERATION_FAILED` with the underlying connection error, or a named
egress refusal if the resolved address turned out to be blocked.
egress refusal if the resolved address turned out to be blocked. The same
refresh also runs automatically for every saved endpoint every 5 minutes
in the background (`refreshAllCustomModelHosts()`, `custom-model-routes.ts`,
started from `server.ts`), so there is no route for triggering "refresh
all" — one endpoint being unreachable on a cycle never blocks the others.
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId } |
{ clear: true }` applies (or clears) the session's selection and
**restarts the session's CLI process in place** — every supported harness
+32 -15
View File
@@ -66,18 +66,26 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
Endpoint management is admin-only in multi-user mode, same as remote/docker
hosts — these are machine-level infra, not per-user settings.
`defaultModelId` names which discovered model the Run-menu picker applies
for that endpoint with no further choice — the settings panel's Edit form
exposes it as a select populated from the endpoint's own discovered
`models`, and the route refuses a value that isn't one of them. Leaving it
unset falls back to the first discovered model; re-discovering drops a
default that no longer appears in the fresh list rather than carrying an
invalid one forward.
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
from the endpoint's own discovered `models`, and the route refuses a value
that isn't one of them. It is applied automatically only when the endpoint
has exactly one discovered model (nothing to choose); with two or more it
is a pre-selection in the model-picker dialog below, never a silent default.
Re-discovering drops a default that no longer appears in the fresh list
rather than carrying an invalid one forward.
**Model lists refresh themselves.** A background sweep (`server.ts`,
`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every
saved endpoint the same way the manual `POST .../discover-models` route
does, best-effort per endpoint — one being unreachable on a given cycle
never blocks the others. Off under `npm test`, same reasoning as the Codex
plan-usage poll it sits beside: no real network to hit, no server instance
to keep the timer alive for.
## The Run-menu picker
With the setting on and at least one endpoint carrying a usable default
model (either an explicit `defaultModelId` or just one discovered model),
With the setting on and at least one endpoint carrying a discovered model,
the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry
per (harness that can redirect to a custom endpoint, saved endpoint) pair,
e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI
@@ -86,13 +94,22 @@ registry's own `capabilities.customModelInjection` at page render
in the frontend — so a CLI whose injection recipe lands later shows up with
no frontend change, and Antigravity (`unsupported`) never does.
Picking an entry runs a single session on that harness exactly the way its
Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`,
`session-ui.js`) rather than trusting anything cached from the dropdown's
own render — the model list can have changed via the 5-minute sweep above
or a settings-panel edit since the menu opened. With exactly one discovered
model it runs straight away; with two or more, a small modal
(`#customModelPickModal`) lists them and asks which one to use for this
launch, with the endpoint's `defaultModelId` marked but not auto-chosen —
the point of asking is letting one launch deliberately differ from the
saved default, not just confirming it. Whichever way the model was decided,
the launch itself runs a single session on that harness exactly the way its
own Run-menu entry would (same case creation, env overrides, everything),
then immediately applies the endpoint's default model to it via the route
below. It is a one-off "try this endpoint" action, not a sticky mode: the
plain Run button still means "this harness, native cloud" afterward.
Entries are hidden entirely for a remote or Docker active case, since the
apply route refuses both (see the next section).
then immediately applies the endpoint and model to it via the route below.
It is a one-off "try this endpoint" action, not a sticky mode: the plain
Run button still means "this harness, native cloud" afterward. Entries are
hidden entirely for a remote or Docker active case, since the apply route
refuses both (see the next section).
## Applying a model to a session
+18 -6
View File
@@ -16,20 +16,32 @@ Still in App Settings → Models → Custom model endpoints:
say). An API key is optional; most local servers don't check one.
2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it.
3. Pick a **default model** from what was discovered. This is the model the Run-menu entry
below applies with no further choice, so set it once you know which one you want.
applies directly when only one model is discovered; with two or more, it's just the one
pre-marked in the picker dialog described below, not a silent default.
Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker
hosts — these are machine-level infra, not a per-user setting.
**Model lists refresh themselves.** Every saved endpoint is re-discovered automatically every
5 minutes in the background, so a model the server starts serving later — or stops serving —
shows up without another manual click of **Discover**. One endpoint being unreachable on a
given cycle (powered off, wrong network) never blocks the others from refreshing.
## Running a session against one
With the setting on and at least one endpoint carrying a usable default model, the **Run**
With the setting on and at least one endpoint carrying a discovered model, the **Run**
dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a
custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a
session on that harness exactly the way its own entry would, then points it at the
endpoint's default model. It is a one-off "try this endpoint" action, not a sticky mode — the
plain **Run** button still means "this harness, native cloud" afterward, and a fresh session
never inherits whatever the last one was pointed at.
session on that harness exactly the way its own entry would. It is a one-off "try this
endpoint" action, not a sticky mode — the plain **Run** button still means "this harness,
native cloud" afterward, and a fresh session never inherits whatever the last one was
pointed at.
**Which model it uses depends on how many the endpoint has discovered.** With exactly one,
the session launches straight away on that model — nothing to choose. With two or more, a
small dialog asks which one to use for this launch before starting the session; the
endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately
use a different one without changing the saved default.
Applying a selection **restarts the harness's process in place** — same tab, same
conversation where the harness supports resuming one, fresh environment. That restart is