feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes

Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:

1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
   apply the endpoint's defaultModelId (or the first discovered model)
   silently. Now, via the new selectCustomModelEntry() (session-ui.js):
   - exactly one discovered model launches straight away, same as before
   - two or more open a new #customModelPickModal listing every discovered
     model; defaultModelId (if set) is marked but never auto-chosen, since
     the point of asking is letting ONE launch deliberately differ from
     the saved default, not just confirming it
   The endpoint is re-fetched at click time rather than trusting anything
   cached from the dropdown's own render, since the model list can have
   changed (the sweep below, or a settings-panel edit) since it opened.
   runCustomModelEntry() itself — the actual launch, routed through run()
   for the in-flight lock, snapshot-guarded against applying to the wrong
   session — is unchanged; it now just always receives an explicit model
   id from one of these two paths instead of computing one itself.

2. Periodic re-discovery. Every saved endpoint's models now refresh
   automatically every 5 minutes in the background
   (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
   as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
   off under testMode), so a model the server starts or stops serving shows
   up without another manual "Discover" click. The manual POST
   .../discover-models route and the new refreshAllCustomModelHosts()
   sweep (custom-model-routes.ts) now share one pure merge step
   (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
   that no longer appears) rather than two copies that could drift. The
   sweep is best-effort per host — one endpoint being unreachable on a
   cycle never blocks the others — and re-reads the store before each
   host's write, keyed by id, so a concurrent edit or delete from the
   settings panel always wins over a sweep that started before it.

Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.

Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 08:50:18 +08:00
co-authored by Claude Sonnet 5
parent 60e1bd52f7
commit 5a9ff07f57
12 changed files with 471 additions and 35 deletions
+1 -1
View File
File diff suppressed because one or more lines are too long
+5 -1
View File
@@ -551,7 +551,11 @@ user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md).
`lastDiscoveredAt`. A `defaultModelId` that no longer appears in the fresh
list is dropped rather than carried forward invalid. Failures answer
`502 OPERATION_FAILED` with the underlying connection error, or a named
egress refusal if the resolved address turned out to be blocked.
egress refusal if the resolved address turned out to be blocked. The same
refresh also runs automatically for every saved endpoint every 5 minutes
in the background (`refreshAllCustomModelHosts()`, `custom-model-routes.ts`,
started from `server.ts`), so there is no route for triggering "refresh
all" — one endpoint being unreachable on a cycle never blocks the others.
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId } |
{ clear: true }` applies (or clears) the session's selection and
**restarts the session's CLI process in place** — every supported harness
+32 -15
View File
@@ -66,18 +66,26 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
Endpoint management is admin-only in multi-user mode, same as remote/docker
hosts — these are machine-level infra, not per-user settings.
`defaultModelId` names which discovered model the Run-menu picker applies
for that endpoint with no further choice — the settings panel's Edit form
exposes it as a select populated from the endpoint's own discovered
`models`, and the route refuses a value that isn't one of them. Leaving it
unset falls back to the first discovered model; re-discovering drops a
default that no longer appears in the fresh list rather than carrying an
invalid one forward.
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
from the endpoint's own discovered `models`, and the route refuses a value
that isn't one of them. It is applied automatically only when the endpoint
has exactly one discovered model (nothing to choose); with two or more it
is a pre-selection in the model-picker dialog below, never a silent default.
Re-discovering drops a default that no longer appears in the fresh list
rather than carrying an invalid one forward.
**Model lists refresh themselves.** A background sweep (`server.ts`,
`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every
saved endpoint the same way the manual `POST .../discover-models` route
does, best-effort per endpoint — one being unreachable on a given cycle
never blocks the others. Off under `npm test`, same reasoning as the Codex
plan-usage poll it sits beside: no real network to hit, no server instance
to keep the timer alive for.
## The Run-menu picker
With the setting on and at least one endpoint carrying a usable default
model (either an explicit `defaultModelId` or just one discovered model),
With the setting on and at least one endpoint carrying a discovered model,
the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry
per (harness that can redirect to a custom endpoint, saved endpoint) pair,
e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI
@@ -86,13 +94,22 @@ registry's own `capabilities.customModelInjection` at page render
in the frontend — so a CLI whose injection recipe lands later shows up with
no frontend change, and Antigravity (`unsupported`) never does.
Picking an entry runs a single session on that harness exactly the way its
Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`,
`session-ui.js`) rather than trusting anything cached from the dropdown's
own render — the model list can have changed via the 5-minute sweep above
or a settings-panel edit since the menu opened. With exactly one discovered
model it runs straight away; with two or more, a small modal
(`#customModelPickModal`) lists them and asks which one to use for this
launch, with the endpoint's `defaultModelId` marked but not auto-chosen —
the point of asking is letting one launch deliberately differ from the
saved default, not just confirming it. Whichever way the model was decided,
the launch itself runs a single session on that harness exactly the way its
own Run-menu entry would (same case creation, env overrides, everything),
then immediately applies the endpoint's default model to it via the route
below. It is a one-off "try this endpoint" action, not a sticky mode: the
plain Run button still means "this harness, native cloud" afterward.
Entries are hidden entirely for a remote or Docker active case, since the
apply route refuses both (see the next section).
then immediately applies the endpoint and model to it via the route below.
It is a one-off "try this endpoint" action, not a sticky mode: the plain
Run button still means "this harness, native cloud" afterward. Entries are
hidden entirely for a remote or Docker active case, since the apply route
refuses both (see the next section).
## Applying a model to a session
+18 -6
View File
@@ -16,20 +16,32 @@ Still in App Settings → Models → Custom model endpoints:
say). An API key is optional; most local servers don't check one.
2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it.
3. Pick a **default model** from what was discovered. This is the model the Run-menu entry
below applies with no further choice, so set it once you know which one you want.
applies directly when only one model is discovered; with two or more, it's just the one
pre-marked in the picker dialog described below, not a silent default.
Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker
hosts — these are machine-level infra, not a per-user setting.
**Model lists refresh themselves.** Every saved endpoint is re-discovered automatically every
5 minutes in the background, so a model the server starts serving later — or stops serving —
shows up without another manual click of **Discover**. One endpoint being unreachable on a
given cycle (powered off, wrong network) never blocks the others from refreshing.
## Running a session against one
With the setting on and at least one endpoint carrying a usable default model, the **Run**
With the setting on and at least one endpoint carrying a discovered model, the **Run**
dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a
custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a
session on that harness exactly the way its own entry would, then points it at the
endpoint's default model. It is a one-off "try this endpoint" action, not a sticky mode — the
plain **Run** button still means "this harness, native cloud" afterward, and a fresh session
never inherits whatever the last one was pointed at.
session on that harness exactly the way its own entry would. It is a one-off "try this
endpoint" action, not a sticky mode — the plain **Run** button still means "this harness,
native cloud" afterward, and a fresh session never inherits whatever the last one was
pointed at.
**Which model it uses depends on how many the endpoint has discovered.** With exactly one,
the session launches straight away on that model — nothing to choose. With two or more, a
small dialog asks which one to use for this launch before starting the session; the
endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately
use a different one without changing the saved default.
Applying a selection **restarts the harness's process in place** — same tab, same
conversation where the harness supports resuming one, fresh environment. That restart is
+3
View File
@@ -311,6 +311,9 @@
'What the Run-menu picker applies for this endpoint. Discover models first.':
'运行菜单选择器会为此端点应用该模型。请先发现可用模型。',
'Custom Endpoints': '自定义端点',
'Choose a model': '选择模型',
'That endpoint no longer exists': '该端点已不存在',
'No models discovered for this endpoint yet': '此端点尚未发现任何模型',
'Subagent Options': '子智能体选项',
'Enable Tracking': '启用跟踪',
'Active Tab Only': '仅活动标签页',
+17
View File
@@ -910,6 +910,23 @@
</div>
</div>
<!-- Custom Model Endpoint Profiles: "which model" picker (docs/custom-model-endpoints-plan.md).
Shown only when the chosen endpoint has more than one discovered model — see
selectCustomModelEntry() in session-ui.js, which skips straight to launch otherwise. -->
<div class="modal" id="customModelPickModal">
<div class="modal-backdrop" onclick="app.closeCustomModelPickModal()"></div>
<div class="modal-content" style="max-width: 380px;">
<div class="modal-header">
<h3 id="customModelPickTitle">Choose a model</h3>
<button class="modal-close" onclick="app.closeCustomModelPickModal()" aria-label="Close model picker">&times;</button>
</div>
<div class="modal-body">
<p class="form-hint" id="customModelPickHint"></p>
<div id="customModelPickList" class="run-mode-custom-models"></div>
</div>
</div>
</div>
<!-- Cron Jobs Modal -->
<div class="modal" id="cronModal">
<div class="modal-backdrop" onclick="app.closeCron()"></div>
+76 -5
View File
@@ -579,8 +579,9 @@ Object.assign(CodemanApp.prototype, {
const rows = [];
for (const host of hosts) {
const modelId = host.defaultModelId || (host.models || [])[0];
if (!modelId) continue; // nothing discovered yet — the settings panel explains why
const models = host.models || [];
if (models.length === 0) continue; // nothing discovered yet — the settings panel explains why
const modelId = host.defaultModelId || models[0];
for (const cli of capableClis) {
// escapeHtml(JSON.stringify(...)) on EVERY arg, not just the untrusted
// one: JSON.stringify's own double quotes would otherwise terminate this
@@ -589,11 +590,11 @@ Object.assign(CodemanApp.prototype, {
// modelId (server-controlled, from the endpoint's own /v1/models reply,
// not this box's) into markup instead of inert data. Same idiom as
// deleteCase's onclick a few hundred lines down.
const args = [cli.id, host.id, modelId].map((v) => escapeHtml(JSON.stringify(v))).join(', ');
const args = [cli.id, host.id].map((v) => escapeHtml(JSON.stringify(v))).join(', ');
rows.push(`
<button class="run-mode-option" data-mode="${escapeHtml(cli.id)}" data-endpoint="${escapeHtml(host.id)}"
onclick="app.runCustomModelEntry(${args})"
title="${escapeHtml(cli.label)} → ${escapeHtml(host.baseUrl)} (${escapeHtml(modelId)})">
onclick="app.selectCustomModelEntry(${args})"
title="${escapeHtml(cli.label)} → ${escapeHtml(host.baseUrl)} (${escapeHtml(modelId)}${models.length > 1 ? `, +${models.length - 1} more` : ''})">
<span class="run-mode-dot ${escapeHtml(cli.id)}"></span>${escapeHtml(cli.label)} (${escapeHtml(host.label)})
</button>`);
}
@@ -604,6 +605,76 @@ Object.assign(CodemanApp.prototype, {
container.innerHTML = rows.join('');
},
/**
* Decides whether picking a Run-menu Custom Endpoint entry can launch
* straight away or needs to ask which model first. Re-fetches the endpoint
* rather than trusting anything cached from the menu render: the models
* list (or the default) could have changed — a re-discovery cycle running
* every 5 minutes in the background, or an edit in the settings panel —
* between opening the dropdown and clicking a row.
*/
async selectCustomModelEntry(mode, endpointId) {
document.getElementById('runModeMenu')?.classList.remove('active');
const hosts = await this._apiJson('/api/model-endpoints');
const host = (hosts || []).find((h) => h.id === endpointId);
if (!host) {
this.showToast('That endpoint no longer exists', 'error');
return;
}
const models = host.models || [];
if (models.length === 0) {
this.showToast('No models discovered for this endpoint yet', 'warning');
return;
}
// Exactly one model: nothing to choose, so asking would just be an extra
// click for the same answer every time. Two or more: always ask, even
// with a defaultModelId set — the point of asking is letting THIS launch
// differ from the default, not just confirming it.
if (models.length === 1) {
return this.runCustomModelEntry(mode, endpointId, models[0]);
}
this._openCustomModelPickModal(mode, host);
},
/** Renders the "which model" picker for a (harness, endpoint) pair with more than one discovered model. */
_openCustomModelPickModal(mode, host) {
const modal = document.getElementById('customModelPickModal');
const list = document.getElementById('customModelPickList');
if (!modal || !list) return;
this._pendingCustomModelPick = { mode, endpointId: host.id };
const cliLabel = (window.__codemanCustomModelClis || []).find((c) => c.id === mode)?.label || mode;
// A static title (translatable by i18n.js's exact-string walker) plus a
// dynamic hint carrying the specifics — same split webviewModalTitle uses,
// since the walker cannot i18n a string a variable is already spliced into.
document.getElementById('customModelPickTitle').textContent = 'Choose a model';
document.getElementById('customModelPickHint').textContent =
`${cliLabel} → ${host.label} — ${(host.models || []).length} models discovered.`;
list.innerHTML = (host.models || [])
.map((m) => {
const isDefault = m === host.defaultModelId;
const arg = escapeHtml(JSON.stringify(m));
return `
<button class="run-mode-option" onclick="app.chooseCustomModelAndRun(${arg})">
<span class="run-mode-dot ${escapeHtml(mode)}"></span>${escapeHtml(m)}${isDefault ? ' <span class="set-scope">Default</span>' : ''}
</button>`;
})
.join('');
modal.classList.add('active');
},
closeCustomModelPickModal() {
document.getElementById('customModelPickModal')?.classList.remove('active');
this._pendingCustomModelPick = null;
},
/** A model row in the picker modal was clicked: close it and launch with that choice. */
chooseCustomModelAndRun(modelId) {
const pending = this._pendingCustomModelPick;
this.closeCustomModelPickModal();
if (!pending) return; // modal reopened/closed from elsewhere between render and click
void this.runCustomModelEntry(pending.mode, pending.endpointId, modelId);
},
/**
* Runs a session on `mode` and immediately applies `endpointId`/`modelId` to it
* via POST /api/sessions/:id/custom-model (see session-routes.ts) — the same
+46 -6
View File
@@ -107,6 +107,51 @@ function describeFetchError(err: unknown): string {
type RedactedHost = ReturnType<typeof redactApiKey>;
/**
* Merges a fresh `GET /v1/models` result into a host record: stamps
* `lastDiscoveredAt`, and drops `defaultModelId` if it no longer appears in
* the fresh list (it would otherwise leave the Run-menu picker applying a
* model id the endpoint just told us it doesn't serve). Pure — no IO, so the
* manual route (which reports a fetch failure's *reason* to the caller) and
* the periodic sweep below (which only cares whether it can move on) can
* each do their own `discoverModels()` + error handling around one shared
* "how to apply a successful result" step.
*/
function applyDiscoveredModels(host: CustomModelHost, models: string[]): CustomModelHost {
const defaultModelId = host.defaultModelId && models.includes(host.defaultModelId) ? host.defaultModelId : undefined;
return { ...host, models, defaultModelId, lastDiscoveredAt: new Date().toISOString() };
}
/**
* Re-discovers every saved endpoint's models, best-effort. One endpoint being
* unreachable (powered off, wrong network) must not stop the others from
* refreshing, and a read-modify-write per host (rather than one batch write
* at the end) means a crash or restart mid-sweep loses at most the endpoints
* not yet reached, never a write already applied. Exported so both the
* periodic timer (server.ts) and a test can drive it directly.
*/
export async function refreshAllCustomModelHosts(): Promise<void> {
const dataDir = getDataDir();
const hosts = await readCustomModelHosts(dataDir);
for (const host of hosts) {
if (isBlockedWebviewUrl(host.baseUrl)) continue;
let models: string[];
try {
models = await discoverModels(host);
} catch {
continue; // unreachable this cycle — try again next tick, not fatal to the sweep
}
// Re-read + splice by id rather than reusing the array captured above: an
// admin editing or deleting an endpoint via the API mid-sweep must win,
// not be silently overwritten by a refresh that started before their change.
const current = await readCustomModelHosts(dataDir);
const index = current.findIndex((item) => item.id === host.id);
if (index === -1) continue; // deleted mid-sweep
current[index] = applyDiscoveredModels(current[index], models);
await writeCustomModelHosts(dataDir, current);
}
}
export function registerCustomModelRoutes(app: FastifyInstance): void {
app.get('/api/model-endpoints', async (req): Promise<RedactedHost[]> => {
if (isMultiUserMode() && !isAdmin(req)) return [];
@@ -179,12 +224,7 @@ export function registerCustomModelRoutes(app: FastifyInstance): void {
try {
const models = await discoverModels(host);
const next = [...hosts];
// A default that no longer appears in the fresh list would leave the Run-menu
// picker applying a model id the endpoint just told us it doesn't serve; drop
// it rather than carry it forward silently invalid.
const defaultModelId =
host.defaultModelId && models.includes(host.defaultModelId) ? host.defaultModelId : undefined;
next[index] = { ...host, models, defaultModelId, lastDiscoveredAt: new Date().toISOString() };
next[index] = applyDiscoveredModels(host, models);
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next);
return { success: true, data: { models } };
} catch (err) {
+1 -1
View File
@@ -27,4 +27,4 @@ export { registerWsRoutes } from './ws-routes.js';
export { registerVoiceRoutes } from './voice-routes.js';
export { registerWebviewRoutes, tryWebviewRefererFallback } from './webview-routes.js';
export { registerTabLayoutRoutes } from './tab-layout-routes.js';
export { registerCustomModelRoutes } from './custom-model-routes.js';
export { registerCustomModelRoutes, refreshAllCustomModelHosts } from './custom-model-routes.js';
+21
View File
@@ -190,6 +190,7 @@ import {
registerWebviewRoutes,
registerTabLayoutRoutes,
registerCustomModelRoutes,
refreshAllCustomModelHosts,
tryWebviewRefererFallback,
} from './routes/index.js';
import { isLostWebviewFrameNavigation } from './webview-proxy.js';
@@ -202,6 +203,7 @@ const __dirname = dirname(fileURLToPath(import.meta.url));
// while capping growth of `sseClientsById` and blocking pathological inputs.
const SSE_CLIENT_ID_RE = /^[A-Za-z0-9_-]{8,64}$/;
const CODEX_USAGE_POLL_INTERVAL_MS = 5 * 60_000;
const CUSTOM_MODEL_REDISCOVER_INTERVAL_MS = 5 * 60_000;
function escapeHtmlText(value: string): string {
return value.replaceAll('&', '&amp;').replaceAll('<', '&lt;').replaceAll('>', '&gt;');
@@ -2734,6 +2736,25 @@ export class WebServer extends EventEmitter {
});
}
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): keeps
// each saved endpoint's discovered model list current with no manual
// "Discover" click, so a model added on the server side (or one that drops
// off) shows up in the Run-menu picker within one cycle. Best-effort per
// endpoint (refreshAllCustomModelHosts skips one that's unreachable rather
// than failing the sweep) and off in tests for the same reason the Codex
// poll above is — no real network to hit, no server instance to keep alive.
if (!this.testMode) {
this.cleanup.setInterval(
() => {
refreshAllCustomModelHosts().catch((err) => {
console.error('[custom-model] periodic re-discovery failed:', getErrorMessage(err));
});
},
CUSTOM_MODEL_REDISCOVER_INTERVAL_MS,
{ description: 'custom model endpoint re-discovery' }
);
}
// Start scheduled runs cleanup timer
this.cleanup.setInterval(
() => {
@@ -0,0 +1,126 @@
/**
* @fileoverview Tests for `refreshAllCustomModelHosts()`, the periodic
* background sweep behind server.ts's "custom model endpoint re-discovery"
* timer (docs/custom-model-endpoints-plan.md). Kept in its own file rather
* than folded into test/routes/custom-model-routes.test.ts: that file's data
* dir is shared across every test in it (one temp HOME per FILE, not per
* test — test/setup.ts), and a sweep that walks every saved host would pick
* up every host any other test in that file happened to create, making an
* exact call-count or exact-host assertion meaningless. A dedicated file
* gets its own clean temp HOME.
*
* Port: N/A (no server; drives readCustomModelHosts/writeCustomModelHosts
* directly plus the mocked webviewFetch dispatcher).
*/
import { describe, it, expect, vi, beforeEach } from 'vitest';
import { getDataDir } from '../src/config/instance.js';
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../src/custom-model-hosts.js';
import { refreshAllCustomModelHosts } from '../src/web/routes/custom-model-routes.js';
import { webviewFetch } from '../src/web/webview-egress.js';
vi.mock('../src/web/webview-egress.js', async () => {
const actual = await vi.importActual<typeof import('../src/web/webview-egress.js')>('../src/web/webview-egress.js');
return { ...actual, webviewFetch: vi.fn() };
});
const fetchMock = vi.mocked(webviewFetch);
function host(overrides: Partial<CustomModelHost> & Pick<CustomModelHost, 'id' | 'baseUrl'>): CustomModelHost {
return { label: overrides.id, ...overrides };
}
beforeEach(() => {
fetchMock.mockReset();
});
describe('refreshAllCustomModelHosts (the periodic re-discovery sweep)', () => {
it('refreshes every saved endpoint, best-effort — one unreachable host does not stop the others', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [
host({ id: 'ok', baseUrl: 'http://localhost:8080' }),
host({ id: 'down', baseUrl: 'http://localhost:8081' }),
]);
fetchMock.mockImplementation(async (url: URL) => {
if (url.href.includes('8081')) throw new TypeError('fetch failed', { cause: new Error('ECONNREFUSED') });
return new Response(JSON.stringify({ data: [{ id: 'qwen3' }] }), { status: 200 });
});
await refreshAllCustomModelHosts();
const hosts = await readCustomModelHosts(dir);
const ok = hosts.find((h) => h.id === 'ok');
const down = hosts.find((h) => h.id === 'down');
expect(ok?.models).toEqual(['qwen3']);
expect(ok?.lastDiscoveredAt).toBeTruthy();
expect(down?.models ?? []).toEqual([]);
expect(down?.lastDiscoveredAt).toBeFalsy();
});
it('skips a host whose baseUrl is blocked, without making a request', async () => {
const dir = getDataDir();
// Written directly rather than through the POST route, which already
// refuses this at save time — this simulates a record that pre-dates the
// guard, or was hand-edited on disk. The sweep must not trust it either.
await writeCustomModelHosts(dir, [host({ id: 'meta', baseUrl: 'http://169.254.169.254/' })]);
fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'x' }] }), { status: 200 }));
await refreshAllCustomModelHosts();
expect(fetchMock).not.toHaveBeenCalled();
});
it('drops a stale default and preserves lastDiscoveredAt semantics, same as manual discovery', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [
host({ id: 'ep', baseUrl: 'http://localhost:8080', models: ['qwen3'], defaultModelId: 'qwen3' }),
]);
fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'llama3' }] }), { status: 200 }));
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.models).toEqual(['llama3']);
expect(updated.defaultModelId).toBeUndefined();
expect(updated.lastDiscoveredAt).toBeTruthy();
});
it('keeps a default that is still present after the sweep', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [
host({ id: 'ep', baseUrl: 'http://localhost:8080', models: ['qwen3'], defaultModelId: 'qwen3' }),
]);
fetchMock.mockResolvedValue(
new Response(JSON.stringify({ data: [{ id: 'qwen3' }, { id: 'llama3' }] }), { status: 200 })
);
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.defaultModelId).toBe('qwen3');
});
it('does not resurrect an endpoint deleted while the sweep was in flight', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'deleted', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockImplementation(async () => {
// Simulate an admin deleting the endpoint between the sweep's fetch and
// its read-modify-write — the delete must win, not be overwritten by a
// refresh that started before it.
const current = await readCustomModelHosts(dir);
await writeCustomModelHosts(
dir,
current.filter((h) => h.id !== 'deleted')
);
return new Response(JSON.stringify({ data: [{ id: 'qwen3' }] }), { status: 200 });
});
await expect(refreshAllCustomModelHosts()).resolves.toBeUndefined();
const hosts = await readCustomModelHosts(dir);
expect(hosts.find((h) => h.id === 'deleted')).toBeUndefined();
});
it('leaves the store untouched when there are no saved endpoints at all', async () => {
await expect(refreshAllCustomModelHosts()).resolves.toBeUndefined();
expect(fetchMock).not.toHaveBeenCalled();
});
});
+125
View File
@@ -49,6 +49,11 @@ function bootApp(
<div id="runModeCustomModelHeader" style="display:none"></div>
<div id="runModeCustomModels"></div>
</div>
<div class="modal" id="customModelPickModal">
<h3 id="customModelPickTitle"></h3>
<p id="customModelPickHint"></p>
<div id="customModelPickList"></div>
</div>
</body>`,
{ url: 'http://localhost/', runScripts: 'dangerously' }
);
@@ -68,6 +73,10 @@ function bootApp(
app.loadAppSettingsFromStorage = () => ({ customModelEndpointsEnabled: options.settingsEnabled ?? true });
app.isCliAvailable = options.cliAvailable ?? (() => true);
app.showToast = () => {};
// Default no-op so a button's onclick (selectCustomModelEntry -> possibly
// straight to runCustomModelEntry for a single-model host) never rejects
// with "this.run is not a function"; tests of the launch itself override it.
app.run = async () => {};
// _apiJson unwraps the {success,data} envelope for real against a live
// server; here it stands in for that, driven from a fixed `hosts` fixture
// so these tests exercise the picker's OWN code, not the envelope helper.
@@ -172,6 +181,122 @@ describe('Custom Model Endpoint Profiles: Run-menu picker generation', () => {
});
});
describe('Custom Model Endpoint Profiles: the "which model" picker', () => {
it('launches straight away for a host with exactly one discovered model, no dialog', async () => {
const { win, app } = bootApp({
hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://localhost:8080', models: ['qwen3'] }],
});
let launched: unknown[] | null = null;
app.runCustomModelEntry = async (...args: unknown[]) => {
launched = args;
};
await app.selectCustomModelEntry('claude', 'llama-box');
expect(launched).toEqual(['claude', 'llama-box', 'qwen3']);
expect(win.document.getElementById('customModelPickModal')!.classList.contains('active')).toBe(false);
});
it('opens the picker for a host with more than one discovered model, rather than launching directly', async () => {
const { win, app } = bootApp({
hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://localhost:8080', models: ['qwen3', 'llama3'] }],
});
let launched = false;
app.runCustomModelEntry = async () => {
launched = true;
};
await app.selectCustomModelEntry('claude', 'llama-box');
expect(launched).toBe(false);
const modal = win.document.getElementById('customModelPickModal')!;
expect(modal.classList.contains('active')).toBe(true);
const list = win.document.getElementById('customModelPickList')!;
expect(list.querySelectorAll('button').length).toBe(2);
expect(list.textContent).toContain('qwen3');
expect(list.textContent).toContain('llama3');
});
it('always asks with 2+ models, even when a defaultModelId is set — the point is letting this launch differ', async () => {
const { win, app } = bootApp({
hosts: [
{
id: 'llama-box',
label: 'llama.cpp',
baseUrl: 'http://localhost:8080',
models: ['qwen3', 'llama3'],
defaultModelId: 'qwen3',
},
],
});
await app.selectCustomModelEntry('claude', 'llama-box');
const modal = win.document.getElementById('customModelPickModal')!;
expect(modal.classList.contains('active')).toBe(true);
// The default is marked, not auto-chosen.
expect(win.document.getElementById('customModelPickList')!.textContent).toContain('Default');
});
it('picking a row in the modal closes it and launches with that exact model', async () => {
const { win, app } = bootApp({
hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://localhost:8080', models: ['qwen3', 'llama3'] }],
});
let launched: unknown[] | null = null;
app.runCustomModelEntry = async (...args: unknown[]) => {
launched = args;
};
win.app = app;
await app.selectCustomModelEntry('claude', 'llama-box');
const buttons = win.document.getElementById('customModelPickList')!.querySelectorAll('button');
const llama3Btn = [...buttons].find((b) => b.textContent?.includes('llama3')) as unknown as HTMLButtonElement & {
onclick: (e: unknown) => void;
};
expect(typeof llama3Btn.onclick).toBe('function');
llama3Btn.onclick(new (win as any).Event('click'));
expect(launched).toEqual(['claude', 'llama-box', 'llama3']);
expect(win.document.getElementById('customModelPickModal')!.classList.contains('active')).toBe(false);
});
it('re-fetches the endpoint at click time rather than trusting anything cached from the menu render', async () => {
// The background re-discovery sweep (server-side, every 5 minutes) or a
// settings-panel edit can change the model list between opening the
// dropdown and clicking a row — the picker must reflect what is current.
let fetchCount = 0;
const { win, app } = bootApp({});
app._apiJson = async (path: string) => {
if (path !== '/api/model-endpoints') return null;
fetchCount += 1;
return [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://x', models: ['qwen3', 'llama3', 'phi4'] }];
};
await app.selectCustomModelEntry('claude', 'llama-box');
expect(fetchCount).toBe(1);
expect(win.document.getElementById('customModelPickList')!.querySelectorAll('button').length).toBe(3);
});
it('toasts and does nothing when the endpoint has vanished by click time', async () => {
const { app } = bootApp({ hosts: [] });
let toastMessage: string | null = null;
app.showToast = (msg: string) => {
toastMessage = msg;
};
await app.selectCustomModelEntry('claude', 'ghost-endpoint');
expect(toastMessage).toMatch(/no longer exists/i);
});
it('toasts and does nothing when the endpoint has zero discovered models by click time', async () => {
const { app } = bootApp({
hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://x', models: [] }],
});
let toastMessage: string | null = null;
app.showToast = (msg: string) => {
toastMessage = msg;
};
await app.selectCustomModelEntry('claude', 'llama-box');
expect(toastMessage).toMatch(/no models discovered/i);
});
});
describe('Custom Model Endpoint Profiles: applying a picked entry', () => {
it('does not apply the endpoint to a session that was already open when the launch fails', async () => {
const { app } = bootApp({});