mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-03 05:59:43 +02:00
feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real llama.cpp server: 1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to apply the endpoint's defaultModelId (or the first discovered model) silently. Now, via the new selectCustomModelEntry() (session-ui.js): - exactly one discovered model launches straight away, same as before - two or more open a new #customModelPickModal listing every discovered model; defaultModelId (if set) is marked but never auto-chosen, since the point of asking is letting ONE launch deliberately differ from the saved default, not just confirming it The endpoint is re-fetched at click time rather than trusting anything cached from the dropdown's own render, since the model list can have changed (the sweep below, or a settings-panel edit) since it opened. runCustomModelEntry() itself — the actual launch, routed through run() for the in-flight lock, snapshot-guarded against applying to the wrong session — is unchanged; it now just always receives an explicit model id from one of these two paths instead of computing one itself. 2. Periodic re-discovery. Every saved endpoint's models now refresh automatically every 5 minutes in the background (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way as the Codex plan-usage poll it sits beside — this.cleanup.setInterval, off under testMode), so a model the server starts or stops serving shows up without another manual "Discover" click. The manual POST .../discover-models route and the new refreshAllCustomModelHosts() sweep (custom-model-routes.ts) now share one pure merge step (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId that no longer appears) rather than two copies that could drift. The sweep is best-effort per host — one endpoint being unreachable on a cycle never blocks the others — and re-reads the store before each host's write, keyed by id, so a concurrent edit or delete from the settings panel always wins over a sweep that started before it. Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated file for the sweep (kept separate from custom-model-routes.test.ts because that file's data dir is shared across every test in it — one temp HOME per FILE, not per test — which would make a sweep-touches-every-host assertion meaningless there). test/custom-model-run-menu-ui.test.ts gained a new describe block driving the real picker modal through JSDOM: single-model bypass, multi-model dialog with the default marked-not-chosen, picking a row closes the modal and launches with that exact model, the endpoint re-fetch, and the two "vanished by click time" toast paths. Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md, docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated — the last of these also caught up two sentences that had gone stale after the draft-review fixes landed (the picker routes through run() now, not a raw run*() call). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
60e1bd52f7
commit
5a9ff07f57
@@ -551,7 +551,11 @@ user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md).
|
||||
`lastDiscoveredAt`. A `defaultModelId` that no longer appears in the fresh
|
||||
list is dropped rather than carried forward invalid. Failures answer
|
||||
`502 OPERATION_FAILED` with the underlying connection error, or a named
|
||||
egress refusal if the resolved address turned out to be blocked.
|
||||
egress refusal if the resolved address turned out to be blocked. The same
|
||||
refresh also runs automatically for every saved endpoint every 5 minutes
|
||||
in the background (`refreshAllCustomModelHosts()`, `custom-model-routes.ts`,
|
||||
started from `server.ts`), so there is no route for triggering "refresh
|
||||
all" — one endpoint being unreachable on a cycle never blocks the others.
|
||||
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId } |
|
||||
{ clear: true }` applies (or clears) the session's selection and
|
||||
**restarts the session's CLI process in place** — every supported harness
|
||||
|
||||
@@ -66,18 +66,26 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
|
||||
Endpoint management is admin-only in multi-user mode, same as remote/docker
|
||||
hosts — these are machine-level infra, not per-user settings.
|
||||
|
||||
`defaultModelId` names which discovered model the Run-menu picker applies
|
||||
for that endpoint with no further choice — the settings panel's Edit form
|
||||
exposes it as a select populated from the endpoint's own discovered
|
||||
`models`, and the route refuses a value that isn't one of them. Leaving it
|
||||
unset falls back to the first discovered model; re-discovering drops a
|
||||
default that no longer appears in the fresh list rather than carrying an
|
||||
invalid one forward.
|
||||
`defaultModelId` names which discovered model the picker pre-marks for that
|
||||
endpoint — the settings panel's Edit form exposes it as a select populated
|
||||
from the endpoint's own discovered `models`, and the route refuses a value
|
||||
that isn't one of them. It is applied automatically only when the endpoint
|
||||
has exactly one discovered model (nothing to choose); with two or more it
|
||||
is a pre-selection in the model-picker dialog below, never a silent default.
|
||||
Re-discovering drops a default that no longer appears in the fresh list
|
||||
rather than carrying an invalid one forward.
|
||||
|
||||
**Model lists refresh themselves.** A background sweep (`server.ts`,
|
||||
`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every
|
||||
saved endpoint the same way the manual `POST .../discover-models` route
|
||||
does, best-effort per endpoint — one being unreachable on a given cycle
|
||||
never blocks the others. Off under `npm test`, same reasoning as the Codex
|
||||
plan-usage poll it sits beside: no real network to hit, no server instance
|
||||
to keep the timer alive for.
|
||||
|
||||
## The Run-menu picker
|
||||
|
||||
With the setting on and at least one endpoint carrying a usable default
|
||||
model (either an explicit `defaultModelId` or just one discovered model),
|
||||
With the setting on and at least one endpoint carrying a discovered model,
|
||||
the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry
|
||||
per (harness that can redirect to a custom endpoint, saved endpoint) pair,
|
||||
e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI
|
||||
@@ -86,13 +94,22 @@ registry's own `capabilities.customModelInjection` at page render
|
||||
in the frontend — so a CLI whose injection recipe lands later shows up with
|
||||
no frontend change, and Antigravity (`unsupported`) never does.
|
||||
|
||||
Picking an entry runs a single session on that harness exactly the way its
|
||||
Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`,
|
||||
`session-ui.js`) rather than trusting anything cached from the dropdown's
|
||||
own render — the model list can have changed via the 5-minute sweep above
|
||||
or a settings-panel edit since the menu opened. With exactly one discovered
|
||||
model it runs straight away; with two or more, a small modal
|
||||
(`#customModelPickModal`) lists them and asks which one to use for this
|
||||
launch, with the endpoint's `defaultModelId` marked but not auto-chosen —
|
||||
the point of asking is letting one launch deliberately differ from the
|
||||
saved default, not just confirming it. Whichever way the model was decided,
|
||||
the launch itself runs a single session on that harness exactly the way its
|
||||
own Run-menu entry would (same case creation, env overrides, everything),
|
||||
then immediately applies the endpoint's default model to it via the route
|
||||
below. It is a one-off "try this endpoint" action, not a sticky mode: the
|
||||
plain Run button still means "this harness, native cloud" afterward.
|
||||
Entries are hidden entirely for a remote or Docker active case, since the
|
||||
apply route refuses both (see the next section).
|
||||
then immediately applies the endpoint and model to it via the route below.
|
||||
It is a one-off "try this endpoint" action, not a sticky mode: the plain
|
||||
Run button still means "this harness, native cloud" afterward. Entries are
|
||||
hidden entirely for a remote or Docker active case, since the apply route
|
||||
refuses both (see the next section).
|
||||
|
||||
## Applying a model to a session
|
||||
|
||||
|
||||
@@ -16,20 +16,32 @@ Still in App Settings → Models → Custom model endpoints:
|
||||
say). An API key is optional; most local servers don't check one.
|
||||
2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it.
|
||||
3. Pick a **default model** from what was discovered. This is the model the Run-menu entry
|
||||
below applies with no further choice, so set it once you know which one you want.
|
||||
applies directly when only one model is discovered; with two or more, it's just the one
|
||||
pre-marked in the picker dialog described below, not a silent default.
|
||||
|
||||
Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker
|
||||
hosts — these are machine-level infra, not a per-user setting.
|
||||
|
||||
**Model lists refresh themselves.** Every saved endpoint is re-discovered automatically every
|
||||
5 minutes in the background, so a model the server starts serving later — or stops serving —
|
||||
shows up without another manual click of **Discover**. One endpoint being unreachable on a
|
||||
given cycle (powered off, wrong network) never blocks the others from refreshing.
|
||||
|
||||
## Running a session against one
|
||||
|
||||
With the setting on and at least one endpoint carrying a usable default model, the **Run**
|
||||
With the setting on and at least one endpoint carrying a discovered model, the **Run**
|
||||
dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a
|
||||
custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a
|
||||
session on that harness exactly the way its own entry would, then points it at the
|
||||
endpoint's default model. It is a one-off "try this endpoint" action, not a sticky mode — the
|
||||
plain **Run** button still means "this harness, native cloud" afterward, and a fresh session
|
||||
never inherits whatever the last one was pointed at.
|
||||
session on that harness exactly the way its own entry would. It is a one-off "try this
|
||||
endpoint" action, not a sticky mode — the plain **Run** button still means "this harness,
|
||||
native cloud" afterward, and a fresh session never inherits whatever the last one was
|
||||
pointed at.
|
||||
|
||||
**Which model it uses depends on how many the endpoint has discovered.** With exactly one,
|
||||
the session launches straight away on that model — nothing to choose. With two or more, a
|
||||
small dialog asks which one to use for this launch before starting the session; the
|
||||
endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately
|
||||
use a different one without changing the saved default.
|
||||
|
||||
Applying a selection **restarts the harness's process in place** — same tab, same
|
||||
conversation where the harness supports resuming one, fresh environment. That restart is
|
||||
|
||||
@@ -311,6 +311,9 @@
|
||||
'What the Run-menu picker applies for this endpoint. Discover models first.':
|
||||
'运行菜单选择器会为此端点应用该模型。请先发现可用模型。',
|
||||
'Custom Endpoints': '自定义端点',
|
||||
'Choose a model': '选择模型',
|
||||
'That endpoint no longer exists': '该端点已不存在',
|
||||
'No models discovered for this endpoint yet': '此端点尚未发现任何模型',
|
||||
'Subagent Options': '子智能体选项',
|
||||
'Enable Tracking': '启用跟踪',
|
||||
'Active Tab Only': '仅活动标签页',
|
||||
|
||||
@@ -910,6 +910,23 @@
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<!-- Custom Model Endpoint Profiles: "which model" picker (docs/custom-model-endpoints-plan.md).
|
||||
Shown only when the chosen endpoint has more than one discovered model — see
|
||||
selectCustomModelEntry() in session-ui.js, which skips straight to launch otherwise. -->
|
||||
<div class="modal" id="customModelPickModal">
|
||||
<div class="modal-backdrop" onclick="app.closeCustomModelPickModal()"></div>
|
||||
<div class="modal-content" style="max-width: 380px;">
|
||||
<div class="modal-header">
|
||||
<h3 id="customModelPickTitle">Choose a model</h3>
|
||||
<button class="modal-close" onclick="app.closeCustomModelPickModal()" aria-label="Close model picker">×</button>
|
||||
</div>
|
||||
<div class="modal-body">
|
||||
<p class="form-hint" id="customModelPickHint"></p>
|
||||
<div id="customModelPickList" class="run-mode-custom-models"></div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<!-- Cron Jobs Modal -->
|
||||
<div class="modal" id="cronModal">
|
||||
<div class="modal-backdrop" onclick="app.closeCron()"></div>
|
||||
|
||||
@@ -579,8 +579,9 @@ Object.assign(CodemanApp.prototype, {
|
||||
|
||||
const rows = [];
|
||||
for (const host of hosts) {
|
||||
const modelId = host.defaultModelId || (host.models || [])[0];
|
||||
if (!modelId) continue; // nothing discovered yet — the settings panel explains why
|
||||
const models = host.models || [];
|
||||
if (models.length === 0) continue; // nothing discovered yet — the settings panel explains why
|
||||
const modelId = host.defaultModelId || models[0];
|
||||
for (const cli of capableClis) {
|
||||
// escapeHtml(JSON.stringify(...)) on EVERY arg, not just the untrusted
|
||||
// one: JSON.stringify's own double quotes would otherwise terminate this
|
||||
@@ -589,11 +590,11 @@ Object.assign(CodemanApp.prototype, {
|
||||
// modelId (server-controlled, from the endpoint's own /v1/models reply,
|
||||
// not this box's) into markup instead of inert data. Same idiom as
|
||||
// deleteCase's onclick a few hundred lines down.
|
||||
const args = [cli.id, host.id, modelId].map((v) => escapeHtml(JSON.stringify(v))).join(', ');
|
||||
const args = [cli.id, host.id].map((v) => escapeHtml(JSON.stringify(v))).join(', ');
|
||||
rows.push(`
|
||||
<button class="run-mode-option" data-mode="${escapeHtml(cli.id)}" data-endpoint="${escapeHtml(host.id)}"
|
||||
onclick="app.runCustomModelEntry(${args})"
|
||||
title="${escapeHtml(cli.label)} → ${escapeHtml(host.baseUrl)} (${escapeHtml(modelId)})">
|
||||
onclick="app.selectCustomModelEntry(${args})"
|
||||
title="${escapeHtml(cli.label)} → ${escapeHtml(host.baseUrl)} (${escapeHtml(modelId)}${models.length > 1 ? `, +${models.length - 1} more` : ''})">
|
||||
<span class="run-mode-dot ${escapeHtml(cli.id)}"></span>${escapeHtml(cli.label)} (${escapeHtml(host.label)})
|
||||
</button>`);
|
||||
}
|
||||
@@ -604,6 +605,76 @@ Object.assign(CodemanApp.prototype, {
|
||||
container.innerHTML = rows.join('');
|
||||
},
|
||||
|
||||
/**
|
||||
* Decides whether picking a Run-menu Custom Endpoint entry can launch
|
||||
* straight away or needs to ask which model first. Re-fetches the endpoint
|
||||
* rather than trusting anything cached from the menu render: the models
|
||||
* list (or the default) could have changed — a re-discovery cycle running
|
||||
* every 5 minutes in the background, or an edit in the settings panel —
|
||||
* between opening the dropdown and clicking a row.
|
||||
*/
|
||||
async selectCustomModelEntry(mode, endpointId) {
|
||||
document.getElementById('runModeMenu')?.classList.remove('active');
|
||||
const hosts = await this._apiJson('/api/model-endpoints');
|
||||
const host = (hosts || []).find((h) => h.id === endpointId);
|
||||
if (!host) {
|
||||
this.showToast('That endpoint no longer exists', 'error');
|
||||
return;
|
||||
}
|
||||
const models = host.models || [];
|
||||
if (models.length === 0) {
|
||||
this.showToast('No models discovered for this endpoint yet', 'warning');
|
||||
return;
|
||||
}
|
||||
// Exactly one model: nothing to choose, so asking would just be an extra
|
||||
// click for the same answer every time. Two or more: always ask, even
|
||||
// with a defaultModelId set — the point of asking is letting THIS launch
|
||||
// differ from the default, not just confirming it.
|
||||
if (models.length === 1) {
|
||||
return this.runCustomModelEntry(mode, endpointId, models[0]);
|
||||
}
|
||||
this._openCustomModelPickModal(mode, host);
|
||||
},
|
||||
|
||||
/** Renders the "which model" picker for a (harness, endpoint) pair with more than one discovered model. */
|
||||
_openCustomModelPickModal(mode, host) {
|
||||
const modal = document.getElementById('customModelPickModal');
|
||||
const list = document.getElementById('customModelPickList');
|
||||
if (!modal || !list) return;
|
||||
this._pendingCustomModelPick = { mode, endpointId: host.id };
|
||||
const cliLabel = (window.__codemanCustomModelClis || []).find((c) => c.id === mode)?.label || mode;
|
||||
// A static title (translatable by i18n.js's exact-string walker) plus a
|
||||
// dynamic hint carrying the specifics — same split webviewModalTitle uses,
|
||||
// since the walker cannot i18n a string a variable is already spliced into.
|
||||
document.getElementById('customModelPickTitle').textContent = 'Choose a model';
|
||||
document.getElementById('customModelPickHint').textContent =
|
||||
`${cliLabel} → ${host.label} — ${(host.models || []).length} models discovered.`;
|
||||
list.innerHTML = (host.models || [])
|
||||
.map((m) => {
|
||||
const isDefault = m === host.defaultModelId;
|
||||
const arg = escapeHtml(JSON.stringify(m));
|
||||
return `
|
||||
<button class="run-mode-option" onclick="app.chooseCustomModelAndRun(${arg})">
|
||||
<span class="run-mode-dot ${escapeHtml(mode)}"></span>${escapeHtml(m)}${isDefault ? ' <span class="set-scope">Default</span>' : ''}
|
||||
</button>`;
|
||||
})
|
||||
.join('');
|
||||
modal.classList.add('active');
|
||||
},
|
||||
|
||||
closeCustomModelPickModal() {
|
||||
document.getElementById('customModelPickModal')?.classList.remove('active');
|
||||
this._pendingCustomModelPick = null;
|
||||
},
|
||||
|
||||
/** A model row in the picker modal was clicked: close it and launch with that choice. */
|
||||
chooseCustomModelAndRun(modelId) {
|
||||
const pending = this._pendingCustomModelPick;
|
||||
this.closeCustomModelPickModal();
|
||||
if (!pending) return; // modal reopened/closed from elsewhere between render and click
|
||||
void this.runCustomModelEntry(pending.mode, pending.endpointId, modelId);
|
||||
},
|
||||
|
||||
/**
|
||||
* Runs a session on `mode` and immediately applies `endpointId`/`modelId` to it
|
||||
* via POST /api/sessions/:id/custom-model (see session-routes.ts) — the same
|
||||
|
||||
@@ -107,6 +107,51 @@ function describeFetchError(err: unknown): string {
|
||||
|
||||
type RedactedHost = ReturnType<typeof redactApiKey>;
|
||||
|
||||
/**
|
||||
* Merges a fresh `GET /v1/models` result into a host record: stamps
|
||||
* `lastDiscoveredAt`, and drops `defaultModelId` if it no longer appears in
|
||||
* the fresh list (it would otherwise leave the Run-menu picker applying a
|
||||
* model id the endpoint just told us it doesn't serve). Pure — no IO, so the
|
||||
* manual route (which reports a fetch failure's *reason* to the caller) and
|
||||
* the periodic sweep below (which only cares whether it can move on) can
|
||||
* each do their own `discoverModels()` + error handling around one shared
|
||||
* "how to apply a successful result" step.
|
||||
*/
|
||||
function applyDiscoveredModels(host: CustomModelHost, models: string[]): CustomModelHost {
|
||||
const defaultModelId = host.defaultModelId && models.includes(host.defaultModelId) ? host.defaultModelId : undefined;
|
||||
return { ...host, models, defaultModelId, lastDiscoveredAt: new Date().toISOString() };
|
||||
}
|
||||
|
||||
/**
|
||||
* Re-discovers every saved endpoint's models, best-effort. One endpoint being
|
||||
* unreachable (powered off, wrong network) must not stop the others from
|
||||
* refreshing, and a read-modify-write per host (rather than one batch write
|
||||
* at the end) means a crash or restart mid-sweep loses at most the endpoints
|
||||
* not yet reached, never a write already applied. Exported so both the
|
||||
* periodic timer (server.ts) and a test can drive it directly.
|
||||
*/
|
||||
export async function refreshAllCustomModelHosts(): Promise<void> {
|
||||
const dataDir = getDataDir();
|
||||
const hosts = await readCustomModelHosts(dataDir);
|
||||
for (const host of hosts) {
|
||||
if (isBlockedWebviewUrl(host.baseUrl)) continue;
|
||||
let models: string[];
|
||||
try {
|
||||
models = await discoverModels(host);
|
||||
} catch {
|
||||
continue; // unreachable this cycle — try again next tick, not fatal to the sweep
|
||||
}
|
||||
// Re-read + splice by id rather than reusing the array captured above: an
|
||||
// admin editing or deleting an endpoint via the API mid-sweep must win,
|
||||
// not be silently overwritten by a refresh that started before their change.
|
||||
const current = await readCustomModelHosts(dataDir);
|
||||
const index = current.findIndex((item) => item.id === host.id);
|
||||
if (index === -1) continue; // deleted mid-sweep
|
||||
current[index] = applyDiscoveredModels(current[index], models);
|
||||
await writeCustomModelHosts(dataDir, current);
|
||||
}
|
||||
}
|
||||
|
||||
export function registerCustomModelRoutes(app: FastifyInstance): void {
|
||||
app.get('/api/model-endpoints', async (req): Promise<RedactedHost[]> => {
|
||||
if (isMultiUserMode() && !isAdmin(req)) return [];
|
||||
@@ -179,12 +224,7 @@ export function registerCustomModelRoutes(app: FastifyInstance): void {
|
||||
try {
|
||||
const models = await discoverModels(host);
|
||||
const next = [...hosts];
|
||||
// A default that no longer appears in the fresh list would leave the Run-menu
|
||||
// picker applying a model id the endpoint just told us it doesn't serve; drop
|
||||
// it rather than carry it forward silently invalid.
|
||||
const defaultModelId =
|
||||
host.defaultModelId && models.includes(host.defaultModelId) ? host.defaultModelId : undefined;
|
||||
next[index] = { ...host, models, defaultModelId, lastDiscoveredAt: new Date().toISOString() };
|
||||
next[index] = applyDiscoveredModels(host, models);
|
||||
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next);
|
||||
return { success: true, data: { models } };
|
||||
} catch (err) {
|
||||
|
||||
@@ -27,4 +27,4 @@ export { registerWsRoutes } from './ws-routes.js';
|
||||
export { registerVoiceRoutes } from './voice-routes.js';
|
||||
export { registerWebviewRoutes, tryWebviewRefererFallback } from './webview-routes.js';
|
||||
export { registerTabLayoutRoutes } from './tab-layout-routes.js';
|
||||
export { registerCustomModelRoutes } from './custom-model-routes.js';
|
||||
export { registerCustomModelRoutes, refreshAllCustomModelHosts } from './custom-model-routes.js';
|
||||
|
||||
@@ -190,6 +190,7 @@ import {
|
||||
registerWebviewRoutes,
|
||||
registerTabLayoutRoutes,
|
||||
registerCustomModelRoutes,
|
||||
refreshAllCustomModelHosts,
|
||||
tryWebviewRefererFallback,
|
||||
} from './routes/index.js';
|
||||
import { isLostWebviewFrameNavigation } from './webview-proxy.js';
|
||||
@@ -202,6 +203,7 @@ const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||
// while capping growth of `sseClientsById` and blocking pathological inputs.
|
||||
const SSE_CLIENT_ID_RE = /^[A-Za-z0-9_-]{8,64}$/;
|
||||
const CODEX_USAGE_POLL_INTERVAL_MS = 5 * 60_000;
|
||||
const CUSTOM_MODEL_REDISCOVER_INTERVAL_MS = 5 * 60_000;
|
||||
|
||||
function escapeHtmlText(value: string): string {
|
||||
return value.replaceAll('&', '&').replaceAll('<', '<').replaceAll('>', '>');
|
||||
@@ -2734,6 +2736,25 @@ export class WebServer extends EventEmitter {
|
||||
});
|
||||
}
|
||||
|
||||
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): keeps
|
||||
// each saved endpoint's discovered model list current with no manual
|
||||
// "Discover" click, so a model added on the server side (or one that drops
|
||||
// off) shows up in the Run-menu picker within one cycle. Best-effort per
|
||||
// endpoint (refreshAllCustomModelHosts skips one that's unreachable rather
|
||||
// than failing the sweep) and off in tests for the same reason the Codex
|
||||
// poll above is — no real network to hit, no server instance to keep alive.
|
||||
if (!this.testMode) {
|
||||
this.cleanup.setInterval(
|
||||
() => {
|
||||
refreshAllCustomModelHosts().catch((err) => {
|
||||
console.error('[custom-model] periodic re-discovery failed:', getErrorMessage(err));
|
||||
});
|
||||
},
|
||||
CUSTOM_MODEL_REDISCOVER_INTERVAL_MS,
|
||||
{ description: 'custom model endpoint re-discovery' }
|
||||
);
|
||||
}
|
||||
|
||||
// Start scheduled runs cleanup timer
|
||||
this.cleanup.setInterval(
|
||||
() => {
|
||||
|
||||
@@ -0,0 +1,126 @@
|
||||
/**
|
||||
* @fileoverview Tests for `refreshAllCustomModelHosts()`, the periodic
|
||||
* background sweep behind server.ts's "custom model endpoint re-discovery"
|
||||
* timer (docs/custom-model-endpoints-plan.md). Kept in its own file rather
|
||||
* than folded into test/routes/custom-model-routes.test.ts: that file's data
|
||||
* dir is shared across every test in it (one temp HOME per FILE, not per
|
||||
* test — test/setup.ts), and a sweep that walks every saved host would pick
|
||||
* up every host any other test in that file happened to create, making an
|
||||
* exact call-count or exact-host assertion meaningless. A dedicated file
|
||||
* gets its own clean temp HOME.
|
||||
*
|
||||
* Port: N/A (no server; drives readCustomModelHosts/writeCustomModelHosts
|
||||
* directly plus the mocked webviewFetch dispatcher).
|
||||
*/
|
||||
import { describe, it, expect, vi, beforeEach } from 'vitest';
|
||||
import { getDataDir } from '../src/config/instance.js';
|
||||
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../src/custom-model-hosts.js';
|
||||
import { refreshAllCustomModelHosts } from '../src/web/routes/custom-model-routes.js';
|
||||
import { webviewFetch } from '../src/web/webview-egress.js';
|
||||
|
||||
vi.mock('../src/web/webview-egress.js', async () => {
|
||||
const actual = await vi.importActual<typeof import('../src/web/webview-egress.js')>('../src/web/webview-egress.js');
|
||||
return { ...actual, webviewFetch: vi.fn() };
|
||||
});
|
||||
|
||||
const fetchMock = vi.mocked(webviewFetch);
|
||||
|
||||
function host(overrides: Partial<CustomModelHost> & Pick<CustomModelHost, 'id' | 'baseUrl'>): CustomModelHost {
|
||||
return { label: overrides.id, ...overrides };
|
||||
}
|
||||
|
||||
beforeEach(() => {
|
||||
fetchMock.mockReset();
|
||||
});
|
||||
|
||||
describe('refreshAllCustomModelHosts (the periodic re-discovery sweep)', () => {
|
||||
it('refreshes every saved endpoint, best-effort — one unreachable host does not stop the others', async () => {
|
||||
const dir = getDataDir();
|
||||
await writeCustomModelHosts(dir, [
|
||||
host({ id: 'ok', baseUrl: 'http://localhost:8080' }),
|
||||
host({ id: 'down', baseUrl: 'http://localhost:8081' }),
|
||||
]);
|
||||
fetchMock.mockImplementation(async (url: URL) => {
|
||||
if (url.href.includes('8081')) throw new TypeError('fetch failed', { cause: new Error('ECONNREFUSED') });
|
||||
return new Response(JSON.stringify({ data: [{ id: 'qwen3' }] }), { status: 200 });
|
||||
});
|
||||
|
||||
await refreshAllCustomModelHosts();
|
||||
|
||||
const hosts = await readCustomModelHosts(dir);
|
||||
const ok = hosts.find((h) => h.id === 'ok');
|
||||
const down = hosts.find((h) => h.id === 'down');
|
||||
expect(ok?.models).toEqual(['qwen3']);
|
||||
expect(ok?.lastDiscoveredAt).toBeTruthy();
|
||||
expect(down?.models ?? []).toEqual([]);
|
||||
expect(down?.lastDiscoveredAt).toBeFalsy();
|
||||
});
|
||||
|
||||
it('skips a host whose baseUrl is blocked, without making a request', async () => {
|
||||
const dir = getDataDir();
|
||||
// Written directly rather than through the POST route, which already
|
||||
// refuses this at save time — this simulates a record that pre-dates the
|
||||
// guard, or was hand-edited on disk. The sweep must not trust it either.
|
||||
await writeCustomModelHosts(dir, [host({ id: 'meta', baseUrl: 'http://169.254.169.254/' })]);
|
||||
|
||||
fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'x' }] }), { status: 200 }));
|
||||
await refreshAllCustomModelHosts();
|
||||
expect(fetchMock).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('drops a stale default and preserves lastDiscoveredAt semantics, same as manual discovery', async () => {
|
||||
const dir = getDataDir();
|
||||
await writeCustomModelHosts(dir, [
|
||||
host({ id: 'ep', baseUrl: 'http://localhost:8080', models: ['qwen3'], defaultModelId: 'qwen3' }),
|
||||
]);
|
||||
fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'llama3' }] }), { status: 200 }));
|
||||
|
||||
await refreshAllCustomModelHosts();
|
||||
|
||||
const [updated] = await readCustomModelHosts(dir);
|
||||
expect(updated.models).toEqual(['llama3']);
|
||||
expect(updated.defaultModelId).toBeUndefined();
|
||||
expect(updated.lastDiscoveredAt).toBeTruthy();
|
||||
});
|
||||
|
||||
it('keeps a default that is still present after the sweep', async () => {
|
||||
const dir = getDataDir();
|
||||
await writeCustomModelHosts(dir, [
|
||||
host({ id: 'ep', baseUrl: 'http://localhost:8080', models: ['qwen3'], defaultModelId: 'qwen3' }),
|
||||
]);
|
||||
fetchMock.mockResolvedValue(
|
||||
new Response(JSON.stringify({ data: [{ id: 'qwen3' }, { id: 'llama3' }] }), { status: 200 })
|
||||
);
|
||||
|
||||
await refreshAllCustomModelHosts();
|
||||
|
||||
const [updated] = await readCustomModelHosts(dir);
|
||||
expect(updated.defaultModelId).toBe('qwen3');
|
||||
});
|
||||
|
||||
it('does not resurrect an endpoint deleted while the sweep was in flight', async () => {
|
||||
const dir = getDataDir();
|
||||
await writeCustomModelHosts(dir, [host({ id: 'deleted', baseUrl: 'http://localhost:8080' })]);
|
||||
|
||||
fetchMock.mockImplementation(async () => {
|
||||
// Simulate an admin deleting the endpoint between the sweep's fetch and
|
||||
// its read-modify-write — the delete must win, not be overwritten by a
|
||||
// refresh that started before it.
|
||||
const current = await readCustomModelHosts(dir);
|
||||
await writeCustomModelHosts(
|
||||
dir,
|
||||
current.filter((h) => h.id !== 'deleted')
|
||||
);
|
||||
return new Response(JSON.stringify({ data: [{ id: 'qwen3' }] }), { status: 200 });
|
||||
});
|
||||
|
||||
await expect(refreshAllCustomModelHosts()).resolves.toBeUndefined();
|
||||
const hosts = await readCustomModelHosts(dir);
|
||||
expect(hosts.find((h) => h.id === 'deleted')).toBeUndefined();
|
||||
});
|
||||
|
||||
it('leaves the store untouched when there are no saved endpoints at all', async () => {
|
||||
await expect(refreshAllCustomModelHosts()).resolves.toBeUndefined();
|
||||
expect(fetchMock).not.toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
@@ -49,6 +49,11 @@ function bootApp(
|
||||
<div id="runModeCustomModelHeader" style="display:none"></div>
|
||||
<div id="runModeCustomModels"></div>
|
||||
</div>
|
||||
<div class="modal" id="customModelPickModal">
|
||||
<h3 id="customModelPickTitle"></h3>
|
||||
<p id="customModelPickHint"></p>
|
||||
<div id="customModelPickList"></div>
|
||||
</div>
|
||||
</body>`,
|
||||
{ url: 'http://localhost/', runScripts: 'dangerously' }
|
||||
);
|
||||
@@ -68,6 +73,10 @@ function bootApp(
|
||||
app.loadAppSettingsFromStorage = () => ({ customModelEndpointsEnabled: options.settingsEnabled ?? true });
|
||||
app.isCliAvailable = options.cliAvailable ?? (() => true);
|
||||
app.showToast = () => {};
|
||||
// Default no-op so a button's onclick (selectCustomModelEntry -> possibly
|
||||
// straight to runCustomModelEntry for a single-model host) never rejects
|
||||
// with "this.run is not a function"; tests of the launch itself override it.
|
||||
app.run = async () => {};
|
||||
// _apiJson unwraps the {success,data} envelope for real against a live
|
||||
// server; here it stands in for that, driven from a fixed `hosts` fixture
|
||||
// so these tests exercise the picker's OWN code, not the envelope helper.
|
||||
@@ -172,6 +181,122 @@ describe('Custom Model Endpoint Profiles: Run-menu picker generation', () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe('Custom Model Endpoint Profiles: the "which model" picker', () => {
|
||||
it('launches straight away for a host with exactly one discovered model, no dialog', async () => {
|
||||
const { win, app } = bootApp({
|
||||
hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://localhost:8080', models: ['qwen3'] }],
|
||||
});
|
||||
let launched: unknown[] | null = null;
|
||||
app.runCustomModelEntry = async (...args: unknown[]) => {
|
||||
launched = args;
|
||||
};
|
||||
|
||||
await app.selectCustomModelEntry('claude', 'llama-box');
|
||||
|
||||
expect(launched).toEqual(['claude', 'llama-box', 'qwen3']);
|
||||
expect(win.document.getElementById('customModelPickModal')!.classList.contains('active')).toBe(false);
|
||||
});
|
||||
|
||||
it('opens the picker for a host with more than one discovered model, rather than launching directly', async () => {
|
||||
const { win, app } = bootApp({
|
||||
hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://localhost:8080', models: ['qwen3', 'llama3'] }],
|
||||
});
|
||||
let launched = false;
|
||||
app.runCustomModelEntry = async () => {
|
||||
launched = true;
|
||||
};
|
||||
|
||||
await app.selectCustomModelEntry('claude', 'llama-box');
|
||||
|
||||
expect(launched).toBe(false);
|
||||
const modal = win.document.getElementById('customModelPickModal')!;
|
||||
expect(modal.classList.contains('active')).toBe(true);
|
||||
const list = win.document.getElementById('customModelPickList')!;
|
||||
expect(list.querySelectorAll('button').length).toBe(2);
|
||||
expect(list.textContent).toContain('qwen3');
|
||||
expect(list.textContent).toContain('llama3');
|
||||
});
|
||||
|
||||
it('always asks with 2+ models, even when a defaultModelId is set — the point is letting this launch differ', async () => {
|
||||
const { win, app } = bootApp({
|
||||
hosts: [
|
||||
{
|
||||
id: 'llama-box',
|
||||
label: 'llama.cpp',
|
||||
baseUrl: 'http://localhost:8080',
|
||||
models: ['qwen3', 'llama3'],
|
||||
defaultModelId: 'qwen3',
|
||||
},
|
||||
],
|
||||
});
|
||||
await app.selectCustomModelEntry('claude', 'llama-box');
|
||||
const modal = win.document.getElementById('customModelPickModal')!;
|
||||
expect(modal.classList.contains('active')).toBe(true);
|
||||
// The default is marked, not auto-chosen.
|
||||
expect(win.document.getElementById('customModelPickList')!.textContent).toContain('Default');
|
||||
});
|
||||
|
||||
it('picking a row in the modal closes it and launches with that exact model', async () => {
|
||||
const { win, app } = bootApp({
|
||||
hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://localhost:8080', models: ['qwen3', 'llama3'] }],
|
||||
});
|
||||
let launched: unknown[] | null = null;
|
||||
app.runCustomModelEntry = async (...args: unknown[]) => {
|
||||
launched = args;
|
||||
};
|
||||
win.app = app;
|
||||
|
||||
await app.selectCustomModelEntry('claude', 'llama-box');
|
||||
const buttons = win.document.getElementById('customModelPickList')!.querySelectorAll('button');
|
||||
const llama3Btn = [...buttons].find((b) => b.textContent?.includes('llama3')) as unknown as HTMLButtonElement & {
|
||||
onclick: (e: unknown) => void;
|
||||
};
|
||||
expect(typeof llama3Btn.onclick).toBe('function');
|
||||
llama3Btn.onclick(new (win as any).Event('click'));
|
||||
|
||||
expect(launched).toEqual(['claude', 'llama-box', 'llama3']);
|
||||
expect(win.document.getElementById('customModelPickModal')!.classList.contains('active')).toBe(false);
|
||||
});
|
||||
|
||||
it('re-fetches the endpoint at click time rather than trusting anything cached from the menu render', async () => {
|
||||
// The background re-discovery sweep (server-side, every 5 minutes) or a
|
||||
// settings-panel edit can change the model list between opening the
|
||||
// dropdown and clicking a row — the picker must reflect what is current.
|
||||
let fetchCount = 0;
|
||||
const { win, app } = bootApp({});
|
||||
app._apiJson = async (path: string) => {
|
||||
if (path !== '/api/model-endpoints') return null;
|
||||
fetchCount += 1;
|
||||
return [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://x', models: ['qwen3', 'llama3', 'phi4'] }];
|
||||
};
|
||||
await app.selectCustomModelEntry('claude', 'llama-box');
|
||||
expect(fetchCount).toBe(1);
|
||||
expect(win.document.getElementById('customModelPickList')!.querySelectorAll('button').length).toBe(3);
|
||||
});
|
||||
|
||||
it('toasts and does nothing when the endpoint has vanished by click time', async () => {
|
||||
const { app } = bootApp({ hosts: [] });
|
||||
let toastMessage: string | null = null;
|
||||
app.showToast = (msg: string) => {
|
||||
toastMessage = msg;
|
||||
};
|
||||
await app.selectCustomModelEntry('claude', 'ghost-endpoint');
|
||||
expect(toastMessage).toMatch(/no longer exists/i);
|
||||
});
|
||||
|
||||
it('toasts and does nothing when the endpoint has zero discovered models by click time', async () => {
|
||||
const { app } = bootApp({
|
||||
hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://x', models: [] }],
|
||||
});
|
||||
let toastMessage: string | null = null;
|
||||
app.showToast = (msg: string) => {
|
||||
toastMessage = msg;
|
||||
};
|
||||
await app.selectCustomModelEntry('claude', 'llama-box');
|
||||
expect(toastMessage).toMatch(/no models discovered/i);
|
||||
});
|
||||
});
|
||||
|
||||
describe('Custom Model Endpoint Profiles: applying a picked entry', () => {
|
||||
it('does not apply the endpoint to a session that was already open when the launch fails', async () => {
|
||||
const { app } = bootApp({});
|
||||
|
||||
Reference in New Issue
Block a user