diff --git a/docs/api-reference.md b/docs/api-reference.md index dcd0d631..3a5dab66 100644 --- a/docs/api-reference.md +++ b/docs/api-reference.md @@ -516,6 +516,53 @@ All four enforce session ownership in multi-user mode; a foreign session id answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the same directory are distinct by construction. +## Custom Model Endpoints + +Points a session's harness at a user-configured OpenAI-compatible endpoint — +local (llama.cpp, vLLM, DGX Spark) or cloud (Azure AI Foundry, OpenRouter) — +instead of its native cloud backend, gated by the opt-in +`customModelEndpointsEnabled` setting (default OFF). Endpoints are +machine-level infra, like remote/docker hosts: writes are admin-only in +multi-user mode. Design: [`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md); +user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md). + +- `GET /api/v1/model-endpoints` -> `CustomModelHost[]`, an unwrapped bare + array like every other list route (still riding the standard `{success, + data}` envelope on the wire — unwrap it the same way). Answers `[]` for a + non-admin in multi-user mode. `apiKey` is never returned; `apiKeySet: + boolean` reports whether one is stored, so a client can render "unchanged + if left blank" without ever holding the real value. +- `POST /api/v1/model-endpoints` with `{ id, label, baseUrl, apiKey?, + authStyle?, defaultModelId? }` creates one. `id` must match + `^[a-zA-Z0-9_-]+$`; `authStyle` is `bearer` (default) or `api-key`, never + both (a real server hung indefinitely when sent both headers on one + request); `baseUrl` must be `http(s)`, carry no embedded credentials, and + is refused if it points at (or resolves to) a link-local or + cloud-metadata address. `409 ALREADY_EXISTS` on a duplicate id. +- `PUT /api/v1/model-endpoints/:id` updates one. An **absent** `apiKey` + keeps the stored one rather than clearing it — the client never receives + the real value to resend deliberately unchanged, so omission is the only + way to say "leave it alone"; there is no way to clear a key back to unset + this way. `defaultModelId`, when set, must be one of that endpoint's own + `models` (`400 INVALID_INPUT` otherwise). +- `DELETE /api/v1/model-endpoints/:id` removes one. +- `POST /api/v1/model-endpoints/:id/discover-models` fetches the endpoint's + own `GET /v1/models` and stores the result as `models`, updating + `lastDiscoveredAt`. A `defaultModelId` that no longer appears in the fresh + list is dropped rather than carried forward invalid. Failures answer + `502 OPERATION_FAILED` with the underlying connection error, or a named + egress refusal if the resolved address turned out to be blocked. +- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId } | + { clear: true }` applies (or clears) the session's selection and + **restarts the session's CLI process in place** — every supported harness + reads its endpoint config at process start, never per turn, so there is + no live hot-swap. A Claude session resumes its existing conversation + across the restart; pi/omp/grok additionally get a forced `--model`/`-m` + value, since for those three the config file alone does not select it. + `400 INVALID_INPUT` for a remote (SSH) or Docker session — both restart + their agent differently under the hood, and applying to one would report + success while changing nothing. + ## Voice dictation Browser dictation transcribed through this server's Claude Code login, i.e. the diff --git a/docs/custom-model-endpoints-plan.md b/docs/custom-model-endpoints-plan.md index 85407adf..2bf5db57 100644 --- a/docs/custom-model-endpoints-plan.md +++ b/docs/custom-model-endpoints-plan.md @@ -208,6 +208,14 @@ extra per-model configuration on Codeman's side at all. ### 4. Toolbar UI +> **Superseded.** This section describes the toolbar-button design as originally +> planned. What actually shipped is a Run-menu picker instead: one generated entry +> per (capable harness, saved endpoint) pair directly in the existing `#runModeMenu` +> dropdown, rather than a separate `#customModelBtn`/`#customModelMenu` surface. See +> [`docs/custom-model-endpoints.md`](custom-model-endpoints.md#the-run-menu-picker) +> for the current design; the sections below (session-restart mechanics, security) +> remain accurate regardless of which UI calls the underlying route. + - New header/toolbar button (e.g. `#customModelBtn`, `btn-toolbar btn-custom-model`), marker-hidden by default (`btn-custom-model--hidden`) and revealed by `applyHeaderVisibilitySettings()` only when diff --git a/docs/wiki/Custom-Model-Endpoints.md b/docs/wiki/Custom-Model-Endpoints.md index f7e2ed28..a989a858 100644 --- a/docs/wiki/Custom-Model-Endpoints.md +++ b/docs/wiki/Custom-Model-Endpoints.md @@ -37,7 +37,9 @@ necessary, not incidental: every supported harness reads its endpoint config at start, never per turn, so there is no live hot-swap while a turn is running. Entries are hidden entirely for a session in a **remote (SSH) or Docker case** — support for -redirecting those hasn't landed yet, see below. +redirecting those hasn't landed yet, see below. The picker also only appears in the desktop +**Run** dropdown; the phone home screen builds its own run picker separately and does not +currently offer these entries. ## Which harnesses actually work @@ -59,6 +61,9 @@ missing entry is the more current answer. (reattaching a durable tmux session rather than relaunching the process), so redirecting them needs its own plumbing that hasn't been built. - **No live hot-swap mid-conversation.** Applying a selection always restarts the process. +- **No button to un-point a session from the UI yet.** Clearing back to native cloud is an + HTTP call (`POST .../custom-model {"clear": true}`) or deleting the session; the settings + panel manages saved endpoints, not what a running session is currently pointed at. - **Nothing is shared with your real cloud credentials.** The endpoint's own key, if any, never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate, explicit choice per session. diff --git a/src/web/public/i18n.js b/src/web/public/i18n.js index bee8fcab..d72cc33a 100644 --- a/src/web/public/i18n.js +++ b/src/web/public/i18n.js @@ -286,6 +286,31 @@ 'Prompt sent': '提示已发送', 'Inserted, press Enter in the terminal to send': '已插入,在终端中按 Enter 发送', 'Could not reach the session': '无法连接到会话', + 'Custom model endpoints': '自定义模型端点', + 'Point a harness at your own OpenAI-compatible server (llama.cpp, vLLM, DGX Spark, Azure AI Foundry, OpenRouter) instead of its native cloud backend. When on, the Run menu offers an extra entry per harness that supports it, per saved endpoint.': + '让工具指向您自己的兼容 OpenAI 服务器(llama.cpp、vLLM、DGX Spark、Azure AI Foundry、OpenRouter),而非其原生云端后端。开启后,"运行"菜单会为每个支持此功能的工具、每个已保存的端点新增一个条目。', + 'Enable custom model endpoints': '启用自定义模型端点', + 'Adds a per-endpoint entry to the Run menu for every harness that can redirect to one.': + '为每个可重定向到端点的工具,在"运行"菜单中添加对应条目。', + 'No endpoints yet. Add one below to point a harness at a local or cloud OpenAI-compatible server.': + '暂无端点。请在下方添加一个,以便将工具指向本地或云端的兼容 OpenAI 服务器。', + Discover: '发现模型', + '+ Add endpoint': '+ 添加端点', + 'Add endpoint': '添加端点', + Id: 'ID', + 'Short, stable — used in URLs, never shown to the CLI.': '简短且固定 — 用于 URL,不会展示给 CLI。', + Label: '标签', + 'Base URL': '基础 URL', + 'API key': 'API 密钥', + 'Optional. Left blank on edit keeps the existing key.': '可选。编辑时留空将保留现有密钥。', + 'Auth header': '认证请求头', + 'Never send both — some servers hang indefinitely.': '切勿同时发送两者 — 部分服务器会因此无限期挂起。', + 'Authorization: Bearer (default)': 'Authorization: Bearer(默认)', + 'api-key header (Azure)': 'api-key 请求头(Azure)', + 'Default model': '默认模型', + 'What the Run-menu picker applies for this endpoint. Discover models first.': + '运行菜单选择器会为此端点应用该模型。请先发现可用模型。', + 'Custom Endpoints': '自定义端点', 'Subagent Options': '子智能体选项', 'Enable Tracking': '启用跟踪', 'Active Tab Only': '仅活动标签页', diff --git a/src/web/public/index.html b/src/web/public/index.html index 2e6fded1..65007ffc 100644 --- a/src/web/public/index.html +++ b/src/web/public/index.html @@ -652,7 +652,7 @@
@@ -2216,42 +2216,46 @@ Enable custom model endpoints Adds a per-endpoint entry to the Run menu for every harness that can redirect to one. - + - - -