From e2034177c5fd978f2a8d6193a1b63b11effb2d8e Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Thu, 17 Sep 2026 13:41:19 +0800 Subject: [PATCH] fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit DeepSeek Harness's own bundled provider module (@deepseek-ai/dsh-llm-deepseek) builds its request URL as `${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its own (its real public API, https://api.deepseek.com, expects the caller's base URL to already carry any needed prefix), while llama-swap/llama.cpp only ever serves the OpenAI-conventional `/v1/chat/completions`. Confirmed two ways: - Installed the real @deepseek-ai/dsh package (all its actual published dependencies) into a scratch dir purely to read dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/ completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL — the same grep-the-real-source bar pi/grok's fixes were held to. - Live against the test-picker's llama-swap: `POST /chat/ completions` -> 404, `POST /v1/chat/completions` -> 200, same endpoint. dsh's own error template ("DeepSeek API error (HTTP ${status})") reproduces the originally-reported "dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly. - New registry field `appendV1Suffix` (env kind only, deepseek's entry alone — claude/gemini must NOT get it, since claude was already confirmed working against the unmodified baseUrl). When set, buildCustomModelInjection runs endpoint.baseUrl through the same withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already use, instead of writing it verbatim. Not yet re-run end-to-end through a real dsh binary — no install available in this environment (not in PATH, and the test-picker container doesn't bundle it) — so this is source-confirmed and live-verified at the HTTP level, not yet promoted to "verified" alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md, the plan doc's confidence table, the wiki page, CLAUDE.md) all updated to reflect this precisely rather than leaving the old "root cause not identified" claim in place. 2 new/updated tests for the /v1 suffix (including idempotency against a baseUrl that already ends in /v1) plus a corrected mock-server contract test. Typecheck/lint clean; full suite shows no new regressions. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG --- CLAUDE.md | 2 +- docs/custom-model-endpoints-plan.md | 22 ++++++------ docs/custom-model-endpoints.md | 19 ++++++++-- docs/wiki/Custom-Model-Endpoints.md | 14 ++++---- src/config/cli-registry/schema.ts | 5 +++ src/config/cli-registry/stock.ts | 21 ++++++++--- src/config/cli-registry/types.ts | 16 +++++++++ src/custom-model-injection.ts | 2 +- test/custom-model-injection-contract.test.ts | 38 +++++++++++--------- test/custom-model-injection.test.ts | 20 +++++++++-- 10 files changed, 114 insertions(+), 45 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 991fefc0..fb83c9f5 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -227,7 +227,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. The session version worked and was still wrong in use: it put a terminal tab on screen next to the web tab the user actually asked for, every single time, and nothing about a long-lived HTTP server needs to be a tab. ⚠️ What a session gave for free now has to be paid for explicitly, and every piece is load-bearing: **exactly one** server (a second click REUSES it rather than racing it for a port, which two sessions structurally could not do), **restarted when the browser authority changes** (`--trusted-host` fences dsh's `/api` against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two, so whoever asks last wins: the asker is by definition the origin about to load the page), **killed on shutdown** (`stopDeepSeekWeb()` in the server teardown, because the child is detached so its whole plugin tree can be signalled at once, which also means it would OUTLIVE Codeman and hold its port against the next start), and **failures returned to the caller**, since with no tab there is nowhere for a stack trace to land. ⚠️ The port search starts at dsh's own default 3080 and walks 40, never fixed: that default is precisely the port most likely to be taken already by the user's own `dsh web`, and hardcoding it killed this feature with EADDRINUSE once. Free-port detection BINDS rather than connects (a connect probe cannot tell "free" from "listening but not answering yet"), so it is racy by nature and the caller still waits for the server to really answer before reporting success. ⚠️ Both `POST` and `DELETE` sit at the **same privilege bar as the profile installer** (`canUsernameRunPrivilegedCommands`) even though the action reads as "open a page": booting a dsh profile executes the plugin code in it, and the server is a single shared instance, so stopping it in multi-user mode takes it out from under other users' tabs. -**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; full stack — settings-panel CRUD + the Run-menu picker, on top of the backend below): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ **Two launch paths, chosen by mechanism, not preference — see the second paragraph below for why**: opencode/codex/gemini/pi/grok/deepseek/omp apply the selection ONE-SHOT, before the session/process ever exists, with no restart at all; claude alone still applies a selection by **restarting the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` kills a WORKING pane, so a CLI whose launch declares a `fallback` chain (claude) gets the live conversation id pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct, and codex only speaks the Responses API since Feb 2026 (`wire_api = "responses"`) — re-verified live against a llama-swap deployment that DOES answer `/v1/responses` (an earlier test's harder failure against a different deployment does not reproduce everywhere): a plain, no-tool-call chat turn gets a real reply, but a real tool-call attempt came back as `agent_message` TEXT (the tool-call JSON printed as the answer) rather than an executable `function_call` item — confirmed via `codex exec --json`'s raw event stream. Tool execution is what makes codex a coding agent, so it remains not usable for real work either way, just with a more precise failure mode than a flat protocol break; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek reaches the server but gets a consistent `HTTP_404` (root cause not identified); antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. ⚠️ **The Run-menu picker generates entries from `window.__codemanCustomModelClis`** (`server.ts`, injected at page render from `enabledClis().filter(kind==='agent' && customModelInjection.kind!=='unsupported')`, JSON-escaped against a literal `` via the exported `escapeScriptJson()` since `label` is a user-`clis.json`-settable string unlike the neighbouring booleans-only `__codemanCliAvailable`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces. One entry per (capable, INSTALLED CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)", filtered through `isCliAvailable()` like the stock entries. Clicking one calls `selectCustomModelEntry(mode, endpointId)` (`session-ui.js`), which re-fetches the endpoint (never trusts anything cached from the dropdown's render — the 5-minute sweep below or a settings edit may have changed it since) and decides the model: exactly one discovered model launches straight away, two or more open `#customModelPickModal` to ask, with `defaultModelId` marked but never auto-chosen (asking exists so ONE launch can deliberately differ from the saved default). Either way the actual launch (`runCustomModelEntry`) routes through `run()` itself via a temporary `_runMode` swap — never `setRunMode()`, which would persist it as the user's new default — rather than a parallel dispatch table, which is what gives a custom-model launch the same `_runInFlight` lock every other Run click gets and means a CLI whose injection recipe lands later needs no update here. It then GETs `/api/sessions/:id/wait?until=idle&timeout=20000` on that session BEFORE applying — measured live, a freshly launched CLI reports itself `busy` for its own startup (boot spinner, workspace-trust check) well before the apply call would otherwise reach it, and the apply route's `isBusy()` guard correctly can't tell that apart from a real turn in progress, so every fresh launch failed with `SESSION_BUSY` until this wait was added. A timeout there is a normal 200 per the wait endpoint's own contract, never an error, so a session still busy after 20s just reaches the apply call anyway and gets that route's own honest error. It then calls `POST /api/sessions/:id/custom-model` on the session `run()` produced, guarded by snapshotting `activeSessionId` before the call and requiring it to have actually changed after — every `run*()` handles its own failure internally and returns normally rather than throwing, so a declined/failed launch must not silently re-point and restart whatever session was already open. ⚠️ The apply call reads the response body itself (`_api()`) rather than `_apiJson()`, which unwraps success but silently discards a failure body — losing the one thing (`error`) that distinguishes "still busy", "not a discovered model", "remote/Docker session" and everything else the route can report; the resulting toast is `type: 'error'`, which `showToast()` now defaults to STICKY (no auto-dismiss, an explicit close button) precisely so a message worth diagnosing survives long enough to be read — a 3s default hid the real reason behind every one of these failures until it was fixed. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to launch with). ⚠️ **Every saved endpoint's models also re-discover themselves automatically**, a `this.cleanup.setInterval` in `server.ts` (`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, 5 minutes, off under `testMode` like the Codex plan-usage poll beside it) calling the exported `refreshAllCustomModelHosts()` (`custom-model-routes.ts`) — one endpoint unreachable on a cycle never blocks the others, and a read-modify-write PER HOST (re-reading the store before each splice, keyed by id) means an admin's concurrent edit or delete wins over a sweep that started before it, never the reverse. +**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; full stack — settings-panel CRUD + the Run-menu picker, on top of the backend below): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ **Two launch paths, chosen by mechanism, not preference — see the second paragraph below for why**: opencode/codex/gemini/pi/grok/deepseek/omp apply the selection ONE-SHOT, before the session/process ever exists, with no restart at all; claude alone still applies a selection by **restarting the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` kills a WORKING pane, so a CLI whose launch declares a `fallback` chain (claude) gets the live conversation id pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct, and codex only speaks the Responses API since Feb 2026 (`wire_api = "responses"`) — re-verified live against a llama-swap deployment that DOES answer `/v1/responses` (an earlier test's harder failure against a different deployment does not reproduce everywhere): a plain, no-tool-call chat turn gets a real reply, but a real tool-call attempt came back as `agent_message` TEXT (the tool-call JSON printed as the answer) rather than an executable `function_call` item — confirmed via `codex exec --json`'s raw event stream. Tool execution is what makes codex a coding agent, so it remains not usable for real work either way, just with a more precise failure mode than a flat protocol break; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek's originally-reported `HTTP_404` is root-caused and fixed — its bundled `@deepseek-ai/dsh-llm-deepseek` module builds `${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` of its own (confirmed by installing the real package and reading its source), so a new `appendV1Suffix` flag on its registry entry (alone — claude/gemini must not get it) runs `endpoint.baseUrl` through `withV1Suffix()` before writing it, live-confirmed against llama-swap (`.../chat/completions` 404s, `.../v1/chat/completions` succeeds) though not yet re-run through an actual `dsh` binary, which isn't installable in this environment; antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. ⚠️ **The Run-menu picker generates entries from `window.__codemanCustomModelClis`** (`server.ts`, injected at page render from `enabledClis().filter(kind==='agent' && customModelInjection.kind!=='unsupported')`, JSON-escaped against a literal `` via the exported `escapeScriptJson()` since `label` is a user-`clis.json`-settable string unlike the neighbouring booleans-only `__codemanCliAvailable`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces. One entry per (capable, INSTALLED CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)", filtered through `isCliAvailable()` like the stock entries. Clicking one calls `selectCustomModelEntry(mode, endpointId)` (`session-ui.js`), which re-fetches the endpoint (never trusts anything cached from the dropdown's render — the 5-minute sweep below or a settings edit may have changed it since) and decides the model: exactly one discovered model launches straight away, two or more open `#customModelPickModal` to ask, with `defaultModelId` marked but never auto-chosen (asking exists so ONE launch can deliberately differ from the saved default). Either way the actual launch (`runCustomModelEntry`) routes through `run()` itself via a temporary `_runMode` swap — never `setRunMode()`, which would persist it as the user's new default — rather than a parallel dispatch table, which is what gives a custom-model launch the same `_runInFlight` lock every other Run click gets and means a CLI whose injection recipe lands later needs no update here. It then GETs `/api/sessions/:id/wait?until=idle&timeout=20000` on that session BEFORE applying — measured live, a freshly launched CLI reports itself `busy` for its own startup (boot spinner, workspace-trust check) well before the apply call would otherwise reach it, and the apply route's `isBusy()` guard correctly can't tell that apart from a real turn in progress, so every fresh launch failed with `SESSION_BUSY` until this wait was added. A timeout there is a normal 200 per the wait endpoint's own contract, never an error, so a session still busy after 20s just reaches the apply call anyway and gets that route's own honest error. It then calls `POST /api/sessions/:id/custom-model` on the session `run()` produced, guarded by snapshotting `activeSessionId` before the call and requiring it to have actually changed after — every `run*()` handles its own failure internally and returns normally rather than throwing, so a declined/failed launch must not silently re-point and restart whatever session was already open. ⚠️ The apply call reads the response body itself (`_api()`) rather than `_apiJson()`, which unwraps success but silently discards a failure body — losing the one thing (`error`) that distinguishes "still busy", "not a discovered model", "remote/Docker session" and everything else the route can report; the resulting toast is `type: 'error'`, which `showToast()` now defaults to STICKY (no auto-dismiss, an explicit close button) precisely so a message worth diagnosing survives long enough to be read — a 3s default hid the real reason behind every one of these failures until it was fixed. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to launch with). ⚠️ **Every saved endpoint's models also re-discover themselves automatically**, a `this.cleanup.setInterval` in `server.ts` (`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, 5 minutes, off under `testMode` like the Codex plan-usage poll beside it) calling the exported `refreshAllCustomModelHosts()` (`custom-model-routes.ts`) — one endpoint unreachable on a cycle never blocks the others, and a read-modify-write PER HOST (re-reading the store before each splice, keyed by id) means an admin's concurrent edit or delete wins over a sweep that started before it, never the reverse. **Everything below landed after the initial backend + picker cut, each confirmed live against a real llama-swap deployment.** ⚠️ **llama-swap runs one model at a time, and switching can disrupt ANOTHER live session** — before applying, both apply routes call llama-swap's own `GET /running` (feature-detected via `getLlamaSwapStatus()`, `custom-model-routes.ts`; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is simply never checked). If a different model is loaded and ready AND another live session's own selection is using it, the apply returns `{requiresConfirmation, currentlyLoadedModel, affectedSessions}` instead of switching silently; retrying with `confirmed: true` skips the check, and switching with nothing else affected proceeds immediately. llama-swap also has no dedicated "switch model" endpoint — the only thing that actually starts a swap is a real inference request naming the model (confirmed live: applying a selection alone never reached llama-swap's own logs, since nothing had asked it to load anything) — so both routes also fire `triggerLlamaSwapLoad()`, a fire-and-forget `POST /v1/chat/completions` with `max_tokens: 1`, whenever the target model isn't already loaded and ready. ⚠️ **That launch-time check cannot catch a swap caused by a DIFFERENT session's LATER, ordinary use** — confirmed live: a second Codex session picking a different model launched with no warning at all (nothing conflicted at that exact instant), yet it silently evicted the first session's model regardless, since llama-swap has no push notification of its own. `detectCustomModelSwapDisplacements()` (`custom-model-routes.ts`) is a separate periodic sweep (`server.ts`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS` = 20s) that compares each live custom-model session's own `modelId` against what `/running` actually reports loaded, broadcasting a `custom-model:swapped-out` SSE event — shown as a global toast, never tied to the displaced session's own tab, since the whole point is telling the user before they type into it — the first time a mismatch appears, via a caller-owned de-dupe `Set` cleared once that session's own model is loaded and ready again so a later, genuinely new displacement notifies again rather than staying silently un-notified forever after the first one. ⚠️ **Context length is read from the REAL launch command, never `/props`** — `/props?model=`'s `default_generation_settings.n_ctx` was confirmed live to report a `--fit-ctx`-launched backend's theoretical/trained maximum rather than the real runtime-configured size (a measured 154112-vs-16384 discrepancy, caught only because the unfixed value still overflowed), so discovery parses the actual configured size straight out of `/running`'s own `cmd` field instead (`parseCtxFromCmd`: `--fit-ctx ` first, then plain llama.cpp `-c`/`--ctx-size`), falling back to `/props` only when `cmd` states no recognizable flag at all. ⚠️ **Claude alone gets a context-window FLOOR check, on top of the ceiling `contextLengthVar` already fixes** — `exceedsSafeContextFloor()` (gated on the registry declaring `contextLengthVar`, so a no-op for every other CLI by construction) compares a model's discovered context against `CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000): confirmed live, twice, that Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on the very first message, before any conversation history exists to compact, so a smaller real context fails outright regardless of what `CLAUDE_CODE_MAX_CONTEXT_TOKENS` says (that var only controls when HISTORY gets compacted, and there is none yet on message one). Below the floor, the apply returns `{requiresContextWarning, modelId, contextLength, minSafeContextTokens}` instead of launching, shown as an in-app dialog naming the actual fix: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on `--fit-ctx` auto-fit, which optimizes for the biggest MODEL that fits rather than the biggest CONTEXT. ⚠️ **A fresh, isolated `CLAUDE_CONFIG_DIR` looks like a brand-new Claude Code profile and replays its ENTIRE first-run sequence on every launch** — the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running with a bypass-permissions flag) a one-time warning about it, confirmed live, none of which a real, already-onboarded profile shows again. `skipFirstRunPrompts` (claude's entry only, requires `apiKeyTrustFile` since it reuses the same file) pre-seeds that same "already been through this" state: `hasCompletedOnboarding` and this session's own `projects[workingDir].hasTrustDialogAccepted` merge into the same `.claude.json` the API-key trust file already writes to, and `skipDangerousModePermissionPrompt` merges into `settings.json` (a different file, same corrupt-tolerant merge). ⚠️ **The loading banner shows the REAL backend log line, not a guess, and has no countdown or auto-timeout at all.** `getLatestLlamaSwapLogLine()` holds one `GET /api/events` SSE connection open per endpoint (confirmed live to stay open indefinitely — read past 220KB over 8s with no `done`; idle-closed after 30s via `pruneIdleLlamaSwapLogTails`, same 20s sweep as the swap-displacement check above), parsing `logData` frames and keeping only `source: "upstream"` (the real `llama-server` process's own stdout) lines, never `source: "proxy"` (llama-swap's own request-access log). ⚠️ `GET /logs` — the endpoint this feature's own first cut targeted, since the name suggested it — was confirmed live to carry ONLY the proxy log and never a single backend line, even seconds after a real, verified model swap; caught and corrected by a live check before merge, not after. The banner itself dropped its size-scaled expected-time estimate and matching auto-timeout (a guess dressed up as a fact that could kill a genuinely slow load partway through on slower hardware) for a generic hardware/model-size disclaimer plus a user-driven **Cancel** button (`_showCenterStatus`'s `onCancel` option, a real button distinct from the plain "×" close glyph an `'error'`-type banner gets) that ends the wait and closes the session on the user's own call rather than a guessed deadline. diff --git a/docs/custom-model-endpoints-plan.md b/docs/custom-model-endpoints-plan.md index bb4ef289..5078f17b 100644 --- a/docs/custom-model-endpoints-plan.md +++ b/docs/custom-model-endpoints-plan.md @@ -104,17 +104,17 @@ declared capability, never an `if (mode === 'claude')` branch. ## Per-CLI injection recipes (confidence-ranked) -| CLI | Mechanism | Confidence | -| ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude | -| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"":{}}}},"model":"custom/"}` | **Verified by user** | -| `codex` | TOML `config.toml`: top-level `model = ""` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol picture more nuanced than a flat break, re-verified live twice on 2026-09-17 against a llama-swap deployment that DOES answer `/v1/responses`** (an earlier test's `Reconnecting...`/`high demand` failure does not reproduce against every llama-swap setup): a plain, no-tool-call chat turn (`codex exec 'reply with just OK'`) returned a real reply. But a real tool-call attempt (`run the shell command: echo hello`) came back as an `agent_message` TEXT item — the tool-call JSON printed as the model's answer, not a `function_call` item codex would actually execute (confirmed via `codex exec --json`'s raw event stream: `item.completed`/`agent_message`, never `function_call`). Since tool execution is what makes codex a coding agent at all, this remains **not usable for real work**, just with a different, more specific failure mode than previously documented — still do not present this as working. Separately, EVERY custom-endpoint codex session also prints `warning: Model metadata for '' not found. Defaulting to fallback metadata...` on launch (confirmed harmless — the successful plain-text reply above still had it): codex's per-model metadata (reasoning tiers, system-prompt templates, context-window figures) comes from `models_cache.json`, a LOCAL CACHE of OpenAI's own hosted model catalog that a custom model can never appear in by construction. No config.toml override exists for it, and the isolated `CODEX_HOME` never gets a `models_cache.json` written into it at all (confirmed: inspected a live, actively-used isolated dir — codex evidently can't reach OpenAI's catalog endpoint for this session and just falls back silently every time, with no file left behind to fix or clean up). Fabricating a fake catalog entry to suppress the warning would mean copying the _shape_ of OpenAI's own proprietary schema — including their real per-model system-prompt content, visible in a genuine `models_cache.json` — for a warning confirmed to have no effect on the actual (broken) tool-calling outcome; not worth building | -| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done | -| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" | -| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly | -| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working | -| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live | -| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism | +| CLI | Mechanism | Confidence | +| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude | +| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"":{}}}},"model":"custom/"}` | **Verified by user** | +| `codex` | TOML `config.toml`: top-level `model = ""` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol picture more nuanced than a flat break, re-verified live twice on 2026-09-17 against a llama-swap deployment that DOES answer `/v1/responses`** (an earlier test's `Reconnecting...`/`high demand` failure does not reproduce against every llama-swap setup): a plain, no-tool-call chat turn (`codex exec 'reply with just OK'`) returned a real reply. But a real tool-call attempt (`run the shell command: echo hello`) came back as an `agent_message` TEXT item — the tool-call JSON printed as the model's answer, not a `function_call` item codex would actually execute (confirmed via `codex exec --json`'s raw event stream: `item.completed`/`agent_message`, never `function_call`). Since tool execution is what makes codex a coding agent at all, this remains **not usable for real work**, just with a different, more specific failure mode than previously documented — still do not present this as working. Separately, EVERY custom-endpoint codex session also prints `warning: Model metadata for '' not found. Defaulting to fallback metadata...` on launch (confirmed harmless — the successful plain-text reply above still had it): codex's per-model metadata (reasoning tiers, system-prompt templates, context-window figures) comes from `models_cache.json`, a LOCAL CACHE of OpenAI's own hosted model catalog that a custom model can never appear in by construction. No config.toml override exists for it, and the isolated `CODEX_HOME` never gets a `models_cache.json` written into it at all (confirmed: inspected a live, actively-used isolated dir — codex evidently can't reach OpenAI's catalog endpoint for this session and just falls back silently every time, with no file left behind to fix or clean up). Fabricating a fake catalog entry to suppress the warning would mean copying the _shape_ of OpenAI's own proprietary schema — including their real per-model system-prompt content, visible in a genuine `models_cache.json` — for a warning confirmed to have no effect on the actual (broken) tool-calling outcome; not worth building | +| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done | +| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" | +| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly | +| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`), now with `appendV1Suffix: true` (see confidence). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Root cause of the original `HTTP_404` found and fixed, by reading dsh's own bundled source — the same bar pi/grok's fixes were held to.** Installed `@deepseek-ai/dsh` (all its real published dependencies) into a scratch directory purely to read `@deepseek-ai/dsh-llm-deepseek/lib/index.js`: it builds its request as `fetch(\`${connection.baseURL}/chat/completions\`, ...)`with`baseURL`read straight from`DEEPSEEK_BASE_URL`(or defaulting to DeepSeek's real public API root,`https://api.deepseek.com`, which also carries no `/v1`) — no `/v1` insertion of dsh's own, unlike the OpenAI-SDK convention this recipe originally assumed. llama-swap/llama.cpp only ever serves the OpenAI-conventional `/v1/chat/completions`. Confirmed live: `POST /chat/completions` → `404`, `POST /v1/chat/completions` → `200`, on the exact same endpoint — and dsh's own error-message template, `DeepSeek API error (HTTP ${status})`, reproduces the originally reported `dsh: HTTP_404: DeepSeek API error (HTTP 404)` precisely. Fixed by adding `appendV1Suffix` (env kind only, deepseek's entry alone — claude/gemini must NOT get it, since claude was already confirmed working against the unmodified `baseUrl`), which runs `endpoint.baseUrl` through the same `withV1Suffix()` helper `configDir`-kind CLIs already use. ⚠️ Not yet re-run end-to-end with a real `dsh` binary — no install available in this environment (no npm-installed CLI binary in `PATH`, and the `codeman-test-picker` container doesn't bundle it either); the fix is source-confirmed and live-verified at the HTTP level, but a genuine "hello world" reply through `dsh` itself is the remaining step before promoting this to **verified** alongside claude/opencode/pi/grok/omp | +| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live | +| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism | Everything web-researched-but-unverified gets implemented but must be smoke-tested against real installs of those CLIs before being called done — diff --git a/docs/custom-model-endpoints.md b/docs/custom-model-endpoints.md index cd42456e..59e5ed56 100644 --- a/docs/custom-model-endpoints.md +++ b/docs/custom-model-endpoints.md @@ -482,8 +482,23 @@ for '' not found. Defaulting to fallback metadata...` on launch — `GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation (several auth workarounds were tried and ruled out); do not rely on Gemini support yet. -- **DeepSeek** — the request reaches the server (env vars are read) but - gets a consistent `HTTP_404`. Root cause not identified; best-effort only. +- **DeepSeek** — root cause of the `HTTP_404` found and fixed. DeepSeek + Harness's own bundled provider module (`@deepseek-ai/dsh-llm-deepseek`) + builds its request URL as `${DEEPSEEK_BASE_URL}/chat/completions` with no + `/v1` insertion of its own (its real public API, `https://api.deepseek.com`, + expects the caller's base URL to already carry any needed prefix) — + confirmed by reading its own source and, live, that + `POST /chat/completions` 404s against llama-swap while + `POST /v1/chat/completions` succeeds; the harness's own error + template (`DeepSeek API error (HTTP ${status})`) matches the originally + reported symptom exactly. `customModelInjection`'s new `appendV1Suffix` + (deepseek's entry only — claude/gemini must NOT get it, since claude was + already confirmed working against the raw `baseUrl`) fixes it by writing + `DEEPSEEK_BASE_URL` with `/v1` appended. Not yet re-run end-to-end with a + real `dsh` binary (no install available in this environment) — the fix + is source-confirmed and live-verified at the HTTP level, but a real + "hello world" reply through `dsh` itself is still outstanding before + calling this fully verified like the harnesses above. - **Antigravity** — no known custom-endpoint mechanism at all; unsupported. See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind diff --git a/docs/wiki/Custom-Model-Endpoints.md b/docs/wiki/Custom-Model-Endpoints.md index 0dbedf91..ebe05fee 100644 --- a/docs/wiki/Custom-Model-Endpoints.md +++ b/docs/wiki/Custom-Model-Endpoints.md @@ -146,13 +146,13 @@ entry. "Launch anyway" is still there if you want to try regardless. ## Which harnesses actually work -| Harness | Status | -| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | -| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. | -| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. | -| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. | -| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. | -| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. | +| Harness | Status | +| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. | +| **Codex** | Config is correct, and plain chat can work against a server that speaks the Responses API — but a real tool-call attempt comes back as inert text instead of running, so it's still not usable for real coding work. | +| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. | +| **DeepSeek** | The original 404 is root-caused and fixed (DeepSeek Harness's own code was missing a `/v1` most local servers require) — not yet re-run against a real `dsh` install to confirm end-to-end. | +| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. | Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry, not a fixed list here, so this table can go stale before this page does — a greyed-out or diff --git a/src/config/cli-registry/schema.ts b/src/config/cli-registry/schema.ts index 2bd75da6..d1249384 100644 --- a/src/config/cli-registry/schema.ts +++ b/src/config/cli-registry/schema.ts @@ -364,6 +364,11 @@ const capabilitiesSchema = z // same file to pre-seed the state a real, already-onboarded profile carries. See // the customModelInjection doc comment in cli-registry/types.ts. skipFirstRunPrompts: z.boolean().optional(), + // DeepSeek-only, confirmed by reading its own bundled SDK source: it concatenates + // "/chat/completions" onto baseUrlVar's value with no "/v1" of its own, while + // llama-swap/llama.cpp only serves the "/v1/..." path — claude/gemini must NOT + // get this. See the customModelInjection doc comment in cli-registry/types.ts. + appendV1Suffix: z.boolean().optional(), }) .strict(), z diff --git a/src/config/cli-registry/stock.ts b/src/config/cli-registry/stock.ts index 37d8f37e..4a357eda 100644 --- a/src/config/cli-registry/stock.ts +++ b/src/config/cli-registry/stock.ts @@ -1103,15 +1103,28 @@ const DEEPSEEK: CliEntry = { // privilege rather than granting it, and clamping it here was a real regression // (test/deepseek-mode.test.ts) fixed before this shipped. privilegedEnvKeys: ['DSH_PERMISSION_MODE', 'DSH_HOME', 'DEEPSEEK_BASE_URL'], - // Web-researched, unverified, partial: reuses the already-existing DEEPSEEK_BASE_URL/ - // DEEPSEEK_API_KEY keys above. No modelVars — dsh's model is a profile-composition - // entry (see `model: { source: 'none' }` above), not an env var, so forcing a specific - // model name may not fully work; verify against a real profile before shipping. + // Reuses the already-existing DEEPSEEK_BASE_URL/DEEPSEEK_API_KEY keys above. No + // modelVars — dsh's model is a profile-composition entry (see `model: { source: 'none' + // }` above), not an env var, so forcing a specific model name may not fully work; + // verify against a real profile before shipping. + // + // ⚠️ appendV1Suffix is REQUIRED, not optional-nice-to-have: without it every request + // 404s. Confirmed live and by reading dsh's own bundled source + // (@deepseek-ai/dsh-llm-deepseek): it builds the request URL as + // `${DEEPSEEK_BASE_URL}/chat/completions` with no "/v1" of its own (its real public + // API, https://api.deepseek.com, expects the caller's base URL to already carry any + // needed prefix), while llama-swap/llama.cpp only serves the OpenAI-conventional + // "/v1/chat/completions" — a bare POST to ".../chat/completions" 404s live, and the + // 404 reported here originally ("dsh: HTTP_404: DeepSeek API error (HTTP 404)") + // matches dsh's own error-message template for exactly this failure. See the + // customModelInjection doc comment in cli-registry/types.ts for the full reasoning, + // including why claude/gemini must NOT get this. customModelInjection: { kind: 'env', baseUrlVar: 'DEEPSEEK_BASE_URL', apiKeyVar: 'DEEPSEEK_API_KEY', modelVars: [], + appendV1Suffix: true, }, }, overlays: { diff --git a/src/config/cli-registry/types.ts b/src/config/cli-registry/types.ts index 8351febf..871a3762 100644 --- a/src/config/cli-registry/types.ts +++ b/src/config/cli-registry/types.ts @@ -535,6 +535,21 @@ export interface CliCapabilities { * in claude's `settings.json` — see `seedFirstRunState`/`seedSkipBypassPermissionsPrompt` * in custom-model-injection-apply.ts. Requires `apiKeyTrustFile` to be set too, since it * reuses that file. + * + * `appendV1Suffix` (env kind only): the raw `endpoint.baseUrl` gets `withV1Suffix()` + * applied before being written to `baseUrlVar`, instead of being used verbatim. + * DeepSeek needs this and claude/gemini must NOT get it — a per-CLI asymmetry confirmed + * by reading each SDK's own request-building source, not assumed: DeepSeek Harness's + * bundled `@deepseek-ai/dsh-llm-deepseek` concatenates `${connection.baseURL}/chat/ + * completions` with no `/v1` insertion of its own (its real public API base, + * `https://api.deepseek.com`, expects the caller's base URL to already carry any + * needed prefix), while llama-swap/llama.cpp only ever serves the OpenAI-conventional + * `/v1/chat/completions` — confirmed live: a bare `POST /chat/completions` + * 404s, `POST /v1/chat/completions` succeeds, and the harness's own error + * message template (`DeepSeek API error (HTTP ${status})`) reproduces the exact + * `HTTP_404` this feature originally shipped with unexplained. Claude Code's own SDK, + * by contrast, was already confirmed working end-to-end against the RAW `baseUrl` with + * no suffix — appending one there would be wrong, not just redundant. */ customModelInjection: | { @@ -547,6 +562,7 @@ export interface CliCapabilities { apiKeyTrustFile?: { relPath: string; shape: 'claude-api-key-responses' }; configDirVar?: string; skipFirstRunPrompts?: boolean; + appendV1Suffix?: boolean; } | { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string } | { diff --git a/src/custom-model-injection.ts b/src/custom-model-injection.ts index 3d678732..9aab9422 100644 --- a/src/custom-model-injection.ts +++ b/src/custom-model-injection.ts @@ -121,7 +121,7 @@ export function buildCustomModelInjection( switch (cap.kind) { case 'env': { const envOverrides: Record = { - [cap.baseUrlVar]: endpoint.baseUrl, + [cap.baseUrlVar]: cap.appendV1Suffix ? withV1Suffix(endpoint.baseUrl) : endpoint.baseUrl, [cap.apiKeyVar]: apiKey, }; for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId; diff --git a/test/custom-model-injection-contract.test.ts b/test/custom-model-injection-contract.test.ts index 85209318..1685504d 100644 --- a/test/custom-model-injection-contract.test.ts +++ b/test/custom-model-injection-contract.test.ts @@ -142,17 +142,19 @@ describe('custom-model-injection contract (mock server)', () => { expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key'); }); - // gemini/deepseek's `env` kind passes the base URL through UNCHANGED (unlike - // opencode/codex/pi/omp/grok, which build a structured config and explicitly append - // /v1) — matching Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the - // SDK appends the path itself. Whether each of these TWO CLIs' own OpenAI-compatible - // client expects the var to already include /v1 (the common OpenAI-SDK convention) or - // appends it itself is genuinely CLI-specific and UNVERIFIED (see the confidence table - // in docs/custom-model-endpoints-plan.md) — these tests model the common OpenAI-SDK convention (base_url - // ends in /v1) since that's the more likely behavior for an OpenAI-compatible client, - // but that assumption should be corrected here the moment it's checked against a real - // binary. (grok WAS in this group too, until live-testing showed the whole `env` recipe - // was wrong for it — see its own test below.) + // gemini's `env` kind still passes the base URL through UNCHANGED (matching + // Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the SDK appends + // the path itself) — whether gemini-cli's own OpenAI-compatible-ish client expects the + // var to already include /v1 or appends it itself remains genuinely UNVERIFIED (it + // fails for an unrelated auth reason before this would even matter — see the + // confidence table in docs/custom-model-endpoints-plan.md); this test models the + // common OpenAI-SDK convention as the best guess, to be corrected the moment it's + // checked against a real client. deepseek WAS in this "passes through unchanged" + // group too, until reading `@deepseek-ai/dsh-llm-deepseek`'s own bundled source + // confirmed it builds its request URL as `${DEEPSEEK_BASE_URL}/chat/completions` with + // no `/v1` of its own — `appendV1Suffix` now fixes that (see its own test below), + // the same way grok's whole `env` recipe turned out to be wrong before live-testing + // corrected it to a `configDir` one. it('gemini: GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY reach the mock', async () => { const injection = buildCustomModelInjection(entryOrThrow('gemini'), endpointFor(mock), 'qwen3'); @@ -186,16 +188,18 @@ describe('custom-model-injection contract (mock server)', () => { expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key'); }); - it('deepseek: DEEPSEEK_BASE_URL/DEEPSEEK_API_KEY reach the mock (base URL/key only, no model var)', async () => { + it('deepseek: DEEPSEEK_BASE_URL already carries the /v1 suffix dsh itself never adds, reaching the mock at the real path dsh requests', async () => { + // Confirmed by reading dsh's own bundled source: it fetches + // `${DEEPSEEK_BASE_URL}/chat/completions` verbatim, no /v1 insertion of its own — so + // this call (unlike gemini's above) passes DEEPSEEK_BASE_URL to callOpenAiCompat + // UNMODIFIED, exactly mirroring what the real harness does, rather than the test + // helping it along. const injection = buildCustomModelInjection(entryOrThrow('deepseek'), endpointFor(mock), 'qwen3'); if (injection.kind !== 'env') throw new Error('unreachable'); expect(Object.keys(injection.envOverrides).sort()).toEqual(['DEEPSEEK_API_KEY', 'DEEPSEEK_BASE_URL']); + expect(injection.envOverrides.DEEPSEEK_BASE_URL).toBe(`${mock.baseUrl}/v1`); - await callOpenAiCompat( - `${injection.envOverrides.DEEPSEEK_BASE_URL}/v1`, - injection.envOverrides.DEEPSEEK_API_KEY, - 'qwen3' - ); + await callOpenAiCompat(injection.envOverrides.DEEPSEEK_BASE_URL, injection.envOverrides.DEEPSEEK_API_KEY, 'qwen3'); expect(mock.requests[0].path).toBe('/v1/chat/completions'); expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key'); diff --git a/test/custom-model-injection.test.ts b/test/custom-model-injection.test.ts index 886a2e35..c99fff49 100644 --- a/test/custom-model-injection.test.ts +++ b/test/custom-model-injection.test.ts @@ -179,15 +179,31 @@ describe('buildCustomModelInjection', () => { expect(result.extraEnv).toEqual({ XAI_API_KEY: 'my-key' }); }); - it('deepseek: env kind sets base URL/key only, no model var', () => { + it('deepseek: env kind sets base URL (with a /v1 suffix appended) and key, no model var', () => { + // appendV1Suffix is REQUIRED here, not cosmetic: confirmed by reading dsh's own + // bundled source (@deepseek-ai/dsh-llm-deepseek) that it builds the request URL as + // `${DEEPSEEK_BASE_URL}/chat/completions` with no "/v1" of its own, while + // llama-swap/llama.cpp only serves "/v1/chat/completions" — without this, every + // request 404s (confirmed live; this is the fix for the originally-reported + // "dsh: HTTP_404: DeepSeek API error (HTTP 404)"). const result = buildCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3'); if (result.kind !== 'env') throw new Error('unreachable'); expect(result.envOverrides).toEqual({ - DEEPSEEK_BASE_URL: 'http://192.168.1.50:8080', + DEEPSEEK_BASE_URL: 'http://192.168.1.50:8080/v1', DEEPSEEK_API_KEY: 'my-key', }); }); + it('deepseek: appending the /v1 suffix is idempotent against a baseUrl that already ends in /v1', () => { + const result = buildCustomModelInjection( + entryOrThrow('deepseek'), + { ...endpoint, baseUrl: 'http://192.168.1.50:8080/v1' }, + 'qwen3' + ); + if (result.kind !== 'env') throw new Error('unreachable'); + expect(result.envOverrides.DEEPSEEK_BASE_URL).toBe('http://192.168.1.50:8080/v1'); + }); + it('antigravity: unsupported', () => { const result = buildCustomModelInjection(entryOrThrow('antigravity'), endpoint, 'qwen3'); expect(result).toEqual({ kind: 'unsupported' });