Merge pull request #430 from opticon454/custom-model-run-menu

This commit is contained in:
Codeman maintainer
2026-09-19 12:25:11 +02:00
43 changed files with 7304 additions and 144 deletions
+99
View File
@@ -582,6 +582,105 @@ All four enforce session ownership in multi-user mode; a foreign session id
answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the
same directory are distinct by construction.
## Custom Model Endpoints
Points a session's harness at a user-configured OpenAI-compatible endpoint —
local (llama.cpp, vLLM, DGX Spark) or cloud (Azure AI Foundry, OpenRouter) —
instead of its native cloud backend, gated by the opt-in
`customModelEndpointsEnabled` setting (default OFF). Endpoints are
machine-level infra, like remote/docker hosts: writes are admin-only in
multi-user mode. Design: [`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md);
user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md).
- `GET /api/v1/model-endpoints` -> `CustomModelHost[]`, an unwrapped bare
array like every other list route (still riding the standard `{success,
data}` envelope on the wire — unwrap it the same way). Answers `[]` for a
non-admin in multi-user mode. `apiKey` is never returned; `apiKeySet:
boolean` reports whether one is stored, so a client can render "unchanged
if left blank" without ever holding the real value.
- `POST /api/v1/model-endpoints` with `{ id, label, baseUrl, apiKey?,
authStyle?, defaultModelId? }` creates one. `id` must match
`^[a-zA-Z0-9_-]+$`; `authStyle` is `bearer` (default) or `api-key`, never
both (a real server hung indefinitely when sent both headers on one
request); `baseUrl` must be `http(s)`, carry no embedded credentials, and
is refused if it points at (or resolves to) a link-local or
cloud-metadata address. `409 ALREADY_EXISTS` on a duplicate id.
- `PUT /api/v1/model-endpoints/:id` updates one. An **absent** `apiKey`
keeps the stored one rather than clearing it — the client never receives
the real value to resend deliberately unchanged, so omission is the only
way to say "leave it alone"; there is no way to clear a key back to unset
this way. `defaultModelId`, when set, must be one of that endpoint's own
`models` (`400 INVALID_INPUT` otherwise).
- `DELETE /api/v1/model-endpoints/:id` removes one.
- `POST /api/v1/model-endpoints/:id/discover-models` fetches the endpoint's
own `GET /v1/models` and stores the result as `models`, updating
`lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own
response already reports loaded) `modelContextLengths` and `modelSizesGB`.
A `defaultModelId` that no longer appears in the fresh list is dropped
rather than carried forward invalid. Failures answer `422 OPERATION_FAILED`
with the underlying connection error, or a named egress refusal if the
resolved address turned out to be blocked. The same refresh also runs
automatically for every saved endpoint every 5 minutes in the background
(`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, started from
`server.ts`), so there is no route for triggering "refresh all" — one
endpoint being unreachable on a cycle never blocks the others.
- `GET /api/v1/model-endpoints/:id/running-status` -> `{ isLlamaSwap,
running: [{model, state, cmd?}], logLine? }`, read-only, no admin gate
(any session owner who could already point a session at this endpoint can
equally ask what it currently has loaded). `isLlamaSwap` is
feature-detected via the endpoint's own `GET /running` — a plain
llama.cpp/OpenAI-compatible server has none and always answers `false`.
`logLine`, present only when `isLlamaSwap` is true, is the most recent
REAL backend `llama-server` process log line (`load_model: ...`,
`llama_server: model loaded`, etc.), sourced from the endpoint's own
`GET /api/events` SSE stream and filtered to `source: "upstream"` frames
only (never llama-swap's own `source: "proxy"` request-access log) — one
connection is held open per endpoint and reused across every poller,
idle-closed after 30s of nobody asking. This is what the Run-menu
picker's loading banner polls once a second while a model is loading.
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId,
confirmed? } | { clear: true }` applies (or clears) the session's
selection and **restarts the session's CLI process in place** — every
supported harness reads its endpoint config at process start, never per
turn, so there is no live hot-swap. (`POST /api/v1/quick-start`'s own
`customModel: { endpointId, modelId, confirmed? }` field is the
no-restart equivalent for a session that doesn't exist yet — see below.)
A Claude session resumes its existing conversation across the restart;
pi/omp/grok additionally get a forced `--model`/`-m` value, since for
those three the config file alone does not select it. `400 INVALID_INPUT`
for a remote (SSH) or Docker session — both restart their agent
differently under the hood, and applying to one would report success
while changing nothing. Two more responses replace the normal
`{customModel, restarted}` shape, neither an error — both require
retrying the same call with `confirmed: true` to proceed anyway, and
neither restarts or creates anything on the first ask:
- `{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}` —
llama.cpp/llama-swap only runs one model at a time, and switching would
unload a model another **live session's own selection** is actively
using. Never returned for a plain (non-llama-swap) server, and never
just because a swap is needed at all — only when it would disrupt
someone else.
- `{requiresContextWarning: true, modelId, contextLength,
minSafeContextTokens}` — Claude Code's own fixed per-turn overhead
(system prompt + tool schemas) can exceed a small model's entire
discovered context on its own, before any conversation history exists
to compact, guaranteeing the very first message fails regardless of
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`. Gated on the CLI registry declaring a
`contextLengthVar` (claude only today), so it never fires for another
harness.
- `POST /api/v1/quick-start`'s `customModel: { endpointId, modelId,
confirmed? }` field (alongside its normal `caseName`/`mode`/etc. body)
computes the same injection **before** the session exists and launches
directly on the endpoint — no restart, because there was never a
native-backend boot to restart away from. Runs the identical checks as
the dedicated route above (`requiresConfirmation`/`requiresContextWarning`,
same shapes, same `confirmed: true` retry), and is refused the same way
for a remote or Docker case. This is what the Run-menu picker uses for
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; Claude still uses the
dedicated restart route above (its `--resume`-based restart is far less
jarring than a full relaunch, and folding it into the one-shot path is
separate work — see `docs/custom-model-endpoints-plan.md`).
## Voice dictation
Browser dictation transcribed through this server's Claude Code login, i.e. the
+25 -13
View File
@@ -104,17 +104,17 @@ declared capability, never an `if (mode === 'claude')` branch.
## Per-CLI injection recipes (confidence-ranked)
| CLI | Mechanism | Confidence |
| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
| CLI | Mechanism | Confidence |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol picture more nuanced than a flat break, re-verified live twice on 2026-09-17 against a llama-swap deployment that DOES answer `/v1/responses`** (an earlier test's `Reconnecting...`/`high demand` failure does not reproduce against every llama-swap setup): a plain, no-tool-call chat turn (`codex exec 'reply with just OK'`) returned a real reply. But a real tool-call attempt (`run the shell command: echo hello`) came back as an `agent_message` TEXT item — the tool-call JSON printed as the model's answer, not a `function_call` item codex would actually execute (confirmed via `codex exec --json`'s raw event stream: `item.completed`/`agent_message`, never `function_call`). Since tool execution is what makes codex a coding agent at all, this remains **not usable for real work**, just with a different, more specific failure mode than previously documented — still do not present this as working. Separately, EVERY custom-endpoint codex session also prints `warning: Model metadata for '<id>' not found. Defaulting to fallback metadata...` on launch (confirmed harmless — the successful plain-text reply above still had it): codex's per-model metadata (reasoning tiers, system-prompt templates, context-window figures) comes from `models_cache.json`, a LOCAL CACHE of OpenAI's own hosted model catalog that a custom model can never appear in by construction. No config.toml override exists for it, and the isolated `CODEX_HOME` never gets a `models_cache.json` written into it at all (confirmed: inspected a live, actively-used isolated dir — codex evidently can't reach OpenAI's catalog endpoint for this session and just falls back silently every time, with no file left behind to fix or clean up). Fabricating a fake catalog entry to suppress the warning would mean copying the _shape_ of OpenAI's own proprietary schema — including their real per-model system-prompt content, visible in a genuine `models_cache.json` — for a warning confirmed to have no effect on the actual (broken) tool-calling outcome; not worth building |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`), now with `appendV1Suffix: true` (see confidence). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Root cause of the original `HTTP_404` found and fixed, by reading dsh's own bundled source — the same bar pi/grok's fixes were held to.** Installed `@deepseek-ai/dsh` (all its real published dependencies) into a scratch directory purely to read `@deepseek-ai/dsh-llm-deepseek/lib/index.js`: it builds its request as `fetch(\`${connection.baseURL}/chat/completions\`, ...)`with`baseURL`read straight from`DEEPSEEK_BASE_URL`(or defaulting to DeepSeek's real public API root,`https://api.deepseek.com`, which also carries no `/v1`) — no `/v1` insertion of dsh's own, unlike the OpenAI-SDK convention this recipe originally assumed. llama-swap/llama.cpp only ever serves the OpenAI-conventional `/v1/chat/completions`. Confirmed live: `POST <baseUrl>/chat/completions` → `404`, `POST <baseUrl>/v1/chat/completions` → `200`, on the exact same endpoint — and dsh's own error-message template, `DeepSeek API error (HTTP ${status})`, reproduces the originally reported `dsh: HTTP_404: DeepSeek API error (HTTP 404)` precisely. Fixed by adding `appendV1Suffix` (env kind only, deepseek's entry alone — claude/gemini must NOT get it, since claude was already confirmed working against the unmodified `baseUrl`), which runs `endpoint.baseUrl` through the same `withV1Suffix()` helper `configDir`-kind CLIs already use. ⚠️ Not yet re-run end-to-end with a real `dsh` binary — no install available in this environment (no npm-installed CLI binary in `PATH`, and the `codeman-test-picker` container doesn't bundle it either); the fix is source-confirmed and live-verified at the HTTP level, but a genuine "hello world" reply through `dsh` itself is the remaining step before promoting this to **verified** alongside claude/opencode/pi/grok/omp |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
Everything web-researched-but-unverified gets implemented but must be
smoke-tested against real installs of those CLIs before being called done —
@@ -208,6 +208,14 @@ extra per-model configuration on Codeman's side at all.
### 4. Toolbar UI
> **Superseded.** This section describes the toolbar-button design as originally
> planned. What actually shipped is a Run-menu picker instead: one generated entry
> per (capable harness, saved endpoint) pair directly in the existing `#runModeMenu`
> dropdown, rather than a separate `#customModelBtn`/`#customModelMenu` surface. See
> [`docs/custom-model-endpoints.md`](custom-model-endpoints.md#the-run-menu-picker)
> for the current design; the sections below (session-restart mechanics, security)
> remain accurate regardless of which UI calls the underlying route.
- New header/toolbar button (e.g. `#customModelBtn`, `btn-toolbar
btn-custom-model`), marker-hidden by default (`btn-custom-model--hidden`)
and revealed by `applyHeaderVisibilitySettings()` only when
@@ -344,8 +352,12 @@ pure unit tests and the live manual checks in Verification:
up automatically with zero edits to the script). Already run to
completion against the author's llama-swap server (a LAN address,
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected**
(Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED**
claude/opencode/pi/grok/omp **PASS**, codex **partially works and still
isn't usable** (plain chat succeeds against a llama-swap deployment that
answers `/v1/responses`, but a real tool-call attempt comes back as
inert text rather than an executable `function_call` — see the
confidence table row for the full, re-verified picture), gemini/deepseek
**UNCONFIRMED**
(reach the server, fail for undiagnosed reasons — see their table rows),
antigravity **SKIP** (no mechanism). Re-run this against a real cloud
endpoint (e.g. an Azure AI Foundry deployment) once one is available, to
+405 -18
View File
@@ -11,20 +11,22 @@ company gateway) — anything answering `GET /v1/models` and
recipe confidence table, and security reasoning:
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
> **Status**: backend is implemented and tested (registry capability, the
> injection engine, the endpoint store + discovery route, the session
> restart route). The toolbar picker / settings UI described below as the
> intended surface is **not yet built** — until it lands, use the HTTP API
> directly (examples below). Antigravity has no known custom-endpoint
> mechanism and is not supported.
> **Status**: fully wired end to end — registry capability, the injection
> engine, the endpoint store + discovery route, both the restart-in-place
> apply route (Claude) and the one-shot quick-start launch path (every
> other supported harness), a settings-panel CRUD surface, and the Run-menu
> picker described below. Antigravity has no known custom-endpoint
> mechanism and is not supported. The HTTP API (examples below) still works
> directly and is what the picker itself calls under the hood.
## Turning it on
App Settings → Agents & CLIs → **Custom Model Endpoints** (synced setting
`customModelEndpointsEnabled`, default **OFF**). Until the toolbar picker
lands, nothing reads this setting: the HTTP routes below work whether it is
on or off, and it exists now only so the picker has a switch to hang off
when it ships. The API equivalent:
App Settings → Models → **Custom model endpoints** (synced setting
`customModelEndpointsEnabled`, default **OFF**). Turning it on does two
things: it reveals the endpoint list/add/edit/discover panel in that same
settings section, and it makes the Run menu offer a generated entry per
(harness, endpoint) pair — see "The Run-menu picker" below. The API
equivalent:
```bash
curl -sk -X PUT https://localhost:3000/api/settings \
@@ -34,6 +36,9 @@ curl -sk -X PUT https://localhost:3000/api/settings \
## Adding an endpoint
Via App Settings → Models → Custom model endpoints → **+ Add endpoint**, or
directly:
```bash
curl -sk -X POST https://localhost:3000/api/model-endpoints \
-H 'Content-Type: application/json' \
@@ -62,7 +67,173 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
Endpoint management is admin-only in multi-user mode, same as remote/docker
hosts — these are machine-level infra, not per-user settings.
## Applying a model to a session
**Context length is discovered too, opportunistically and safely.** The plain
`GET /v1/models` response has no context-window field. Discovery only ever
looks for one for a model llama-swap's own response already reports
`status.value === "loaded"` for — never for an unloaded one, because
llama-swap treats `?model=` as a routing hint and asking about a model that
isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side
effect of what should be read-only discovery. A server with no `status` field
on any entry at all (not llama-swap) gets no context-length enrichment,
rather than guessing. A model's previously-learned context length survives a
later cycle where it wasn't the loaded one; it's dropped only once the model
disappears from the endpoint's list entirely. Stored per model in
`modelContextLengths` and applied automatically (see "Applying a model to a
session" below) so a CLI that would otherwise assume a large default context
window for an unrecognized model id stops silently overflowing a much
smaller real one.
**Where that number actually comes from matters, and got this wrong once
already.** The first cut read it from llama.cpp's own
`GET /props?model=<id>` (`n_ctx`) — plausible, and it worked in testing, but
confirmed live to be actively WRONG for a `--fit-ctx`-launched llama-swap
backend: `/props` reported `n_ctx: 154112` for a model llama-swap itself had
launched with `--fit-ctx 16384`, and the real server then refused a request
right at that real 16384-token limit — `/props`'s `n_ctx` appears to report
the model's theoretical/trained maximum there, not the runtime-configured
one. Discovery now parses the REAL configured size straight out of
llama-swap's own launch command instead (`GET /running`'s `cmd` field —
`--fit-ctx <N>` first, then the plain llama.cpp `-c`/`--ctx-size` a
hand-written command might use), and only falls back to the `/props` probe
when `cmd` states no recognizable flag at all.
**File size is discovered too, when the server states one.** llama-swap
writes a GB figure into an auto-discovered model's own `description`
(`"Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"`), parsed
into `modelSizesGB` — unlike context length, this needs no `/props` probe
(the figure is right there in the `/v1/models` response) and so is populated
for every model regardless of loaded state. A hand-configured profile's own
description has no such figure and correctly gets no entry, never a guess.
Used only to label the Run-menu picker's "loading model" banner (e.g.
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB) on llama-swap..."); never anything
a server-side check relies on.
**The loading banner is unbounded by design, and says so — no countdown, no
automatic give-up.** An earlier version scaled an expected-time estimate and
a timeout off the model's file size and auto-closed the session once that
elapsed, but a real load's actual duration depends on hardware this feature
has no way to know (VRAM, storage speed, whatever else is contending for the
GPU) — any fixed number was a guess dressed up as a fact, and a model that
genuinely takes 10+ minutes on slower hardware would just get killed
mid-load by its own display. The banner now says outright that it can take a
while depending on hardware and model size, polls
`GET /api/model-endpoints/:id/running-status` every second for as long as it
takes, and carries a **Cancel** button (rendered on the banner itself) that
ends the wait and closes the session the load was for — the user's own call
on when it's taking too long, not a fixed number baked into the client.
**The banner's second line is the real backend log line, not a guess.**
llama-swap's `GET /api/events` SSE stream carries the actual `llama-server`
process's own stdout — `load_model: loading model '<path>'`,
`llama_server: model loaded`, tokenizer warnings, all of it — tagged
`source: "upstream"`, distinct from llama-swap's own `source: "proxy"`
request-access lines. `running-status`'s response now includes `logLine`
(via `getLatestLlamaSwapLogLine`), and the banner shows it on its own line
under the disclaimer, e.g. "llama.cpp: load_model: loading model '...'" —
confirmed live end-to-end through a real forced swap, sequentially showing
the model path, a tokenizer warning, then staying on whatever llama.cpp last
printed once the load goes quiet (never cleared back to blank). ⚠️
**`GET /logs` — the endpoint this feature's own first cut was built
against — turns out to carry ONLY llama-swap's own proxy request-access
log.** Confirmed live it never showed a single backend line, even seconds
after a real, verified model swap; `/api/events`'s `logData` frames are the
only source that actually has it, and its own `source` field (`upstream` vs
`proxy`) is what `getLatestLlamaSwapLogLine` filters on. One `/api/events`
connection is held open per endpoint and reused across every session
watching a load on it (confirmed live to stay open indefinitely, unlike
`/logs`, which closes after a fixed ~100KB), idle-closed after 30s of nobody
polling it (`pruneIdleLlamaSwapLogTails`, same 20s sweep as the
swap-displacement check below).
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
from the endpoint's own discovered `models`, and the route refuses a value
that isn't one of them. It is applied automatically only when the endpoint
has exactly one discovered model (nothing to choose); with two or more it
is a pre-selection in the model-picker dialog below, never a silent default.
Re-discovering drops a default that no longer appears in the fresh list
rather than carrying an invalid one forward.
**Model lists refresh themselves.** A background sweep (`server.ts`,
`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every
saved endpoint the same way the manual `POST .../discover-models` route
does, best-effort per endpoint — one being unreachable on a given cycle
never blocks the others. Off under `npm test`, same reasoning as the Codex
plan-usage poll it sits beside: no real network to hit, no server instance
to keep the timer alive for.
## The Run-menu picker
With the setting on and at least one endpoint carrying a discovered model,
the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry
per (harness that can redirect to a custom endpoint, saved endpoint) pair,
e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI
registry's own `capabilities.customModelInjection` at page render
(`window.__codemanCustomModelClis`, `server.ts`) — never a hardcoded id list
in the frontend — so a CLI whose injection recipe lands later shows up with
no frontend change, and Antigravity (`unsupported`) never does.
Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`,
`session-ui.js`) rather than trusting anything cached from the dropdown's
own render — the model list can have changed via the 5-minute sweep above
or a settings-panel edit since the menu opened. With exactly one discovered
model it runs straight away; with two or more, a small modal
(`#customModelPickModal`) lists them and asks which one to use for this
launch, with the endpoint's `defaultModelId` marked but not auto-chosen —
the point of asking is letting one launch deliberately differ from the
saved default, not just confirming it.
**How the launch itself applies the endpoint depends on the harness.** For
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP (`runCustomModelEntry` →
`_runCustomModelEntryOneShot`), the endpoint/model is folded into the SAME
`POST /api/quick-start` call that creates the session (`customModel` field),
so the session launches directly on the endpoint — no restart, no visible
relaunch. Claude (`_runCustomModelEntryViaRestart`) still uses the original
two-step design: the launch runs a single native session exactly the way its
own Run-menu entry would, then **waits for the new session to go idle**
(`GET .../wait?until=idle`, bounded at 20s — a normal 200 either way, never
an error, per the wait endpoint's own contract) before applying the endpoint
via the restart route below. That wait exists because a freshly launched CLI
reports itself as `busy` for its own startup (a boot spinner, a
workspace-trust check) well before the apply call would otherwise reach it,
and the apply route correctly refuses to restart a session mid-turn — a
fresh boot looks exactly like one from the outside. A session still busy
after the wait reaches the apply call anyway and gets that route's own
honest `SESSION_BUSY` error, now visible as a sticky toast with a close
button rather than a generic message that vanished in three seconds. Claude
stays on this path because its own restart (`--resume`-based, keeping the
conversation) is far less jarring than the other seven's, and `runClaude()`'s
multi-tab launch and docker-config-drift confirm/retry loop make folding it
into the one-shot path separate work. It is a
one-off "try this endpoint" action, not a sticky mode: the plain Run button
still means "this harness, native cloud" afterward. Entries are hidden
entirely for a remote or Docker active case, since the apply route refuses
both (see the next section).
## Launching directly on an endpoint (no restart)
```bash
curl -sk -X POST https://localhost:3000/api/quick-start \
-H 'Content-Type: application/json' \
-d '{"caseName": "myapp", "mode": "codex", "customModel": {"endpointId": "llama-box", "modelId": "qwen3"}}'
```
`POST /api/quick-start`'s `customModel` field (`{endpointId, modelId,
confirmed?}`) computes the same injection the restart route below does, but
BEFORE the session exists — the session is minted its own id up front
(`crypto.randomUUID()`), the injection (env vars, and for a `configDir`-kind
CLI, the written config file) targets that real id, and the session launches
already pointed at the endpoint. No restart, because there was never a
native-backend launch to restart away from. Runs the same llama-swap
conflict check as the restart route (below) — a `409`-shaped
`{requiresConfirmation, currentlyLoadedModel, affectedSessions}` response
with no session created, resolved by retrying with `confirmed: true` — and
is refused the same way for a remote or Docker case. This is what the
Run-menu picker uses for opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP;
Claude still uses the restart route below (see "The Run-menu picker" above
for why).
## Applying a model to an ALREADY-RUNNING session
```bash
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
@@ -84,6 +255,178 @@ since for those three the config file alone does not switch the model.
reattaches the durable remote/in-container tmux rather than relaunching the
agent, so the selection would report success and change nothing.
**Claude gets two more env vars when known/applicable, both declared on its
registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is set to `modelId`'s discovered context
length (see the discovery section above) whenever one is known. Without
it, Claude Code assumes a large (200k) window for any unrecognized custom
model id and never compacts, which reliably overflows a much smaller real
local context — confirmed live: a stock ~33.7K-token system prompt against
a 16384-token llama-swap model failed with `exceeds the available context
size`. No entry for the model in `modelContextLengths` means the var is
simply omitted, never a guess. ⚠️ **This var only affects when Claude
Code compacts conversation _history_ — it cannot fix a model whose real
context is smaller than Claude Code's own fixed per-turn overhead**
(system prompt + tool schemas, empirically ~36.4K tokens, confirmed live
via an `in:0 out:0` failure on the very first message, before any
history exists to compact). No context-length declaration changes that
fixed overhead, so a model below the safe floor fails outright on
message one regardless of what this var says. See "Context-window floor
warning" below for how Codeman catches this case before launching
instead of after.
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
the `configDir`-kind CLIs use (empty, no files written into it), so the
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
claude.ai OAuth login. Claude Code still prints "Both claude.ai and
ANTHROPIC_API_KEY set" when the two coexist in the same config directory —
cosmetic (confirmed live: the API key wins for actual requests either way,
visible in the terminal's own `API Usage Billing` line) but worth
eliminating rather than living with. The directory's `projects`
subdirectory is symlinked (a junction on Windows) back to the real
`~/.claude/projects` so the response viewer, subagent windows and Read My
Mind keep working for that session — the same trade-off and fix documented
for a manually-set `CLAUDE_CONFIG_DIR` in
[`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied
automatically here. Best-effort: a platform that refuses the symlink keeps
the pre-existing blind-response-viewer side effect rather than failing the
whole custom-model apply over it. ⚠️ **This relocates the whole `.claude`
tree, not just transcripts**: a custom-model Claude session also loses the
user's global `settings.json`, user-level skills (the codeman agent skill
included), user-level agents and commands, and the MCP servers configured
in `~/.claude.json` — none of those are symlinked back, only `projects` is.
A fine trade for "point this session at my local llama.cpp," but worth
knowing before it surprises you mid-session.
**That isolated directory needed one more fix to actually be usable
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
real profile's prior "Detected a custom API key — use it?" approvals, so
without more, Claude Code stops and asks that on _every single launch_ —
confirmed live, and with nobody at a TTY to answer, its own default answer
("No") silently refuses the very key this feature just injected, which
looks like the endpoint being ignored entirely. `customModelInjection`'s
`apiKeyTrustFile` (`{ relPath: '.claude.json', shape:
'claude-api-key-responses' }` on claude's entry) pre-seeds that exact
approval: the apply step merges `customApiKeyResponses.approved: [apiKey]`
into `<configDir>/.claude.json`, the same field a real answered prompt
itself writes to (confirmed against a real file after answering by hand
once) — this answers the prompt in advance rather than bypassing it. The
merge preserves whatever else the CLI already wrote into that file on an
earlier launch in the same isolated directory (`userID`, `numStartups`,
earlier approved keys), and a missing or corrupt file is treated as empty
rather than failing the apply.
**A fresh `CLAUDE_CONFIG_DIR` isn't just missing that one approval — Claude
Code treats it as a brand-new profile and replays its ENTIRE first-run
sequence on every launch: the theme picker, the security-notes screen, the
per-project "trust this folder?" dialog, and (running with
`--dangerously-skip-permissions`) a one-time warning about bypassing
permissions.** Confirmed live: none of these show up again for a real,
already-onboarded profile, but every custom-model session gets a fresh,
otherwise-empty isolated directory, so it saw all four every single time.
`customModelInjection`'s `skipFirstRunPrompts` (`true` on claude's entry,
requires `apiKeyTrustFile` since it reuses the same file) pre-seeds the
state a real profile accumulates from answering all of that once:
`hasCompletedOnboarding: true` and the launching session's own
`projects[workingDir].hasTrustDialogAccepted: true` go into the same
`<configDir>/.claude.json` the API-key approval above already merges into
(other projects, and other fields on this session's own project entry, are
left untouched), and `skipDangerousModePermissionPrompt: true` goes into
`<configDir>/settings.json` — a different file, merged the same
corrupt-tolerant way. `workingDir` is used exactly as the session was
launched with as its cwd, never realpath'd or slash-normalized, since
that's the literal string Claude Code itself uses as the project key.
**llama-swap gets two more fixes on top of the context-length/config-dir
ones above, both from watching a real switch live.** llama.cpp only ever
runs one model at a time; llama-swap swaps the backing process on demand,
which can take anywhere from a few seconds to well over a minute:
- **The conflict check.** Both apply routes (the restart one here and the
one-shot `POST /api/quick-start` above) call llama-swap's own
`GET /running` first — feature-detected, so a plain llama.cpp/OpenAI-
compatible server (no such endpoint) is simply never checked. If a
_different_ model is currently loaded and ready, and another **live
session's own selection** is using it, the apply returns
`{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}`
instead of silently switching — nothing is applied or created yet.
Retrying with `confirmed: true` skips the check. Switching with nothing
else affected proceeds immediately; this is a warning about disrupting
another session, never a gate on the switch itself.
- **Actually starting the load.** llama-swap has no "switch model" admin
call — the only thing that starts a swap is a real inference request
naming the model, and confirmed live: applying a selection alone never
reached llama-swap at all (nothing in its own server logs), since nothing
had actually asked it to load anything yet. Both apply routes now also
send the smallest real request that will —
`POST <baseUrl>/v1/chat/completions` with `max_tokens: 1` and one
throwaway message — whenever the
target model isn't already the one loaded and ready, fire-and-forget (its
response is never read; `GET /api/model-endpoints/:id/running-status`,
polled client-side, is what actually confirms readiness). The response
also carries `modelSwapInProgress: true` in that case, which is what
drives the Run-menu picker's own "loading model" status banner.
## Catching a swap after the fact
The conflict check above only runs at the moment a session is created or a
model is applied — it has no way to catch a swap that happens **later**.
Confirmed live: a session created while nothing else conflicted at that
exact instant can still get silently displaced afterward, once a
_different_ session's own normal use (or its own create-time load trigger)
asks llama-swap to load something else. llama-swap has no push
notification of its own for this, so a background sweep
(`detectCustomModelSwapDisplacements`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS`
= 20s in `server.ts`) polls `GET /running` once per distinct endpoint that
has at least one live custom-model session, and compares each such
session's own `modelId` against what is actually loaded. A session whose
model is no longer in that list gets a `custom-model:swapped-out` SSE event
(`{sessionId, sessionName, endpointId, previousModel, currentlyLoadedModel}`),
shown as a global toast — global rather than tied to that session's tab,
since the whole point is telling the user before they type into it
expecting the model they picked. Notifies **once per displacement**: the
same de-dupe `Set` clears a session's flag once its own model is loaded and
ready again, so a later, genuinely new displacement notifies again rather
than the session staying silently un-notified forever after the first one.
## Context-window floor warning
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
empirically ~36.4K tokens) can exceed a small local model's _entire_ real
context on its own, before any conversation history exists to fill it —
confirmed live twice, both as an `in:0 out:0` failure on the very first
message sent. `CLAUDE_CODE_MAX_CONTEXT_TOKENS` (above) cannot fix this: it
only governs when Claude Code compacts conversation history, and there is
no history yet on message one. Applying such a model would look like the
endpoint being ignored, or the wrong model being used, when in fact the
endpoint applied correctly and the model is simply too small for this CLI.
Both apply routes (the restart route and the one-shot `POST
/api/quick-start`) now check for this **before** launching or restarting
anything, gated on the CLI's registry entry declaring a `contextLengthVar`
(currently only claude — the check is a no-op for every other CLI by
construction, never a hardcoded mode check). If the model's discovered
context (`modelContextLengths`, from discovery above) is below
`CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000, comfortably above the measured
~36.4K overhead), the response is `{requiresContextWarning: true, modelId,
contextLength, minSafeContextTokens}` instead of applying — nothing is
restarted or created yet. A context length that was never discovered at
all skips the check entirely (nothing to compare, so it fails open rather
than warning on every model an endpoint hasn't reported a size for).
Retrying with `confirmed: true` launches anyway.
The Run-menu picker shows this as an in-app modal
(`#customModelContextWarningModal`, matching the llama-swap conflict
modal's look) naming the model, its discovered context, and the safe
floor, and explaining the fix: reconfigure llama-swap to give that model
(or a smaller one) an explicit larger context instead of relying on
auto-fit (`--fit-ctx`), which optimizes for the biggest _model_ that fits
rather than the biggest _context_ — e.g. adding `-c 65536` (or as large a
`--ctx-size` as the hardware holds) to that model's llama-swap config
entry. A smaller model at a much larger explicit context often fits in
the same VRAM a bigger model's auto-fit context gets shrunk to make room
for.
Clear back to the harness's native cloud default with:
```bash
@@ -101,6 +444,14 @@ id, model and injected key NAMES are persisted, the values are re-derived
from the endpoint store on recovery, and the pane keeps running against the
endpoint in between because tmux retains its environment.
⚠️ Clearing removes injected keys **by name**, and `CLAUDE_CONFIG_DIR` is one
of the names claude's selection injects — so a session that ALSO had
`CLAUDE_CONFIG_DIR` set through the generic `envOverrides` field (the
per-client-account case) loses that override on clear too, and silently
falls back to the server's default Claude account. If you route a session
to a specific account this way, re-apply the override after clearing a
custom-model selection from it.
**New sessions always default back to the harness's native backend.** A
custom-endpoint selection is a per-session choice, never a sticky global
default — starting a fresh session doesn't inherit whatever the last one was
@@ -115,17 +466,53 @@ automatically). Results:
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
came back through the endpoint.
- **Codex** — the config is structurally correct, but Codex only speaks the
Responses API since Feb 2026, which llama.cpp/llama-swap don't implement.
This is a real protocol incompatibility, not a bug here; Codex support
needs a Responses-API-compatible endpoint.
- **Codex** — the config is structurally correct, and against a llama-swap
server that DOES answer `/v1/responses` (confirmed live: a plain,
no-tool-call chat turn returned a real reply), the picture is more
nuanced than a flat failure. A real tool-call attempt (`run the shell
command: echo hello`) came back as `agent_message` TEXT — literally the
tool-call JSON printed as the model's answer — instead of a
`function_call` item Codex would actually execute (confirmed via `codex
exec --json`'s raw event stream). So plain chat can work while the thing
that makes Codex a coding agent — actually running commands and editing
files — does not; treat Codex as still unreliable for real work against a
llama.cpp/llama-swap endpoint, tool-calling gap included, not just the
earlier-documented `wire_api` mismatch (which not every deployment hits
the same way — some legitimately have no `/v1/responses` route at all).
Separately, EVERY custom-endpoint Codex session prints `Model metadata
for '<id>' not found. Defaulting to fallback metadata...` on launch —
confirmed harmless (the reply above still came back correctly): Codex's
model metadata (reasoning-tier options, per-model system-prompt
templates, context-window figures) comes from `models_cache.json`, a
local cache of OpenAI's own hosted model catalog that a custom local
model can never appear in by construction, since it isn't one of
OpenAI's models. There's no config.toml override for a model's metadata,
and fabricating a fake catalog entry would mean copying the _shape_ of
OpenAI's own proprietary schema (their per-model system-prompt content
included) for a warning that doesn't otherwise affect behavior — not
something to build into discovery.
- **Gemini** — fails with `Invalid auth method selected`, traced to an
undocumented `GATEWAY` auth path gemini-cli selects once
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
(several auth workarounds were tried and ruled out); do not rely on
Gemini support yet.
- **DeepSeek** — the request reaches the server (env vars are read) but
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
- **DeepSeek** — root cause of the `HTTP_404` found and fixed. DeepSeek
Harness's own bundled provider module (`@deepseek-ai/dsh-llm-deepseek`)
builds its request URL as `${DEEPSEEK_BASE_URL}/chat/completions` with no
`/v1` insertion of its own (its real public API, `https://api.deepseek.com`,
expects the caller's base URL to already carry any needed prefix) —
confirmed by reading its own source and, live, that
`POST <baseUrl>/chat/completions` 404s against llama-swap while
`POST <baseUrl>/v1/chat/completions` succeeds; the harness's own error
template (`DeepSeek API error (HTTP ${status})`) matches the originally
reported symptom exactly. `customModelInjection`'s new `appendV1Suffix`
(deepseek's entry only — claude/gemini must NOT get it, since claude was
already confirmed working against the raw `baseUrl`) fixes it by writing
`DEEPSEEK_BASE_URL` with `/v1` appended. Not yet re-run end-to-end with a
real `dsh` binary (no install available in this environment) — the fix
is source-confirmed and live-verified at the HTTP level, but a real
"hello world" reply through `dsh` itself is still outstanding before
calling this fully verified like the harnesses above.
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind
+7
View File
@@ -273,9 +273,16 @@ into the case's `.claude/settings.local.json` so that `/model` keeps working.
- **Shell** for the times you want a terminal on your phone with no agent at all. It is a
genuinely useful mode, not a fallback.
## Pointing one at your own server
Most of these harnesses can also run against a custom OpenAI-compatible endpoint instead of
their native cloud backend, for one session at a time, an opt-in feature covered in full on
[Custom Model Endpoints](Custom-Model-Endpoints).
## Read next
- [Core Concepts](Core-Concepts) - run modes versus location overlays.
- [Custom Model Endpoints](Custom-Model-Endpoints) - run a harness against your own server.
- [Settings Reference](Settings-Reference) - model, effort, and permission-mode settings.
- [Keeping Agents Running](Keeping-Agents-Running) - what idle detection does per mode.
- [Security](Security) - what skipping permission prompts actually means.
+183
View File
@@ -0,0 +1,183 @@
# Custom Model Endpoints
Point a harness at your own OpenAI-compatible server instead of its native cloud backend, for
one session at a time. "Custom endpoint" covers **local** hardware (llama.cpp, Ollama, vLLM,
a home GPU rig, DGX Spark, Strix Halo) and **cloud** services (Azure AI Foundry's
OpenAI-compatible endpoint, OpenRouter, a company gateway) alike, anything answering
`GET /v1/models` and `POST /v1/chat/completions` in the standard shape.
**Off by default.** Turn it on in App Settings → Models → **Custom model endpoints**.
## Adding an endpoint
Still in App Settings → Models → Custom model endpoints:
1. **+ Add endpoint** — give it an id, a label, and the base URL (`http://192.168.1.50:8080`,
say). An API key is optional; most local servers don't check one.
2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it.
3. Pick a **default model** from what was discovered. This is the model the Run-menu entry
applies directly when only one model is discovered; with two or more, it's just the one
pre-marked in the picker dialog described below, not a silent default.
Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker
hosts — these are machine-level infra, not a per-user setting.
**Model lists refresh themselves.** Every saved endpoint is re-discovered automatically every
5 minutes in the background, so a model the server starts serving later — or stops serving —
shows up without another manual click of **Discover**. One endpoint being unreachable on a
given cycle (powered off, wrong network) never blocks the others from refreshing.
**Context length is picked up automatically where it can be, safely.** Against a
llama.cpp/llama-swap server, discovery also learns each _currently loaded_ model's real
context window and applies it to the launched session (Claude Code today — see below), so
the harness stops assuming a large default window for a model name it doesn't recognise and
overflowing a much smaller real one. It's deliberately never probed for a model that isn't
already loaded, since asking a llama-swap server about an unloaded model can trigger an
actual, slow model swap as a side effect — a model just not currently loaded keeps whatever
context length an earlier cycle already learned for it instead.
## Running a session against one
With the setting on and at least one endpoint carrying a discovered model, the **Run**
dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a
custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a
session on that harness exactly the way its own entry would. It is a one-off "try this
endpoint" action, not a sticky mode — the plain **Run** button still means "this harness,
native cloud" afterward, and a fresh session never inherits whatever the last one was
pointed at.
**Which model it uses depends on how many the endpoint has discovered.** With exactly one,
the session launches straight away on that model — nothing to choose. With two or more, a
small dialog asks which one to use for this launch before starting the session; the
endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately
use a different one without changing the saved default.
**For opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP, picking an entry launches
straight onto the endpoint** — no restart, because the endpoint is applied before the
session's process ever starts. **Claude still restarts the harness's process in place** —
same tab, same conversation (`--resume`) — after a normal native launch, since that restart
is far less jarring for Claude than for the other seven, whose own TUI can fully
reinitialize on a restart. Either way, every supported harness reads its endpoint config at
process start, never per turn, so there is no live hot-swap while a turn is running.
Picking an entry that launches a **brand-new** Claude session waits (up to 20 seconds) for it to
finish its own startup before applying — a freshly started CLI reports itself as busy for its
boot sequence, and applying to a genuinely busy session is refused so a real, in-progress
turn is never interrupted out from under you. A session that is still busy after that wait
(a very slow-starting CLI, or one you started typing into right away) surfaces that refusal
as an ordinary error, which now stays on screen with a close button instead of vanishing
after a few seconds — read it, it names the actual reason rather than a generic failure.
Entries are hidden entirely for a session in a **remote (SSH) or Docker case** — support for
redirecting those hasn't landed yet, see below. The picker also only appears in the desktop
**Run** dropdown; the phone home screen builds its own run picker separately and does not
currently offer these entries.
**Against llama-swap, applying a selection also starts the actual model load, rather than
waiting on your first prompt to do it.** llama-swap has no "switch model" button of its own
— the only thing that starts a swap is a real request naming the model, and confirmed live:
just applying a selection never reached llama-swap's own logs at all until something asked
it to load. Picking an entry now also sends the smallest real request that will trigger
that load, in the background, the moment the target model isn't already loaded and ready.
**The centred loading banner has no countdown and no automatic timeout — it waits as long as
it takes, and tells you so.** When it knows the model's discovered file size (its GB figure,
when llama-swap states one) it's shown too, e.g. "Loading qwen3.8-27b (16.4 GB) on
llama-swap — this can take a while depending on your hardware and the model size." An
earlier version tried to estimate and enforce a time limit, but real load time depends on
hardware this feature has no way to know, so a fixed number was always a guess — worse, one
that could kill a genuinely slow load partway through. If it really is taking too long, a
**Cancel** button right on the banner ends the wait and **closes the session that load was
for**, on your own call rather than a guessed deadline.
**The banner also shows a real, live second line of what llama.cpp itself is doing** — not
a made-up progress phase, the actual next line the `llama-server` process printed, e.g.
"llama.cpp: load_model: loading model '/models/.../Qwen3.8-27B.gguf'" then later
"llama.cpp: llama_server: model loaded". It comes straight from llama-swap's own event
feed, filtered down to just the backend process's own output (not llama-swap's own request
logging), and stays on whatever it last said once the load goes quiet, rather than
clearing back to nothing.
**You'll also be told if a session's model gets swapped out from under it later, not just
at launch.** The conflict warning above only fires at the moment you launch or apply a
model — llama.cpp only runs one model at a time, so if a DIFFERENT session using the same
endpoint later triggers its own load, whatever was loaded before (including a session you
already had running) gets silently evicted, with no warning at that instant since nothing
conflicted when it was first set up. A background check (every 20 seconds) catches this
after the fact and shows a toast naming which session lost its model and what's loaded now
— so you know before typing into that session that it's about to reload (and, in turn,
evict whatever displaced it).
**Claude Code specifically gets three extra fixes applied automatically:**
- Its discovered context length (see above) is passed through as
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`, so it doesn't send a full-size prompt against a much
smaller real local context and overflow it.
- Its session runs with an isolated `CLAUDE_CONFIG_DIR`, so the injected API key never sits
in the same directory as a stored claude.ai login — that combination is harmless for actual
requests (the API key wins) but the CLI still prints a "both claude.ai and
ANTHROPIC_API_KEY set" warning about it, which this avoids entirely. The isolated directory
keeps a link back to your real session history so the response viewer and similar features
still work for that session. That isolated directory starts with no prior approvals of its
own, so Codeman also pre-approves the injected key the same way answering Claude Code's own
"Detected a custom API key" prompt once would — without it, that prompt would otherwise
reappear on every single launch with nobody there to answer it.
- **That same fresh isolated directory also looks like a brand-new Claude Code profile**, so
without this fix it replayed the WHOLE first-run sequence every single launch: the theme
picker, the security-notes screen, the "trust this folder?" dialog, and a one-time warning
about running with permissions bypassed — none of which a real, already-used profile shows
again. Codeman now pre-seeds that same "already been through this once" state (onboarding
completed, this session's own project marked trusted, the bypass-permissions warning
acknowledged) so a custom-model launch reaches the actual conversation exactly as fast as a
native cloud one does, instead of stopping at a wizard with nobody there to click through it.
**If a model's real context is too small for Claude Code to even get started, you get a
warning instead of a confusing failure.** Claude Code's own system prompt and tools take up
roughly 40K tokens on their own, before you've typed anything — a small local model with a
smaller real context than that fails outright on the very first message, no matter what
context size Codeman tells it to expect (raising the declared context only changes when
Claude Code trims _conversation history_, and there is none yet on message one). Picking
such a model now shows an in-app dialog naming the model, its discovered context and what's
needed, before anything launches or restarts, with the fix spelled out: reconfigure
llama-swap to give that model (or a smaller one) an explicit larger context instead of
relying on auto-fit (`--fit-ctx`), which sizes the context around fitting the biggest model
rather than the biggest context — for example adding `-c 65536` to that model's llama-swap
entry. "Launch anyway" is still there if you want to try regardless.
## Which harnesses actually work
| Harness | Status |
| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
| **Codex** | Config is correct, and plain chat can work against a server that speaks the Responses API — but a real tool-call attempt comes back as inert text instead of running, so it's still not usable for real coding work. |
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
| **DeepSeek** | The original 404 is root-caused and fixed (DeepSeek Harness's own code was missing a `/v1` most local servers require) — not yet re-run against a real `dsh` install to confirm end-to-end. |
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry,
not a fixed list here, so this table can go stale before this page does — a greyed-out or
missing entry is the more current answer.
## What it does not do
- **No remote or Docker sessions yet.** Both restart their agent differently under the hood
(reattaching a durable tmux session rather than relaunching the process), so redirecting
them needs its own plumbing that hasn't been built.
- **No live hot-swap mid-conversation.** Applying a selection always restarts the process.
- **No button to un-point a session from the UI yet.** Clearing back to native cloud is an
HTTP call (`POST .../custom-model {"clear": true}`) or deleting the session; the settings
panel manages saved endpoints, not what a running session is currently pointed at.
- **Nothing is shared with your real cloud credentials.** The endpoint's own key, if any,
never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate,
explicit choice per session.
## Security
An endpoint's base URL can't point at a link-local or cloud-metadata address (both at save
time and against the address it actually resolves to), the same guard Web Tabs uses for
saved dashboards. Endpoint records and any per-session config files a harness needs are
written with owner-only permissions. See
[custom-model-endpoints-plan.md](https://github.com/Ark0N/Codeman/blob/master/docs/custom-model-endpoints-plan.md)
in the repository for the full design reasoning, including why this feature closed a
pre-existing gap in how session environment overrides were guarded rather than opening a new
one.
+4
View File
@@ -92,6 +92,10 @@ Model and effort are both **soft defaults**: the model is written into the case'
`.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort`
inside a session override them at any time.
**Custom model endpoints** (off by default) adds a saved-endpoint list plus a matching
section to the Run dropdown, for pointing a harness at your own OpenAI-compatible server
instead of its native cloud backend. See [Custom Model Endpoints](Custom-Model-Endpoints).
### Agents & CLIs
| Setting | Notes |
+1
View File
@@ -12,6 +12,7 @@
- [The Dashboard](The-Dashboard)
- [Agent CLIs](Agent-CLIs)
- [Custom Model Endpoints](Custom-Model-Endpoints)
- [Working With Files](Working-With-Files)
- [Input And Voice](Input-And-Voice)
- [Mobile Guide](Mobile-Guide)