mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. Covers local hardware (llama.cpp, Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter). Off by default (customModelEndpointsEnabled, synced, default OFF). - Registry: capabilities.customModelInjection per CLI entry (env / configContentEnv / configDir / unsupported kinds) - Pure injection builder (custom-model-injection.ts) turning an endpoint + model id into the real env vars / config content per CLI - Endpoint store + CRUD routes (custom-model-hosts.ts, custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded - Session integration: Session.setCustomModel()/restartCli() (POST /api/sessions/:id/custom-model), reusing the existing respawn-pane -k primitive to restart the CLI process with new env - Multi-user hardening: every new redirect-capable env var added to its CLI's privilegedEnvKeys, closing a pre-existing gap where several were already reachable via the generic envOverrides field's prefix allowlist - Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries against a real endpoint outside the web UI, independent of tmux/sessions - Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying every CLI's injected values through a real HTTP shape Real end-to-end validation against a live llama-swap server (inside a codeman/agent:llm-test Docker image with all 9 CLI binaries) found and fixed three real bugs before they shipped: - Codex's config.toml schema was wrong ([model].default table instead of a top-level model string + [model_providers.custom]); fixing it then surfaced a genuine, documented protocol incompatibility (Codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement) - Claude Code's async session-title-generation call validates ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and hangs the whole -p invocation on an unrecognized name; documented for chunk 6, worked around in the standalone script only (--bare is NOT safe for a real interactive session, which needs hooks) - The discovery route's authStyle: 'both' option (send both Authorization and api-key headers) reliably hung a real server; removed the option entirely rather than just changing the default Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see PR.md and deployment_plan.md for the full chunk breakdown and confidence table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
a017e9a8e0
commit
41416566aa
@@ -105,3 +105,7 @@ readme-preview.mjs
|
||||
|
||||
# Uploaded images land here under each session working dir (runtime artifact)
|
||||
.claude-images/
|
||||
|
||||
# Local-LLM harness smoke-test config (real IPs/keys) — see the .example.json
|
||||
# alongside it in scripts/, which IS tracked as the template.
|
||||
scripts/local-llm-test.config.json
|
||||
|
||||
@@ -225,6 +225,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
|
||||
|
||||
**DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. The session version worked and was still wrong in use: it put a terminal tab on screen next to the web tab the user actually asked for, every single time, and nothing about a long-lived HTTP server needs to be a tab. ⚠️ What a session gave for free now has to be paid for explicitly, and every piece is load-bearing: **exactly one** server (a second click REUSES it rather than racing it for a port, which two sessions structurally could not do), **restarted when the browser authority changes** (`--trusted-host` fences dsh's `/api` against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two, so whoever asks last wins: the asker is by definition the origin about to load the page), **killed on shutdown** (`stopDeepSeekWeb()` in the server teardown, because the child is detached so its whole plugin tree can be signalled at once, which also means it would OUTLIVE Codeman and hold its port against the next start), and **failures returned to the caller**, since with no tab there is nowhere for a stack trace to land. ⚠️ The port search starts at dsh's own default 3080 and walks 40, never fixed: that default is precisely the port most likely to be taken already by the user's own `dsh web`, and hardcoding it killed this feature with EADDRINUSE once. Free-port detection BINDS rather than connects (a connect probe cannot tell "free" from "listening but not answering yet"), so it is racy by nature and the caller still waits for the server to really answer before reporting success. ⚠️ Both `POST` and `DELETE` sit at the **same privilege bar as the profile installer** (`canUsernameRunPrivilegedCommands`) even though the action reads as "open a page": booting a dsh profile executes the plugin code in it, and the server is a single shared instance, so stopping it in multi-user mode takes it out from under other users' tabs.
|
||||
|
||||
**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `deployment_plan.md`): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET <baseUrl>/v1/models`; `CustomModelHost.authStyle` (default `'both'`) sends BOTH `Authorization: Bearer` and `api-key` headers on discovery since cloud gateways (Azure) and local servers (llama.cpp) disagree on the convention and there is no way to know in advance which one a given endpoint wants. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/grok/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi` config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `PI_CONFIG_DIR`, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. Confidence is per-CLI: claude/opencode/codex are hand-verified against a real llama.cpp/llama-swap server; gemini/pi/grok/deepseek/omp have their ONE-SHOT INVOCATION flags confirmed against real installed binaries' `--help` output, but their custom-endpoint env/config conventions remain unverified — see the confidence table in `deployment_plan.md`. The standalone `scripts/test-local-llm-harnesses.mjs` (reads a gitignored `scripts/local-llm-test.config.json`, template `.example.json` tracked) smoke-tests real CLI binaries against a real endpoint outside the web UI entirely, independent of tmux/sessions.
|
||||
|
||||
**Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w<n>-<case>` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. ⚠️ **Closing has the mirror-image race and one owner**: `closeSession()` reads `wasActive` BEFORE its `await` and announces the delete via `_closingSessions`, while `_onSessionDeleted` skips the active-session handoff for an id in that set. Both used to read `activeSessionId` after the fact, so the `session_deleted` broadcast for your own delete could null it first and closing the tab you were on landed on the welcome screen instead of the next session, on the same build, depending on timing. The fallback also picks the first order entry that is still in `sessions` (a dead id can linger in `sessionOrder`, same reason Alt+N indexes a live-filtered list). A delete from ANOTHER client still shows the welcome screen, which is the honest answer when what you were looking at was taken away. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization)
|
||||
|
||||
**Session lineage lines** (tab → tab it spawned, `sessionLineageLines`, per-device, desktop default ON): a create request may name the session that spawned it, as a `parentSessionId` body field on `POST /api/sessions` / `POST /api/quick-start` or the `X-Codeman-Parent-Session` header (the agent skill sets that once on its shared curl invocation, so every spawn recipe carries it). `resolveParentSessionId()` (route-helpers.ts) **resolves rather than trusts** it: exact id, else a UNIQUE ≥8-char prefix (ids reach agents truncated), it must be a live session the caller can see AND carry the same owner, and **anything unresolvable is DROPPED, never a 400** — a cosmetic field must not be able to fail a worker spawn. It rides `toState()` into `session_created`, so there is no new SSE event. ⚠️ Rendering is an ADDITIONAL LAYER on the existing SVG pass (`_appendLineageConnectionLines` called at the tail of `_updateConnectionLinesImmediate()`, exactly like ultracode), sharing one batched read→write reflow and the `tab:<id>` rect cache; geometry is pure in `computeLineagePath()` (constants.js). ⚠️ **ONE shape, and the second one was the bug**: every pair (flat strip or wrapped) gets a U-bridge hanging below the strip, anchored on both tabs' BOTTOM edges. A wrapped strip used to get a parent-bottom → child-TOP bezier with a ~14px row gap to bend in, which drew a flat line hidden in the gap with siblings overprinting. ⚠️ The dip is a **mis-tuned-in-both-directions corridor** (44px cap = straight thread at strip-wide spans, #285; 104px cap + full row offset = ~106px over-bow into the terminal, 2026-08-15): it now hangs from the **STRIP's bottom edge** (fallback: lower tab bottom), capped at 64px, with NO per-row offsets stacked on top — the strip-bottom baseline is also what keeps a row-1 pair's arc from drawing through row 2's tab labels. ⚠️ **Colors are keyed on the SPAWNING tab, not per child**: every arc leaving one tab is the same color however many workers it spawns, so the strip reads as "these five came from w1, those two came from w2" — per-child coloring gave one tab's own children a different color each, which is the distinction the colors exist to make. A child that spawns in turn is a parent in its own right and gets its own color for the arcs below it, so a chain changes color at each generation while each generation's fan-out stays uniform. Assignment cycles `CodemanLineage.COLORS` in first-seen order per parent id (first entry empty = the skin-tuned `--session-blue`, so the first spawning tab keeps it; the rest vivid fixed hexes), memoized rather than derived from draw index (the SVG is wiped and rebuilt constantly, so an index-based color would flicker), and set inline as `--lineage-color` so styles.css keeps owning opacity/glow/dash. `test/session-lineage-lines.test.ts` drives the real `_appendLineageConnectionLines()` and asserts the painted property, since testing the color function alone would pass just as happily with the child id passed back in. ⚠️ **Desktop only**: the overlay is `z-index: 999` and the desktop header is 100 (arcs paint over it, which is what lets them touch tab bottoms), but under 1024px mobile.css makes the header `fixed; z-index: 1200` and would bury them. ⚠️ Paths carry `data-agent-id="lineage:<childId>"` because that is what `_applyLineEntrances()` queries — that one attribute is what gives them the entrance animation and its negative-`animation-delay` resume across `svg.innerHTML=''`. ⚠️ `.session-tabs` is `overflow-x: auto`, so a scrolled-out tab still HAS a rect (over the logo); edges with an endpoint outside the strip are skipped, and a passive `scroll` listener re-anchors the rest.
|
||||
|
||||
@@ -0,0 +1,218 @@
|
||||
# feat: Custom Model Endpoint Profiles (local or cloud, all harnesses)
|
||||
|
||||
> **⭐ Shout-out up front:** this feature was partly inspired by — and is a
|
||||
> great fit for — **[Ark0N/Qwen5090](https://github.com/Ark0N/Qwen5090)**,
|
||||
> the maintainer's other project: a one-click Windows / one-command Linux
|
||||
> installer that stands up Qwen3.8-27B locally on an RTX 5090 behind an
|
||||
> OpenAI-compatible API (vLLM / NInfer / llama.cpp). Once this feature lands,
|
||||
> pointing Codeman at a Qwen5090 box is just adding one endpoint entry — no
|
||||
> extra code, no special-casing. Qwen5090 already wires up DeepSeek Harness
|
||||
> and Claude Code as local coding agents itself, which is basically this
|
||||
> feature's idea in miniature. 🙂
|
||||
|
||||
**Status: draft / work-in-progress.** This PR is not ready to merge — see
|
||||
[Status](#status) below for exactly what's done and what's still open.
|
||||
|
||||
---
|
||||
|
||||
## What
|
||||
|
||||
Adds a settings-gated (default **OFF**) way to point any Codeman-supported
|
||||
harness — Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP, or
|
||||
Antigravity — at a **custom OpenAI-compatible endpoint** instead of its
|
||||
native cloud backend, for a given session. "Custom endpoint" covers both:
|
||||
|
||||
- **Local hardware**: llama.cpp, Ollama, vLLM, a home GPU rig, or
|
||||
purpose-built on-prem boxes like NVIDIA DGX Spark, AMD Strix Halo
|
||||
(Ryzen AI Max) mini-PCs, or the [Qwen5090](https://github.com/Ark0N/Qwen5090)
|
||||
setup above.
|
||||
- **Cloud**: Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
|
||||
company's self-hosted gateway.
|
||||
|
||||
The user adds an endpoint by base URL (+ optional API key), Codeman
|
||||
discovers its available models via `GET /v1/models`, and a new toolbar
|
||||
picker lets them apply one of those models to a session — which then
|
||||
restarts that session's CLI process pointed at the endpoint.
|
||||
|
||||
**Why discovery instead of asking the user to type a model name:** it turns
|
||||
"go read your inference server's docs to find the exact model identifier it
|
||||
expects" into "pick from a list Codeman already fetched" — one less place
|
||||
for a user to get a name/casing wrong and have a harness fail with an
|
||||
opaque "model not found." It also means this feature works unmodified
|
||||
against **multi-model hosting setups**, not just a single-model server: a
|
||||
gateway like **[llama-swap](https://github.com/mostlygeek/llama-swap)**
|
||||
(or vLLM/LiteLLM/Ollama serving several loaded/loadable models behind one
|
||||
`/v1/models` list) already advertises every model it can hot-swap to, so
|
||||
the toolbar picker becomes a live menu of everything that endpoint can
|
||||
serve — no per-model endpoint entries, no separate configuration step,
|
||||
just "add the gateway once, everything behind it shows up."
|
||||
|
||||
## Why
|
||||
|
||||
The maintainer pays for a Claude Code subscription but also runs a capable
|
||||
local model. Every harness Codeman drives already *has* its own mechanism
|
||||
for pointing at a custom endpoint (env vars for Claude, a JSON config blob
|
||||
for opencode, a TOML file for Codex, etc.) — Codeman just never exposed a
|
||||
UI for it. Full motivation, the per-CLI recipe table, and the on-prem
|
||||
hardware use cases are written up in **[`deployment_plan.md`](deployment_plan.md)**.
|
||||
|
||||
## How
|
||||
|
||||
- **`src/config/cli-registry/{types,schema,stock}.ts`** — new
|
||||
`capabilities.customModelInjection` field per CLI entry, one of four
|
||||
kinds: `env` (Claude, Gemini, Grok, DeepSeek), `configContentEnv`
|
||||
(opencode, reusing its existing `OPENCODE_CONFIG_CONTENT` mechanism),
|
||||
`configDir` (Codex/Pi/OMP — writes an isolated config file, never touches
|
||||
the user's real one), or `unsupported` (Antigravity — no known mechanism,
|
||||
toolbar entry stays disabled). Declared data-driven per the repo's
|
||||
existing "never branch on CLI id" rule.
|
||||
- **`src/custom-model-injection.ts`** — pure function turning
|
||||
`(CliEntry, endpoint, modelId)` into the real env vars / config content.
|
||||
No IO; a caller writes `configDir` files to disk.
|
||||
- **`src/custom-model-hosts.ts`** + **`src/web/routes/custom-model-routes.ts`** —
|
||||
read/write-array endpoint store (`~/.codeman/custom-model-hosts.json`,
|
||||
same shape as `remote-hosts.ts`) and `GET/POST/PUT/DELETE
|
||||
/api/model-endpoints` + `POST /:id/discover-models`, admin-gated in
|
||||
multi-user mode, SSRF-guarded via the same `isBlockedWebviewUrl()` check
|
||||
web tabs use.
|
||||
- **`src/web/schemas.ts`** — `customModelEndpointsEnabled` (synced, default
|
||||
OFF) + the endpoint payload schema.
|
||||
- **`scripts/test-local-llm-harnesses.mjs`** — standalone smoke-test script
|
||||
that spawns each real CLI binary one-shot against a real endpoint and
|
||||
checks it can answer "hello world," independent of the web UI. Reads
|
||||
defaults from a gitignored `scripts/local-llm-test.config.json` (see the
|
||||
committed `.example.json`) so real IPs/keys never land in git.
|
||||
|
||||
### A finding along the way: multi-user privilege hardening
|
||||
|
||||
Building this surfaced that several env vars (`GOOGLE_GEMINI_BASE_URL`,
|
||||
`GROK_BASE_URL`, `CODEX_HOME`, `PI_CONFIG_DIR`, `OPENCODE_CONFIG_CONTENT`,
|
||||
etc.) were **already** reachable via the generic `envOverrides` API today,
|
||||
pre-existing this PR, because Codeman's env allowlist is prefix-based and
|
||||
global. A non-granted multi-user owner could already redirect a session's
|
||||
endpoint/credentials via a plain `envOverrides` field. This PR adds all of
|
||||
them to their CLI's `privilegedEnvKeys` (the existing clamp mechanism
|
||||
`DEEPSEEK_BASE_URL` already used), closing that gap rather than widening it.
|
||||
`CODEX_HOME` and `PI_CONFIG_DIR` are flagged as extra-sensitive: a
|
||||
redirected config dir can restate approval/sandbox policy or, for Pi,
|
||||
redirect to a dir Pi will execute `.pi/extensions` TypeScript from.
|
||||
|
||||
Claude is the deliberate exception: `ANTHROPIC_*` is **not** added to
|
||||
Claude's allowed env prefixes at all, so it stays reachable only through
|
||||
the dedicated, admin-configured, SSRF-guarded custom-model route — never
|
||||
through a plain client-supplied `envOverrides`.
|
||||
|
||||
## Status
|
||||
|
||||
Built in reviewable chunks; ✅ = done and verified (typecheck + lint +
|
||||
format + tests green), ⬜ = not started.
|
||||
|
||||
- ✅ **1. Registry types** — `customModelInjection` capability shape
|
||||
- ✅ **2. Pure injection builder** — `custom-model-injection.ts` + 15 unit tests
|
||||
- ✅ **3. Endpoint store + CRUD routes** — `custom-model-hosts.ts`,
|
||||
`custom-model-routes.ts`, discovery + SSRF guard, 7 route tests
|
||||
- ✅ **4. Settings + security hardening** — `customModelEndpointsEnabled`
|
||||
flag, `privilegedEnvKeys` additions across 7 CLI entries
|
||||
- ✅ **5. Session integration** — `session.customModel` state field,
|
||||
`session.setCustomModel()`/`session.restartCli()` (a generalized,
|
||||
de-restricted `reattachRemote()` reusing the existing `respawn-pane -k`
|
||||
primitive), `POST /api/sessions/:id/custom-model` restart route. 5 new
|
||||
route tests; the existing `session.test.ts`/`session-cleanup.test.ts`
|
||||
suites can't run at all on this Windows dev box (no local `tmux` —
|
||||
confirmed identical on unmodified `master`, not a regression), which is
|
||||
exactly why the container test below matters.
|
||||
- ⬜ **6. Frontend** — settings group, toolbar picker, tab badge
|
||||
- ✅ **7. Mock-server contract tests** — `test/fixtures/mock-openai-server.ts`
|
||||
and `test/custom-model-injection-contract.test.ts`, 10 tests replaying
|
||||
every CLI's real injected values through an HTTP call shaped the way that
|
||||
CLI sends it, against an in-process fake server
|
||||
- ✅ **8. Docs** — `docs/custom-model-endpoints.md` (user guide, HTTP-API-only
|
||||
until chunk 6 lands) + a CLAUDE.md pointer bullet
|
||||
|
||||
Also done outside the chunk list: the standalone
|
||||
`scripts/test-local-llm-harnesses.mjs` smoke-test script + its gitignored
|
||||
config file, the on-prem-hardware use-case writeup in `deployment_plan.md`
|
||||
(DGX Spark, Strix Halo, Qwen5090), and a `codeman/agent:llm-test` Docker
|
||||
image (all 9 CLI binaries, built from `docker/agent.Dockerfile`) for the
|
||||
real end-to-end test against a live llama-swap server.
|
||||
|
||||
## Testing performed so far
|
||||
|
||||
- `npm run typecheck` — clean after every chunk
|
||||
- `npm run lint` / `npx prettier --check` — clean
|
||||
- `npm test -- test/cli-registry test/custom-model-injection.test.ts
|
||||
test/custom-model-injection-contract.test.ts test/routes/custom-model-routes.test.ts
|
||||
test/routes/session-custom-model.test.ts test/routes/external-cli-bypass-clamp.test.ts` —
|
||||
245+ tests passing, including the existing multi-user clamp suite (no
|
||||
regressions from the `privilegedEnvKeys` additions)
|
||||
- `node --check scripts/test-local-llm-harnesses.mjs` + manual `--help` run
|
||||
- **Real end-to-end run against the maintainer's live llama-swap server**
|
||||
(`http://10.10.11.241:8080`), inside `codeman/agent:llm-test` (all 9 CLI
|
||||
binaries, built via `docker/agent.Dockerfile`), against the smallest
|
||||
available model (`qwen3.5-0.8b-ud-q8_k_xl`, 1.1GB — picked by parsing the
|
||||
server's own reported model sizes). Real findings, not simulated:
|
||||
- **opencode: PASS.** Genuinely round-tripped a "hello world" reply
|
||||
through the real endpoint.
|
||||
- **codex: real bug found and fixed.** The recipe's TOML shape
|
||||
(`[model].default`) was rejected by a real codex binary ("invalid
|
||||
type: map, expected a string") — codex wants a top-level `model`
|
||||
string plus `[model_providers.custom]`, and the API key rides as an
|
||||
`env_key`-named env var, never a literal TOML field. Fixed in
|
||||
`custom-model-injection.ts`, the standalone script, and both test
|
||||
suites. **Then a second, deeper finding**: codex only speaks the
|
||||
Responses API now (`wire_api = "responses"`, the only value it accepts
|
||||
since dropping `"chat"` support in Feb 2026) — a real run against the
|
||||
now-correctly-shaped config still failed (`Reconnecting...` × 5, then
|
||||
"high demand" errors) because llama-swap doesn't implement
|
||||
`/v1/responses`. This is a genuine, currently-unresolved protocol
|
||||
incompatibility, not a bug in this PR's code — documented prominently
|
||||
in `deployment_plan.md`'s confidence table.
|
||||
- **claude: PASS, after two real bugs found and fixed.** (1) Claude
|
||||
Code's async session-title-generation call also uses
|
||||
`ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude's own
|
||||
internal recognized-model list, printing `[claude-code:unrecognized_model]`
|
||||
and, in `-p` mode, hanging the whole invocation rather than just
|
||||
warning. `--settings '{"autoTitle":false}'` does NOT
|
||||
stop it (confirmed); `--bare` does — the warning still prints, but the
|
||||
real prompt now runs and returns the real answer. ⚠️ `--bare` is only
|
||||
safe for this standalone one-shot test script — it also disables hooks,
|
||||
LSP, plugin sync, and CLAUDE.md auto-discovery, so it must NEVER be
|
||||
applied to a real interactive Codeman session (which depends on hooks
|
||||
for idle detection, trust-dialog auto-accept, etc.). Whether an
|
||||
INTERACTIVE session with a custom model hits the same hang (vs. just a
|
||||
background warning, which would be harmless) is untested — flagged as
|
||||
an open item for chunk 5/6, not assumed either way. (2) A separate,
|
||||
genuinely nasty bug in the test script itself: a `POST` issued right
|
||||
after a `GET` in the same Node process reliably HUNG indefinitely
|
||||
against this real server (reproduced repeatedly: GET alone ~30ms, POST
|
||||
alone ~1-2s, GET-then-immediate-POST times out completely; a 2s pause
|
||||
between them fixed it every time) — looks like Node's fetch/undici
|
||||
reusing a pooled keep-alive connection the server doesn't handle
|
||||
cleanly for a second request right behind a first. Fixed with a 2s
|
||||
pause between the script's discovery GET and its baseline POST. This
|
||||
is a tooling-correctness fix (affects the script's own baseline check),
|
||||
not a claim about how any CLI's own HTTP client behaves.
|
||||
- **Also found and fixed**: an earlier design sent BOTH `Authorization:
|
||||
Bearer` and `api-key` auth header conventions on every discovery/
|
||||
baseline request, on the theory that an unused header is harmless.
|
||||
Live-tested against the real server, sending both reliably HUNG the
|
||||
request (reproduced 3×: either header alone ~500-600ms, both together
|
||||
no response inside 15s). Removed the `'both'` option entirely from
|
||||
`CustomModelAuthStyle` (was `'bearer' | 'api-key' | 'both'`, now just
|
||||
the first two, default `'bearer'`) — in the schema, the store type, the
|
||||
discovery route, and the standalone script (`--auth-style` flag added).
|
||||
This was a real, currently-shipped-in-this-PR bug fixed before it ever
|
||||
reached anyone, not a pre-existing one.
|
||||
|
||||
## Not yet done / open questions for review
|
||||
|
||||
- Six of nine per-CLI recipes (Gemini, Pi, Grok, DeepSeek, OMP) are
|
||||
**web-researched, not verified** against real binaries — see the
|
||||
confidence table in `deployment_plan.md`. Antigravity has no known
|
||||
mechanism at all and stays unsupported.
|
||||
- Chunk 5's session-restart design needs a careful look before
|
||||
implementation: switching a session's endpoint restarts its CLI process
|
||||
in place (confirmed acceptable with the maintainer — these harnesses
|
||||
read endpoint config at process start, not per-turn).
|
||||
|
||||
🤖 Generated with [Claude Code](https://claude.com/claude-code)
|
||||
@@ -0,0 +1,348 @@
|
||||
# Custom Model Endpoint Profiles (all harnesses, local or cloud)
|
||||
|
||||
## Context
|
||||
|
||||
Devvyn pays for Claude Code but also runs a capable local model behind an
|
||||
OpenAI-compatible server (llama.cpp) — and wants the same mechanism to work
|
||||
against a **cloud** OpenAI-compatible endpoint too (e.g. Azure AI Foundry's
|
||||
OpenAI-compatible inference endpoint, OpenRouter, a self-hosted gateway).
|
||||
Right now every Codeman session mode defaults to its native cloud backend
|
||||
with no way to redirect a session at any other endpoint from the UI — the
|
||||
closest existing precedent is DeepSeek's server-env-sourced
|
||||
`DEEPSEEK_BASE_URL`, which isn't user-facing.
|
||||
|
||||
**Scope note**: this plan originally said "local LLM." It now covers any
|
||||
OpenAI-compatible endpoint the user configures — local (llama.cpp, Ollama,
|
||||
vLLM) or cloud (Azure AI Foundry, OpenRouter, a company gateway). The
|
||||
mechanism is identical (a base URL Codeman probes via `GET /v1/models`); the
|
||||
only real differences are auth-header convention (cloud endpoints often want
|
||||
an `api-key` header, e.g. Azure, rather than `Authorization: Bearer`) and
|
||||
that a cloud "model" may actually be a deployment name distinct from the
|
||||
underlying model family (Azure AI Foundry deployments) — both are called out
|
||||
where they matter below. Naming throughout this plan is **"custom model
|
||||
endpoint,"** not "local model," to keep that scope explicit.
|
||||
|
||||
### Additional use case: on-premises AI hardware
|
||||
|
||||
"Local" isn't limited to a desktop running llama.cpp — a growing category of
|
||||
purpose-built, on-premises AI hardware exists specifically to run a serious
|
||||
model on-site with an OpenAI-compatible server, and this feature is exactly
|
||||
the on-ramp for pointing Codeman at one:
|
||||
|
||||
- **NVIDIA DGX Spark** (and the DGX Spark-class "Spark" mini-supercomputer
|
||||
line) — a compact on-prem inference/training box aimed at running large
|
||||
local models with an OpenAI-compatible API surface.
|
||||
- **AMD "Strix Halo" (Ryzen AI Max)** on-prem AI mini-PCs — unified-memory
|
||||
APU hardware marketed for local LLM inference, typically fronted by
|
||||
llama.cpp/Ollama/vLLM the same way a home server would be.
|
||||
|
||||
Neither needs anything new from this design: both present a standard
|
||||
`/v1/models` + `/v1/chat/completions` OpenAI-compatible surface once the
|
||||
inference server is running, so they're just another `baseUrl` entry in the
|
||||
custom-model-hosts store, same as llama.cpp or a cloud endpoint. The
|
||||
justification for building this generically (rather than hardcoding "point
|
||||
Claude at my llama.cpp box") is precisely this: **the same endpoint registry
|
||||
and per-CLI injection mechanism should work unmodified for any current or
|
||||
future OpenAI-compatible box or service** — a home GPU rig today, a Spark or
|
||||
Strix Halo appliance tomorrow, a company's on-prem inference cluster after
|
||||
that — without Codeman needing to know or care what's actually serving the
|
||||
model on the other end of that URL.
|
||||
|
||||
A concrete example worth naming: **[Ark0N/Qwen5090](https://github.com/Ark0N/Qwen5090)**
|
||||
(from the same GitHub account as this project's owner) is a one-click
|
||||
Windows / one-command Linux installer that stands up Qwen3.8-27B locally on
|
||||
an RTX 5090 (or another RTX 50-series card with ≥24GB) behind an
|
||||
OpenAI-compatible API, served by any of vLLM, NInfer, or llama.cpp — MIT-
|
||||
licensed tooling over Apache-2.0 Qwen weights. It's a direct, ready-made
|
||||
target for this feature: point a custom-model-hosts entry at whichever
|
||||
backend it's running, and it needs nothing further from Codeman's side. It's
|
||||
also notable for already wiring up DeepSeek Harness and Claude Code as
|
||||
coding agents against that local server itself, which is effectively the
|
||||
same "point a Codeman-supported harness at a local endpoint" idea this
|
||||
feature is generalizing — worth using as a real-world reference/test target
|
||||
once chunk 5 (session integration) exists, alongside Devvyn's own llama.cpp
|
||||
box.
|
||||
|
||||
Each harness has its own (different-shaped) mechanism for pointing at a
|
||||
custom OpenAI-compatible base URL + model — env vars for Claude, a JSON
|
||||
config blob for opencode, a TOML file for Codex, etc. Devvyn gave the
|
||||
starting recipes for those three; the rest (Gemini, Pi, Grok, DeepSeek, OMP,
|
||||
Antigravity) were researched for this plan and are flagged by confidence
|
||||
below. A real end-to-end pass against Devvyn's own llama-swap server
|
||||
(`scripts/test-local-llm-harnesses.mjs`, inside a `codeman/agent:llm-test`
|
||||
Docker image with all 9 CLIs installed) then confirmed **claude and
|
||||
opencode work end-to-end**, corrected a real Codex config.toml schema bug
|
||||
the given recipe had (see the Codex row below), and surfaced that Codex's
|
||||
*protocol* — not just its config shape — does not work against a plain
|
||||
OpenAI-Chat-Completions server like llama.cpp/llama-swap at all. Confidence
|
||||
below reflects what was actually observed, not just what was planned.
|
||||
|
||||
The feature must be:
|
||||
|
||||
- **Off by default**, one settings toggle turns it on.
|
||||
- Endpoint entry: user gives a base URL — a LAN address or a cloud URL —
|
||||
plus an optional API key, and Codeman calls `GET <baseUrl>/v1/models` to
|
||||
discover and store the available model (or deployment) list.
|
||||
- A **new toolbar selector** (separate from the existing Run-mode menu, since
|
||||
it's a modifier on top of whichever harness is already selected/running)
|
||||
lets the user pick "Cloud (default)" — the harness's own native backend —
|
||||
or a model discovered from one of the configured custom endpoints.
|
||||
- Picking a custom-endpoint model for an **already-running session restarts
|
||||
that session's CLI process** with the injected env/config pointed at that
|
||||
endpoint (confirmed with Devvyn — these harnesses read endpoint config at
|
||||
process start, not per-turn, so a live hot-swap isn't possible).
|
||||
- **New sessions always default back to the harness's native cloud backend.**
|
||||
A custom-endpoint selection is a per-session override, not a sticky global
|
||||
default — starting a fresh CLI (any mode) always launches against its
|
||||
native backend unless the user explicitly picks a custom endpoint for that
|
||||
new session too. The toolbar selector is scoped to "this session," never
|
||||
carried forward as the default for future sessions.
|
||||
|
||||
This follows the repo's existing data-driven CLI-registry philosophy
|
||||
(`test/cli-registry-no-id-branching.test.ts`): per-CLI behavior is a
|
||||
declared capability, never an `if (mode === 'claude')` branch.
|
||||
|
||||
## Per-CLI injection recipes (confidence-ranked)
|
||||
|
||||
| CLI | Mechanism | Confidence |
|
||||
|---|---|---|
|
||||
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
|
||||
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
|
||||
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
|
||||
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` (or `GOOGLE_VERTEX_BASE_URL`) + `GEMINI_API_KEY`; CLI needs a restart to pick them up (matches our restart-on-switch design). Model selection via `--model`/`GEMINI_MODEL`-style override — verify exact var name against the installed `gemini-cli` version before shipping | Web-researched, unverified |
|
||||
| `pi` | Config file `~/.pi/agent/models.json` (hot-reloadable) with a custom provider block: `baseUrl`, `apiKey`, `api:"openai-completions"`. Redirect via `PI_CONFIG_DIR` (already allowlisted per CLAUDE.md) pointed at an isolated dir containing just this file, rather than overwriting the user's real one | Web-researched, unverified |
|
||||
| `grok` | Env vars `GROK_BASE_URL`, `XAI_API_KEY` (dummy ok for local; a real key for most cloud endpoints), `GROK_MODEL`. All three already fit inside the existing `XAI_*`/CLI-specific allowlist shape | Web-researched, unverified |
|
||||
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts:879-913`, already in `privilegedEnvKeys`). Model selection is murkier — CLAUDE.md notes dsh model is "a profile composition entry," not a flag/env var, so redirecting the endpoint is solid but forcing a specific model name may not fully work; document as best-effort and verify against a real profile | Web-researched, unverified, partial |
|
||||
| `omp` | Config file `~/.omp/agent/models.yml`-equivalent with a custom provider `baseUrl`. CLAUDE.md notes omp's config tree is itself relocatable via `PI_CONFIG_DIR` — reuse the same isolated-dir-redirect approach as `pi` | Web-researched, unverified |
|
||||
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
|
||||
|
||||
Everything web-researched-but-unverified gets implemented but must be
|
||||
smoke-tested against real installs of those CLIs before being called done —
|
||||
call this out explicitly when implementing, don't just ship on faith.
|
||||
|
||||
**Cloud-endpoint specifics** to keep in mind per recipe above: an Azure AI
|
||||
Foundry-style endpoint typically wants the API key in an `api-key` header
|
||||
rather than (or in addition to) `Authorization: Bearer`, and its "model" is
|
||||
often a deployment name rather than the underlying model family name — the
|
||||
discovery step (`GET /v1/models`) still works the same way against Azure AI
|
||||
Foundry's OpenAI-compatible endpoint shape, but a user may need to type the
|
||||
deployment name manually if it isn't returned as expected.
|
||||
|
||||
## Architecture
|
||||
|
||||
### 1. Registry: new `capabilities.customModelInjection` field
|
||||
|
||||
Extend `src/config/cli-registry/types.ts` / `schema.ts` with a discriminated
|
||||
union on each `CliEntry.capabilities`:
|
||||
|
||||
```ts
|
||||
type CustomModelInjection =
|
||||
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[] }
|
||||
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json' }
|
||||
| { kind: 'configDir'; dirEnvVar: string; fileName: string; template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' }
|
||||
| { kind: 'unsupported' }
|
||||
```
|
||||
|
||||
Declared per stock.ts entry per the table above. A pure function in a new
|
||||
`src/custom-model-injection.ts` (`buildCustomModelInjection(entry, endpoint, modelId)`)
|
||||
turns `(CliEntry, endpoint, modelId)` into either an `envOverrides` object
|
||||
(kind `env`/`configContentEnv`) or a `{ dirEnvVar, files: [{path, content}] }`
|
||||
descriptor (kind `configDir`) — unit-testable with no IO, mirroring how
|
||||
`session-cli-builder.ts` is pure. The `configDir` kind additionally needs an
|
||||
IO wrapper that writes those files under
|
||||
`dataPath('custom-model-configs/<sessionId>/')` (new dir, cleaned up on
|
||||
session delete — same lifecycle as other per-session generated state).
|
||||
|
||||
### 2. Endpoint registry: `src/custom-model-hosts.ts`
|
||||
|
||||
Same read-array/write-array shape as `src/remote-hosts.ts` /
|
||||
`src/webview-store.ts`: `~/.codeman/custom-model-hosts.json` holding
|
||||
`CustomModelEndpoint[] = { id, label, baseUrl, apiKey?, authStyle?: 'bearer'|'api-key'|'both', models?: string[], lastDiscoveredAt? }`.
|
||||
`authStyle` defaults to `'both'` (send both header conventions on the
|
||||
discovery probe, same approach the smoke-test script below uses) so one
|
||||
endpoint entry works whether it's llama.cpp or Azure without the user having
|
||||
to know which header their box wants in advance.
|
||||
|
||||
New route file `src/web/routes/custom-model-routes.ts` (registered in the
|
||||
routes barrel), mirroring `case-routes.ts`'s remote/docker-host CRUD
|
||||
(`GET/POST/PUT/DELETE /api/model-endpoints`, admin-gated in multi-user mode
|
||||
the same way) plus:
|
||||
|
||||
- `POST /api/model-endpoints/:id/discover-models` — fetches
|
||||
`${baseUrl}/v1/models`, stores the `data[].id` list, returns it. Bounded
|
||||
timeout, and run the target through the **same SSRF egress guard already
|
||||
used for web tabs** (`webview-egress-policy.ts` — reject link-local/cloud
|
||||
metadata addresses) — this still matters for a cloud URL too, since the
|
||||
guard is about preventing a redirect to internal infra, not about
|
||||
local-vs-cloud.
|
||||
|
||||
**Why discovery rather than a free-text model field**: it removes the one
|
||||
piece of configuration most likely to trip a user up — hand-typing the
|
||||
exact model identifier a given inference server expects, which varies by
|
||||
server and is an easy source of a silent "model not found" failure with no
|
||||
useful error surfaced back through a CLI's own startup. Discovery also
|
||||
means this design is not limited to a single-model box: a **multi-model
|
||||
gateway** such as **[llama-swap](https://github.com/mostlygeek/llama-swap)**
|
||||
(hot-swaps between several loaded llama.cpp model configs behind one
|
||||
OpenAI-compatible endpoint) or a vLLM/LiteLLM/Ollama instance serving
|
||||
several models advertises ALL of them through the same `/v1/models` call —
|
||||
so one endpoint entry surfaces every model that gateway can serve, with no
|
||||
extra per-model configuration on Codeman's side at all.
|
||||
|
||||
### 3. Settings
|
||||
|
||||
- New synced boolean `customModelEndpointsEnabled` in `SettingsUpdateSchema`
|
||||
(`src/web/schemas.ts`), default `false`, documented inline like
|
||||
`readMyMindEnabled`/`workspaceHooksEnabled`.
|
||||
- New `.set-group` "Custom Model Endpoints" inside the **Agents & CLIs**
|
||||
section (`settings-clis`, `index.html:2150+`) with the enable toggle plus
|
||||
a list-editor (add/refresh-models/delete rows) for endpoints — closest
|
||||
existing precedent is the respawn-presets array editor
|
||||
(`schemas.ts:1285-1305`, `index.html:1243-1244`) for add/apply/delete-by-id
|
||||
semantics, backed by the new CRUD routes above.
|
||||
|
||||
### 4. Toolbar UI
|
||||
|
||||
- New header/toolbar button (e.g. `#customModelBtn`, `btn-toolbar
|
||||
btn-custom-model`), marker-hidden by default (`btn-custom-model--hidden`)
|
||||
and revealed by `applyHeaderVisibilitySettings()` only when
|
||||
`customModelEndpointsEnabled` is on — same pattern as the File
|
||||
Viewer/Cron buttons.
|
||||
- Clicking opens a dropdown (`#customModelMenu`, same `.run-mode-menu`-style
|
||||
markup as the existing Run-mode gear menu) listing "Cloud (default)" plus
|
||||
every discovered model, grouped by endpoint. An entry is disabled with a
|
||||
tooltip when the active session's CLI has `customModelInjection.kind ===
|
||||
'unsupported'` (Antigravity) or none declared.
|
||||
- Selecting an entry calls a new route:
|
||||
`POST /api/sessions/:id/custom-model { endpointId, modelId } | { clear: true }`.
|
||||
Server: resolve the CLI entry for `session.mode`, build the injection via
|
||||
§1, persist it as a new `session.customModel` state field (surfaced in
|
||||
`toState()`/SSE so the tab can show a small badge, e.g. "🖥 qwen3 (local)"
|
||||
or "☁ gpt-4o-mini (azure)", and the choice survives reload), merge into
|
||||
the session's `envOverrides`, and **respawn the pane's CLI process**
|
||||
through the same respawn/interactive-restart path
|
||||
`session.ts`/`tmux-manager.ts` already use for effort/model changes
|
||||
(`_configureCliEnv()` + `applyEnvOverrides()` at spawn time) — reuse,
|
||||
don't reinvent, the existing kill-and-relaunch-in-pane machinery.
|
||||
- New-session creation deliberately does **not** inherit a prior custom-
|
||||
endpoint choice: `buildEnvOverrides()` (session-ui.js) never carries the
|
||||
toolbar selection forward to the next `run()` call. Every new session
|
||||
starts on its native backend; picking a custom endpoint in the toolbar for
|
||||
a session applies only to that session (and, if done before Run is
|
||||
clicked, to the one session about to be created — not to sessions created
|
||||
afterward).
|
||||
|
||||
### 5. Multi-user security clamp
|
||||
|
||||
Every new env var this feature introduces that can redirect a session's
|
||||
traffic (and thus wherever its credentials go) — `ANTHROPIC_BASE_URL`,
|
||||
`GOOGLE_GEMINI_BASE_URL`, `GROK_BASE_URL`, the `CODEX_HOME`/`PI_CONFIG_DIR`
|
||||
dir-redirects, plus the already-privileged `DEEPSEEK_BASE_URL` — must be
|
||||
added to each CLI's `capabilities.privilegedEnvKeys` so
|
||||
`clampEnvOverridesForOwner()` strips them for a non-granted multi-user
|
||||
owner, exactly the precedent already documented for `DEEPSEEK_BASE_URL`/
|
||||
`OMP_AUTH_BROKER_URL`. This matters *more*, not less, now that endpoints can
|
||||
be cloud URLs: redirecting a non-granted user's session to an attacker's
|
||||
cloud endpoint is a credential-exfiltration path, not just a mischief
|
||||
redirect to a LAN box. Endpoint CRUD itself stays admin-only in multi-user
|
||||
mode, same as remote/docker hosts.
|
||||
|
||||
## Files touched (representative, not exhaustive)
|
||||
|
||||
- `src/config/cli-registry/types.ts`, `schema.ts`, `stock.ts` — new capability + per-entry declarations
|
||||
- `src/custom-model-injection.ts` (new) — pure per-CLI descriptor builder + unit tests
|
||||
- `src/custom-model-hosts.ts` (new) — endpoint store
|
||||
- `src/web/routes/custom-model-routes.ts` (new) — CRUD + discovery route
|
||||
- `src/web/routes/session-routes.ts` — `POST /api/sessions/:id/custom-model`, clamp wiring
|
||||
- `src/web/schemas.ts` — `customModelEndpointsEnabled`, endpoint/discover payload schemas, privileged-key updates
|
||||
- `src/session.ts` — `customModel` state field, `toState()` surface
|
||||
- `src/web/public/index.html`, `settings-ui.js`, `session-ui.js`, `styles.css` — settings group, toolbar button/menu, badge, accent CSS
|
||||
- `src/web/sse-events.ts` + `constants.js` — if a dedicated SSE event is warranted for the badge (or just ride existing session-update broadcasts)
|
||||
- `test/fixtures/mock-openai-server.ts` (new) + `test/custom-model-injection-contract.test.ts` (new) — see Mock-server validation below
|
||||
- `scripts/test-local-llm-harnesses.mjs` (already added, this branch) — the standalone real-CLI-and-real-endpoint smoke test; despite the filename (kept for continuity with when it was written) it already supports any `--base-url`, local or cloud
|
||||
- `docs/custom-model-endpoints.md` (new) + a CLAUDE.md pointer bullet under External CLI modes / envOverrides
|
||||
|
||||
## Mock-server validation strategy (CI-runnable, no real CLI binaries needed)
|
||||
|
||||
Spawning nine real CLI binaries in CI isn't realistic, and neither Devvyn's
|
||||
llama.cpp box nor a real cloud subscription can be a CI dependency. So the
|
||||
injection *logic* gets a tier of automated coverage that sits between the
|
||||
pure unit tests and the live manual checks in Verification:
|
||||
|
||||
1. **`test/fixtures/mock-openai-server.ts`** — a small in-process HTTP
|
||||
server (plain `http.createServer`, no external deps, port picked per the
|
||||
existing `const PORT = 3150+` convention) that:
|
||||
- Serves `GET /v1/models` → a fixed fake model list (`{data:[{id:'qwen3'},...]}`),
|
||||
for testing the discovery route.
|
||||
- Serves `POST /v1/chat/completions` (OpenAI shape) **and**
|
||||
`POST /v1/messages` (Anthropic Messages-API shape, since that's what
|
||||
`ANTHROPIC_BASE_URL` traffic looks like) and records every request it
|
||||
receives (headers, body, path) into an array the test can assert on —
|
||||
including which auth header style it saw, so the `authStyle: 'both'`
|
||||
default and Azure's `api-key` convention both get real coverage.
|
||||
- Returns a minimal valid completion so a client library doesn't choke
|
||||
on the response shape.
|
||||
|
||||
2. **`test/custom-model-injection-contract.test.ts`** — for every CLI with a
|
||||
`customModelInjection` capability (i.e. every row in the table above
|
||||
except `antigravity`):
|
||||
- Point a fixture `CustomModelEndpoint` at the mock server's URL.
|
||||
- Call `buildCustomModelInjection(entry, endpoint, modelId)` (the pure
|
||||
function from §1) to get the real env vars / config-file content that
|
||||
would be injected into that CLI's session.
|
||||
- Replay those exact values through a minimal HTTP request shaped the
|
||||
way that CLI is documented to send it (Anthropic Messages shape for
|
||||
claude; OpenAI chat-completions shape for opencode/codex/pi/grok/omp;
|
||||
`GOOGLE_GEMINI_BASE_URL`'s OpenAI-compat shape for gemini; dsh's
|
||||
provider call for deepseek) against the mock server.
|
||||
- Assert the mock server received the request **at the injected
|
||||
`baseUrl`**, with **the injected API key** in the expected header, and
|
||||
**the injected model id** in the body/path — i.e. prove the values
|
||||
Codeman computes are internally consistent and would reach the right
|
||||
place with the right identifiers, end to end, in CI, on every push.
|
||||
- Also cover the `configDir` kind (codex/pi/omp): assert the written
|
||||
`config.toml`/`models.json`/`models.yml` file parses and contains the
|
||||
same base URL/key/model, and that it's written under the isolated
|
||||
per-session dir rather than the user's real config path.
|
||||
|
||||
3. **Explicit, stated limitation** (goes in the test file's `@fileoverview`
|
||||
and in this doc, not left implicit): this proves *"if the CLI honors its
|
||||
documented env/config contract, it will hit the right endpoint with the
|
||||
right model."* It does **not** prove the real CLI binary actually reads
|
||||
that env var / config file the way its docs say — that's still the job
|
||||
of the live manual checks in Verification step 4-5 below, and is exactly
|
||||
why the confidence table above stays "unverified" for six of the nine
|
||||
CLIs until someone runs those binaries for real. The mock-server suite
|
||||
catches regressions in Codeman's own logic; it cannot catch a CLI
|
||||
changing its env-var name in a future release, or a real cloud endpoint
|
||||
behaving differently from the mock.
|
||||
|
||||
## Verification
|
||||
|
||||
1. `npm run typecheck && npm test` after each slice — this now includes the
|
||||
mock-server contract suite from above, so injection-logic regressions
|
||||
are caught automatically without touching real infrastructure.
|
||||
2. Unit tests for `buildCustomModelInjection()` per CLI kind (pure, no IO).
|
||||
3. Route tests (`app.inject`) for the new CRUD + discover-models endpoint
|
||||
(mock `fetch` for `/v1/models`), and for the multi-user clamp on the new
|
||||
privileged keys (mirror `test/routes/external-cli-bypass-clamp.test.ts`).
|
||||
4. **Standalone real-binary smoke test**: `scripts/test-local-llm-harnesses.mjs`
|
||||
(already written on this branch) exercises every harness against a real
|
||||
`--base-url` — local or cloud — outside of Codeman's UI entirely. Run it
|
||||
against Devvyn's llama.cpp server first (`claude`/`opencode`/`codex`
|
||||
should PASS, since those recipes are verified; the rest report
|
||||
UNCONFIRMED/SKIP until their guessed flags are corrected via
|
||||
`--probe-help`), then again against a real cloud endpoint (e.g. an Azure
|
||||
AI Foundry deployment) once one is available, to prove the `authStyle`/
|
||||
deployment-name handling holds up outside llama.cpp.
|
||||
5. Once the full feature (not just the standalone script) is built: add an
|
||||
endpoint via the real UI, hit discover-models, confirm the returned model
|
||||
list, pick Claude + the model on a real session, confirm via
|
||||
`tmux -L codeman capture-pane`/`tmux showenv -t <pane>` that
|
||||
`ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY`/`ANTHROPIC_DEFAULT_*_MODEL` are
|
||||
set post-restart, and confirm the endpoint's own logs show the next
|
||||
prompt actually landing there. Repeat for opencode and Codex at minimum
|
||||
before considering this shippable; spot-check the web-researched CLIs
|
||||
and correct the plan's confidence table with what's actually observed.
|
||||
6. `npm run lint && npm run format:check`.
|
||||
7. Update `CHANGELOG.md`/changeset per the COM workflow when shipping.
|
||||
@@ -0,0 +1,105 @@
|
||||
# Custom Model Endpoint Profiles
|
||||
|
||||
Point any Codeman-supported harness — Claude, opencode, Codex, Gemini, Pi,
|
||||
Grok, DeepSeek, or OMP — at a custom OpenAI-compatible endpoint instead of
|
||||
its native cloud backend, for a given session. "Custom endpoint" covers both
|
||||
**local** hardware (llama.cpp, Ollama, vLLM, a home GPU rig, or purpose-built
|
||||
boxes like NVIDIA DGX Spark or AMD Strix Halo mini-PCs) and **cloud**
|
||||
services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
|
||||
company gateway) — anything answering `GET /v1/models` and
|
||||
`POST /v1/chat/completions` in the standard shape. Design doc, per-CLI
|
||||
recipe confidence table, and security reasoning:
|
||||
[`deployment_plan.md`](../deployment_plan.md).
|
||||
|
||||
> **Status**: backend is implemented and tested (registry capability, the
|
||||
> injection engine, the endpoint store + discovery route, the session
|
||||
> restart route). The toolbar picker / settings UI described below as the
|
||||
> intended surface is **not yet built** — until it lands, use the HTTP API
|
||||
> directly (examples below). Antigravity has no known custom-endpoint
|
||||
> mechanism and is not supported.
|
||||
|
||||
## Turning it on
|
||||
|
||||
App Settings → Agents & CLIs → **Custom Model Endpoints** (synced setting
|
||||
`customModelEndpointsEnabled`, default **OFF**). The API equivalent:
|
||||
|
||||
```bash
|
||||
curl -sk -X PUT https://localhost:3000/api/settings \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"customModelEndpointsEnabled": true}'
|
||||
```
|
||||
|
||||
## Adding an endpoint
|
||||
|
||||
```bash
|
||||
curl -sk -X POST https://localhost:3000/api/model-endpoints \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"id": "llama-box", "label": "Home llama.cpp", "baseUrl": "http://192.168.1.50:8080"}'
|
||||
```
|
||||
|
||||
`apiKey` is optional (most local servers don't check it). `authStyle`
|
||||
(`bearer` | `api-key` | `both`, default `both`) controls which auth header
|
||||
convention discovery uses — `both` works whether the endpoint is llama.cpp
|
||||
(ignores the header) or a cloud gateway like Azure (wants `api-key`).
|
||||
|
||||
Discover its available models:
|
||||
|
||||
```bash
|
||||
curl -sk -X POST https://localhost:3000/api/model-endpoints/llama-box/discover-models
|
||||
```
|
||||
|
||||
This calls the endpoint's own `GET /v1/models` and stores the returned list
|
||||
on the endpoint record; `GET /api/model-endpoints` lists everything
|
||||
configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
|
||||
Endpoint management is admin-only in multi-user mode, same as remote/docker
|
||||
hosts — these are machine-level infra, not per-user settings.
|
||||
|
||||
## Applying a model to a session
|
||||
|
||||
```bash
|
||||
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"endpointId": "llama-box", "modelId": "qwen3"}'
|
||||
```
|
||||
|
||||
This computes the CLI-specific env vars / config for that session's mode
|
||||
(see the recipe table in `deployment_plan.md`) and **restarts the session's
|
||||
CLI process in place** — same pane, same tmux session, fresh env. That
|
||||
restart is necessary, not incidental: every supported harness reads its
|
||||
endpoint config at process start, not per-turn, so there is no live
|
||||
hot-swap. Clear back to the harness's native cloud default with:
|
||||
|
||||
```bash
|
||||
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
||||
-H 'Content-Type: application/json' -d '{"clear": true}'
|
||||
```
|
||||
|
||||
**New sessions always default back to the harness's native backend.** A
|
||||
custom-endpoint selection is a per-session choice, never a sticky global
|
||||
default — starting a fresh session doesn't inherit whatever the last one was
|
||||
pointed at.
|
||||
|
||||
## Confidence per harness
|
||||
|
||||
Only Claude, opencode, and Codex have been verified against a real
|
||||
llama.cpp server by hand. Gemini, Pi, Grok, DeepSeek, and OMP's recipes are
|
||||
correct on their one-shot invocation flags (confirmed against real
|
||||
installed binaries' own `--help` output) but their env-var/config
|
||||
conventions for a _custom_ endpoint are still web-researched, not verified
|
||||
end-to-end — see the confidence table in `deployment_plan.md` before relying
|
||||
on one of those five in production. `scripts/test-local-llm-harnesses.mjs`
|
||||
is the standalone script used to check a harness against a real endpoint
|
||||
outside the web UI entirely; see its own `--help` for usage.
|
||||
|
||||
## Security note
|
||||
|
||||
Every env var this feature can set that redirects a session's traffic
|
||||
(`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, etc.) is
|
||||
listed in that CLI's `privilegedEnvKeys` in the CLI registry, so a
|
||||
non-granted multi-user owner cannot set one directly via the generic
|
||||
`envOverrides` API field — only through this feature's own route, which
|
||||
computes the value from an admin-configured, SSRF-guarded endpoint rather
|
||||
than trusting arbitrary client input. See the "Multi-user security
|
||||
hardening" section of `deployment_plan.md` for the full reasoning; several
|
||||
of these were reachable via the generic `envOverrides` field even before
|
||||
this feature existed, and building this surfaced and closed that gap.
|
||||
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"_comment": "Copy this file to local-llm-test.config.json (gitignored) and fill in your own values. CLI flags on scripts/test-local-llm-harnesses.mjs always override these. Any field can be omitted. apiKey is OPTIONAL — omit it entirely (or delete this line) for an endpoint like llama.cpp that doesn't check one; it defaults to a harmless placeholder either way.",
|
||||
"baseUrl": "http://192.168.1.50:8080",
|
||||
"model": "qwen3",
|
||||
"apiKey": "",
|
||||
"prompt": "Reply with exactly: hello world",
|
||||
"timeout": 30000,
|
||||
"only": []
|
||||
}
|
||||
@@ -0,0 +1,631 @@
|
||||
#!/usr/bin/env node
|
||||
/**
|
||||
* Standalone smoke-test for pointing each Codeman-supported harness CLI at a
|
||||
* custom OpenAI-compatible endpoint — local (llama.cpp, Ollama, vLLM, ...) or
|
||||
* cloud (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
|
||||
* self-hosted gateway, ...). Anything that answers GET /v1/models and POST
|
||||
* /v1/chat/completions in the standard shape qualifies; --base-url is not
|
||||
* assumed to be a LAN address.
|
||||
*
|
||||
* This is intentionally OUTSIDE the npm test suite and outside Codeman's own
|
||||
* session/tmux machinery: it spawns each real CLI binary directly, one-shot,
|
||||
* with the env vars / config files that CLI's own docs say redirect it to a
|
||||
* custom endpoint, and checks it can answer "hello world".
|
||||
*
|
||||
* Cloud endpoints often differ from a bare llama.cpp box in two ways this
|
||||
* script accounts for: (1) auth may be an `api-key` header (Azure's
|
||||
* convention) rather than `Authorization: Bearer` — the baseline check in
|
||||
* Step 0 sends both, since an extra header is harmless to servers that
|
||||
* ignore it; each CLI's OWN auth convention (set via its env vars/config,
|
||||
* not this script) still needs to match what that endpoint expects. (2) a
|
||||
* cloud endpoint's "model" may actually be a deployment name distinct from
|
||||
* the model family (Azure AI Foundry deployments) — always pass --model
|
||||
* explicitly for those rather than relying on GET /v1/models discovery.
|
||||
*
|
||||
* IMPORTANT CONFIDENCE NOTE: only claude/opencode/codex recipes are verified
|
||||
* (Devvyn confirmed them by hand). gemini/pi/grok/deepseek/omp are best
|
||||
* guesses from public docs, not verified against this repo or against real
|
||||
* binaries. antigravity has no known CLI/env mechanism at all and is always
|
||||
* skipped. Read a harness's UNCONFIRMED/FAIL output before trusting it — use
|
||||
* --probe-help to read that binary's real --help and fix the guessed flag.
|
||||
*
|
||||
* Usage:
|
||||
* node scripts/test-local-llm-harnesses.mjs --base-url http://192.168.1.50:8080 [options]
|
||||
* node scripts/test-local-llm-harnesses.mjs --base-url https://<resource>.services.ai.azure.com/openai/v1 --model <deployment-name> --api-key $AZURE_AI_KEY
|
||||
*
|
||||
* Options:
|
||||
* --base-url <url> Required. Root URL of the OpenAI-compatible endpoint (local or cloud).
|
||||
* --model <name> Model/deployment id to request. Default: first from GET /v1/models.
|
||||
* --api-key <key> API key to send. Default: local-dummy-key (fine for llama.cpp; required for most cloud endpoints).
|
||||
* --auth-style <style> "bearer" (default, Authorization: Bearer) or "api-key" (the
|
||||
* `api-key` header some cloud gateways, e.g. Azure, want).
|
||||
* NEVER send both — live-tested against a real server, doing
|
||||
* so reliably HANGS the request indefinitely.
|
||||
* --prompt <text> Prompt to send. Default: "Reply with exactly: hello world".
|
||||
* --only <id,id,...> Restrict to these harness ids (comma-separated).
|
||||
* --timeout <ms> Per-harness spawn timeout. Default: 30000.
|
||||
* --probe-help Instead of testing, resolve each installed binary and print --help.
|
||||
* --keep-temp Don't delete generated per-harness config dirs afterward.
|
||||
* --list Dry run: print the resolved plan per harness, execute nothing.
|
||||
* -h, --help Show this help.
|
||||
*/
|
||||
|
||||
import { execFileSync, spawn } from 'node:child_process';
|
||||
import { mkdtempSync, mkdirSync, writeFileSync, rmSync, readFileSync, existsSync } from 'node:fs';
|
||||
import { tmpdir, homedir } from 'node:os';
|
||||
import { join, delimiter, dirname } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
|
||||
const TAG = '[test-local-llm-harnesses]';
|
||||
const SCRIPT_DIR = dirname(fileURLToPath(import.meta.url));
|
||||
const CONFIG_PATH = join(SCRIPT_DIR, 'local-llm-test.config.json');
|
||||
const CONFIG_EXAMPLE_PATH = join(SCRIPT_DIR, 'local-llm-test.config.example.json');
|
||||
|
||||
/**
|
||||
* Loads scripts/local-llm-test.config.json (gitignored — real IP/model/key,
|
||||
* per-machine) if present, so you don't have to retype --base-url every run.
|
||||
* See local-llm-test.config.example.json (tracked) for the shape. CLI flags
|
||||
* always override whatever this file sets; this only supplies defaults.
|
||||
*/
|
||||
function loadConfigFile() {
|
||||
if (!existsSync(CONFIG_PATH)) return {};
|
||||
try {
|
||||
const raw = JSON.parse(readFileSync(CONFIG_PATH, 'utf8'));
|
||||
return {
|
||||
baseUrl: raw.baseUrl ?? null,
|
||||
model: raw.model ?? null,
|
||||
apiKey: raw.apiKey || undefined, // empty string counts as "not set", not a real key
|
||||
authStyle: raw.authStyle === 'api-key' ? 'api-key' : undefined, // never 'both'
|
||||
prompt: raw.prompt ?? undefined,
|
||||
only: Array.isArray(raw.only) && raw.only.length ? raw.only : null,
|
||||
timeout: typeof raw.timeout === 'number' ? raw.timeout : undefined,
|
||||
};
|
||||
} catch (err) {
|
||||
console.error(`${TAG} failed to parse ${CONFIG_PATH}: ${err.message} (ignoring it)`);
|
||||
return {};
|
||||
}
|
||||
}
|
||||
|
||||
function parseArgs(argv, configDefaults) {
|
||||
const opts = {
|
||||
baseUrl: configDefaults.baseUrl ?? null,
|
||||
model: configDefaults.model ?? null,
|
||||
apiKey: configDefaults.apiKey ?? 'local-dummy-key',
|
||||
authStyle: configDefaults.authStyle ?? 'bearer',
|
||||
prompt: configDefaults.prompt ?? 'Reply with exactly: hello world',
|
||||
only: configDefaults.only ?? null,
|
||||
timeout: configDefaults.timeout ?? 30000,
|
||||
probeHelp: false,
|
||||
keepTemp: false,
|
||||
list: false,
|
||||
help: false,
|
||||
};
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
const a = argv[i];
|
||||
switch (a) {
|
||||
case '--base-url':
|
||||
opts.baseUrl = argv[++i];
|
||||
break;
|
||||
case '--model':
|
||||
opts.model = argv[++i];
|
||||
break;
|
||||
case '--api-key':
|
||||
opts.apiKey = argv[++i];
|
||||
break;
|
||||
case '--auth-style':
|
||||
opts.authStyle = argv[++i];
|
||||
if (opts.authStyle !== 'bearer' && opts.authStyle !== 'api-key') {
|
||||
console.error(`${TAG} --auth-style must be "bearer" or "api-key"`);
|
||||
opts.help = true;
|
||||
}
|
||||
break;
|
||||
case '--prompt':
|
||||
opts.prompt = argv[++i];
|
||||
break;
|
||||
case '--only':
|
||||
opts.only = argv[++i].split(',').map((s) => s.trim()).filter(Boolean);
|
||||
break;
|
||||
case '--timeout':
|
||||
opts.timeout = Number(argv[++i]);
|
||||
break;
|
||||
case '--probe-help':
|
||||
opts.probeHelp = true;
|
||||
break;
|
||||
case '--keep-temp':
|
||||
opts.keepTemp = true;
|
||||
break;
|
||||
case '--list':
|
||||
opts.list = true;
|
||||
break;
|
||||
case '-h':
|
||||
case '--help':
|
||||
opts.help = true;
|
||||
break;
|
||||
default:
|
||||
console.error(`${TAG} unknown argument: ${a}`);
|
||||
opts.help = true;
|
||||
}
|
||||
}
|
||||
return opts;
|
||||
}
|
||||
|
||||
function printUsage() {
|
||||
console.log(`Usage: node scripts/test-local-llm-harnesses.mjs [--base-url <url>] [options]
|
||||
|
||||
Reads defaults from scripts/local-llm-test.config.json if it exists (copy
|
||||
scripts/local-llm-test.config.example.json to create it — gitignored, since
|
||||
it holds a real IP/model/key). CLI flags always override the config file.
|
||||
--base-url becomes optional once that file supplies one.
|
||||
|
||||
Works against any custom OpenAI-compatible endpoint, local or cloud
|
||||
(llama.cpp, Ollama, vLLM, Azure AI Foundry, OpenRouter, a self-hosted
|
||||
gateway, ...) — anything answering GET /v1/models and POST
|
||||
/v1/chat/completions in the standard shape.
|
||||
|
||||
Options:
|
||||
--base-url <url> Required. Root URL of the OpenAI-compatible endpoint.
|
||||
--model <name> Model/deployment id to request. Default: first from GET /v1/models.
|
||||
--api-key <key> API key to send. Default: local-dummy-key (required for most cloud endpoints).
|
||||
--auth-style <style> "bearer" (default) or "api-key" (Azure-style). Never both — sending
|
||||
both headers together reliably hangs some real servers.
|
||||
--prompt <text> Prompt to send. Default: "Reply with exactly: hello world".
|
||||
--only <id,id,...> Restrict to these harness ids.
|
||||
--timeout <ms> Per-harness spawn timeout. Default: 30000.
|
||||
--probe-help Print each installed binary's --help instead of testing.
|
||||
--keep-temp Keep generated per-harness config dirs afterward.
|
||||
--list Dry run: print the resolved plan, execute nothing.
|
||||
-h, --help Show this help.
|
||||
|
||||
Harness ids: claude, opencode, codex, gemini, pi, grok, deepseek, omp, antigravity
|
||||
|
||||
Examples:
|
||||
node scripts/test-local-llm-harnesses.mjs --base-url http://192.168.1.50:8080
|
||||
node scripts/test-local-llm-harnesses.mjs --base-url https://<resource>.services.ai.azure.com/openai/v1 --model <deployment-name> --api-key $AZURE_AI_KEY`);
|
||||
}
|
||||
|
||||
const HOME = homedir();
|
||||
const EXTRA_SEARCH_DIRS = [
|
||||
join(HOME, '.local', 'bin'),
|
||||
join(HOME, '.opencode', 'bin'),
|
||||
join(HOME, '.codex', 'bin'),
|
||||
join(HOME, '.gemini', 'bin'),
|
||||
join(HOME, '.antigravity', 'bin'),
|
||||
join(HOME, '.grok', 'bin'),
|
||||
join(HOME, '.omp', 'bin'),
|
||||
join(HOME, '.bun', 'bin'),
|
||||
join(HOME, '.npm-global', 'bin'),
|
||||
join(HOME, 'bin'),
|
||||
'/usr/local/bin',
|
||||
];
|
||||
|
||||
function pathWithExtraDirs() {
|
||||
return [...EXTRA_SEARCH_DIRS, process.env.PATH ?? ''].join(delimiter);
|
||||
}
|
||||
|
||||
/** Resolve a binary by trying `<bin> --version` with extra search dirs prefixed onto PATH. */
|
||||
function resolveBinary(bin) {
|
||||
try {
|
||||
execFileSync(bin, ['--version'], {
|
||||
timeout: 5000,
|
||||
stdio: 'pipe',
|
||||
env: { ...process.env, PATH: pathWithExtraDirs() },
|
||||
});
|
||||
return bin;
|
||||
} catch (err) {
|
||||
// Some CLIs (e.g. dsh) don't support --version cleanly for identity but
|
||||
// still exist on PATH; a non-ENOENT failure still counts as "found".
|
||||
if (err && err.code === 'ENOENT') return null;
|
||||
return bin;
|
||||
}
|
||||
}
|
||||
|
||||
function printHelp(bin) {
|
||||
try {
|
||||
const out = execFileSync(bin, ['--help'], {
|
||||
timeout: 5000,
|
||||
stdio: 'pipe',
|
||||
env: { ...process.env, PATH: pathWithExtraDirs() },
|
||||
});
|
||||
console.log(out.toString());
|
||||
} catch (err) {
|
||||
console.log((err.stdout ?? err.message ?? String(err)).toString());
|
||||
}
|
||||
}
|
||||
|
||||
// --- per-harness definitions -----------------------------------------------
|
||||
|
||||
/** kind: 'env' | 'configContentEnv' | 'configDir' | 'unsupported' */
|
||||
const HARNESSES = {
|
||||
claude: {
|
||||
binary: 'claude',
|
||||
confidence: 'verified',
|
||||
buildEnv: (baseUrl, apiKey, model) => ({
|
||||
ANTHROPIC_BASE_URL: baseUrl,
|
||||
ANTHROPIC_API_KEY: apiKey,
|
||||
ANTHROPIC_DEFAULT_SONNET_MODEL: model,
|
||||
ANTHROPIC_DEFAULT_HAIKU_MODEL: model,
|
||||
ANTHROPIC_DEFAULT_OPUS_MODEL: model,
|
||||
}),
|
||||
// Claude Code's async session-title-generation call also uses
|
||||
// ANTHROPIC_DEFAULT_HAIKU_MODEL and validates it against Claude's OWN internal
|
||||
// recognized-model list, printing [claude-code:unrecognized_model] to stderr for
|
||||
// a local model name. Confirmed live: `--settings '{"autoTitle":false}'` does NOT
|
||||
// stop it (still hung the whole run); `--bare` does — the warning still prints,
|
||||
// but the actual prompt now runs and returns the real answer. Confirmed against
|
||||
// a real llama-swap server.
|
||||
buildArgv: (prompt) => ['--dangerously-skip-permissions', '--bare', '-p', prompt],
|
||||
},
|
||||
opencode: {
|
||||
binary: 'opencode',
|
||||
confidence: 'verified',
|
||||
buildEnv: (baseUrl, apiKey, model) => ({
|
||||
OPENCODE_CONFIG_CONTENT: JSON.stringify({
|
||||
$schema: 'https://opencode.ai/config.json',
|
||||
provider: {
|
||||
local: {
|
||||
options: { baseURL: `${baseUrl}/v1`, apiKey },
|
||||
models: { [model]: {} },
|
||||
},
|
||||
},
|
||||
model: `local/${model}`,
|
||||
}),
|
||||
}),
|
||||
buildArgv: (prompt) => ['run', prompt],
|
||||
},
|
||||
codex: {
|
||||
binary: 'codex',
|
||||
confidence: 'verified',
|
||||
configDir: {
|
||||
dirEnvVar: 'CODEX_HOME',
|
||||
fileName: 'config.toml',
|
||||
// Verified against a real codex binary: `model` must be a top-level STRING
|
||||
// (an earlier `[model].default` table was rejected with "invalid type: map,
|
||||
// expected a string"). The API key is NEVER a literal TOML field — codex only
|
||||
// supports `env_key`, the NAME of an env var it reads the value from, so the
|
||||
// real key rides as an extra env var (see extraEnv below), never in the file.
|
||||
// ⚠️ `wire_api = "responses"` is the only value codex still accepts (support
|
||||
// for "chat" was dropped Feb 2026) — a plain OpenAI Chat-Completions server
|
||||
// (llama.cpp, llama-swap) does NOT implement the Responses API, so this may
|
||||
// still fail at the PROTOCOL level even with a correctly-shaped file.
|
||||
content: (baseUrl, _apiKey, model) =>
|
||||
`model = "${model}"\nmodel_provider = "custom"\n\n[model_providers.custom]\nname = "Custom Endpoint"\nbase_url = "${baseUrl}/v1"\nenv_key = "CODEMAN_CUSTOM_MODEL_API_KEY"\nwire_api = "responses"\n`,
|
||||
extraEnv: (_baseUrl, apiKey) => ({ CODEMAN_CUSTOM_MODEL_API_KEY: apiKey }),
|
||||
},
|
||||
buildArgv: (prompt) => ['exec', '--dangerously-bypass-approvals-and-sandbox', prompt],
|
||||
},
|
||||
gemini: {
|
||||
binary: 'gemini',
|
||||
confidence: 'researched',
|
||||
buildEnv: (baseUrl, apiKey, model) => ({
|
||||
GOOGLE_GEMINI_BASE_URL: baseUrl,
|
||||
GEMINI_API_KEY: apiKey,
|
||||
GEMINI_MODEL: model,
|
||||
}),
|
||||
buildArgv: (prompt) => ['-p', prompt, '--approval-mode', 'yolo'],
|
||||
},
|
||||
pi: {
|
||||
binary: 'pi',
|
||||
confidence: 'researched',
|
||||
configDir: {
|
||||
dirEnvVar: 'PI_CONFIG_DIR',
|
||||
fileName: join('agent', 'models.json'),
|
||||
content: (baseUrl, apiKey, model) =>
|
||||
JSON.stringify(
|
||||
{
|
||||
providers: {
|
||||
local: {
|
||||
baseUrl: `${baseUrl}/v1`,
|
||||
apiKey,
|
||||
api: 'openai-completions',
|
||||
models: { [model]: {} },
|
||||
},
|
||||
},
|
||||
},
|
||||
null,
|
||||
2
|
||||
),
|
||||
},
|
||||
buildArgv: (prompt) => ['--approve', '-p', prompt],
|
||||
},
|
||||
grok: {
|
||||
binary: 'grok',
|
||||
confidence: 'researched',
|
||||
buildEnv: (baseUrl, apiKey, model) => ({
|
||||
GROK_BASE_URL: baseUrl,
|
||||
XAI_API_KEY: apiKey,
|
||||
GROK_MODEL: model,
|
||||
}),
|
||||
buildArgv: (prompt) => ['--always-approve', '-p', prompt],
|
||||
},
|
||||
deepseek: {
|
||||
binary: 'dsh',
|
||||
confidence: 'unknown',
|
||||
note: 'dsh is a profile launcher, not a documented one-shot prompt flag. Best-effort only.',
|
||||
buildEnv: (baseUrl, apiKey) => ({
|
||||
DEEPSEEK_BASE_URL: baseUrl,
|
||||
DEEPSEEK_API_KEY: apiKey,
|
||||
DSH_PERMISSION_MODE: 'danger-full-access',
|
||||
}),
|
||||
buildArgv: (prompt) => ['--profile', 'headless', prompt],
|
||||
},
|
||||
omp: {
|
||||
binary: 'omp',
|
||||
confidence: 'researched',
|
||||
configDir: {
|
||||
dirEnvVar: 'PI_CONFIG_DIR', // omp's ~/.omp tree is relocatable via PI_CONFIG_DIR per CLAUDE.md
|
||||
fileName: join('agent', 'models.yml'),
|
||||
content: (baseUrl, apiKey, model) =>
|
||||
`providers:\n local:\n baseUrl: ${baseUrl}/v1\n apiKey: ${apiKey}\n models:\n - ${model}\n`,
|
||||
},
|
||||
buildArgv: (prompt) => ['-p', prompt],
|
||||
},
|
||||
antigravity: {
|
||||
binary: 'agy',
|
||||
confidence: 'unsupported',
|
||||
note: 'No known CLI/env/config mechanism for a custom endpoint (GUI-only per public docs). Always skipped.',
|
||||
buildEnv: null,
|
||||
buildArgv: null,
|
||||
},
|
||||
};
|
||||
|
||||
// --- baseline server check ---------------------------------------------------
|
||||
|
||||
async function baselineCheck(baseUrl, apiKey, authStyle, model, prompt, timeoutMs) {
|
||||
console.log(`\n=== Step 0: baseline check against ${baseUrl} (auth: ${authStyle}) ===`);
|
||||
|
||||
// Exactly ONE header, never both. An earlier version sent both auth conventions
|
||||
// (Bearer + api-key) on the theory that an unused header is harmless — live-
|
||||
// tested against a real llama-swap server, sending both reliably HUNG the
|
||||
// request indefinitely (reproduced 3x: Bearer alone ~500ms, api-key alone
|
||||
// ~600ms, both together no response inside a 15s timeout). Use --auth-style
|
||||
// api-key for endpoints that specifically want that header (e.g. Azure AI
|
||||
// Foundry); default 'bearer' covers everything else.
|
||||
const authHeaders = authStyle === 'api-key' ? { 'api-key': apiKey } : { Authorization: `Bearer ${apiKey}` };
|
||||
|
||||
let discoveredModel = model;
|
||||
try {
|
||||
const res = await fetch(`${baseUrl}/v1/models`, {
|
||||
headers: authHeaders,
|
||||
signal: AbortSignal.timeout(timeoutMs),
|
||||
});
|
||||
if (!res.ok) throw new Error(`HTTP ${res.status}`);
|
||||
const body = await res.json();
|
||||
const ids = (body.data ?? []).map((m) => m.id);
|
||||
console.log(`GET /v1/models -> ${ids.length ? ids.join(', ') : '(empty list)'}`);
|
||||
if (!discoveredModel && ids.length) discoveredModel = ids[0];
|
||||
} catch (err) {
|
||||
console.error(`${TAG} GET /v1/models failed: ${err.message}`);
|
||||
console.error(`${TAG} Is the server actually running at ${baseUrl}? Aborting.`);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
if (!discoveredModel) {
|
||||
console.error(`${TAG} No --model given and none discovered from /v1/models. Aborting.`);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
// Live-tested against a real llama-swap server: a POST issued right after a GET on
|
||||
// the same Node process reliably HANGS indefinitely (reproduced repeatedly — GET
|
||||
// alone ~30ms, POST alone ~1-2s, GET-then-immediate-POST times out completely; a
|
||||
// 2s pause between them fixed it every time). This looks like Node's fetch (undici)
|
||||
// reusing a pooled keep-alive connection the server doesn't handle cleanly for a
|
||||
// second request right behind a first. A short pause is the simplest portable fix
|
||||
// (no extra deps, no need for undici's Agent/dispatcher API).
|
||||
await new Promise((resolve) => setTimeout(resolve, 2000));
|
||||
|
||||
try {
|
||||
const res = await fetch(`${baseUrl}/v1/chat/completions`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json', ...authHeaders },
|
||||
body: JSON.stringify({
|
||||
model: discoveredModel,
|
||||
messages: [{ role: 'user', content: prompt }],
|
||||
}),
|
||||
signal: AbortSignal.timeout(timeoutMs),
|
||||
});
|
||||
if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
|
||||
const body = await res.json();
|
||||
const reply = body.choices?.[0]?.message?.content ?? '';
|
||||
if (!reply.trim()) throw new Error('empty reply');
|
||||
console.log(`POST /v1/chat/completions -> "${reply.trim().slice(0, 200)}"`);
|
||||
console.log('Server baseline: PASS\n');
|
||||
} catch (err) {
|
||||
console.error(`${TAG} POST /v1/chat/completions failed: ${err.message}`);
|
||||
console.error(`${TAG} Server responded to /v1/models but not to a chat request. Aborting.`);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
return discoveredModel;
|
||||
}
|
||||
|
||||
// --- per-harness run ----------------------------------------------------------
|
||||
|
||||
function makeTempConfigDir(id) {
|
||||
const dir = mkdtempSync(join(tmpdir(), `codeman-local-llm-test-${id}-`));
|
||||
return dir;
|
||||
}
|
||||
|
||||
function runChild(bin, argv, env, timeoutMs) {
|
||||
return new Promise((resolve) => {
|
||||
let stdout = '';
|
||||
let stderr = '';
|
||||
let settled = false;
|
||||
const child = spawn(bin, argv, {
|
||||
env: { ...process.env, ...env, PATH: pathWithExtraDirs() },
|
||||
stdio: ['ignore', 'pipe', 'pipe'],
|
||||
});
|
||||
const timer = setTimeout(() => {
|
||||
if (settled) return;
|
||||
settled = true;
|
||||
child.kill('SIGKILL');
|
||||
resolve({ code: null, stdout, stderr, timedOut: true });
|
||||
}, timeoutMs);
|
||||
child.stdout.on('data', (d) => (stdout += d.toString()));
|
||||
child.stderr.on('data', (d) => (stderr += d.toString()));
|
||||
child.on('error', (err) => {
|
||||
if (settled) return;
|
||||
settled = true;
|
||||
clearTimeout(timer);
|
||||
resolve({ code: null, stdout, stderr: `${stderr}\n${err.message}`, timedOut: false });
|
||||
});
|
||||
child.on('close', (code) => {
|
||||
if (settled) return;
|
||||
settled = true;
|
||||
clearTimeout(timer);
|
||||
resolve({ code, stdout, stderr, timedOut: false });
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
async function runHarness(id, def, opts, model) {
|
||||
const result = { id, confidence: def.confidence, status: 'SKIP', detail: '' };
|
||||
|
||||
if (def.confidence === 'unsupported') {
|
||||
result.status = 'SKIP';
|
||||
result.detail = def.note ?? 'no known mechanism';
|
||||
return result;
|
||||
}
|
||||
|
||||
const resolved = resolveBinary(def.binary);
|
||||
if (!resolved) {
|
||||
result.status = 'SKIP';
|
||||
result.detail = `binary "${def.binary}" not found on PATH or search dirs`;
|
||||
return result;
|
||||
}
|
||||
|
||||
let env = def.buildEnv ? def.buildEnv(opts.baseUrl, opts.apiKey, model) : {};
|
||||
let tempDir = null;
|
||||
|
||||
if (def.configDir) {
|
||||
tempDir = makeTempConfigDir(id);
|
||||
const filePath = join(tempDir, def.configDir.fileName);
|
||||
mkdirSync(join(filePath, '..'), { recursive: true });
|
||||
writeFileSync(filePath, def.configDir.content(opts.baseUrl, opts.apiKey, model), 'utf8');
|
||||
const extraEnv = def.configDir.extraEnv ? def.configDir.extraEnv(opts.baseUrl, opts.apiKey, model) : {};
|
||||
env = { ...env, [def.configDir.dirEnvVar]: tempDir, ...extraEnv };
|
||||
}
|
||||
|
||||
const argv = def.buildArgv(opts.prompt);
|
||||
|
||||
if (opts.list) {
|
||||
result.status = 'LIST';
|
||||
result.detail = `${def.binary} ${argv.join(' ')} | env: ${Object.keys(env).join(', ')}${
|
||||
tempDir ? ` | configDir: ${tempDir}` : ''
|
||||
}`;
|
||||
if (tempDir && !opts.keepTemp) rmSync(tempDir, { recursive: true, force: true });
|
||||
return result;
|
||||
}
|
||||
|
||||
const { code, stdout, stderr, timedOut } = await runChild(def.binary, argv, env, opts.timeout);
|
||||
|
||||
if (tempDir && !opts.keepTemp) rmSync(tempDir, { recursive: true, force: true });
|
||||
else if (tempDir) result.detail += ` [config kept at ${tempDir}]`;
|
||||
|
||||
if (timedOut) {
|
||||
result.status = 'FAIL';
|
||||
result.detail = `timed out after ${opts.timeout}ms. stderr: ${stderr.slice(-300)}`;
|
||||
return result;
|
||||
}
|
||||
|
||||
const reply = stdout.trim();
|
||||
const matched = /hello/i.test(reply) && /world/i.test(reply);
|
||||
|
||||
if (code !== 0) {
|
||||
result.status = def.confidence === 'verified' ? 'FAIL' : 'UNCONFIRMED';
|
||||
result.detail = `exit ${code}. stderr: ${stderr.trim().slice(-300) || '(empty)'}`;
|
||||
return result;
|
||||
}
|
||||
|
||||
if (!reply) {
|
||||
result.status = def.confidence === 'verified' ? 'FAIL' : 'UNCONFIRMED';
|
||||
result.detail = 'exit 0 but empty stdout';
|
||||
return result;
|
||||
}
|
||||
|
||||
if (matched) {
|
||||
result.status = 'PASS';
|
||||
result.detail = reply.slice(0, 200);
|
||||
} else {
|
||||
result.status = 'UNCONFIRMED';
|
||||
result.detail = `reply didn't match heuristic, judge by eye: "${reply.slice(0, 300)}"`;
|
||||
}
|
||||
return result;
|
||||
}
|
||||
|
||||
// --- main ---------------------------------------------------------------------
|
||||
|
||||
async function main() {
|
||||
const configDefaults = loadConfigFile();
|
||||
const opts = parseArgs(process.argv.slice(2), configDefaults);
|
||||
if (opts.help) {
|
||||
printUsage();
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
const ids = opts.only ?? Object.keys(HARNESSES);
|
||||
const unknownIds = ids.filter((id) => !HARNESSES[id]);
|
||||
if (unknownIds.length) {
|
||||
console.error(`${TAG} unknown harness id(s): ${unknownIds.join(', ')}`);
|
||||
console.error(`${TAG} known ids: ${Object.keys(HARNESSES).join(', ')}`);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
// --probe-help never touches the network — no --base-url needed for it.
|
||||
if (opts.probeHelp) {
|
||||
for (const id of ids) {
|
||||
const def = HARNESSES[id];
|
||||
const resolved = resolveBinary(def.binary);
|
||||
console.log(`\n=== ${id} (${def.binary}) ===`);
|
||||
if (!resolved) {
|
||||
console.log('(not found on PATH or search dirs)');
|
||||
continue;
|
||||
}
|
||||
printHelp(def.binary);
|
||||
}
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
if (!opts.baseUrl) {
|
||||
console.error(`${TAG} --base-url is required (pass it, or set "baseUrl" in ${CONFIG_PATH}).`);
|
||||
console.error(`${TAG} See ${CONFIG_EXAMPLE_PATH} for the config file shape.\n`);
|
||||
printUsage();
|
||||
process.exit(1);
|
||||
}
|
||||
opts.baseUrl = opts.baseUrl.replace(/\/+$/, '');
|
||||
|
||||
// --list is a pure dry run: never touch the network, even if --model was given.
|
||||
let model = opts.model;
|
||||
if (opts.list) {
|
||||
model = opts.model ?? 'local-model';
|
||||
console.log(`\n=== Step 0 skipped (--list never hits the network; using placeholder "${model}") ===\n`);
|
||||
} else {
|
||||
model = await baselineCheck(opts.baseUrl, opts.apiKey, opts.authStyle, opts.model, opts.prompt, opts.timeout);
|
||||
}
|
||||
|
||||
console.log(`=== Testing ${ids.length} harness(es) ===`);
|
||||
const results = [];
|
||||
for (const id of ids) {
|
||||
process.stdout.write(`\n--- ${id} ---\n`);
|
||||
const result = await runHarness(id, HARNESSES[id], opts, model);
|
||||
results.push(result);
|
||||
console.log(`${result.status}: ${result.detail}`);
|
||||
}
|
||||
|
||||
console.log('\n=== Summary ===');
|
||||
const width = Math.max(...results.map((r) => r.id.length)) + 2;
|
||||
for (const r of results) {
|
||||
console.log(`${r.id.padEnd(width)} [${r.confidence.padEnd(11)}] ${r.status.padEnd(11)} ${r.detail.slice(0, 100)}`);
|
||||
}
|
||||
|
||||
const hardFail = results.some((r) => r.status === 'FAIL' && r.confidence === 'verified');
|
||||
if (hardFail) {
|
||||
console.error(`\n${TAG} at least one VERIFIED harness FAILed — that's a real regression, not just an unconfirmed guess.`);
|
||||
process.exit(1);
|
||||
}
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
main().catch((err) => {
|
||||
console.error(`${TAG} unexpected error:`, err);
|
||||
process.exit(1);
|
||||
});
|
||||
@@ -310,6 +310,34 @@ const capabilitiesSchema = z
|
||||
privilegedEnvKeys: z.array(envName).max(8),
|
||||
gates: z.record(z.string(), z.object({ minVersion: z.string().max(20), failClosed: z.boolean() }).strict()),
|
||||
maxFrameBytes: z.number().int().positive().optional(),
|
||||
customModelInjection: z.discriminatedUnion('kind', [
|
||||
z
|
||||
.object({
|
||||
kind: z.literal('env'),
|
||||
baseUrlVar: envName,
|
||||
apiKeyVar: envName,
|
||||
// Empty is valid: deepseek's model routing is a profile-composition concern, not
|
||||
// an env var, so it declares baseUrl/apiKey injection with no model var at all.
|
||||
modelVars: z.array(envName).max(8),
|
||||
})
|
||||
.strict(),
|
||||
z
|
||||
.object({
|
||||
kind: z.literal('configContentEnv'),
|
||||
envVar: envName,
|
||||
template: z.literal('opencode-json'),
|
||||
})
|
||||
.strict(),
|
||||
z
|
||||
.object({
|
||||
kind: z.literal('configDir'),
|
||||
dirEnvVar: envName,
|
||||
fileName: z.string().min(1).max(80),
|
||||
template: z.enum(['codex-toml', 'pi-models-json', 'omp-models-yml']),
|
||||
})
|
||||
.strict(),
|
||||
z.object({ kind: z.literal('unsupported') }).strict(),
|
||||
]),
|
||||
})
|
||||
.strict();
|
||||
|
||||
|
||||
@@ -178,6 +178,12 @@ const CLAUDE: CliEntry = {
|
||||
unset: ['CLAUDECODE', 'COLORTERM'],
|
||||
tmuxSetenvKeys: [],
|
||||
dockerExecEnvNames: [],
|
||||
// Deliberately excludes ANTHROPIC_* (base URL / API key / default-model overrides):
|
||||
// custom-model-injection.ts's claude recipe uses those names, but they must reach a
|
||||
// session ONLY through the admin-configured, SSRF-guarded custom-model route, never
|
||||
// through a plain client-supplied envOverrides field. Widening this prefix would let
|
||||
// any session-create caller redirect a session's Anthropic traffic and credentials to
|
||||
// an arbitrary, unvalidated URL.
|
||||
allowedPrefixes: ['CLAUDE_CODE_'],
|
||||
allowedKeys: ['CLAUDE_CONFIG_DIR'],
|
||||
},
|
||||
@@ -209,8 +215,28 @@ const CLAUDE: CliEntry = {
|
||||
statusLineTelemetry: true,
|
||||
model: { source: 'claude-settings-file' },
|
||||
privilegedParams: [],
|
||||
privilegedEnvKeys: [],
|
||||
// ANTHROPIC_* is NOT in allowedPrefixes/allowedKeys above (deliberately — see the
|
||||
// allowedPrefixes comment nearby), so these are unreachable via plain envOverrides
|
||||
// today; listed here only so the dedicated custom-model route (deployment_plan.md
|
||||
// chunk 5) clamps them for a non-granted multi-user owner the same way every other
|
||||
// CLI's injection vars are clamped, the day that route widens who can set them.
|
||||
privilegedEnvKeys: [
|
||||
'ANTHROPIC_BASE_URL',
|
||||
'ANTHROPIC_API_KEY',
|
||||
'ANTHROPIC_DEFAULT_SONNET_MODEL',
|
||||
'ANTHROPIC_DEFAULT_HAIKU_MODEL',
|
||||
'ANTHROPIC_DEFAULT_OPUS_MODEL',
|
||||
],
|
||||
gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } },
|
||||
// Custom Model Endpoint Profiles (deployment_plan.md) — verified by hand against a real
|
||||
// llama.cpp server. Claude reads these at process start only, so switching requires a
|
||||
// respawn, never a live hot-swap.
|
||||
customModelInjection: {
|
||||
kind: 'env',
|
||||
baseUrlVar: 'ANTHROPIC_BASE_URL',
|
||||
apiKeyVar: 'ANTHROPIC_API_KEY',
|
||||
modelVars: ['ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL'],
|
||||
},
|
||||
},
|
||||
overlays: {
|
||||
// Mirrors the local default so the remote/in-container agent runs non-interactively
|
||||
@@ -276,6 +302,7 @@ const SHELL: CliEntry = {
|
||||
privilegedParams: [],
|
||||
privilegedEnvKeys: [],
|
||||
gates: {},
|
||||
customModelInjection: { kind: 'unsupported' }, // a raw shell has no "model" concept
|
||||
},
|
||||
overlays: {
|
||||
// No `remote` entry: defaultRemoteCommandForMode special-cases kind==='shell' directly
|
||||
@@ -355,6 +382,15 @@ const OPENCODE: CliEntry = {
|
||||
...agentDefaults(),
|
||||
altScreen: 'strip-mux-only',
|
||||
echo: { policy: 'buffer', anchor: { kind: 'cursor' }, predictProfile: undefined },
|
||||
// Verified by hand against a real llama.cpp server. Reuses the SAME env var opencode's
|
||||
// own `env.configContentVar` already declares — the builder in custom-model-injection.ts
|
||||
// must merge into whatever opencode config Codeman would otherwise send, not clobber it.
|
||||
customModelInjection: { kind: 'configContentEnv', envVar: 'OPENCODE_CONFIG_CONTENT', template: 'opencode-json' },
|
||||
// OPENCODE_CONFIG_CONTENT already matches the OPENCODE_ allowedPrefix above, so it was
|
||||
// ALREADY reachable via plain envOverrides before this feature existed — it replaces
|
||||
// opencode's whole config, provider api keys included, so a non-granted multi-user owner
|
||||
// sending it is a pre-existing credential-redirection gap, not one this feature opens.
|
||||
privilegedEnvKeys: ['OPENCODE_CONFIG_CONTENT'],
|
||||
},
|
||||
overlays: {
|
||||
credStore: { rel: '.config/opencode', seedWhole: true },
|
||||
@@ -444,6 +480,23 @@ const CODEX: CliEntry = {
|
||||
// `dangerouslyBypassApprovals` on the wire), so it is the one that would have caught a
|
||||
// regression; `schema.ts` now rejects a name that is not a declared param.
|
||||
privilegedParams: [{ param: 'bypassApprovals', clampTo: false }],
|
||||
// Verified by hand against a real llama.cpp server. Written to an isolated CODEX_HOME
|
||||
// so the user's real ~/.codex/config.toml is never touched.
|
||||
customModelInjection: {
|
||||
kind: 'configDir',
|
||||
dirEnvVar: 'CODEX_HOME',
|
||||
fileName: 'config.toml',
|
||||
template: 'codex-toml',
|
||||
},
|
||||
// CODEX_HOME already matches the CODEX_ allowedPrefix above, so it was ALREADY
|
||||
// reachable via plain envOverrides before this feature existed. It is arguably
|
||||
// MORE sensitive than a bare base-url var: a redirected CODEX_HOME points codex at a
|
||||
// config.toml a non-granted owner fully controls, which can restate sandbox/approval
|
||||
// policy INSIDE that file — a path the argv-level `bypassApprovals` clamp above
|
||||
// cannot see or stop.
|
||||
// CODEMAN_CUSTOM_MODEL_API_KEY: the credential config.toml's env_key references
|
||||
// (see custom-model-injection.ts) — same reasoning as CODEX_HOME above.
|
||||
privilegedEnvKeys: ['CODEX_HOME', 'CODEMAN_CUSTOM_MODEL_API_KEY'],
|
||||
},
|
||||
overlays: {
|
||||
credStore: {
|
||||
@@ -527,6 +580,20 @@ const GEMINI: CliEntry = {
|
||||
// MATERIALIZE a config (not just touch an already-sent one) or a non-granted owner who
|
||||
// sends no geminiConfig at all would still get yolo for free.
|
||||
privilegedParams: [{ param: 'approvalMode', clampTo: 'auto_edit', materializeWhenAbsent: true }],
|
||||
// Web-researched, unverified — needs a restart to pick up (CLI reads these at process
|
||||
// start). Confirm the exact model-override env var name against the installed
|
||||
// gemini-cli version before shipping.
|
||||
customModelInjection: {
|
||||
kind: 'env',
|
||||
baseUrlVar: 'GOOGLE_GEMINI_BASE_URL',
|
||||
apiKeyVar: 'GEMINI_API_KEY',
|
||||
modelVars: ['GEMINI_MODEL'],
|
||||
},
|
||||
// All three already match the GEMINI_/GOOGLE_ allowedPrefixes above, so they were
|
||||
// ALREADY reachable via plain envOverrides before this feature existed — a non-granted
|
||||
// multi-user owner redirecting a gemini session's endpoint/credentials is a
|
||||
// pre-existing gap this feature's analysis surfaced, not one it opens.
|
||||
privilegedEnvKeys: ['GOOGLE_GEMINI_BASE_URL', 'GEMINI_API_KEY', 'GEMINI_MODEL'],
|
||||
},
|
||||
overlays: {
|
||||
credStore: { rel: '.gemini', seedWhole: true }, // also covers antigravity — see its own entry
|
||||
@@ -592,6 +659,10 @@ const ANTIGRAVITY: CliEntry = {
|
||||
// Like codex: an ABSENT config already defaults safe (no bypass flag), so only a
|
||||
// SENT config needs the flag forced off — nothing is materialized.
|
||||
privilegedParams: [{ param: 'dangerouslySkipPermissions', clampTo: false }],
|
||||
// No known CLI/env/config mechanism — Antigravity's own docs describe a GUI-only
|
||||
// custom-endpoint setting and explicitly say it "cannot currently" become the core
|
||||
// reasoning model. Toolbar entry stays disabled for this mode.
|
||||
customModelInjection: { kind: 'unsupported' },
|
||||
},
|
||||
overlays: {
|
||||
// No credStore of its own: agy nests its whole state under ~/.gemini/antigravity-cli/,
|
||||
@@ -678,6 +749,20 @@ const PI: CliEntry = {
|
||||
// just answer "yes" to, so omitting --approve is not itself a clamp — MATERIALIZE
|
||||
// approveProjectTrust:false so buildPiCommand emits --no-approve outright.
|
||||
privilegedParams: [{ param: 'approveProjectTrust', clampTo: false, materializeWhenAbsent: true }],
|
||||
// Web-researched, unverified. pi's models.json hot-reloads, but this feature always
|
||||
// restarts the CLI on switch for consistency with the other 8 harnesses. Written to an
|
||||
// isolated PI_CONFIG_DIR so the user's real ~/.pi/agent/models.json is never touched.
|
||||
customModelInjection: {
|
||||
kind: 'configDir',
|
||||
dirEnvVar: 'PI_CONFIG_DIR',
|
||||
fileName: 'agent/models.json',
|
||||
template: 'pi-models-json',
|
||||
},
|
||||
// PI_CONFIG_DIR already matches the PI_ allowedPrefix above, so it was ALREADY
|
||||
// reachable via plain envOverrides before this feature existed — and pi executes
|
||||
// repo-local .pi/extensions TypeScript (see the External CLI modes note in CLAUDE.md),
|
||||
// so redirecting this dir is a code-execution surface, not just a config swap.
|
||||
privilegedEnvKeys: ['PI_CONFIG_DIR'],
|
||||
},
|
||||
overlays: {
|
||||
credStore: {
|
||||
@@ -773,6 +858,16 @@ const GROK: CliEntry = {
|
||||
// already its safe interactive ask-mode, so the multi-user clamp only needs to force an
|
||||
// EXPLICITLY-SENT bypass flag back off — nothing is materialized when config is absent.
|
||||
privilegedParams: [{ param: 'alwaysApprove', clampTo: false }],
|
||||
// Web-researched, unverified.
|
||||
customModelInjection: {
|
||||
kind: 'env',
|
||||
baseUrlVar: 'GROK_BASE_URL',
|
||||
apiKeyVar: 'XAI_API_KEY',
|
||||
modelVars: ['GROK_MODEL'],
|
||||
},
|
||||
// All three already match the GROK_/XAI_ allowedPrefixes above, so they were ALREADY
|
||||
// reachable via plain envOverrides before this feature existed.
|
||||
privilegedEnvKeys: ['GROK_BASE_URL', 'XAI_API_KEY', 'GROK_MODEL'],
|
||||
},
|
||||
overlays: {
|
||||
// ~/.grok also holds sessions/, memory/, downloads/ (the ~160MB binary), completions/,
|
||||
@@ -924,7 +1019,19 @@ const DEEPSEEK: CliEntry = {
|
||||
// The half no other CLI needs. `DSH_*` is an allowlisted envOverrides prefix and
|
||||
// applyEnvOverrides() runs LAST, so without this a non-granted owner could send
|
||||
// DSH_PERMISSION_MODE on the same request and land after the config clamp.
|
||||
privilegedEnvKeys: ['DSH_PERMISSION_MODE', 'DSH_HOME', 'DEEPSEEK_BASE_URL'],
|
||||
// DEEPSEEK_API_KEY added alongside DEEPSEEK_BASE_URL for the custom-model-injection.ts
|
||||
// recipe (deployment_plan.md) — the pair travels together, same reasoning as base URL.
|
||||
privilegedEnvKeys: ['DSH_PERMISSION_MODE', 'DSH_HOME', 'DEEPSEEK_BASE_URL', 'DEEPSEEK_API_KEY'],
|
||||
// Web-researched, unverified, partial: reuses the already-existing DEEPSEEK_BASE_URL/
|
||||
// DEEPSEEK_API_KEY keys above. No modelVars — dsh's model is a profile-composition
|
||||
// entry (see `model: { source: 'none' }` above), not an env var, so forcing a specific
|
||||
// model name may not fully work; verify against a real profile before shipping.
|
||||
customModelInjection: {
|
||||
kind: 'env',
|
||||
baseUrlVar: 'DEEPSEEK_BASE_URL',
|
||||
apiKeyVar: 'DEEPSEEK_API_KEY',
|
||||
modelVars: [],
|
||||
},
|
||||
},
|
||||
overlays: {
|
||||
// No credStore: dsh keeps everything under $DSH_HOME (default ~/.dsh), which is
|
||||
@@ -1026,7 +1133,20 @@ const OMP: CliEntry = {
|
||||
// Where omp resolves its auth from. No known concrete exfiltration path today (omp
|
||||
// forwards no operator-held key into a pane), but a non-granted owner redirecting where
|
||||
// a shared multi-tenant deployment resolves auth is not something to allow silently.
|
||||
privilegedEnvKeys: ['OMP_AUTH_BROKER_URL', 'OMP_AUTH_BROKER_TOKEN'],
|
||||
// PI_CONFIG_DIR added for custom-model-injection.ts's omp recipe, which reuses pi's
|
||||
// dir-redirect mechanism (see the customModelInjection comment below) — already
|
||||
// reachable via the PI_ allowedPrefix (pi's own entry), so this closes the same
|
||||
// pre-existing gap for an omp session that PI's own entry closes for a pi session.
|
||||
privilegedEnvKeys: ['OMP_AUTH_BROKER_URL', 'OMP_AUTH_BROKER_TOKEN', 'PI_CONFIG_DIR'],
|
||||
// Web-researched, unverified. omp's ~/.omp tree is itself relocatable via PI_CONFIG_DIR
|
||||
// (see the DeepSeek/OMP note in CLAUDE.md), so this reuses that same redirect rather
|
||||
// than inventing an OMP-specific dir env var.
|
||||
customModelInjection: {
|
||||
kind: 'configDir',
|
||||
dirEnvVar: 'PI_CONFIG_DIR',
|
||||
fileName: 'agent/models.yml',
|
||||
template: 'omp-models-yml',
|
||||
},
|
||||
},
|
||||
overlays: {
|
||||
// `~/.omp/agent` also holds agent.db/history.db/models.db (SQLite caches) and
|
||||
|
||||
@@ -441,6 +441,37 @@ export interface CliCapabilities {
|
||||
gates: Record<string, { minVersion: string; failClosed: boolean }>;
|
||||
/** Cap on a single terminal frame, when this CLI needs a tighter one than the default. */
|
||||
maxFrameBytes?: number;
|
||||
/**
|
||||
* How this CLI is pointed at a user-supplied custom OpenAI-compatible
|
||||
* endpoint (local, e.g. llama.cpp, or cloud, e.g. Azure AI Foundry) — the
|
||||
* Custom Model Endpoint Profiles feature (`deployment_plan.md`). Declared
|
||||
* per entry, never branched on id, same as every other capability here.
|
||||
*
|
||||
* `env`: plain env vars (claude's `ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY`/
|
||||
* `ANTHROPIC_DEFAULT_*_MODEL`). `configContentEnv`: a full config blob
|
||||
* carried in one env var (opencode's `OPENCODE_CONFIG_CONTENT`).
|
||||
* `configDir`: a generated config file under an isolated, dir-redirect-env-
|
||||
* pointed directory so the user's real CLI config is never touched
|
||||
* (codex's `CODEX_HOME`/`config.toml`, pi/omp's `PI_CONFIG_DIR`).
|
||||
* `unsupported`: no known mechanism (antigravity) — the toolbar entry
|
||||
* stays disabled for this CLI.
|
||||
*
|
||||
* Every env var name this introduces that can redirect a session's
|
||||
* traffic MUST also appear in `privilegedEnvKeys` above, exactly like
|
||||
* `DEEPSEEK_BASE_URL` — a non-granted multi-user owner redirecting a
|
||||
* session to their own endpoint is a credential-exfiltration path, not
|
||||
* just a mischief redirect.
|
||||
*/
|
||||
customModelInjection:
|
||||
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[] }
|
||||
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json' }
|
||||
| {
|
||||
kind: 'configDir';
|
||||
dirEnvVar: string;
|
||||
fileName: string;
|
||||
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml';
|
||||
}
|
||||
| { kind: 'unsupported' };
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
/**
|
||||
* @fileoverview Read/write-array store for user-configured custom OpenAI-compatible
|
||||
* model endpoints (local or cloud — deployment_plan.md). Same shape as
|
||||
* `remote-hosts.ts` / `webview-store.ts`: `~/.codeman/custom-model-hosts.json`
|
||||
* holding a plain array, read/written whole.
|
||||
*/
|
||||
|
||||
import { existsSync, mkdirSync } from 'node:fs';
|
||||
import fs from 'node:fs/promises';
|
||||
import { join } from 'node:path';
|
||||
|
||||
const CUSTOM_MODEL_HOSTS_FILE = 'custom-model-hosts.json';
|
||||
|
||||
export type CustomModelAuthStyle = 'bearer' | 'api-key';
|
||||
|
||||
export interface CustomModelHost {
|
||||
id: string;
|
||||
label: string;
|
||||
/** Root URL, local or cloud — e.g. "http://192.168.1.50:8080" or an Azure AI Foundry URL. */
|
||||
baseUrl: string;
|
||||
apiKey?: string;
|
||||
/**
|
||||
* Defaults to 'bearer' (the common `Authorization: Bearer` convention — matches
|
||||
* llama.cpp, OpenAI-compatible servers, and most gateways). Pick 'api-key' for
|
||||
* endpoints that specifically want the `api-key` header, e.g. Azure AI Foundry.
|
||||
*
|
||||
* ⚠️ There is deliberately NO 'both' option. An earlier design sent BOTH headers
|
||||
* on every discovery request on the theory that an unused header is harmless —
|
||||
* live-tested against a real llama-swap server, sending both reliably HUNG the
|
||||
* request indefinitely (reproduced 3× — Bearer alone: ~500ms, api-key alone:
|
||||
* ~600ms, both together: no response inside a 15s timeout). Whatever auth
|
||||
* middleware some servers run apparently does not handle two simultaneous
|
||||
* credential conventions gracefully, so "send everything and let the server
|
||||
* ignore what it doesn't need" is not a safe default — it can silently turn a
|
||||
* working endpoint into one that always times out.
|
||||
*/
|
||||
authStyle?: CustomModelAuthStyle;
|
||||
models?: string[];
|
||||
lastDiscoveredAt?: string;
|
||||
}
|
||||
|
||||
export function customModelHostsPath(configDir: string): string {
|
||||
return join(configDir, CUSTOM_MODEL_HOSTS_FILE);
|
||||
}
|
||||
|
||||
export async function readCustomModelHosts(configDir: string): Promise<CustomModelHost[]> {
|
||||
try {
|
||||
const raw = await fs.readFile(customModelHostsPath(configDir), 'utf-8');
|
||||
const parsed = JSON.parse(raw);
|
||||
return Array.isArray(parsed) ? (parsed as CustomModelHost[]) : [];
|
||||
} catch {
|
||||
return [];
|
||||
}
|
||||
}
|
||||
|
||||
export async function writeCustomModelHosts(configDir: string, hosts: CustomModelHost[]): Promise<void> {
|
||||
if (!existsSync(configDir)) mkdirSync(configDir, { recursive: true });
|
||||
await fs.writeFile(customModelHostsPath(configDir), JSON.stringify(hosts, null, 2));
|
||||
}
|
||||
@@ -0,0 +1,188 @@
|
||||
/**
|
||||
* @fileoverview Pure builder for the Custom Model Endpoint Profiles feature
|
||||
* (deployment_plan.md): turns a CLI registry entry's
|
||||
* `capabilities.customModelInjection` declaration, a configured endpoint,
|
||||
* and a chosen model id into the concrete env vars / config-file content
|
||||
* that would redirect that CLI's session at the endpoint.
|
||||
*
|
||||
* No IO here on purpose (mirrors `session-cli-builder.ts`) — a caller
|
||||
* writes `ConfigDirInjection.files` to disk under an isolated per-session
|
||||
* directory and points `dirEnvVar` at it; this module only computes what
|
||||
* those files/env vars should contain.
|
||||
*
|
||||
* Confidence: `claude` and `opencode` are verified end-to-end against a real
|
||||
* llama-swap server (a real "hello world" reply came back). `codex`'s
|
||||
* config.toml STRUCTURE is now verified (an earlier `[model].default` table
|
||||
* shape was rejected by a real codex binary with "invalid type: map,
|
||||
* expected a string" — caught by `scripts/test-local-llm-harnesses.mjs`),
|
||||
* but `wire_api = "responses"` is the only value codex still accepts
|
||||
* (support for `"chat"` was dropped in Feb 2026), and a plain OpenAI
|
||||
* Chat-Completions server (llama.cpp, llama-swap, most local setups) does
|
||||
* NOT implement the Responses API — so codex may still fail at the
|
||||
* PROTOCOL level even with a correctly-shaped config file. That gap is
|
||||
* real and current, not a stale warning; see deployment_plan.md. The rest
|
||||
* (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION flags
|
||||
* confirmed against real installed binaries' own `--help` output, but
|
||||
* their custom-endpoint env/config conventions remain web-researched,
|
||||
* unverified.
|
||||
*/
|
||||
|
||||
import type { CliEntry } from './config/cli-registry/types.js';
|
||||
|
||||
export interface CustomModelEndpoint {
|
||||
id: string;
|
||||
label: string;
|
||||
/** Root URL, no trailing slash required — e.g. "http://192.168.1.50:8080" or an Azure AI Foundry URL. */
|
||||
baseUrl: string;
|
||||
/** Falls back to a harmless placeholder for endpoints (llama.cpp) that don't check it. */
|
||||
apiKey?: string;
|
||||
}
|
||||
|
||||
export interface EnvInjection {
|
||||
kind: 'env';
|
||||
/** Ready to merge into a session's envOverrides. */
|
||||
envOverrides: Record<string, string>;
|
||||
}
|
||||
|
||||
export interface ConfigDirInjection {
|
||||
kind: 'configDir';
|
||||
/** Env var that must be set to the directory the caller writes `files` under. */
|
||||
dirEnvVar: string;
|
||||
files: Array<{ relPath: string; content: string }>;
|
||||
/**
|
||||
* Env vars the written config file REFERENCES by name rather than embedding a
|
||||
* literal value (codex's `env_key = "..."` convention: config.toml never carries
|
||||
* the API key itself, only the name of an env var codex reads it from). Merge
|
||||
* these into the session's envOverrides alongside `dirEnvVar` — never skip them,
|
||||
* or the config points at a credential that was never actually set.
|
||||
*/
|
||||
extraEnv?: Record<string, string>;
|
||||
}
|
||||
|
||||
export interface UnsupportedInjection {
|
||||
kind: 'unsupported';
|
||||
}
|
||||
|
||||
export type CustomModelInjectionResult = EnvInjection | ConfigDirInjection | UnsupportedInjection;
|
||||
|
||||
const DEFAULT_API_KEY = 'local-dummy-key';
|
||||
|
||||
/** Normalizes a base URL to end in exactly one trailing `/v1`, for CLIs whose config expects the OpenAI-style suffix. */
|
||||
export function withV1Suffix(baseUrl: string): string {
|
||||
const trimmed = baseUrl.replace(/\/+$/, '');
|
||||
return /\/v1$/.test(trimmed) ? trimmed : `${trimmed}/v1`;
|
||||
}
|
||||
|
||||
/** JSON-escapes a string for embedding in a TOML/YAML double-quoted scalar — a safe superset of both grammars' basic escapes. */
|
||||
function quoted(value: string): string {
|
||||
return JSON.stringify(value);
|
||||
}
|
||||
|
||||
export function buildCustomModelInjection(
|
||||
entry: Pick<CliEntry, 'capabilities'>,
|
||||
endpoint: CustomModelEndpoint,
|
||||
modelId: string
|
||||
): CustomModelInjectionResult {
|
||||
const cap = entry.capabilities.customModelInjection;
|
||||
const apiKey = endpoint.apiKey?.trim() || DEFAULT_API_KEY;
|
||||
|
||||
switch (cap.kind) {
|
||||
case 'env': {
|
||||
const envOverrides: Record<string, string> = {
|
||||
[cap.baseUrlVar]: endpoint.baseUrl,
|
||||
[cap.apiKeyVar]: apiKey,
|
||||
};
|
||||
for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId;
|
||||
return { kind: 'env', envOverrides };
|
||||
}
|
||||
|
||||
case 'configContentEnv': {
|
||||
const content = renderConfigContent(cap.template, endpoint, modelId, apiKey);
|
||||
return { kind: 'env', envOverrides: { [cap.envVar]: content } };
|
||||
}
|
||||
|
||||
case 'configDir': {
|
||||
const { content, extraEnv } = renderConfigFile(cap.template, endpoint, modelId, apiKey);
|
||||
return { kind: 'configDir', dirEnvVar: cap.dirEnvVar, files: [{ relPath: cap.fileName, content }], extraEnv };
|
||||
}
|
||||
|
||||
case 'unsupported':
|
||||
return { kind: 'unsupported' };
|
||||
}
|
||||
}
|
||||
|
||||
function renderConfigContent(
|
||||
template: 'opencode-json',
|
||||
endpoint: CustomModelEndpoint,
|
||||
modelId: string,
|
||||
apiKey: string
|
||||
): string {
|
||||
switch (template) {
|
||||
case 'opencode-json':
|
||||
return JSON.stringify({
|
||||
$schema: 'https://opencode.ai/config.json',
|
||||
provider: {
|
||||
custom: {
|
||||
options: { baseURL: withV1Suffix(endpoint.baseUrl), apiKey },
|
||||
models: { [modelId]: {} },
|
||||
},
|
||||
},
|
||||
model: `custom/${modelId}`,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
const CODEX_API_KEY_ENV_VAR = 'CODEMAN_CUSTOM_MODEL_API_KEY';
|
||||
|
||||
function renderConfigFile(
|
||||
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml',
|
||||
endpoint: CustomModelEndpoint,
|
||||
modelId: string,
|
||||
apiKey: string
|
||||
): { content: string; extraEnv?: Record<string, string> } {
|
||||
const baseUrl = withV1Suffix(endpoint.baseUrl);
|
||||
switch (template) {
|
||||
case 'codex-toml': {
|
||||
// Verified against real codex (>= Feb 2026): `model` is a top-level STRING, never
|
||||
// a `[model].default` table — codex rejects that with "invalid type: map, expected
|
||||
// a string" (caught by scripts/test-local-llm-harnesses.mjs against a real llama-swap
|
||||
// server). The API key is NEVER a literal TOML field: codex's schema only supports
|
||||
// `env_key`, the NAME of an env var it reads the credential from at runtime, so the
|
||||
// actual value must ride along as an extra env var, never embedded in the file.
|
||||
// ⚠️ `wire_api = "responses"` is the only value codex still accepts (it dropped
|
||||
// `"chat"` support in Feb 2026) — a plain OpenAI Chat-Completions server (llama.cpp,
|
||||
// llama-swap, most local setups) does NOT implement the Responses API, so this
|
||||
// recipe may still fail at the PROTOCOL level even though the file now parses
|
||||
// correctly. That is a real, currently-unresolved compatibility gap, not a syntax
|
||||
// bug — track it before calling codex support done.
|
||||
const content = [
|
||||
`model = ${quoted(modelId)}`,
|
||||
`model_provider = "custom"`,
|
||||
'',
|
||||
'[model_providers.custom]',
|
||||
`name = "Custom Endpoint"`,
|
||||
`base_url = ${quoted(baseUrl)}`,
|
||||
`env_key = ${quoted(CODEX_API_KEY_ENV_VAR)}`,
|
||||
`wire_api = "responses"`,
|
||||
'',
|
||||
].join('\n');
|
||||
return { content, extraEnv: { [CODEX_API_KEY_ENV_VAR]: apiKey } };
|
||||
}
|
||||
case 'pi-models-json':
|
||||
return {
|
||||
content: JSON.stringify(
|
||||
{
|
||||
providers: {
|
||||
custom: { baseUrl, apiKey, api: 'openai-completions', models: { [modelId]: {} } },
|
||||
},
|
||||
},
|
||||
null,
|
||||
2
|
||||
),
|
||||
};
|
||||
case 'omp-models-yml':
|
||||
return {
|
||||
content: `providers:\n custom:\n baseUrl: ${quoted(baseUrl)}\n apiKey: ${quoted(apiKey)}\n models:\n - ${quoted(modelId)}\n`,
|
||||
};
|
||||
}
|
||||
}
|
||||
+83
-1
@@ -577,6 +577,14 @@ export class Session extends EventEmitter {
|
||||
// the CLAUDE_CODE_EFFORT_LEVEL env var, which would hard-lock the session.
|
||||
private _effort: EffortLevel | undefined;
|
||||
|
||||
// Custom Model Endpoint Profiles (deployment_plan.md). `envKeys` and `configDir` are
|
||||
// internal bookkeeping ONLY (never surfaced via toState()/customModel getter): they are
|
||||
// what setCustomModel() needs to undo a previous injection (remove exactly the env keys
|
||||
// it added, delete a previous isolated config dir) without guessing what it once wrote.
|
||||
private _customModel:
|
||||
| { endpointId: string; modelId: string; label?: string; envKeys: string[]; configDir?: string }
|
||||
| undefined;
|
||||
|
||||
// tmux history-limit (scrollback lines) allocated when this session's pane is created.
|
||||
private readonly _tmuxHistoryLimit: number;
|
||||
|
||||
@@ -1238,6 +1246,40 @@ export class Session extends EventEmitter {
|
||||
}
|
||||
}
|
||||
|
||||
// Custom Model Endpoint Profiles (deployment_plan.md) — public-safe subset only
|
||||
// (never envKeys/configDir, which are internal bookkeeping for setCustomModel below).
|
||||
get customModel(): { endpointId: string; modelId: string; label?: string } | undefined {
|
||||
if (!this._customModel) return undefined;
|
||||
const { endpointId, modelId, label } = this._customModel;
|
||||
return { endpointId, modelId, label };
|
||||
}
|
||||
|
||||
/**
|
||||
* Update this session's custom-model selection and merge the endpoint's injected env
|
||||
* vars into `_envOverrides` — first UNDOING whatever the previous selection injected
|
||||
* (removing exactly those env keys), so switching endpoints, or clearing back to the
|
||||
* harness's native cloud default, never leaves a stale key behind. Synchronous and
|
||||
* side-effect-free beyond mutating state, matching `setNice`/`setColor` above — this
|
||||
* class does no file IO, so it returns the PREVIOUS `configDir` (if any) for the
|
||||
* caller to clean up on disk (custom-model-injection.ts's configDir kind).
|
||||
*/
|
||||
setCustomModel(
|
||||
next: { endpointId: string; modelId: string; label?: string; envKeys: string[]; configDir?: string } | undefined,
|
||||
envOverrides?: Record<string, string>
|
||||
): string | undefined {
|
||||
const previousConfigDir = this._customModel?.configDir;
|
||||
if (this._customModel) {
|
||||
for (const key of this._customModel.envKeys) {
|
||||
if (this._envOverrides) delete this._envOverrides[key];
|
||||
}
|
||||
}
|
||||
this._customModel = next;
|
||||
if (envOverrides && Object.keys(envOverrides).length > 0) {
|
||||
this._envOverrides = { ...(this._envOverrides ?? {}), ...envOverrides };
|
||||
}
|
||||
return previousConfigDir;
|
||||
}
|
||||
|
||||
// Token tracking getters and setters
|
||||
get totalTokens(): number {
|
||||
return this._totalInputTokens + this._totalOutputTokens;
|
||||
@@ -1478,6 +1520,7 @@ export class Session extends EventEmitter {
|
||||
ompConfig: this._ompConfig,
|
||||
resumeSessionId: this._resumeSessionId,
|
||||
effort: this._effort,
|
||||
customModel: this.customModel,
|
||||
// COD-118: runtime-only — surfaced so the frontend can require explicit user
|
||||
// intent before restarting a crash-looped session. Deliberately NOT restored
|
||||
// by the constructor: a Codeman restart starts with a fresh breaker so boot
|
||||
@@ -1710,10 +1753,49 @@ export class Session extends EventEmitter {
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Kill and relaunch this session's CLI process IN PLACE — same pane, same tmux
|
||||
* session, fresh env/args from current state. Custom Model Endpoint Profiles
|
||||
* (deployment_plan.md) is the first caller: after `setCustomModel()` merges new
|
||||
* env vars into `_envOverrides`, the running CLI process still has the OLD env
|
||||
* (inherited at its own process start, not live-reloaded), so switching a
|
||||
* session's model/endpoint requires this restart to actually take effect.
|
||||
*
|
||||
* A GENERALIZED {@link reattachRemote} with the `!this._remote` guard dropped —
|
||||
* `_buildRespawnPaneOptions()` already passes `remote: this._remote` through
|
||||
* unconditionally, so `mux.respawnPane()` builds the right command either way
|
||||
* (a local session gets `respawn-pane -k` + the real launch line, which is the
|
||||
* kill-and-relaunch this method exists for; a remote session gets the existing
|
||||
* reattach-to-durable-tmux behavior). Deliberately does NOT check `isBusy()` —
|
||||
* that's the caller's job (mirrors `/interactive`'s guard), since a raw restart
|
||||
* primitive shouldn't itself decide when it's safe to use.
|
||||
*
|
||||
* @returns true if the pane was respawned, false otherwise (no mux session, or
|
||||
* the mux session is gone — see {@link reattachRemote} for that reasoning).
|
||||
*/
|
||||
async restartCli(): Promise<boolean> {
|
||||
if (!this._useMux || !this._mux || !this._muxSession) return false;
|
||||
const mux = this._mux;
|
||||
|
||||
if (!mux.muxSessionExists(this._muxSession.muxName)) {
|
||||
console.log('[Session] restartCli: mux session gone, skipping:', this._muxSession.muxName);
|
||||
return false;
|
||||
}
|
||||
|
||||
this._pinOmpRespawnId();
|
||||
const newPid = await mux.respawnPane(this._buildRespawnPaneOptions());
|
||||
if (!newPid) {
|
||||
console.error('[Session] restartCli: respawnPane failed for', this._muxSession.muxName);
|
||||
return false;
|
||||
}
|
||||
console.log('[Session] restartCli: restarted CLI for', this._muxSession.muxName, 'pid', newPid);
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Assemble the {@link RespawnPaneOptions} for this session. Single source of
|
||||
* truth shared by interactive start, shell start (via their inline copies),
|
||||
* and {@link reattachRemote} so the remote reattach path can never drift from
|
||||
* {@link reattachRemote}, and {@link restartCli} so no respawn path can drift from
|
||||
* the spawn path.
|
||||
*/
|
||||
private _buildRespawnPaneOptions(): import('./mux-interface.js').RespawnPaneOptions {
|
||||
|
||||
@@ -677,6 +677,13 @@ export interface SessionState {
|
||||
resumeSessionId?: string;
|
||||
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
|
||||
effort?: EffortLevel;
|
||||
/**
|
||||
* Custom Model Endpoint Profiles (deployment_plan.md): the custom OpenAI-compatible
|
||||
* endpoint (local or cloud) this session's CLI is currently pointed at, if any.
|
||||
* Undefined = the harness's native cloud default. No secrets here — the endpoint's
|
||||
* base URL/api key live only in Session._envOverrides, never in this public state.
|
||||
*/
|
||||
customModel?: { endpointId: string; modelId: string; label?: string };
|
||||
/** Sanitized per-session attachment history. */
|
||||
attachmentHistory?: SessionAttachmentHistoryItem[];
|
||||
/**
|
||||
|
||||
@@ -0,0 +1,128 @@
|
||||
/**
|
||||
* @fileoverview Custom Model Endpoint Profiles CRUD + discovery
|
||||
* (deployment_plan.md). Endpoints are machine-level infra, like remote/docker
|
||||
* hosts, so writes are admin-only in multi-user mode
|
||||
* (`case-routes.ts`'s `/api/remote-hosts` is the pattern this mirrors).
|
||||
*
|
||||
* Discovery (`POST /:id/discover-models`) fetches `${baseUrl}/v1/models`.
|
||||
* `isBlockedWebviewUrl()` is the same synchronous hostname/link-local/cloud-
|
||||
* metadata check `webview-egress-policy.ts` uses for saved dashboard URLs —
|
||||
* reused here as a save-time and discover-time guard. It does NOT re-check
|
||||
* the DNS-RESOLVED address the way `webviewFetch()`'s undici lookup hook
|
||||
* does; wiring that dispatcher-level guard here is a followup, not done in
|
||||
* this pass, since this route is already admin-only in multi-user mode.
|
||||
*/
|
||||
|
||||
import type { FastifyInstance, FastifyRequest } from 'fastify';
|
||||
import { ApiErrorCode, createErrorResponse, type ApiResponse } from '../../types.js';
|
||||
import { isAdmin, parseBody } from '../route-helpers.js';
|
||||
import { isMultiUserMode } from '../../config/multiuser.js';
|
||||
import { getDataDir } from '../../config/instance.js';
|
||||
import { isBlockedWebviewUrl } from '../webview-egress-policy.js';
|
||||
import { CustomModelHostSchema } from '../schemas.js';
|
||||
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../../custom-model-hosts.js';
|
||||
|
||||
const CODEMAN_CONFIG_DIR = getDataDir();
|
||||
const DISCOVER_TIMEOUT_MS = 8000;
|
||||
|
||||
function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }): ApiResponse<never> | null {
|
||||
if (!isMultiUserMode() || isAdmin(req)) return null;
|
||||
reply.code(403);
|
||||
return createErrorResponse(ApiErrorCode.FORBIDDEN, 'Admin only in multi-user mode');
|
||||
}
|
||||
|
||||
async function discoverModels(host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' | 'authStyle'>): Promise<string[]> {
|
||||
const headers: Record<string, string> = {};
|
||||
const apiKey = host.apiKey?.trim();
|
||||
// Exactly ONE header, never both — see custom-model-hosts.ts's CustomModelAuthStyle
|
||||
// doc comment for why: sending both reliably HANGS some real servers.
|
||||
const style = host.authStyle ?? 'bearer';
|
||||
if (apiKey && style === 'bearer') headers.Authorization = `Bearer ${apiKey}`;
|
||||
if (apiKey && style === 'api-key') headers['api-key'] = apiKey;
|
||||
|
||||
const res = await fetch(`${host.baseUrl.replace(/\/+$/, '')}/v1/models`, {
|
||||
headers,
|
||||
signal: AbortSignal.timeout(DISCOVER_TIMEOUT_MS),
|
||||
});
|
||||
if (!res.ok) throw new Error(`HTTP ${res.status}`);
|
||||
const body = (await res.json()) as { data?: Array<{ id?: unknown }> };
|
||||
return (body.data ?? []).map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0);
|
||||
}
|
||||
|
||||
export function registerCustomModelRoutes(app: FastifyInstance): void {
|
||||
app.get('/api/model-endpoints', async (req) =>
|
||||
isMultiUserMode() && !isAdmin(req) ? [] : readCustomModelHosts(CODEMAN_CONFIG_DIR)
|
||||
);
|
||||
|
||||
app.post('/api/model-endpoints', async (req, reply): Promise<ApiResponse<{ host: CustomModelHost }>> => {
|
||||
const denied = adminOnly(req, reply);
|
||||
if (denied) return denied;
|
||||
const host = parseBody(CustomModelHostSchema, req.body);
|
||||
if (isBlockedWebviewUrl(host.baseUrl)) {
|
||||
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
|
||||
}
|
||||
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
|
||||
if (hosts.some((item) => item.id === host.id)) {
|
||||
return createErrorResponse(ApiErrorCode.ALREADY_EXISTS, 'Model endpoint already exists');
|
||||
}
|
||||
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, [...hosts, host]);
|
||||
return { success: true, data: { host } };
|
||||
});
|
||||
|
||||
app.put('/api/model-endpoints/:id', async (req, reply): Promise<ApiResponse<{ host: CustomModelHost }>> => {
|
||||
const denied = adminOnly(req, reply);
|
||||
if (denied) return denied;
|
||||
const { id } = req.params as { id: string };
|
||||
const host = parseBody(CustomModelHostSchema, { ...(req.body as object), id });
|
||||
if (isBlockedWebviewUrl(host.baseUrl)) {
|
||||
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
|
||||
}
|
||||
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
|
||||
const index = hosts.findIndex((item) => item.id === id);
|
||||
if (index === -1) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
|
||||
const next = [...hosts];
|
||||
next[index] = host;
|
||||
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next);
|
||||
return { success: true, data: { host } };
|
||||
});
|
||||
|
||||
app.delete('/api/model-endpoints/:id', async (req, reply): Promise<ApiResponse<{ id: string }>> => {
|
||||
const denied = adminOnly(req, reply);
|
||||
if (denied) return denied;
|
||||
const { id } = req.params as { id: string };
|
||||
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
|
||||
await writeCustomModelHosts(
|
||||
CODEMAN_CONFIG_DIR,
|
||||
hosts.filter((item) => item.id !== id)
|
||||
);
|
||||
return { success: true, data: { id } };
|
||||
});
|
||||
|
||||
app.post(
|
||||
'/api/model-endpoints/:id/discover-models',
|
||||
async (req, reply): Promise<ApiResponse<{ models: string[] }>> => {
|
||||
const denied = adminOnly(req, reply);
|
||||
if (denied) return denied;
|
||||
const { id } = req.params as { id: string };
|
||||
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
|
||||
const index = hosts.findIndex((item) => item.id === id);
|
||||
if (index === -1) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
|
||||
const host = hosts[index];
|
||||
if (isBlockedWebviewUrl(host.baseUrl)) {
|
||||
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
|
||||
}
|
||||
try {
|
||||
const models = await discoverModels(host);
|
||||
const next = [...hosts];
|
||||
next[index] = { ...host, models, lastDiscoveredAt: new Date().toISOString() };
|
||||
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next);
|
||||
return { success: true, data: { models } };
|
||||
} catch (err) {
|
||||
return createErrorResponse(
|
||||
ApiErrorCode.OPERATION_FAILED,
|
||||
`Could not reach endpoint: ${err instanceof Error ? err.message : String(err)}`
|
||||
);
|
||||
}
|
||||
}
|
||||
);
|
||||
}
|
||||
@@ -27,3 +27,4 @@ export { registerWsRoutes } from './ws-routes.js';
|
||||
export { registerVoiceRoutes } from './voice-routes.js';
|
||||
export { registerWebviewRoutes, tryWebviewRefererFallback } from './webview-routes.js';
|
||||
export { registerTabLayoutRoutes } from './tab-layout-routes.js';
|
||||
export { registerCustomModelRoutes } from './custom-model-routes.js';
|
||||
|
||||
@@ -8,7 +8,7 @@ import { FastifyInstance, type FastifyReply } from 'fastify';
|
||||
import { z } from 'zod';
|
||||
import { join, dirname, extname, basename } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
import { existsSync, statSync, mkdirSync, writeFileSync } from 'node:fs';
|
||||
import { existsSync, statSync, mkdirSync, writeFileSync, rmSync } from 'node:fs';
|
||||
import { execFile } from 'node:child_process';
|
||||
import fs from 'node:fs/promises';
|
||||
import { randomBytes } from 'node:crypto';
|
||||
@@ -51,7 +51,10 @@ import {
|
||||
SessionOrderUpdateSchema,
|
||||
SessionWaitQuerySchema,
|
||||
SessionWaitOutputQuerySchema,
|
||||
CustomModelSelectionSchema,
|
||||
} from '../schemas.js';
|
||||
import { readCustomModelHosts } from '../../custom-model-hosts.js';
|
||||
import { buildCustomModelInjection } from '../../custom-model-injection.js';
|
||||
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
|
||||
import { TabLayoutValidationError } from '../../tab-layout.js';
|
||||
import {
|
||||
@@ -1161,6 +1164,88 @@ export function registerSessionRoutes(
|
||||
return { color: session.color };
|
||||
});
|
||||
|
||||
// ========== Custom Model Endpoint Profiles (deployment_plan.md) ==========
|
||||
//
|
||||
// Applies (or clears) a session's custom OpenAI-compatible endpoint selection and
|
||||
// RESTARTS the pane's CLI process — these harnesses read endpoint config at process
|
||||
// start, not per-turn, so a live hot-swap isn't possible (confirmed with the
|
||||
// maintainer). Endpoints come from the admin-configured custom-model-hosts store
|
||||
// (chunk 3's CRUD routes), never raw client-supplied env — that's what keeps this
|
||||
// route safe to let any session owner call for their own session, unlike the
|
||||
// generic envOverrides field the privilegedEnvKeys clamp exists to guard.
|
||||
app.post('/api/sessions/:id/custom-model', async (req) => {
|
||||
const { id } = req.params as { id: string };
|
||||
const body = parseBody(CustomModelSelectionSchema, req.body, 'Invalid request body');
|
||||
const session = findSessionOrFail(ctx, id, req);
|
||||
|
||||
if (session.isBusy()) {
|
||||
return createErrorResponse(ApiErrorCode.SESSION_BUSY, 'Session is busy');
|
||||
}
|
||||
|
||||
if ('clear' in body) {
|
||||
const previousConfigDir = session.setCustomModel(undefined);
|
||||
if (previousConfigDir) rmSync(previousConfigDir, { recursive: true, force: true });
|
||||
const restarted = await session.restartCli();
|
||||
persistAndBroadcastSession(ctx, session);
|
||||
return { customModel: session.customModel, restarted };
|
||||
}
|
||||
|
||||
const entry = getCli(session.mode);
|
||||
if (!entry) {
|
||||
return createErrorResponse(ApiErrorCode.INVALID_INPUT, `No CLI registry entry for mode ${session.mode}`);
|
||||
}
|
||||
if (entry.capabilities.customModelInjection.kind === 'unsupported') {
|
||||
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
|
||||
}
|
||||
|
||||
const hosts = await readCustomModelHosts(getDataDir());
|
||||
const endpoint = hosts.find((h) => h.id === body.endpointId);
|
||||
if (!endpoint) {
|
||||
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
|
||||
}
|
||||
|
||||
const injection = buildCustomModelInjection(entry, endpoint, body.modelId);
|
||||
|
||||
let envOverrides: Record<string, string>;
|
||||
let envKeys: string[];
|
||||
let configDir: string | undefined;
|
||||
|
||||
if (injection.kind === 'env') {
|
||||
envOverrides = injection.envOverrides;
|
||||
envKeys = Object.keys(injection.envOverrides);
|
||||
} else if (injection.kind === 'configDir') {
|
||||
// Isolated per-session dir — never the user's real CLI config path.
|
||||
configDir = join(dataPath('custom-model-configs'), session.id);
|
||||
for (const file of injection.files) {
|
||||
const filePath = join(configDir, file.relPath);
|
||||
mkdirSync(dirname(filePath), { recursive: true });
|
||||
writeFileSync(filePath, file.content, 'utf8');
|
||||
}
|
||||
// extraEnv: vars the written config file REFERENCES by name (codex's `env_key`
|
||||
// convention) rather than embedding a literal value — must ride alongside
|
||||
// dirEnvVar or the config points at a credential that was never actually set.
|
||||
envOverrides = { [injection.dirEnvVar]: configDir, ...injection.extraEnv };
|
||||
envKeys = [injection.dirEnvVar, ...Object.keys(injection.extraEnv ?? {})];
|
||||
} else {
|
||||
// 'unsupported' is already handled above; this keeps the switch exhaustive.
|
||||
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
|
||||
}
|
||||
|
||||
const previousConfigDir = session.setCustomModel(
|
||||
{ endpointId: endpoint.id, modelId: body.modelId, label: endpoint.label, envKeys, configDir },
|
||||
envOverrides
|
||||
);
|
||||
// Clean up the OLD config dir on disk, unless the new one happens to reuse the same
|
||||
// path (same session, configDir kind again) — never delete the dir we just wrote.
|
||||
if (previousConfigDir && previousConfigDir !== configDir) {
|
||||
rmSync(previousConfigDir, { recursive: true, force: true });
|
||||
}
|
||||
|
||||
const restarted = await session.restartCli();
|
||||
persistAndBroadcastSession(ctx, session);
|
||||
return { customModel: session.customModel, restarted };
|
||||
});
|
||||
|
||||
// ========== Delete Session ==========
|
||||
|
||||
app.delete('/api/sessions/:id', async (req) => {
|
||||
|
||||
@@ -741,6 +741,29 @@ export const RemoteHostSchema = z.object({
|
||||
commands: RemoteCommandOverridesSchema,
|
||||
});
|
||||
|
||||
// Custom Model Endpoint Profiles (deployment_plan.md) — a user-configured custom
|
||||
// OpenAI-compatible endpoint, local (llama.cpp) or cloud (Azure AI Foundry, etc.).
|
||||
export const CustomModelHostSchema = z.object({
|
||||
id: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
|
||||
label: z.string().min(1).max(100),
|
||||
baseUrl: z.string().url().max(2048),
|
||||
apiKey: z.string().max(4096).optional(),
|
||||
// No 'both': live-tested against a real server, sending both auth header
|
||||
// conventions on one request reliably HANGS it — see custom-model-hosts.ts.
|
||||
authStyle: z.enum(['bearer', 'api-key']).optional(),
|
||||
models: z.array(z.string().max(200)).max(200).optional(),
|
||||
lastDiscoveredAt: z.string().max(64).optional(),
|
||||
});
|
||||
|
||||
/** POST /api/sessions/:id/custom-model — apply or clear a session's custom-model selection. */
|
||||
export const CustomModelSelectionSchema = z.union([
|
||||
z.object({
|
||||
endpointId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
|
||||
modelId: z.string().min(1).max(200),
|
||||
}),
|
||||
z.object({ clear: z.literal(true) }),
|
||||
]);
|
||||
|
||||
export const RemoteCaseLinkSchema = z.object({
|
||||
name: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid case name format'),
|
||||
hostId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid remote host id'),
|
||||
@@ -1239,6 +1262,13 @@ export const SettingsUpdateSchema = z
|
||||
* stored profiles stay until DELETE /api/sessions/:id/intent.
|
||||
*/
|
||||
readMyMindEnabled: z.boolean().optional(),
|
||||
/**
|
||||
* Custom Model Endpoint Profiles (deployment_plan.md): the toolbar picker that lets a
|
||||
* session point at a user-configured custom OpenAI-compatible endpoint (local or
|
||||
* cloud) instead of its native cloud backend. SYNCED, default OFF — endpoint entry,
|
||||
* discovery, and the extra toolbar surface are all opt-in.
|
||||
*/
|
||||
customModelEndpointsEnabled: z.boolean().optional(),
|
||||
/**
|
||||
* Read My Mind predictor model override. Empty/absent = the AI-checker
|
||||
* default (opus: prediction quality is the product and it runs only on an
|
||||
|
||||
@@ -179,6 +179,7 @@ import {
|
||||
registerVoiceRoutes,
|
||||
registerWebviewRoutes,
|
||||
registerTabLayoutRoutes,
|
||||
registerCustomModelRoutes,
|
||||
tryWebviewRefererFallback,
|
||||
} from './routes/index.js';
|
||||
import { CronService } from '../cron/cron-service.js';
|
||||
@@ -1051,6 +1052,7 @@ export class WebServer extends EventEmitter {
|
||||
registerOrchestratorRoutes(this.app, ctx);
|
||||
registerWebviewRoutes(this.app, ctx, this.basePath);
|
||||
registerTabLayoutRoutes(this.app, ctx);
|
||||
registerCustomModelRoutes(this.app);
|
||||
|
||||
// Cron: build the service from the same context, recompute
|
||||
// due times for any persisted jobs, then expose it to its routes.
|
||||
|
||||
@@ -0,0 +1,211 @@
|
||||
/**
|
||||
* @fileoverview Contract tests for Custom Model Endpoint Profiles
|
||||
* (deployment_plan.md chunk 7): for every CLI with a `customModelInjection`
|
||||
* capability, build the real injection via `buildCustomModelInjection()`,
|
||||
* then replay those exact values through an HTTP request shaped the way that
|
||||
* CLI is documented to send it, against the in-process mock server
|
||||
* (`test/fixtures/mock-openai-server.ts`). Asserts the mock received the
|
||||
* request at the injected base URL, with the injected API key in the
|
||||
* expected header, and the injected model id in the body.
|
||||
*
|
||||
* LIMITATION (stated here and in deployment_plan.md, not left implicit): this
|
||||
* proves "if the CLI honors its documented env/config contract, it will hit
|
||||
* the right endpoint with the right model." It does NOT prove the real CLI
|
||||
* binary actually reads that env var / config file the way its docs say —
|
||||
* that's still the job of `scripts/test-local-llm-harnesses.mjs` against a
|
||||
* real endpoint and real binaries. This suite catches regressions in
|
||||
* Codeman's own injection logic; it cannot catch a CLI changing its env-var
|
||||
* name in a future release.
|
||||
*
|
||||
* Port: N/A (mock server binds a random free port, not a fixed one)
|
||||
*/
|
||||
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
|
||||
import { getCli } from '../src/config/cli-registry/index.js';
|
||||
import { buildCustomModelInjection, type CustomModelEndpoint } from '../src/custom-model-injection.js';
|
||||
import { startMockOpenAiServer, type MockOpenAiServer } from './fixtures/mock-openai-server.js';
|
||||
|
||||
let mock: MockOpenAiServer;
|
||||
|
||||
beforeEach(async () => {
|
||||
mock = await startMockOpenAiServer();
|
||||
});
|
||||
|
||||
afterEach(async () => {
|
||||
await mock.close();
|
||||
});
|
||||
|
||||
function entryOrThrow(id: string) {
|
||||
const entry = getCli(id);
|
||||
if (!entry) throw new Error(`missing CLI registry entry: ${id}`);
|
||||
return entry;
|
||||
}
|
||||
|
||||
function endpointFor(mock: MockOpenAiServer): CustomModelEndpoint {
|
||||
return { id: 'ep1', label: 'mock', baseUrl: mock.baseUrl, apiKey: 'contract-test-key' };
|
||||
}
|
||||
|
||||
/** Replays an OpenAI-shaped chat-completions call using the given base URL/key/model. */
|
||||
async function callOpenAiCompat(baseUrl: string, apiKey: string, model: string) {
|
||||
return fetch(`${baseUrl}/chat/completions`, {
|
||||
method: 'POST',
|
||||
headers: { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` },
|
||||
body: JSON.stringify({ model, messages: [{ role: 'user', content: 'hello world' }] }),
|
||||
});
|
||||
}
|
||||
|
||||
describe('custom-model-injection contract (mock server)', () => {
|
||||
it('claude: ANTHROPIC_BASE_URL/API_KEY reach a real Anthropic-shaped /v1/messages call', async () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('claude'), endpointFor(mock), 'qwen3');
|
||||
if (injection.kind !== 'env') throw new Error('unreachable');
|
||||
|
||||
await fetch(`${injection.envOverrides.ANTHROPIC_BASE_URL}/v1/messages`, {
|
||||
method: 'POST',
|
||||
headers: { 'content-type': 'application/json', 'x-api-key': injection.envOverrides.ANTHROPIC_API_KEY },
|
||||
body: JSON.stringify({
|
||||
model: injection.envOverrides.ANTHROPIC_DEFAULT_SONNET_MODEL,
|
||||
messages: [{ role: 'user', content: 'hello world' }],
|
||||
}),
|
||||
});
|
||||
|
||||
expect(mock.requests).toHaveLength(1);
|
||||
expect(mock.requests[0].path).toBe('/v1/messages');
|
||||
expect(mock.requests[0].headers['x-api-key']).toBe('contract-test-key');
|
||||
expect((mock.requests[0].body as { model: string }).model).toBe('qwen3');
|
||||
});
|
||||
|
||||
it('opencode: OPENCODE_CONFIG_CONTENT decodes to a baseURL/apiKey that reach the mock', async () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('opencode'), endpointFor(mock), 'qwen3');
|
||||
if (injection.kind !== 'env') throw new Error('unreachable');
|
||||
const config = JSON.parse(injection.envOverrides.OPENCODE_CONFIG_CONTENT);
|
||||
const { baseURL, apiKey } = config.provider.custom.options;
|
||||
expect(baseURL).toBe(`${mock.baseUrl}/v1`);
|
||||
|
||||
await callOpenAiCompat(baseURL, apiKey, 'qwen3');
|
||||
|
||||
expect(mock.requests[0].path).toBe('/v1/chat/completions');
|
||||
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
|
||||
expect((mock.requests[0].body as { model: string }).model).toBe('qwen3');
|
||||
});
|
||||
|
||||
it('codex: config.toml decodes to a base_url/model, and env_key/extraEnv reach the mock over /v1/responses', async () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('codex'), endpointFor(mock), 'qwen3');
|
||||
if (injection.kind !== 'configDir') throw new Error('unreachable');
|
||||
const toml = injection.files[0].content;
|
||||
const baseUrl = /base_url = "([^"]+)"/.exec(toml)?.[1];
|
||||
const model = /^model = "([^"]+)"/m.exec(toml)?.[1];
|
||||
const envKeyName = /env_key = "([^"]+)"/.exec(toml)?.[1];
|
||||
expect(baseUrl).toBe(`${mock.baseUrl}/v1`);
|
||||
expect(model).toBe('qwen3');
|
||||
expect(toml).toContain('wire_api = "responses"');
|
||||
expect(toml).not.toContain('api_key ='); // never a literal TOML field
|
||||
expect(envKeyName).toBe('CODEMAN_CUSTOM_MODEL_API_KEY');
|
||||
expect(injection.extraEnv).toEqual({ CODEMAN_CUSTOM_MODEL_API_KEY: 'contract-test-key' });
|
||||
|
||||
// The real credential rides as an env var (env_key names it) — replay it, not a
|
||||
// value read from the file, since the file itself never carries the secret.
|
||||
const apiKey = injection.extraEnv!.CODEMAN_CUSTOM_MODEL_API_KEY;
|
||||
await fetch(`${baseUrl}/responses`, {
|
||||
method: 'POST',
|
||||
headers: { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` },
|
||||
body: JSON.stringify({ model, input: 'hello world' }),
|
||||
});
|
||||
|
||||
expect(mock.requests[0].path).toBe('/v1/responses');
|
||||
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
|
||||
});
|
||||
|
||||
it('pi: models.json decodes to a baseUrl/apiKey that reach the mock', async () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('pi'), endpointFor(mock), 'qwen3');
|
||||
if (injection.kind !== 'configDir') throw new Error('unreachable');
|
||||
const parsed = JSON.parse(injection.files[0].content);
|
||||
const { baseUrl, apiKey } = parsed.providers.custom;
|
||||
expect(baseUrl).toBe(`${mock.baseUrl}/v1`);
|
||||
|
||||
await callOpenAiCompat(baseUrl, apiKey, 'qwen3');
|
||||
|
||||
expect(mock.requests[0].path).toBe('/v1/chat/completions');
|
||||
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
|
||||
});
|
||||
|
||||
it('omp: models.yml decodes to a baseUrl/apiKey that reach the mock', async () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('omp'), endpointFor(mock), 'qwen3');
|
||||
if (injection.kind !== 'configDir') throw new Error('unreachable');
|
||||
const yml = injection.files[0].content;
|
||||
const baseUrl = JSON.parse(/baseUrl: (".*")\n/.exec(yml)![1]);
|
||||
const apiKey = JSON.parse(/apiKey: (".*")\n/.exec(yml)![1]);
|
||||
expect(baseUrl).toBe(`${mock.baseUrl}/v1`);
|
||||
|
||||
await callOpenAiCompat(baseUrl, apiKey, 'qwen3');
|
||||
|
||||
expect(mock.requests[0].path).toBe('/v1/chat/completions');
|
||||
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
|
||||
});
|
||||
|
||||
// gemini/grok/deepseek's `env` kind passes the base URL through UNCHANGED (unlike
|
||||
// opencode/codex/pi/omp, which build a structured config and explicitly append /v1) —
|
||||
// matching Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the SDK
|
||||
// appends the path itself. Whether each of these THREE CLIs' own OpenAI-compatible
|
||||
// client expects the var to already include /v1 (the common OpenAI-SDK convention) or
|
||||
// appends it itself is genuinely CLI-specific and UNVERIFIED (see the confidence table
|
||||
// in deployment_plan.md) — these tests model the common OpenAI-SDK convention (base_url
|
||||
// ends in /v1) since that's the more likely behavior for an OpenAI-compatible client,
|
||||
// but that assumption should be corrected here the moment it's checked against a real
|
||||
// binary.
|
||||
|
||||
it('gemini: GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY reach the mock', async () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('gemini'), endpointFor(mock), 'qwen3');
|
||||
if (injection.kind !== 'env') throw new Error('unreachable');
|
||||
|
||||
await callOpenAiCompat(
|
||||
`${injection.envOverrides.GOOGLE_GEMINI_BASE_URL}/v1`,
|
||||
injection.envOverrides.GEMINI_API_KEY,
|
||||
injection.envOverrides.GEMINI_MODEL
|
||||
);
|
||||
|
||||
expect(mock.requests[0].path).toBe('/v1/chat/completions');
|
||||
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
|
||||
expect((mock.requests[0].body as { model: string }).model).toBe('qwen3');
|
||||
});
|
||||
|
||||
it('grok: GROK_BASE_URL/XAI_API_KEY reach the mock', async () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('grok'), endpointFor(mock), 'qwen3');
|
||||
if (injection.kind !== 'env') throw new Error('unreachable');
|
||||
|
||||
await callOpenAiCompat(
|
||||
`${injection.envOverrides.GROK_BASE_URL}/v1`,
|
||||
injection.envOverrides.XAI_API_KEY,
|
||||
injection.envOverrides.GROK_MODEL
|
||||
);
|
||||
|
||||
expect(mock.requests[0].path).toBe('/v1/chat/completions');
|
||||
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
|
||||
});
|
||||
|
||||
it('deepseek: DEEPSEEK_BASE_URL/DEEPSEEK_API_KEY reach the mock (base URL/key only, no model var)', async () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('deepseek'), endpointFor(mock), 'qwen3');
|
||||
if (injection.kind !== 'env') throw new Error('unreachable');
|
||||
expect(Object.keys(injection.envOverrides).sort()).toEqual(['DEEPSEEK_API_KEY', 'DEEPSEEK_BASE_URL']);
|
||||
|
||||
await callOpenAiCompat(
|
||||
`${injection.envOverrides.DEEPSEEK_BASE_URL}/v1`,
|
||||
injection.envOverrides.DEEPSEEK_API_KEY,
|
||||
'qwen3'
|
||||
);
|
||||
|
||||
expect(mock.requests[0].path).toBe('/v1/chat/completions');
|
||||
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
|
||||
});
|
||||
|
||||
it('antigravity: unsupported, never reaches the mock', () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('antigravity'), endpointFor(mock), 'qwen3');
|
||||
expect(injection).toEqual({ kind: 'unsupported' });
|
||||
expect(mock.requests).toHaveLength(0);
|
||||
});
|
||||
|
||||
it('mock server also answers GET /v1/models for the discovery route', async () => {
|
||||
const res = await fetch(`${mock.baseUrl}/v1/models`);
|
||||
const body = await res.json();
|
||||
expect(body.data.map((m: { id: string }) => m.id)).toEqual(['qwen3', 'llama3']);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,149 @@
|
||||
/**
|
||||
* @fileoverview Tests for the Custom Model Endpoint Profiles pure builder.
|
||||
* Uses the real CLI registry entries (getCli) rather than hand-rolled
|
||||
* fixtures, so a change to a real entry's customModelInjection declaration
|
||||
* is exercised here automatically instead of silently diverging.
|
||||
*
|
||||
* Port: N/A (no server needed)
|
||||
*/
|
||||
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import { getCli } from '../src/config/cli-registry/index.js';
|
||||
import { buildCustomModelInjection, withV1Suffix, type CustomModelEndpoint } from '../src/custom-model-injection.js';
|
||||
|
||||
const endpoint: CustomModelEndpoint = {
|
||||
id: 'ep1',
|
||||
label: 'llama.cpp box',
|
||||
baseUrl: 'http://192.168.1.50:8080',
|
||||
apiKey: 'my-key',
|
||||
};
|
||||
|
||||
function entryOrThrow(id: string) {
|
||||
const entry = getCli(id);
|
||||
if (!entry) throw new Error(`missing CLI registry entry: ${id}`);
|
||||
return entry;
|
||||
}
|
||||
|
||||
describe('withV1Suffix', () => {
|
||||
it('appends /v1 when missing', () => {
|
||||
expect(withV1Suffix('http://host:8080')).toBe('http://host:8080/v1');
|
||||
});
|
||||
|
||||
it('is idempotent when already present', () => {
|
||||
expect(withV1Suffix('http://host:8080/v1')).toBe('http://host:8080/v1');
|
||||
expect(withV1Suffix('http://host:8080/v1/')).toBe('http://host:8080/v1');
|
||||
});
|
||||
|
||||
it('strips a trailing slash with no /v1', () => {
|
||||
expect(withV1Suffix('http://host:8080/')).toBe('http://host:8080/v1');
|
||||
});
|
||||
});
|
||||
|
||||
describe('buildCustomModelInjection', () => {
|
||||
it('claude: env kind sets base URL, api key, and all three tier model vars', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3');
|
||||
expect(result.kind).toBe('env');
|
||||
if (result.kind !== 'env') throw new Error('unreachable');
|
||||
expect(result.envOverrides).toEqual({
|
||||
ANTHROPIC_BASE_URL: 'http://192.168.1.50:8080',
|
||||
ANTHROPIC_API_KEY: 'my-key',
|
||||
ANTHROPIC_DEFAULT_SONNET_MODEL: 'qwen3',
|
||||
ANTHROPIC_DEFAULT_HAIKU_MODEL: 'qwen3',
|
||||
ANTHROPIC_DEFAULT_OPUS_MODEL: 'qwen3',
|
||||
});
|
||||
});
|
||||
|
||||
it('claude: falls back to a dummy key when the endpoint has none', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('claude'), { ...endpoint, apiKey: undefined }, 'qwen3');
|
||||
if (result.kind !== 'env') throw new Error('unreachable');
|
||||
expect(result.envOverrides.ANTHROPIC_API_KEY).toBe('local-dummy-key');
|
||||
});
|
||||
|
||||
it('opencode: configContentEnv carries a JSON blob in OPENCODE_CONFIG_CONTENT', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('opencode'), endpoint, 'qwen3');
|
||||
expect(result.kind).toBe('env');
|
||||
if (result.kind !== 'env') throw new Error('unreachable');
|
||||
const parsed = JSON.parse(result.envOverrides.OPENCODE_CONFIG_CONTENT);
|
||||
expect(parsed.model).toBe('custom/qwen3');
|
||||
expect(parsed.provider.custom.options.baseURL).toBe('http://192.168.1.50:8080/v1');
|
||||
expect(parsed.provider.custom.options.apiKey).toBe('my-key');
|
||||
expect(parsed.provider.custom.models.qwen3).toEqual({});
|
||||
});
|
||||
|
||||
it('codex: configDir writes an isolated config.toml with model/base_url, and the key rides as extraEnv (never a literal TOML field)', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('codex'), endpoint, 'qwen3');
|
||||
expect(result.kind).toBe('configDir');
|
||||
if (result.kind !== 'configDir') throw new Error('unreachable');
|
||||
expect(result.dirEnvVar).toBe('CODEX_HOME');
|
||||
expect(result.files).toHaveLength(1);
|
||||
expect(result.files[0].relPath).toBe('config.toml');
|
||||
expect(result.files[0].content).toContain('model = "qwen3"');
|
||||
expect(result.files[0].content).toContain('base_url = "http://192.168.1.50:8080/v1"');
|
||||
expect(result.files[0].content).toContain('wire_api = "responses"');
|
||||
expect(result.files[0].content).not.toContain('api_key ='); // never a literal TOML field
|
||||
expect(result.files[0].content).toContain('env_key = "CODEMAN_CUSTOM_MODEL_API_KEY"');
|
||||
expect(result.extraEnv).toEqual({ CODEMAN_CUSTOM_MODEL_API_KEY: 'my-key' });
|
||||
});
|
||||
|
||||
it('codex: escapes a quote in the model id so it cannot break out of the TOML string', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('codex'), endpoint, 'weird"model');
|
||||
if (result.kind !== 'configDir') throw new Error('unreachable');
|
||||
expect(result.files[0].content).toContain('model = "weird\\"model"');
|
||||
});
|
||||
|
||||
it('pi: configDir writes models.json under agent/', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('pi'), endpoint, 'qwen3');
|
||||
if (result.kind !== 'configDir') throw new Error('unreachable');
|
||||
expect(result.dirEnvVar).toBe('PI_CONFIG_DIR');
|
||||
expect(result.files[0].relPath).toBe('agent/models.json');
|
||||
const parsed = JSON.parse(result.files[0].content);
|
||||
expect(parsed.providers.custom.baseUrl).toBe('http://192.168.1.50:8080/v1');
|
||||
});
|
||||
|
||||
it('omp: configDir writes models.yml under agent/, redirected via PI_CONFIG_DIR', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('omp'), endpoint, 'qwen3');
|
||||
if (result.kind !== 'configDir') throw new Error('unreachable');
|
||||
expect(result.dirEnvVar).toBe('PI_CONFIG_DIR');
|
||||
expect(result.files[0].relPath).toBe('agent/models.yml');
|
||||
expect(result.files[0].content).toContain('baseUrl: "http://192.168.1.50:8080/v1"');
|
||||
});
|
||||
|
||||
it('gemini: env kind sets GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY/GEMINI_MODEL', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('gemini'), endpoint, 'qwen3');
|
||||
if (result.kind !== 'env') throw new Error('unreachable');
|
||||
expect(result.envOverrides).toEqual({
|
||||
GOOGLE_GEMINI_BASE_URL: 'http://192.168.1.50:8080',
|
||||
GEMINI_API_KEY: 'my-key',
|
||||
GEMINI_MODEL: 'qwen3',
|
||||
});
|
||||
});
|
||||
|
||||
it('grok: env kind sets GROK_BASE_URL/XAI_API_KEY/GROK_MODEL', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('grok'), endpoint, 'qwen3');
|
||||
if (result.kind !== 'env') throw new Error('unreachable');
|
||||
expect(result.envOverrides).toEqual({
|
||||
GROK_BASE_URL: 'http://192.168.1.50:8080',
|
||||
XAI_API_KEY: 'my-key',
|
||||
GROK_MODEL: 'qwen3',
|
||||
});
|
||||
});
|
||||
|
||||
it('deepseek: env kind sets base URL/key only, no model var', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3');
|
||||
if (result.kind !== 'env') throw new Error('unreachable');
|
||||
expect(result.envOverrides).toEqual({
|
||||
DEEPSEEK_BASE_URL: 'http://192.168.1.50:8080',
|
||||
DEEPSEEK_API_KEY: 'my-key',
|
||||
});
|
||||
});
|
||||
|
||||
it('antigravity: unsupported', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('antigravity'), endpoint, 'qwen3');
|
||||
expect(result).toEqual({ kind: 'unsupported' });
|
||||
});
|
||||
|
||||
it('shell: unsupported', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('shell'), endpoint, 'qwen3');
|
||||
expect(result).toEqual({ kind: 'unsupported' });
|
||||
});
|
||||
});
|
||||
Vendored
+118
@@ -0,0 +1,118 @@
|
||||
/**
|
||||
* @fileoverview In-process fake OpenAI/Anthropic-compatible HTTP server for the
|
||||
* Custom Model Endpoint Profiles contract tests (deployment_plan.md chunk 7).
|
||||
*
|
||||
* No external deps — plain `node:http`. Captures every request it receives
|
||||
* (method, path, headers, parsed JSON body) so a test can assert the injected
|
||||
* base URL / API key / model actually reached the right place, with the right
|
||||
* auth header, in the shape a real llama.cpp/Azure/etc. endpoint would see it.
|
||||
*
|
||||
* Serves the request shapes this feature's recipes produce: OpenAI-style
|
||||
* `POST /v1/chat/completions` (opencode/pi/grok/omp/gemini's compat
|
||||
* endpoint), Anthropic-style `POST /v1/messages` (claude's ANTHROPIC_BASE_URL
|
||||
* traffic), OpenAI's newer `POST /v1/responses` (codex's actual wire protocol
|
||||
* as of Feb 2026 — it dropped chat-completions support), plus `GET /v1/models`
|
||||
* for the discovery route's own tests.
|
||||
*/
|
||||
|
||||
import { createServer, type IncomingMessage, type Server } from 'node:http';
|
||||
import { AddressInfo } from 'node:net';
|
||||
|
||||
export interface CapturedRequest {
|
||||
method: string;
|
||||
path: string;
|
||||
headers: Record<string, string | string[] | undefined>;
|
||||
body: unknown;
|
||||
}
|
||||
|
||||
export interface MockOpenAiServer {
|
||||
baseUrl: string;
|
||||
requests: CapturedRequest[];
|
||||
close(): Promise<void>;
|
||||
}
|
||||
|
||||
function readJsonBody(req: IncomingMessage): Promise<unknown> {
|
||||
return new Promise((resolve) => {
|
||||
const chunks: Buffer[] = [];
|
||||
req.on('data', (c) => chunks.push(c));
|
||||
req.on('end', () => {
|
||||
const raw = Buffer.concat(chunks).toString('utf8');
|
||||
if (!raw) return resolve(undefined);
|
||||
try {
|
||||
resolve(JSON.parse(raw));
|
||||
} catch {
|
||||
resolve(raw);
|
||||
}
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
/** Starts the mock server on a random free port and resolves once it's listening. */
|
||||
export async function startMockOpenAiServer(): Promise<MockOpenAiServer> {
|
||||
const requests: CapturedRequest[] = [];
|
||||
|
||||
const server: Server = createServer((req, res) => {
|
||||
void (async () => {
|
||||
const body = await readJsonBody(req);
|
||||
const path = (req.url ?? '').split('?')[0];
|
||||
requests.push({ method: req.method ?? 'GET', path, headers: req.headers, body });
|
||||
|
||||
res.setHeader('content-type', 'application/json');
|
||||
|
||||
if (path === '/v1/models' && req.method === 'GET') {
|
||||
res.writeHead(200);
|
||||
res.end(JSON.stringify({ data: [{ id: 'qwen3' }, { id: 'llama3' }] }));
|
||||
return;
|
||||
}
|
||||
|
||||
if (path === '/v1/chat/completions' && req.method === 'POST') {
|
||||
res.writeHead(200);
|
||||
res.end(
|
||||
JSON.stringify({
|
||||
id: 'mock-completion',
|
||||
choices: [{ index: 0, message: { role: 'assistant', content: 'hello world' } }],
|
||||
})
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
if (path === '/v1/messages' && req.method === 'POST') {
|
||||
res.writeHead(200);
|
||||
res.end(
|
||||
JSON.stringify({
|
||||
id: 'mock-message',
|
||||
role: 'assistant',
|
||||
content: [{ type: 'text', text: 'hello world' }],
|
||||
})
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
// Codex's real wire protocol (verified against a live binary: it dropped
|
||||
// wire_api="chat" support in Feb 2026, so its config.toml always says
|
||||
// wire_api="responses") — a different shape from OpenAI's chat-completions.
|
||||
if (path === '/v1/responses' && req.method === 'POST') {
|
||||
res.writeHead(200);
|
||||
res.end(
|
||||
JSON.stringify({
|
||||
id: 'mock-response',
|
||||
output: [{ type: 'message', role: 'assistant', content: [{ type: 'output_text', text: 'hello world' }] }],
|
||||
})
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
res.writeHead(404);
|
||||
res.end(JSON.stringify({ error: 'not found in mock server', path }));
|
||||
})();
|
||||
});
|
||||
|
||||
await new Promise<void>((resolve) => server.listen(0, '127.0.0.1', resolve));
|
||||
const { port } = server.address() as AddressInfo;
|
||||
|
||||
return {
|
||||
baseUrl: `http://127.0.0.1:${port}`,
|
||||
requests,
|
||||
close: () => new Promise<void>((resolve, reject) => server.close((err) => (err ? reject(err) : resolve()))),
|
||||
};
|
||||
}
|
||||
@@ -318,6 +318,22 @@ export class MockSession extends EventEmitter {
|
||||
this.color = c;
|
||||
});
|
||||
|
||||
/** Custom Model Endpoint Profiles (deployment_plan.md) */
|
||||
customModel: { endpointId: string; modelId: string; label?: string } | undefined = undefined;
|
||||
private _mockCustomModelConfigDir: string | undefined;
|
||||
setCustomModel = vi.fn(
|
||||
(
|
||||
next: { endpointId: string; modelId: string; label?: string; envKeys: string[]; configDir?: string } | undefined,
|
||||
_envOverrides?: Record<string, string>
|
||||
): string | undefined => {
|
||||
const previous = this._mockCustomModelConfigDir;
|
||||
this._mockCustomModelConfigDir = next?.configDir;
|
||||
this.customModel = next ? { endpointId: next.endpointId, modelId: next.modelId, label: next.label } : undefined;
|
||||
return previous;
|
||||
}
|
||||
);
|
||||
restartCli = vi.fn(async () => true);
|
||||
|
||||
/** Stub for sendInput */
|
||||
sendInput = vi.fn();
|
||||
|
||||
|
||||
@@ -0,0 +1,151 @@
|
||||
/**
|
||||
* @fileoverview Route tests for Custom Model Endpoint Profiles CRUD + discovery.
|
||||
* Port: N/A (app.inject, no real port needed)
|
||||
*/
|
||||
import { describe, it, expect, vi, afterEach } from 'vitest';
|
||||
import { registerCustomModelRoutes } from '../../src/web/routes/custom-model-routes.js';
|
||||
import { createRouteTestHarness } from './_route-test-utils.js';
|
||||
|
||||
async function setup() {
|
||||
return createRouteTestHarness(registerCustomModelRoutes);
|
||||
}
|
||||
|
||||
describe('custom model endpoint CRUD', () => {
|
||||
afterEach(() => {
|
||||
vi.unstubAllGlobals();
|
||||
});
|
||||
|
||||
it('starts empty', async () => {
|
||||
const { app } = await setup();
|
||||
const res = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
|
||||
expect(res.json()).toEqual([]);
|
||||
});
|
||||
|
||||
it('creates, lists, updates, and deletes an endpoint', async () => {
|
||||
const { app } = await setup();
|
||||
|
||||
const create = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/model-endpoints',
|
||||
payload: { id: 'ep1', label: 'llama.cpp box', baseUrl: 'http://192.168.1.50:8080' },
|
||||
});
|
||||
expect(create.statusCode).toBe(200);
|
||||
expect(create.json().data.host.id).toBe('ep1');
|
||||
|
||||
const list = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
|
||||
expect(list.json()).toHaveLength(1);
|
||||
|
||||
const update = await app.inject({
|
||||
method: 'PUT',
|
||||
url: '/api/model-endpoints/ep1',
|
||||
payload: { label: 'Renamed', baseUrl: 'http://192.168.1.50:8080' },
|
||||
});
|
||||
expect(update.statusCode).toBe(200);
|
||||
expect(update.json().data.host.label).toBe('Renamed');
|
||||
|
||||
const del = await app.inject({ method: 'DELETE', url: '/api/model-endpoints/ep1' });
|
||||
expect(del.statusCode).toBe(200);
|
||||
|
||||
const listAfter = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
|
||||
expect(listAfter.json()).toEqual([]);
|
||||
});
|
||||
|
||||
it('rejects a duplicate id on create', async () => {
|
||||
const { app } = await setup();
|
||||
const payload = { id: 'dup', label: 'A', baseUrl: 'http://localhost:8080' };
|
||||
await app.inject({ method: 'POST', url: '/api/model-endpoints', payload });
|
||||
const second = await app.inject({ method: 'POST', url: '/api/model-endpoints', payload });
|
||||
expect(second.json().success).toBe(false);
|
||||
expect(second.json().errorCode).toBe('ALREADY_EXISTS');
|
||||
});
|
||||
|
||||
it('404s updating/deleting an id that does not exist', async () => {
|
||||
const { app } = await setup();
|
||||
const update = await app.inject({
|
||||
method: 'PUT',
|
||||
url: '/api/model-endpoints/ghost',
|
||||
payload: { label: 'A', baseUrl: 'http://localhost:8080' },
|
||||
});
|
||||
expect(update.json().errorCode).toBe('NOT_FOUND');
|
||||
});
|
||||
|
||||
it('rejects a link-local/cloud-metadata base URL', async () => {
|
||||
const { app } = await setup();
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/model-endpoints',
|
||||
payload: { id: 'meta', label: 'A', baseUrl: 'http://169.254.169.254/' },
|
||||
});
|
||||
expect(res.json().success).toBe(false);
|
||||
expect(res.json().errorCode).toBe('INVALID_INPUT');
|
||||
});
|
||||
|
||||
it('discovers models via GET /v1/models and stores the result', async () => {
|
||||
const { app } = await setup();
|
||||
await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/model-endpoints',
|
||||
payload: { id: 'ep1', label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'k' },
|
||||
});
|
||||
|
||||
const fetchMock = vi.fn(async (url: string, init?: RequestInit) => {
|
||||
expect(url).toBe('http://localhost:8080/v1/models');
|
||||
const headers = init?.headers as Record<string, string>;
|
||||
// Exactly ONE auth header — never both (a real server hung when sent both).
|
||||
expect(headers.Authorization).toBe('Bearer k');
|
||||
expect(headers['api-key']).toBeUndefined();
|
||||
return new Response(JSON.stringify({ data: [{ id: 'qwen3' }, { id: 'llama3' }] }), { status: 200 });
|
||||
});
|
||||
vi.stubGlobal('fetch', fetchMock);
|
||||
|
||||
const res = await app.inject({ method: 'POST', url: '/api/model-endpoints/ep1/discover-models' });
|
||||
expect(res.json().data.models).toEqual(['qwen3', 'llama3']);
|
||||
|
||||
// Data dir is shared across this WHOLE test file (one temp HOME per file, not per
|
||||
// test — test/setup.ts), so find by id rather than assuming index 0.
|
||||
const list = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
|
||||
const stored = (list.json() as Array<{ id: string }>).find((h) => h.id === 'ep1');
|
||||
expect(stored?.models).toEqual(['qwen3', 'llama3']);
|
||||
expect(stored?.lastDiscoveredAt).toBeTruthy();
|
||||
});
|
||||
|
||||
it('discovers models with authStyle "api-key" using only that header, never Authorization', async () => {
|
||||
const { app } = await setup();
|
||||
await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/model-endpoints',
|
||||
payload: { id: 'ep-azure', label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'k', authStyle: 'api-key' },
|
||||
});
|
||||
|
||||
const fetchMock = vi.fn(async (_url: string, init?: RequestInit) => {
|
||||
const headers = init?.headers as Record<string, string>;
|
||||
expect(headers['api-key']).toBe('k');
|
||||
expect(headers.Authorization).toBeUndefined();
|
||||
return new Response(JSON.stringify({ data: [] }), { status: 200 });
|
||||
});
|
||||
vi.stubGlobal('fetch', fetchMock);
|
||||
|
||||
await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-azure/discover-models' });
|
||||
expect(fetchMock).toHaveBeenCalledTimes(1);
|
||||
});
|
||||
|
||||
it('reports a clear error when the endpoint is unreachable', async () => {
|
||||
const { app } = await setup();
|
||||
await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/model-endpoints',
|
||||
payload: { id: 'ep-err', label: 'A', baseUrl: 'http://localhost:8080' },
|
||||
});
|
||||
vi.stubGlobal(
|
||||
'fetch',
|
||||
vi.fn(async () => {
|
||||
throw new Error('connect ECONNREFUSED');
|
||||
})
|
||||
);
|
||||
|
||||
const res = await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-err/discover-models' });
|
||||
expect(res.json().success).toBe(false);
|
||||
expect(res.json().errorCode).toBe('OPERATION_FAILED');
|
||||
expect(res.json().error).toContain('ECONNREFUSED');
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,121 @@
|
||||
/**
|
||||
* @fileoverview Tests for POST /api/sessions/:id/custom-model (deployment_plan.md
|
||||
* chunk 5 — applying/clearing a session's custom model endpoint + CLI restart).
|
||||
* Port: N/A (app.inject, no real port needed)
|
||||
*/
|
||||
import { describe, it, expect, beforeEach } from 'vitest';
|
||||
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
|
||||
import { createRouteTestHarness } from './_route-test-utils.js';
|
||||
import { getDataDir } from '../../src/config/instance.js';
|
||||
import { writeCustomModelHosts, type CustomModelHost } from '../../src/custom-model-hosts.js';
|
||||
|
||||
const CLAUDE_ENDPOINT: CustomModelHost = {
|
||||
id: 'ep1',
|
||||
label: 'llama.cpp box',
|
||||
baseUrl: 'http://192.168.1.50:8080',
|
||||
apiKey: 'k',
|
||||
};
|
||||
|
||||
async function setup() {
|
||||
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT]);
|
||||
return createRouteTestHarness(registerSessionRoutes);
|
||||
}
|
||||
|
||||
describe('POST /api/sessions/:id/custom-model', () => {
|
||||
beforeEach(async () => {
|
||||
await writeCustomModelHosts(getDataDir(), []);
|
||||
});
|
||||
|
||||
it('applies an endpoint/model to a claude-mode session and restarts the CLI', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
const session = ctx.sessions.get('test-session-1')!;
|
||||
session.mode = 'claude';
|
||||
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/sessions/test-session-1/custom-model',
|
||||
payload: { endpointId: 'ep1', modelId: 'qwen3' },
|
||||
});
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
const body = res.json();
|
||||
expect(body.customModel).toEqual({ endpointId: 'ep1', modelId: 'qwen3', label: 'llama.cpp box' });
|
||||
expect(body.restarted).toBe(true);
|
||||
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
|
||||
expect(session.restartCli).toHaveBeenCalledTimes(1);
|
||||
|
||||
// Verify the actual injected env vars via setCustomModel's captured call args.
|
||||
const [next, envOverrides] = session.setCustomModel.mock.calls[0];
|
||||
expect(next.envKeys).toEqual([
|
||||
'ANTHROPIC_BASE_URL',
|
||||
'ANTHROPIC_API_KEY',
|
||||
'ANTHROPIC_DEFAULT_SONNET_MODEL',
|
||||
'ANTHROPIC_DEFAULT_HAIKU_MODEL',
|
||||
'ANTHROPIC_DEFAULT_OPUS_MODEL',
|
||||
]);
|
||||
expect(envOverrides.ANTHROPIC_BASE_URL).toBe('http://192.168.1.50:8080');
|
||||
expect(envOverrides.ANTHROPIC_API_KEY).toBe('k');
|
||||
});
|
||||
|
||||
it('clears back to the native default', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
const session = ctx.sessions.get('test-session-1')!;
|
||||
session.mode = 'claude';
|
||||
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/sessions/test-session-1/custom-model',
|
||||
payload: { clear: true },
|
||||
});
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json().customModel).toBeUndefined();
|
||||
expect(session.setCustomModel).toHaveBeenCalledWith(undefined);
|
||||
expect(session.restartCli).toHaveBeenCalledTimes(1);
|
||||
});
|
||||
|
||||
it('404s for an unknown endpoint id', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
ctx.sessions.get('test-session-1')!.mode = 'claude';
|
||||
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/sessions/test-session-1/custom-model',
|
||||
payload: { endpointId: 'ghost', modelId: 'qwen3' },
|
||||
});
|
||||
|
||||
expect(res.json().success).toBe(false);
|
||||
expect(res.json().errorCode).toBe('NOT_FOUND');
|
||||
});
|
||||
|
||||
it('refuses a mode with no known custom-model mechanism (antigravity)', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
ctx.sessions.get('test-session-1')!.mode = 'antigravity';
|
||||
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/sessions/test-session-1/custom-model',
|
||||
payload: { endpointId: 'ep1', modelId: 'qwen3' },
|
||||
});
|
||||
|
||||
expect(res.json().success).toBe(false);
|
||||
expect(res.json().errorCode).toBe('OPERATION_FAILED');
|
||||
});
|
||||
|
||||
it('refuses to touch a busy session', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
const session = ctx.sessions.get('test-session-1')!;
|
||||
session.mode = 'claude';
|
||||
session.isBusy = () => true;
|
||||
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/sessions/test-session-1/custom-model',
|
||||
payload: { endpointId: 'ep1', modelId: 'qwen3' },
|
||||
});
|
||||
|
||||
expect(res.json().success).toBe(false);
|
||||
expect(res.json().errorCode).toBe('SESSION_BUSY');
|
||||
expect(session.setCustomModel).not.toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
Reference in New Issue
Block a user