test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end

Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.

Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.

Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:

- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
  speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
  don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
  undocumented GATEWAY AuthType gemini-cli selects once
  GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
  tried
- deepseek: reaches the server (env vars are read) but gets a consistent
  HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)

Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).

deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
This commit is contained in:
Devvyn
2026-09-13 17:42:35 +08:00
co-authored by Claude Sonnet 5
parent 41416566aa
commit 61779745aa
14 changed files with 1061 additions and 770 deletions
+1 -1
View File
@@ -225,7 +225,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. The session version worked and was still wrong in use: it put a terminal tab on screen next to the web tab the user actually asked for, every single time, and nothing about a long-lived HTTP server needs to be a tab. ⚠️ What a session gave for free now has to be paid for explicitly, and every piece is load-bearing: **exactly one** server (a second click REUSES it rather than racing it for a port, which two sessions structurally could not do), **restarted when the browser authority changes** (`--trusted-host` fences dsh's `/api` against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two, so whoever asks last wins: the asker is by definition the origin about to load the page), **killed on shutdown** (`stopDeepSeekWeb()` in the server teardown, because the child is detached so its whole plugin tree can be signalled at once, which also means it would OUTLIVE Codeman and hold its port against the next start), and **failures returned to the caller**, since with no tab there is nowhere for a stack trace to land. ⚠️ The port search starts at dsh's own default 3080 and walks 40, never fixed: that default is precisely the port most likely to be taken already by the user's own `dsh web`, and hardcoding it killed this feature with EADDRINUSE once. Free-port detection BINDS rather than connects (a connect probe cannot tell "free" from "listening but not answering yet"), so it is racy by nature and the caller still waits for the server to really answer before reporting success. ⚠️ Both `POST` and `DELETE` sit at the **same privilege bar as the profile installer** (`canUsernameRunPrivilegedCommands`) even though the action reads as "open a page": booting a dsh profile executes the plugin code in it, and the server is a single shared instance, so stopping it in multi-user mode takes it out from under other users' tabs.
**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `deployment_plan.md`): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET <baseUrl>/v1/models`; `CustomModelHost.authStyle` (default `'both'`) sends BOTH `Authorization: Bearer` and `api-key` headers on discovery since cloud gateways (Azure) and local servers (llama.cpp) disagree on the convention and there is no way to know in advance which one a given endpoint wants. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/grok/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi` config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `PI_CONFIG_DIR`, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. Confidence is per-CLI: claude/opencode/codex are hand-verified against a real llama.cpp/llama-swap server; gemini/pi/grok/deepseek/omp have their ONE-SHOT INVOCATION flags confirmed against real installed binaries' `--help` output, but their custom-endpoint env/config conventions remain unverified — see the confidence table in `deployment_plan.md`. The standalone `scripts/test-local-llm-harnesses.mjs` (reads a gitignored `scripts/local-llm-test.config.json`, template `.example.json` tracked) smoke-tests real CLI binaries against a real endpoint outside the web UI entirely, independent of tmux/sessions.
**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `deployment_plan.md`): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET <baseUrl>/v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.<name>]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct but codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement — a confirmed protocol gap, not a bug; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek reaches the server but gets a consistent `HTTP_404` (root cause not identified); antigravity has no known mechanism at all. See the confidence table in `deployment_plan.md` for the full detail on each.
**Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w<n>-<case>` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. ⚠️ **Closing has the mirror-image race and one owner**: `closeSession()` reads `wasActive` BEFORE its `await` and announces the delete via `_closingSessions`, while `_onSessionDeleted` skips the active-session handoff for an id in that set. Both used to read `activeSessionId` after the fact, so the `session_deleted` broadcast for your own delete could null it first and closing the tab you were on landed on the welcome screen instead of the next session, on the same build, depending on timing. The fallback also picks the first order entry that is still in `sessions` (a dead id can linger in `sessionOrder`, same reason Alt+N indexes a live-filtered list). A delete from ANOTHER client still shows the welcome screen, which is the honest answer when what you were looking at was taken away. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization)
+93 -18
View File
@@ -50,7 +50,7 @@ just "add the gateway once, everything behind it shows up."
## Why
The maintainer pays for a Claude Code subscription but also runs a capable
local model. Every harness Codeman drives already *has* its own mechanism
local model. Every harness Codeman drives already _has_ its own mechanism
for pointing at a custom endpoint (env vars for Claude, a JSON config blob
for opencode, a TOML file for Codex, etc.) — Codeman just never exposed a
UI for it. Full motivation, the per-CLI recipe table, and the on-prem
@@ -77,7 +77,7 @@ hardware use cases are written up in **[`deployment_plan.md`](deployment_plan.md
web tabs use.
- **`src/web/schemas.ts`** — `customModelEndpointsEnabled` (synced, default
OFF) + the endpoint payload schema.
- **`scripts/test-local-llm-harnesses.mjs`** — standalone smoke-test script
- **`scripts/test-local-llm-harnesses.ts`** — standalone smoke-test script
that spawns each real CLI binary one-shot against a real endpoint and
checks it can answer "hello world," independent of the web UI. Reads
defaults from a gitignored `scripts/local-llm-test.config.json` (see the
@@ -130,11 +130,13 @@ format + tests green), ⬜ = not started.
until chunk 6 lands) + a CLAUDE.md pointer bullet
Also done outside the chunk list: the standalone
`scripts/test-local-llm-harnesses.mjs` smoke-test script + its gitignored
config file, the on-prem-hardware use-case writeup in `deployment_plan.md`
(DGX Spark, Strix Halo, Qwen5090), and a `codeman/agent:llm-test` Docker
image (all 9 CLI binaries, built from `docker/agent.Dockerfile`) for the
real end-to-end test against a live llama-swap server.
`scripts/test-local-llm-harnesses.ts` smoke-test script (now dynamic —
reads the live CLI registry rather than a hand-maintained harness list) +
its gitignored config file, the on-prem-hardware use-case writeup in
`deployment_plan.md` (DGX Spark, Strix Halo, Qwen5090), a
`codeman/agent:llm-test` Docker image (all 9 CLI binaries, built from
`docker/agent.Dockerfile`), and a **completed real end-to-end run of all 9
harnesses** against a live llama-swap server — see Testing below.
## Testing performed so far
@@ -145,14 +147,67 @@ test/custom-model-injection-contract.test.ts test/routes/custom-model-routes.tes
test/routes/session-custom-model.test.ts test/routes/external-cli-bypass-clamp.test.ts` —
245+ tests passing, including the existing multi-user clamp suite (no
regressions from the `privilegedEnvKeys` additions)
- `node --check scripts/test-local-llm-harnesses.mjs` + manual `--help` run
- Refactored `scripts/test-local-llm-harnesses.ts` (now `npx tsx`-run, was
plain `.mjs`) to import `enabledClis()` and `buildCustomModelInjection()`
directly from source instead of keeping a second hand-maintained copy of
every CLI's env/config shape — a registry change now needs zero edits to
the test script. Extracted the config-dir-write logic shared with the
production route into `custom-model-injection-apply.ts` so both places
call exactly one implementation.
- **Real end-to-end run against the maintainer's live llama-swap server**
(`http://10.10.11.241:8080`), inside `codeman/agent:llm-test` (all 9 CLI
binaries, built via `docker/agent.Dockerfile`), against the smallest
available model (`qwen3.5-0.8b-ud-q8_k_xl`, 1.1GB — picked by parsing the
server's own reported model sizes). Real findings, not simulated:
server's own reported model sizes). **Full 9-harness result: claude,
opencode, pi, grok, omp all PASS with a genuine "hello world" reply
round-tripped through the real endpoint; codex FAILs for a confirmed
protocol reason (not a bug — see below); gemini and deepseek reach the
server but fail for reasons not yet root-caused; antigravity SKIPs (no
known mechanism); all correctly classified by the now-dynamic
`scripts/test-local-llm-harnesses.ts`, which reads the live CLI registry
rather than a hand-maintained harness list.** Real findings, not
simulated:
- **opencode: PASS.** Genuinely round-tripped a "hello world" reply
through the real endpoint.
- **pi: PASS, after two real bugs found and fixed.** `PI_CONFIG_DIR` does
nothing for pi at all (grepped pi's entire bundled JS source — the
string appears nowhere); the real redirect is the child process's own
`HOME`, since pi hardcodes `~/.pi/agent/models.json` with no dedicated
override. Separately, pi's `models` field must be an **array** of
`{id}` objects, not an object keyed by id (confirmed against pi's own
bundled `docs/models.md`) — the object shape silently loaded zero
models. Also needs an explicit `--model custom/<id>` on invocation.
- **grok: PASS, after the original recipe turned out to be flat-out
wrong**, not just unverified — the env-var recipe in this table's first
draft (`GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) produced "Not signed
in" against a real binary. Researched xAI's actual docs and corrected
to a `config.toml` with a `[model.<name>]` block redirected via
`GROK_HOME`, with the key riding as an `env_key`-named env var — then
confirmed working end-to-end.
- **omp: PASS**, after the same two fixes as pi (array-shaped `models`,
`HOME`-redirect instead of `PI_CONFIG_DIR`) plus `--model custom/<id>`.
Unverified against omp's own official docs (none are bundled in the
install), but empirically confirmed working live.
- **gemini: confirmed broken, unresolved after real investigation.**
Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an
undocumented `AuthType.GATEWAY` path with validation requirements a
live run never satisfies (`Invalid auth method selected`, regardless of
key format). Tried and ruled out: a Google-format dummy key,
`GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE`
override, and a hand-written `settings.json`. `--skip-trust` is a real,
separate fix for a different symptom (an untrusted-folder check
silently overriding `--approval-mode yolo`) and is kept, but does not
touch this auth failure. Left as an open, documented gap rather than
claimed as working.
- **deepseek: confirmed reaching the server, still failing, unresolved.**
A real run returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)`
consistently — the env vars are read (the request reaches the network
rather than failing locally), but the root cause was not identified in
the time available. By analogy with codex's Responses-API gap, `dsh`
may expect DeepSeek's own API response shape rather than a generic
OpenAI-compatible one, but this was not confirmed by reading dsh's own
bundled source the way the pi/grok questions were resolved. Documented
as best-effort/unknown, matching its pre-existing lowest confidence tag.
- **codex: real bug found and fixed.** The recipe's TOML shape
(`[model].default`) was rejected by a real codex binary ("invalid
type: map, expected a string") — codex wants a top-level `model`
@@ -193,7 +248,7 @@ test/routes/session-custom-model.test.ts test/routes/external-cli-bypass-clamp.t
is a tooling-correctness fix (affects the script's own baseline check),
not a claim about how any CLI's own HTTP client behaves.
- **Also found and fixed**: an earlier design sent BOTH `Authorization:
Bearer` and `api-key` auth header conventions on every discovery/
Bearer` and `api-key` auth header conventions on every discovery/
baseline request, on the theory that an unused header is harmless.
Live-tested against the real server, sending both reliably HUNG the
request (reproduced 3×: either header alone ~500-600ms, both together
@@ -206,13 +261,33 @@ test/routes/session-custom-model.test.ts test/routes/external-cli-bypass-clamp.t
## Not yet done / open questions for review
- Six of nine per-CLI recipes (Gemini, Pi, Grok, DeepSeek, OMP) are
**web-researched, not verified** against real binaries — see the
confidence table in `deployment_plan.md`. Antigravity has no known
mechanism at all and stays unsupported.
- Chunk 5's session-restart design needs a careful look before
implementation: switching a session's endpoint restarts its CLI process
in place (confirmed acceptable with the maintainer — these harnesses
read endpoint config at process start, not per-turn).
- **Chunk 6 (frontend)** — settings group, toolbar picker, tab badge — is
still entirely unbuilt; the feature is currently HTTP-API-only (see
`docs/custom-model-endpoints.md`).
- **Gemini is confirmed broken end-to-end** (`Invalid auth method
selected`, traced to an undocumented `GATEWAY` AuthType gemini-cli
selects once `GOOGLE_GEMINI_BASE_URL` is set) — needs upstream
investigation before it can be called supported. Documented in full in
`deployment_plan.md`'s confidence table rather than silently shipped as
working.
- **DeepSeek is confirmed reaching the server but failing** with a
consistent `HTTP_404`, root cause not identified — documented as
best-effort/unknown, same as its pre-existing lowest confidence tag.
- **Codex cannot work against a plain OpenAI-Chat-Completions server**
(llama.cpp/llama-swap/Ollama/vLLM's default) — it only speaks the
Responses API since Feb 2026. This is an external protocol
incompatibility, not something this PR can fix; codex support is real
only against a Responses-API-compatible endpoint.
- Antigravity has no known mechanism at all and stays unsupported.
- Chunk 5's session-restart design needs a careful look before merge:
switching a session's endpoint restarts its CLI process in place
(confirmed acceptable with the maintainer — these harnesses read
endpoint config at process start, not per-turn). Whether an INTERACTIVE
claude session with a custom model hits the same async-title-generation
hang the standalone script worked around with `--bare` (vs. just a
harmless background warning) is untested and should be checked before
calling claude's chunk 5 support done — `--bare` itself must never be
applied to a real interactive session, since it disables hooks Codeman
depends on.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
+49 -34
View File
@@ -69,11 +69,11 @@ config blob for opencode, a TOML file for Codex, etc. Devvyn gave the
starting recipes for those three; the rest (Gemini, Pi, Grok, DeepSeek, OMP,
Antigravity) were researched for this plan and are flagged by confidence
below. A real end-to-end pass against Devvyn's own llama-swap server
(`scripts/test-local-llm-harnesses.mjs`, inside a `codeman/agent:llm-test`
(`scripts/test-local-llm-harnesses.ts`, inside a `codeman/agent:llm-test`
Docker image with all 9 CLIs installed) then confirmed **claude and
opencode work end-to-end**, corrected a real Codex config.toml schema bug
the given recipe had (see the Codex row below), and surfaced that Codex's
*protocol* — not just its config shape — does not work against a plain
_protocol_ — not just its config shape — does not work against a plain
OpenAI-Chat-Completions server like llama.cpp/llama-swap at all. Confidence
below reflects what was actually observed, not just what was planned.
@@ -104,17 +104,17 @@ declared capability, never an `if (mode === 'claude')` branch.
## Per-CLI injection recipes (confidence-ranked)
| CLI | Mechanism | Confidence |
|---|---|---|
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` (or `GOOGLE_VERTEX_BASE_URL`) + `GEMINI_API_KEY`; CLI needs a restart to pick them up (matches our restart-on-switch design). Model selection via `--model`/`GEMINI_MODEL`-style override — verify exact var name against the installed `gemini-cli` version before shipping | Web-researched, unverified |
| `pi` | Config file `~/.pi/agent/models.json` (hot-reloadable) with a custom provider block: `baseUrl`, `apiKey`, `api:"openai-completions"`. Redirect via `PI_CONFIG_DIR` (already allowlisted per CLAUDE.md) pointed at an isolated dir containing just this file, rather than overwriting the user's real one | Web-researched, unverified |
| `grok` | Env vars `GROK_BASE_URL`, `XAI_API_KEY` (dummy ok for local; a real key for most cloud endpoints), `GROK_MODEL`. All three already fit inside the existing `XAI_*`/CLI-specific allowlist shape | Web-researched, unverified |
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts:879-913`, already in `privilegedEnvKeys`). Model selection is murkier — CLAUDE.md notes dsh model is "a profile composition entry," not a flag/env var, so redirecting the endpoint is solid but forcing a specific model name may not fully work; document as best-effort and verify against a real profile | Web-researched, unverified, partial |
| `omp` | Config file `~/.omp/agent/models.yml`-equivalent with a custom provider `baseUrl`. CLAUDE.md notes omp's config tree is itself relocatable via `PI_CONFIG_DIR` — reuse the same isolated-dir-redirect approach as `pi` | Web-researched, unverified |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
| CLI | Mechanism | Confidence |
| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`, already in `privilegedEnvKeys`). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
Everything web-researched-but-unverified gets implemented but must be
smoke-tested against real installs of those CLIs before being called done —
@@ -139,8 +139,13 @@ union on each `CliEntry.capabilities`:
type CustomModelInjection =
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[] }
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json' }
| { kind: 'configDir'; dirEnvVar: string; fileName: string; template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' }
| { kind: 'unsupported' }
| {
kind: 'configDir';
dirEnvVar: string;
fileName: string;
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml';
}
| { kind: 'unsupported' };
```
Declared per stock.ts entry per the table above. A pure function in a new
@@ -241,7 +246,7 @@ dir-redirects, plus the already-privileged `DEEPSEEK_BASE_URL` — must be
added to each CLI's `capabilities.privilegedEnvKeys` so
`clampEnvOverridesForOwner()` strips them for a non-granted multi-user
owner, exactly the precedent already documented for `DEEPSEEK_BASE_URL`/
`OMP_AUTH_BROKER_URL`. This matters *more*, not less, now that endpoints can
`OMP_AUTH_BROKER_URL`. This matters _more_, not less, now that endpoints can
be cloud URLs: redirecting a non-granted user's session to an attacker's
cloud endpoint is a credential-exfiltration path, not just a mischief
redirect to a LAN box. Endpoint CRUD itself stays admin-only in multi-user
@@ -259,14 +264,14 @@ mode, same as remote/docker hosts.
- `src/web/public/index.html`, `settings-ui.js`, `session-ui.js`, `styles.css` — settings group, toolbar button/menu, badge, accent CSS
- `src/web/sse-events.ts` + `constants.js` — if a dedicated SSE event is warranted for the badge (or just ride existing session-update broadcasts)
- `test/fixtures/mock-openai-server.ts` (new) + `test/custom-model-injection-contract.test.ts` (new) — see Mock-server validation below
- `scripts/test-local-llm-harnesses.mjs` (already added, this branch) — the standalone real-CLI-and-real-endpoint smoke test; despite the filename (kept for continuity with when it was written) it already supports any `--base-url`, local or cloud
- `scripts/test-local-llm-harnesses.ts` (already added, this branch; run via `npx tsx`) — the standalone real-CLI-and-real-endpoint smoke test, supporting any `--base-url` (local or cloud). Dynamic: derives its harness list and every env var/config it injects from the live CLI registry + `buildCustomModelInjection()` rather than a second hand-maintained copy — only the one-shot invocation flags (`ONE_SHOT` table) are CLI-specific info the registry doesn't model and stay hand-maintained
- `docs/custom-model-endpoints.md` (new) + a CLAUDE.md pointer bullet under External CLI modes / envOverrides
## Mock-server validation strategy (CI-runnable, no real CLI binaries needed)
Spawning nine real CLI binaries in CI isn't realistic, and neither Devvyn's
llama.cpp box nor a real cloud subscription can be a CI dependency. So the
injection *logic* gets a tier of automated coverage that sits between the
injection _logic_ gets a tier of automated coverage that sits between the
pure unit tests and the live manual checks in Verification:
1. **`test/fixtures/mock-openai-server.ts`** — a small in-process HTTP
@@ -306,16 +311,21 @@ pure unit tests and the live manual checks in Verification:
per-session dir rather than the user's real config path.
3. **Explicit, stated limitation** (goes in the test file's `@fileoverview`
and in this doc, not left implicit): this proves *"if the CLI honors its
and in this doc, not left implicit): this proves _"if the CLI honors its
documented env/config contract, it will hit the right endpoint with the
right model."* It does **not** prove the real CLI binary actually reads
right model."_ It does **not** prove the real CLI binary actually reads
that env var / config file the way its docs say — that's still the job
of the live manual checks in Verification step 4-5 below, and is exactly
why the confidence table above stays "unverified" for six of the nine
CLIs until someone runs those binaries for real. The mock-server suite
catches regressions in Codeman's own logic; it cannot catch a CLI
changing its env-var name in a future release, or a real cloud endpoint
behaving differently from the mock.
why the confidence table above did not stop at "researched" — every CLI
except antigravity (no mechanism at all) has since been run against a
real llama-swap server via `scripts/test-local-llm-harnesses.ts`:
claude/opencode/pi/grok/omp are confirmed PASS end-to-end, codex is
confirmed FAIL for a real documented protocol reason (Responses-API-only
since Feb 2026), and gemini/deepseek are confirmed reaching the server
but failing for reasons not yet root-caused (see their table rows). The
mock-server suite catches regressions in Codeman's own logic; it cannot
catch a CLI changing its env-var name in a future release, or a real
cloud endpoint behaving differently from a local llama.cpp box.
## Verification
@@ -326,15 +336,20 @@ pure unit tests and the live manual checks in Verification:
3. Route tests (`app.inject`) for the new CRUD + discover-models endpoint
(mock `fetch` for `/v1/models`), and for the multi-user clamp on the new
privileged keys (mirror `test/routes/external-cli-bypass-clamp.test.ts`).
4. **Standalone real-binary smoke test**: `scripts/test-local-llm-harnesses.mjs`
(already written on this branch) exercises every harness against a real
`--base-url` — local or cloud — outside of Codeman's UI entirely. Run it
against Devvyn's llama.cpp server first (`claude`/`opencode`/`codex`
should PASS, since those recipes are verified; the rest report
UNCONFIRMED/SKIP until their guessed flags are corrected via
`--probe-help`), then again against a real cloud endpoint (e.g. an Azure
AI Foundry deployment) once one is available, to prove the `authStyle`/
deployment-name handling holds up outside llama.cpp.
4. **Standalone real-binary smoke test**: `scripts/test-local-llm-harnesses.ts`
exercises every harness the CLI registry declares `customModelInjection`
support for against a real `--base-url` — local or cloud — outside of
Codeman's UI entirely, and is DYNAMIC (reads `enabledClis()` + calls the
real `buildCustomModelInjection()`, so a future registry change is picked
up automatically with zero edits to the script). Already run to
completion against Devvyn's llama-swap server (`http://10.10.11.241:8080`,
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected**
(Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED**
(reach the server, fail for undiagnosed reasons — see their table rows),
antigravity **SKIP** (no mechanism). Re-run this against a real cloud
endpoint (e.g. an Azure AI Foundry deployment) once one is available, to
prove the `authStyle`/deployment-name handling holds up outside llama.cpp.
5. Once the full feature (not just the standalone script) is built: add an
endpoint via the real UI, hit discover-models, confirm the returned model
list, pick Claude + the model on a real session, confirm via
+24 -9
View File
@@ -81,15 +81,30 @@ pointed at.
## Confidence per harness
Only Claude, opencode, and Codex have been verified against a real
llama.cpp server by hand. Gemini, Pi, Grok, DeepSeek, and OMP's recipes are
correct on their one-shot invocation flags (confirmed against real
installed binaries' own `--help` output) but their env-var/config
conventions for a _custom_ endpoint are still web-researched, not verified
end-to-end — see the confidence table in `deployment_plan.md` before relying
on one of those five in production. `scripts/test-local-llm-harnesses.mjs`
is the standalone script used to check a harness against a real endpoint
outside the web UI entirely; see its own `--help` for usage.
Every harness except Antigravity has now been run end-to-end against a real
llama-swap server via `scripts/test-local-llm-harnesses.ts` (a dynamic
script that reads the live CLI registry, so a registry change is picked up
automatically). Results:
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
came back through the endpoint.
- **Codex** — the config is structurally correct, but Codex only speaks the
Responses API since Feb 2026, which llama.cpp/llama-swap don't implement.
This is a real protocol incompatibility, not a bug here; Codex support
needs a Responses-API-compatible endpoint.
- **Gemini** — fails with `Invalid auth method selected`, traced to an
undocumented `GATEWAY` auth path gemini-cli selects once
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
(several auth workarounds were tried and ruled out); do not rely on
Gemini support yet.
- **DeepSeek** — the request reaches the server (env vars are read) but
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
See the confidence table in `deployment_plan.md` for the full detail behind
each result. `scripts/test-local-llm-harnesses.ts` is the standalone script
used to check a harness against a real endpoint outside the web UI
entirely; see its own `--help` for usage.
## Security note
-631
View File
@@ -1,631 +0,0 @@
#!/usr/bin/env node
/**
* Standalone smoke-test for pointing each Codeman-supported harness CLI at a
* custom OpenAI-compatible endpoint — local (llama.cpp, Ollama, vLLM, ...) or
* cloud (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
* self-hosted gateway, ...). Anything that answers GET /v1/models and POST
* /v1/chat/completions in the standard shape qualifies; --base-url is not
* assumed to be a LAN address.
*
* This is intentionally OUTSIDE the npm test suite and outside Codeman's own
* session/tmux machinery: it spawns each real CLI binary directly, one-shot,
* with the env vars / config files that CLI's own docs say redirect it to a
* custom endpoint, and checks it can answer "hello world".
*
* Cloud endpoints often differ from a bare llama.cpp box in two ways this
* script accounts for: (1) auth may be an `api-key` header (Azure's
* convention) rather than `Authorization: Bearer` — the baseline check in
* Step 0 sends both, since an extra header is harmless to servers that
* ignore it; each CLI's OWN auth convention (set via its env vars/config,
* not this script) still needs to match what that endpoint expects. (2) a
* cloud endpoint's "model" may actually be a deployment name distinct from
* the model family (Azure AI Foundry deployments) — always pass --model
* explicitly for those rather than relying on GET /v1/models discovery.
*
* IMPORTANT CONFIDENCE NOTE: only claude/opencode/codex recipes are verified
* (Devvyn confirmed them by hand). gemini/pi/grok/deepseek/omp are best
* guesses from public docs, not verified against this repo or against real
* binaries. antigravity has no known CLI/env mechanism at all and is always
* skipped. Read a harness's UNCONFIRMED/FAIL output before trusting it — use
* --probe-help to read that binary's real --help and fix the guessed flag.
*
* Usage:
* node scripts/test-local-llm-harnesses.mjs --base-url http://192.168.1.50:8080 [options]
* node scripts/test-local-llm-harnesses.mjs --base-url https://<resource>.services.ai.azure.com/openai/v1 --model <deployment-name> --api-key $AZURE_AI_KEY
*
* Options:
* --base-url <url> Required. Root URL of the OpenAI-compatible endpoint (local or cloud).
* --model <name> Model/deployment id to request. Default: first from GET /v1/models.
* --api-key <key> API key to send. Default: local-dummy-key (fine for llama.cpp; required for most cloud endpoints).
* --auth-style <style> "bearer" (default, Authorization: Bearer) or "api-key" (the
* `api-key` header some cloud gateways, e.g. Azure, want).
* NEVER send both — live-tested against a real server, doing
* so reliably HANGS the request indefinitely.
* --prompt <text> Prompt to send. Default: "Reply with exactly: hello world".
* --only <id,id,...> Restrict to these harness ids (comma-separated).
* --timeout <ms> Per-harness spawn timeout. Default: 30000.
* --probe-help Instead of testing, resolve each installed binary and print --help.
* --keep-temp Don't delete generated per-harness config dirs afterward.
* --list Dry run: print the resolved plan per harness, execute nothing.
* -h, --help Show this help.
*/
import { execFileSync, spawn } from 'node:child_process';
import { mkdtempSync, mkdirSync, writeFileSync, rmSync, readFileSync, existsSync } from 'node:fs';
import { tmpdir, homedir } from 'node:os';
import { join, delimiter, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
const TAG = '[test-local-llm-harnesses]';
const SCRIPT_DIR = dirname(fileURLToPath(import.meta.url));
const CONFIG_PATH = join(SCRIPT_DIR, 'local-llm-test.config.json');
const CONFIG_EXAMPLE_PATH = join(SCRIPT_DIR, 'local-llm-test.config.example.json');
/**
* Loads scripts/local-llm-test.config.json (gitignored — real IP/model/key,
* per-machine) if present, so you don't have to retype --base-url every run.
* See local-llm-test.config.example.json (tracked) for the shape. CLI flags
* always override whatever this file sets; this only supplies defaults.
*/
function loadConfigFile() {
if (!existsSync(CONFIG_PATH)) return {};
try {
const raw = JSON.parse(readFileSync(CONFIG_PATH, 'utf8'));
return {
baseUrl: raw.baseUrl ?? null,
model: raw.model ?? null,
apiKey: raw.apiKey || undefined, // empty string counts as "not set", not a real key
authStyle: raw.authStyle === 'api-key' ? 'api-key' : undefined, // never 'both'
prompt: raw.prompt ?? undefined,
only: Array.isArray(raw.only) && raw.only.length ? raw.only : null,
timeout: typeof raw.timeout === 'number' ? raw.timeout : undefined,
};
} catch (err) {
console.error(`${TAG} failed to parse ${CONFIG_PATH}: ${err.message} (ignoring it)`);
return {};
}
}
function parseArgs(argv, configDefaults) {
const opts = {
baseUrl: configDefaults.baseUrl ?? null,
model: configDefaults.model ?? null,
apiKey: configDefaults.apiKey ?? 'local-dummy-key',
authStyle: configDefaults.authStyle ?? 'bearer',
prompt: configDefaults.prompt ?? 'Reply with exactly: hello world',
only: configDefaults.only ?? null,
timeout: configDefaults.timeout ?? 30000,
probeHelp: false,
keepTemp: false,
list: false,
help: false,
};
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
switch (a) {
case '--base-url':
opts.baseUrl = argv[++i];
break;
case '--model':
opts.model = argv[++i];
break;
case '--api-key':
opts.apiKey = argv[++i];
break;
case '--auth-style':
opts.authStyle = argv[++i];
if (opts.authStyle !== 'bearer' && opts.authStyle !== 'api-key') {
console.error(`${TAG} --auth-style must be "bearer" or "api-key"`);
opts.help = true;
}
break;
case '--prompt':
opts.prompt = argv[++i];
break;
case '--only':
opts.only = argv[++i].split(',').map((s) => s.trim()).filter(Boolean);
break;
case '--timeout':
opts.timeout = Number(argv[++i]);
break;
case '--probe-help':
opts.probeHelp = true;
break;
case '--keep-temp':
opts.keepTemp = true;
break;
case '--list':
opts.list = true;
break;
case '-h':
case '--help':
opts.help = true;
break;
default:
console.error(`${TAG} unknown argument: ${a}`);
opts.help = true;
}
}
return opts;
}
function printUsage() {
console.log(`Usage: node scripts/test-local-llm-harnesses.mjs [--base-url <url>] [options]
Reads defaults from scripts/local-llm-test.config.json if it exists (copy
scripts/local-llm-test.config.example.json to create it — gitignored, since
it holds a real IP/model/key). CLI flags always override the config file.
--base-url becomes optional once that file supplies one.
Works against any custom OpenAI-compatible endpoint, local or cloud
(llama.cpp, Ollama, vLLM, Azure AI Foundry, OpenRouter, a self-hosted
gateway, ...) — anything answering GET /v1/models and POST
/v1/chat/completions in the standard shape.
Options:
--base-url <url> Required. Root URL of the OpenAI-compatible endpoint.
--model <name> Model/deployment id to request. Default: first from GET /v1/models.
--api-key <key> API key to send. Default: local-dummy-key (required for most cloud endpoints).
--auth-style <style> "bearer" (default) or "api-key" (Azure-style). Never both — sending
both headers together reliably hangs some real servers.
--prompt <text> Prompt to send. Default: "Reply with exactly: hello world".
--only <id,id,...> Restrict to these harness ids.
--timeout <ms> Per-harness spawn timeout. Default: 30000.
--probe-help Print each installed binary's --help instead of testing.
--keep-temp Keep generated per-harness config dirs afterward.
--list Dry run: print the resolved plan, execute nothing.
-h, --help Show this help.
Harness ids: claude, opencode, codex, gemini, pi, grok, deepseek, omp, antigravity
Examples:
node scripts/test-local-llm-harnesses.mjs --base-url http://192.168.1.50:8080
node scripts/test-local-llm-harnesses.mjs --base-url https://<resource>.services.ai.azure.com/openai/v1 --model <deployment-name> --api-key $AZURE_AI_KEY`);
}
const HOME = homedir();
const EXTRA_SEARCH_DIRS = [
join(HOME, '.local', 'bin'),
join(HOME, '.opencode', 'bin'),
join(HOME, '.codex', 'bin'),
join(HOME, '.gemini', 'bin'),
join(HOME, '.antigravity', 'bin'),
join(HOME, '.grok', 'bin'),
join(HOME, '.omp', 'bin'),
join(HOME, '.bun', 'bin'),
join(HOME, '.npm-global', 'bin'),
join(HOME, 'bin'),
'/usr/local/bin',
];
function pathWithExtraDirs() {
return [...EXTRA_SEARCH_DIRS, process.env.PATH ?? ''].join(delimiter);
}
/** Resolve a binary by trying `<bin> --version` with extra search dirs prefixed onto PATH. */
function resolveBinary(bin) {
try {
execFileSync(bin, ['--version'], {
timeout: 5000,
stdio: 'pipe',
env: { ...process.env, PATH: pathWithExtraDirs() },
});
return bin;
} catch (err) {
// Some CLIs (e.g. dsh) don't support --version cleanly for identity but
// still exist on PATH; a non-ENOENT failure still counts as "found".
if (err && err.code === 'ENOENT') return null;
return bin;
}
}
function printHelp(bin) {
try {
const out = execFileSync(bin, ['--help'], {
timeout: 5000,
stdio: 'pipe',
env: { ...process.env, PATH: pathWithExtraDirs() },
});
console.log(out.toString());
} catch (err) {
console.log((err.stdout ?? err.message ?? String(err)).toString());
}
}
// --- per-harness definitions -----------------------------------------------
/** kind: 'env' | 'configContentEnv' | 'configDir' | 'unsupported' */
const HARNESSES = {
claude: {
binary: 'claude',
confidence: 'verified',
buildEnv: (baseUrl, apiKey, model) => ({
ANTHROPIC_BASE_URL: baseUrl,
ANTHROPIC_API_KEY: apiKey,
ANTHROPIC_DEFAULT_SONNET_MODEL: model,
ANTHROPIC_DEFAULT_HAIKU_MODEL: model,
ANTHROPIC_DEFAULT_OPUS_MODEL: model,
}),
// Claude Code's async session-title-generation call also uses
// ANTHROPIC_DEFAULT_HAIKU_MODEL and validates it against Claude's OWN internal
// recognized-model list, printing [claude-code:unrecognized_model] to stderr for
// a local model name. Confirmed live: `--settings '{"autoTitle":false}'` does NOT
// stop it (still hung the whole run); `--bare` does — the warning still prints,
// but the actual prompt now runs and returns the real answer. Confirmed against
// a real llama-swap server.
buildArgv: (prompt) => ['--dangerously-skip-permissions', '--bare', '-p', prompt],
},
opencode: {
binary: 'opencode',
confidence: 'verified',
buildEnv: (baseUrl, apiKey, model) => ({
OPENCODE_CONFIG_CONTENT: JSON.stringify({
$schema: 'https://opencode.ai/config.json',
provider: {
local: {
options: { baseURL: `${baseUrl}/v1`, apiKey },
models: { [model]: {} },
},
},
model: `local/${model}`,
}),
}),
buildArgv: (prompt) => ['run', prompt],
},
codex: {
binary: 'codex',
confidence: 'verified',
configDir: {
dirEnvVar: 'CODEX_HOME',
fileName: 'config.toml',
// Verified against a real codex binary: `model` must be a top-level STRING
// (an earlier `[model].default` table was rejected with "invalid type: map,
// expected a string"). The API key is NEVER a literal TOML field — codex only
// supports `env_key`, the NAME of an env var it reads the value from, so the
// real key rides as an extra env var (see extraEnv below), never in the file.
// ⚠️ `wire_api = "responses"` is the only value codex still accepts (support
// for "chat" was dropped Feb 2026) — a plain OpenAI Chat-Completions server
// (llama.cpp, llama-swap) does NOT implement the Responses API, so this may
// still fail at the PROTOCOL level even with a correctly-shaped file.
content: (baseUrl, _apiKey, model) =>
`model = "${model}"\nmodel_provider = "custom"\n\n[model_providers.custom]\nname = "Custom Endpoint"\nbase_url = "${baseUrl}/v1"\nenv_key = "CODEMAN_CUSTOM_MODEL_API_KEY"\nwire_api = "responses"\n`,
extraEnv: (_baseUrl, apiKey) => ({ CODEMAN_CUSTOM_MODEL_API_KEY: apiKey }),
},
buildArgv: (prompt) => ['exec', '--dangerously-bypass-approvals-and-sandbox', prompt],
},
gemini: {
binary: 'gemini',
confidence: 'researched',
buildEnv: (baseUrl, apiKey, model) => ({
GOOGLE_GEMINI_BASE_URL: baseUrl,
GEMINI_API_KEY: apiKey,
GEMINI_MODEL: model,
}),
buildArgv: (prompt) => ['-p', prompt, '--approval-mode', 'yolo'],
},
pi: {
binary: 'pi',
confidence: 'researched',
configDir: {
dirEnvVar: 'PI_CONFIG_DIR',
fileName: join('agent', 'models.json'),
content: (baseUrl, apiKey, model) =>
JSON.stringify(
{
providers: {
local: {
baseUrl: `${baseUrl}/v1`,
apiKey,
api: 'openai-completions',
models: { [model]: {} },
},
},
},
null,
2
),
},
buildArgv: (prompt) => ['--approve', '-p', prompt],
},
grok: {
binary: 'grok',
confidence: 'researched',
buildEnv: (baseUrl, apiKey, model) => ({
GROK_BASE_URL: baseUrl,
XAI_API_KEY: apiKey,
GROK_MODEL: model,
}),
buildArgv: (prompt) => ['--always-approve', '-p', prompt],
},
deepseek: {
binary: 'dsh',
confidence: 'unknown',
note: 'dsh is a profile launcher, not a documented one-shot prompt flag. Best-effort only.',
buildEnv: (baseUrl, apiKey) => ({
DEEPSEEK_BASE_URL: baseUrl,
DEEPSEEK_API_KEY: apiKey,
DSH_PERMISSION_MODE: 'danger-full-access',
}),
buildArgv: (prompt) => ['--profile', 'headless', prompt],
},
omp: {
binary: 'omp',
confidence: 'researched',
configDir: {
dirEnvVar: 'PI_CONFIG_DIR', // omp's ~/.omp tree is relocatable via PI_CONFIG_DIR per CLAUDE.md
fileName: join('agent', 'models.yml'),
content: (baseUrl, apiKey, model) =>
`providers:\n local:\n baseUrl: ${baseUrl}/v1\n apiKey: ${apiKey}\n models:\n - ${model}\n`,
},
buildArgv: (prompt) => ['-p', prompt],
},
antigravity: {
binary: 'agy',
confidence: 'unsupported',
note: 'No known CLI/env/config mechanism for a custom endpoint (GUI-only per public docs). Always skipped.',
buildEnv: null,
buildArgv: null,
},
};
// --- baseline server check ---------------------------------------------------
async function baselineCheck(baseUrl, apiKey, authStyle, model, prompt, timeoutMs) {
console.log(`\n=== Step 0: baseline check against ${baseUrl} (auth: ${authStyle}) ===`);
// Exactly ONE header, never both. An earlier version sent both auth conventions
// (Bearer + api-key) on the theory that an unused header is harmless — live-
// tested against a real llama-swap server, sending both reliably HUNG the
// request indefinitely (reproduced 3x: Bearer alone ~500ms, api-key alone
// ~600ms, both together no response inside a 15s timeout). Use --auth-style
// api-key for endpoints that specifically want that header (e.g. Azure AI
// Foundry); default 'bearer' covers everything else.
const authHeaders = authStyle === 'api-key' ? { 'api-key': apiKey } : { Authorization: `Bearer ${apiKey}` };
let discoveredModel = model;
try {
const res = await fetch(`${baseUrl}/v1/models`, {
headers: authHeaders,
signal: AbortSignal.timeout(timeoutMs),
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = await res.json();
const ids = (body.data ?? []).map((m) => m.id);
console.log(`GET /v1/models -> ${ids.length ? ids.join(', ') : '(empty list)'}`);
if (!discoveredModel && ids.length) discoveredModel = ids[0];
} catch (err) {
console.error(`${TAG} GET /v1/models failed: ${err.message}`);
console.error(`${TAG} Is the server actually running at ${baseUrl}? Aborting.`);
process.exit(1);
}
if (!discoveredModel) {
console.error(`${TAG} No --model given and none discovered from /v1/models. Aborting.`);
process.exit(1);
}
// Live-tested against a real llama-swap server: a POST issued right after a GET on
// the same Node process reliably HANGS indefinitely (reproduced repeatedly — GET
// alone ~30ms, POST alone ~1-2s, GET-then-immediate-POST times out completely; a
// 2s pause between them fixed it every time). This looks like Node's fetch (undici)
// reusing a pooled keep-alive connection the server doesn't handle cleanly for a
// second request right behind a first. A short pause is the simplest portable fix
// (no extra deps, no need for undici's Agent/dispatcher API).
await new Promise((resolve) => setTimeout(resolve, 2000));
try {
const res = await fetch(`${baseUrl}/v1/chat/completions`, {
method: 'POST',
headers: { 'Content-Type': 'application/json', ...authHeaders },
body: JSON.stringify({
model: discoveredModel,
messages: [{ role: 'user', content: prompt }],
}),
signal: AbortSignal.timeout(timeoutMs),
});
if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
const body = await res.json();
const reply = body.choices?.[0]?.message?.content ?? '';
if (!reply.trim()) throw new Error('empty reply');
console.log(`POST /v1/chat/completions -> "${reply.trim().slice(0, 200)}"`);
console.log('Server baseline: PASS\n');
} catch (err) {
console.error(`${TAG} POST /v1/chat/completions failed: ${err.message}`);
console.error(`${TAG} Server responded to /v1/models but not to a chat request. Aborting.`);
process.exit(1);
}
return discoveredModel;
}
// --- per-harness run ----------------------------------------------------------
function makeTempConfigDir(id) {
const dir = mkdtempSync(join(tmpdir(), `codeman-local-llm-test-${id}-`));
return dir;
}
function runChild(bin, argv, env, timeoutMs) {
return new Promise((resolve) => {
let stdout = '';
let stderr = '';
let settled = false;
const child = spawn(bin, argv, {
env: { ...process.env, ...env, PATH: pathWithExtraDirs() },
stdio: ['ignore', 'pipe', 'pipe'],
});
const timer = setTimeout(() => {
if (settled) return;
settled = true;
child.kill('SIGKILL');
resolve({ code: null, stdout, stderr, timedOut: true });
}, timeoutMs);
child.stdout.on('data', (d) => (stdout += d.toString()));
child.stderr.on('data', (d) => (stderr += d.toString()));
child.on('error', (err) => {
if (settled) return;
settled = true;
clearTimeout(timer);
resolve({ code: null, stdout, stderr: `${stderr}\n${err.message}`, timedOut: false });
});
child.on('close', (code) => {
if (settled) return;
settled = true;
clearTimeout(timer);
resolve({ code, stdout, stderr, timedOut: false });
});
});
}
async function runHarness(id, def, opts, model) {
const result = { id, confidence: def.confidence, status: 'SKIP', detail: '' };
if (def.confidence === 'unsupported') {
result.status = 'SKIP';
result.detail = def.note ?? 'no known mechanism';
return result;
}
const resolved = resolveBinary(def.binary);
if (!resolved) {
result.status = 'SKIP';
result.detail = `binary "${def.binary}" not found on PATH or search dirs`;
return result;
}
let env = def.buildEnv ? def.buildEnv(opts.baseUrl, opts.apiKey, model) : {};
let tempDir = null;
if (def.configDir) {
tempDir = makeTempConfigDir(id);
const filePath = join(tempDir, def.configDir.fileName);
mkdirSync(join(filePath, '..'), { recursive: true });
writeFileSync(filePath, def.configDir.content(opts.baseUrl, opts.apiKey, model), 'utf8');
const extraEnv = def.configDir.extraEnv ? def.configDir.extraEnv(opts.baseUrl, opts.apiKey, model) : {};
env = { ...env, [def.configDir.dirEnvVar]: tempDir, ...extraEnv };
}
const argv = def.buildArgv(opts.prompt);
if (opts.list) {
result.status = 'LIST';
result.detail = `${def.binary} ${argv.join(' ')} | env: ${Object.keys(env).join(', ')}${
tempDir ? ` | configDir: ${tempDir}` : ''
}`;
if (tempDir && !opts.keepTemp) rmSync(tempDir, { recursive: true, force: true });
return result;
}
const { code, stdout, stderr, timedOut } = await runChild(def.binary, argv, env, opts.timeout);
if (tempDir && !opts.keepTemp) rmSync(tempDir, { recursive: true, force: true });
else if (tempDir) result.detail += ` [config kept at ${tempDir}]`;
if (timedOut) {
result.status = 'FAIL';
result.detail = `timed out after ${opts.timeout}ms. stderr: ${stderr.slice(-300)}`;
return result;
}
const reply = stdout.trim();
const matched = /hello/i.test(reply) && /world/i.test(reply);
if (code !== 0) {
result.status = def.confidence === 'verified' ? 'FAIL' : 'UNCONFIRMED';
result.detail = `exit ${code}. stderr: ${stderr.trim().slice(-300) || '(empty)'}`;
return result;
}
if (!reply) {
result.status = def.confidence === 'verified' ? 'FAIL' : 'UNCONFIRMED';
result.detail = 'exit 0 but empty stdout';
return result;
}
if (matched) {
result.status = 'PASS';
result.detail = reply.slice(0, 200);
} else {
result.status = 'UNCONFIRMED';
result.detail = `reply didn't match heuristic, judge by eye: "${reply.slice(0, 300)}"`;
}
return result;
}
// --- main ---------------------------------------------------------------------
async function main() {
const configDefaults = loadConfigFile();
const opts = parseArgs(process.argv.slice(2), configDefaults);
if (opts.help) {
printUsage();
process.exit(0);
}
const ids = opts.only ?? Object.keys(HARNESSES);
const unknownIds = ids.filter((id) => !HARNESSES[id]);
if (unknownIds.length) {
console.error(`${TAG} unknown harness id(s): ${unknownIds.join(', ')}`);
console.error(`${TAG} known ids: ${Object.keys(HARNESSES).join(', ')}`);
process.exit(1);
}
// --probe-help never touches the network — no --base-url needed for it.
if (opts.probeHelp) {
for (const id of ids) {
const def = HARNESSES[id];
const resolved = resolveBinary(def.binary);
console.log(`\n=== ${id} (${def.binary}) ===`);
if (!resolved) {
console.log('(not found on PATH or search dirs)');
continue;
}
printHelp(def.binary);
}
process.exit(0);
}
if (!opts.baseUrl) {
console.error(`${TAG} --base-url is required (pass it, or set "baseUrl" in ${CONFIG_PATH}).`);
console.error(`${TAG} See ${CONFIG_EXAMPLE_PATH} for the config file shape.\n`);
printUsage();
process.exit(1);
}
opts.baseUrl = opts.baseUrl.replace(/\/+$/, '');
// --list is a pure dry run: never touch the network, even if --model was given.
let model = opts.model;
if (opts.list) {
model = opts.model ?? 'local-model';
console.log(`\n=== Step 0 skipped (--list never hits the network; using placeholder "${model}") ===\n`);
} else {
model = await baselineCheck(opts.baseUrl, opts.apiKey, opts.authStyle, opts.model, opts.prompt, opts.timeout);
}
console.log(`=== Testing ${ids.length} harness(es) ===`);
const results = [];
for (const id of ids) {
process.stdout.write(`\n--- ${id} ---\n`);
const result = await runHarness(id, HARNESSES[id], opts, model);
results.push(result);
console.log(`${result.status}: ${result.detail}`);
}
console.log('\n=== Summary ===');
const width = Math.max(...results.map((r) => r.id.length)) + 2;
for (const r of results) {
console.log(`${r.id.padEnd(width)} [${r.confidence.padEnd(11)}] ${r.status.padEnd(11)} ${r.detail.slice(0, 100)}`);
}
const hardFail = results.some((r) => r.status === 'FAIL' && r.confidence === 'verified');
if (hardFail) {
console.error(`\n${TAG} at least one VERIFIED harness FAILed — that's a real regression, not just an unconfirmed guess.`);
process.exit(1);
}
process.exit(0);
}
main().catch((err) => {
console.error(`${TAG} unexpected error:`, err);
process.exit(1);
});
+699
View File
@@ -0,0 +1,699 @@
#!/usr/bin/env -S npx tsx
/**
* Standalone smoke-test for pointing each Codeman-supported harness CLI at a
* custom OpenAI-compatible endpoint — local (llama.cpp, Ollama, vLLM, ...) or
* cloud (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
* self-hosted gateway, ...). Anything that answers GET /v1/models and POST
* /v1/chat/completions in the standard shape qualifies; --base-url is not
* assumed to be a LAN address.
*
* This is intentionally OUTSIDE the npm test suite and outside Codeman's own
* session/tmux machinery: it spawns each real CLI binary directly, one-shot,
* with the env vars / config files that CLI's own docs say redirect it to a
* custom endpoint, and checks it can answer "hello world".
*
* DYNAMIC BY DESIGN: this file imports the SAME `enabledClis()` registry and
* `buildCustomModelInjection()` builder the production feature uses (see
* ../src/config/cli-registry/, ../src/custom-model-injection.ts,
* ../src/custom-model-injection-apply.ts) rather than keeping a second,
* hand-maintained copy of each CLI's env vars/config shape. A registry
* change (a new CLI, an edited env var name, a fixed config template) is
* picked up here automatically with zero edits to this file. Only the
* ONE-SHOT INVOCATION FLAGS (how to make each CLI answer one prompt and
* exit — information the registry doesn't model at all, since it only knows
* how to launch the interactive TUI) stay in the small ONE_SHOT table below;
* a CLI newly added to the registry with no ONE_SHOT entry is reported
* UNKNOWN rather than silently skipped or guessed at.
*
* Cloud endpoints often differ from a bare llama.cpp box in two ways this
* script accounts for: (1) auth may be an `api-key` header (Azure's
* convention) rather than `Authorization: Bearer` — see --auth-style below.
* (2) a cloud endpoint's "model" may actually be a deployment name distinct
* from the model family (Azure AI Foundry deployments) — always pass
* --model explicitly for those rather than relying on GET /v1/models
* discovery.
*
* IMPORTANT CONFIDENCE NOTE: claude and opencode are verified end-to-end
* against a real llama-swap server. codex's config STRUCTURE is verified,
* but it only speaks the Responses API (dropped Chat-Completions support
* Feb 2026) — expect it to fail against a plain OpenAI-compatible server,
* that's a real protocol gap, not a bug here. gemini/pi/grok/omp have their
* ONE-SHOT INVOCATION flags confirmed against real installed binaries'
* `--help` output, but their custom-endpoint env/config conventions remain
* web-researched, unverified. deepseek (dsh) is a profile launcher with no
* documented one-shot prompt flag at all — best-effort only. antigravity
* has no known CLI/env/config mechanism (GUI-only per public docs) — its
* registry entry declares `customModelInjection: { kind: 'unsupported' }`,
* which this script picks up dynamically and always skips.
*
* Usage:
* npx tsx scripts/test-local-llm-harnesses.ts --base-url http://192.168.1.50:8080 [options]
* npx tsx scripts/test-local-llm-harnesses.ts --base-url https://<resource>.services.ai.azure.com/openai/v1 --model <deployment-name> --api-key $AZURE_AI_KEY
*
* Options:
* --base-url <url> Required. Root URL of the OpenAI-compatible endpoint (local or cloud).
* --model <name> Model/deployment id to request. Default: first from GET /v1/models.
* --api-key <key> API key to send. Default: local-dummy-key (fine for llama.cpp; required for most cloud endpoints).
* --auth-style <style> "bearer" (default, Authorization: Bearer) or "api-key" (the
* `api-key` header some cloud gateways, e.g. Azure, want).
* NEVER send both — live-tested against a real server, doing
* so reliably HANGS the request indefinitely.
* --prompt <text> Prompt to send. Default: "Reply with exactly: hello world".
* --only <id,id,...> Restrict to these harness ids (comma-separated).
* --timeout <ms> Per-harness spawn timeout. Default: 30000.
* --probe-help Instead of testing, resolve each installed binary and print --help.
* --keep-temp Don't delete generated per-harness config dirs afterward.
* --list Dry run: print the resolved plan per harness, execute nothing.
* -h, --help Show this help.
*/
import { execFileSync, spawn } from 'node:child_process';
import { mkdtempSync, rmSync, readFileSync, existsSync } from 'node:fs';
import { tmpdir, homedir } from 'node:os';
import { join, delimiter, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
import { enabledClis } from '../src/config/cli-registry/index.js';
import type { CliEntry } from '../src/config/cli-registry/types.js';
import {
buildCustomModelInjection,
GROK_CUSTOM_MODEL_NAME,
type CustomModelEndpoint,
} from '../src/custom-model-injection.js';
import { applyConfigDirInjection } from '../src/custom-model-injection-apply.js';
const TAG = '[test-local-llm-harnesses]';
const SCRIPT_DIR = dirname(fileURLToPath(import.meta.url));
const CONFIG_PATH = join(SCRIPT_DIR, 'local-llm-test.config.json');
const CONFIG_EXAMPLE_PATH = join(SCRIPT_DIR, 'local-llm-test.config.example.json');
type AuthStyle = 'bearer' | 'api-key';
interface ConfigDefaults {
baseUrl?: string | null;
model?: string | null;
apiKey?: string;
authStyle?: AuthStyle;
prompt?: string;
only?: string[] | null;
timeout?: number;
}
/**
* Loads scripts/local-llm-test.config.json (gitignored — real IP/model/key,
* per-machine) if present, so you don't have to retype --base-url every run.
* See local-llm-test.config.example.json (tracked) for the shape. CLI flags
* always override whatever this file sets; this only supplies defaults.
*/
function loadConfigFile(): ConfigDefaults {
if (!existsSync(CONFIG_PATH)) return {};
try {
const raw = JSON.parse(readFileSync(CONFIG_PATH, 'utf8'));
return {
baseUrl: raw.baseUrl ?? null,
model: raw.model ?? null,
apiKey: raw.apiKey || undefined, // empty string counts as "not set", not a real key
authStyle: raw.authStyle === 'api-key' ? 'api-key' : undefined, // never 'both'
prompt: raw.prompt ?? undefined,
only: Array.isArray(raw.only) && raw.only.length ? raw.only : null,
timeout: typeof raw.timeout === 'number' ? raw.timeout : undefined,
};
} catch (err) {
console.error(`${TAG} failed to parse ${CONFIG_PATH}: ${(err as Error).message} (ignoring it)`);
return {};
}
}
interface Opts {
baseUrl: string | null;
model: string | null;
apiKey: string;
authStyle: AuthStyle;
prompt: string;
only: string[] | null;
timeout: number;
probeHelp: boolean;
keepTemp: boolean;
list: boolean;
help: boolean;
}
function parseArgs(argv: string[], configDefaults: ConfigDefaults): Opts {
const opts: Opts = {
baseUrl: configDefaults.baseUrl ?? null,
model: configDefaults.model ?? null,
apiKey: configDefaults.apiKey ?? 'local-dummy-key',
authStyle: configDefaults.authStyle ?? 'bearer',
prompt: configDefaults.prompt ?? 'Reply with exactly: hello world',
only: configDefaults.only ?? null,
timeout: configDefaults.timeout ?? 30000,
probeHelp: false,
keepTemp: false,
list: false,
help: false,
};
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
switch (a) {
case '--base-url':
opts.baseUrl = argv[++i];
break;
case '--model':
opts.model = argv[++i];
break;
case '--api-key':
opts.apiKey = argv[++i];
break;
case '--auth-style':
opts.authStyle = argv[++i] as AuthStyle;
if (opts.authStyle !== 'bearer' && opts.authStyle !== 'api-key') {
console.error(`${TAG} --auth-style must be "bearer" or "api-key"`);
opts.help = true;
}
break;
case '--prompt':
opts.prompt = argv[++i];
break;
case '--only':
opts.only = argv[++i]
.split(',')
.map((s) => s.trim())
.filter(Boolean);
break;
case '--timeout':
opts.timeout = Number(argv[++i]);
break;
case '--probe-help':
opts.probeHelp = true;
break;
case '--keep-temp':
opts.keepTemp = true;
break;
case '--list':
opts.list = true;
break;
case '-h':
case '--help':
opts.help = true;
break;
default:
console.error(`${TAG} unknown argument: ${a}`);
opts.help = true;
}
}
return opts;
}
function printUsage(): void {
console.log(`Usage: npx tsx scripts/test-local-llm-harnesses.ts [--base-url <url>] [options]
Reads defaults from scripts/local-llm-test.config.json if it exists (copy
scripts/local-llm-test.config.example.json to create it — gitignored, since
it holds a real IP/model/key). CLI flags always override the config file.
--base-url becomes optional once that file supplies one.
Works against any custom OpenAI-compatible endpoint, local or cloud
(llama.cpp, Ollama, vLLM, Azure AI Foundry, OpenRouter, a self-hosted
gateway, ...) — anything answering GET /v1/models and POST
/v1/chat/completions in the standard shape.
Options:
--base-url <url> Required. Root URL of the OpenAI-compatible endpoint.
--model <name> Model/deployment id to request. Default: first from GET /v1/models.
--api-key <key> API key to send. Default: local-dummy-key (required for most cloud endpoints).
--auth-style <style> "bearer" (default) or "api-key" (Azure-style). Never both — sending
both headers together reliably hangs some real servers.
--prompt <text> Prompt to send. Default: "Reply with exactly: hello world".
--only <id,id,...> Restrict to these harness ids.
--timeout <ms> Per-harness spawn timeout. Default: 30000.
--probe-help Print each installed binary's --help instead of testing.
--keep-temp Keep generated per-harness config dirs afterward.
--list Dry run: print the resolved plan, execute nothing.
-h, --help Show this help.
Harness ids are read from the CLI registry at run time — pass an unknown
one and the error message lists what's actually enabled right now.
Examples:
npx tsx scripts/test-local-llm-harnesses.ts --base-url http://192.168.1.50:8080
npx tsx scripts/test-local-llm-harnesses.ts --base-url https://<resource>.services.ai.azure.com/openai/v1 --model <deployment-name> --api-key $AZURE_AI_KEY`);
}
const HOME = homedir();
/** Expands a leading `~` the way the CLI registry's own search dirs are written. */
function expandHome(p: string): string {
if (p === '~') return HOME;
if (p.startsWith('~/')) return join(HOME, p.slice(2));
return p;
}
function pathWithExtraDirs(extraDirs: string[]): string {
return [...extraDirs.map(expandHome), '/usr/local/bin', process.env.PATH ?? ''].join(delimiter);
}
/** Resolve a binary by trying `<bin> --version` with the CLI's own registry search dirs prefixed onto PATH. */
function resolveBinary(bin: string, searchDirs: string[]): string | null {
try {
execFileSync(bin, ['--version'], {
timeout: 5000,
stdio: 'pipe',
env: { ...process.env, PATH: pathWithExtraDirs(searchDirs) },
});
return bin;
} catch (err) {
// Some CLIs (e.g. dsh) don't support --version cleanly for identity but
// still exist on PATH; a non-ENOENT failure still counts as "found".
if (err && (err as NodeJS.ErrnoException).code === 'ENOENT') return null;
return bin;
}
}
function printHelp(bin: string, searchDirs: string[]): void {
try {
const out = execFileSync(bin, ['--help'], {
timeout: 5000,
stdio: 'pipe',
env: { ...process.env, PATH: pathWithExtraDirs(searchDirs) },
});
console.log(out.toString());
} catch (err) {
const e = err as { stdout?: Buffer; message?: string };
console.log((e.stdout ?? e.message ?? String(err)).toString());
}
}
// --- one-shot invocation table (NOT in the registry — genuinely separate info) ---
type Confidence = 'verified' | 'researched' | 'unknown';
interface OneShot {
/** `modelId` is the RAW model/deployment id (e.g. "qwen3.5-0.8b-...") — CLIs whose
* config wraps it under a provider/block name (pi/omp's "custom/<id>", grok's fixed
* block name) build the full `--model` value here, not in the injection layer. */
argv: (prompt: string, modelId: string) => string[];
confidence: Confidence;
note?: string;
}
/**
* How to make each CLI answer ONE prompt and exit. The registry has no concept
* of this (it only knows the interactive TUI launch line), so this table is
* necessarily hand-maintained — but it is the ONLY hand-maintained part left;
* everything about WHERE the prompt goes (env vars, config files) comes from
* the real registry + `buildCustomModelInjection()` above.
*
* A CLI enabled in the registry with no entry here reports UNKNOWN rather
* than being silently skipped or guessed at — see `resolveOneShot()`.
*/
const ONE_SHOT: Record<string, OneShot> = {
claude: {
confidence: 'verified',
// Claude Code's async session-title-generation call also uses
// ANTHROPIC_DEFAULT_HAIKU_MODEL and validates it against Claude's OWN internal
// recognized-model list, printing [claude-code:unrecognized_model] to stderr for
// a local model name. Confirmed live: `--settings '{"autoTitle":false}'` does NOT
// stop it (still hung the whole run); `--bare` does — the warning still prints,
// but the actual prompt now runs and returns the real answer. Confirmed against
// a real llama-swap server. ⚠️ `--bare` also disables hooks/LSP/plugin sync/
// CLAUDE.md auto-discovery — fine for this ISOLATED one-shot test, never safe to
// apply to a real interactive Codeman session (which needs hooks).
argv: (prompt) => ['--dangerously-skip-permissions', '--bare', '-p', prompt],
},
opencode: { confidence: 'verified', argv: (prompt) => ['run', prompt] },
codex: {
confidence: 'verified',
note: 'config STRUCTURE verified; codex only speaks the Responses API (dropped Chat-Completions Feb 2026) — expect FAIL against a plain OpenAI-compatible server, that is a protocol gap, not a bug here.',
argv: (prompt) => ['exec', '--dangerously-bypass-approvals-and-sandbox', prompt],
},
gemini: {
confidence: 'researched',
// --skip-trust: without it, an untrusted-folder check silently overrides
// --approval-mode yolo back to 'default' (confirmed live: "Approval mode
// overridden to 'default' because the current folder is not trusted").
argv: (prompt) => ['-p', prompt, '--approval-mode', 'yolo', '--skip-trust'],
},
pi: {
confidence: 'verified',
// --model custom/<id>: without an explicit --model, pi uses its own default
// provider (not our injected "custom" one) and fails with "No API key found
// for the selected model" — confirmed live. "custom" matches the provider name
// pi-models-json writes in custom-model-injection.ts. Verified end-to-end
// against a real llama-swap server after two real bugs were found and fixed:
// pi's `models` field must be an ARRAY of `{id}` objects (an object keyed by
// id silently loaded zero models), and PI_CONFIG_DIR does nothing for pi at
// all (grepped pi's own bundled source — not present anywhere); the actual
// working redirect is the CHILD PROCESS's `HOME` itself, since pi hardcodes
// `~/.pi/agent/models.json` with no dedicated override.
argv: (prompt, modelId) => ['--approve', '--model', `custom/${modelId}`, '-p', prompt],
},
grok: {
confidence: 'verified',
// -m <block name>: grok's config.toml (grok-toml template) declares the custom
// model under a fixed [model.<name>] block; GROK_CUSTOM_MODEL_NAME is that same
// name, imported from custom-model-injection.ts so the two can never drift apart.
// Verified end-to-end against a real llama-swap server after correcting the
// ORIGINAL recipe, which was wrong (env vars, not a config file — see the
// customModelInjection comment on grok's registry entry).
argv: (prompt) => ['--always-approve', '-m', GROK_CUSTOM_MODEL_NAME, '-p', prompt],
},
deepseek: {
confidence: 'unknown',
note: 'dsh is a profile launcher, not a documented one-shot prompt flag. Best-effort only.',
argv: (prompt) => ['--profile', 'headless', prompt],
},
omp: {
confidence: 'verified',
// --model custom/<id>: same reasoning as pi — omp's own default model has no
// credential, so without an explicit --model it never reaches our injected
// provider at all. Verified end-to-end against a real llama-swap server after
// the same two fixes as pi (array-shaped `models`, HOME-redirect instead of
// PI_CONFIG_DIR — omp hardcodes `~/.omp/agent/models.yml`).
argv: (prompt, modelId) => ['--model', `custom/${modelId}`, '-p', prompt],
},
};
// --- baseline server check ---------------------------------------------------
async function baselineCheck(
baseUrl: string,
apiKey: string,
authStyle: AuthStyle,
model: string | null,
prompt: string,
timeoutMs: number
): Promise<string> {
console.log(`\n=== Step 0: baseline check against ${baseUrl} (auth: ${authStyle}) ===`);
// Exactly ONE header, never both. An earlier version sent both auth conventions
// (Bearer + api-key) on the theory that an unused header is harmless — live-
// tested against a real llama-swap server, sending both reliably HUNG the
// request indefinitely (reproduced 3x: Bearer alone ~500ms, api-key alone
// ~600ms, both together no response inside a 15s timeout). Use --auth-style
// api-key for endpoints that specifically want that header (e.g. Azure AI
// Foundry); default 'bearer' covers everything else.
const authHeaders: Record<string, string> =
authStyle === 'api-key' ? { 'api-key': apiKey } : { Authorization: `Bearer ${apiKey}` };
let discoveredModel = model;
try {
const res = await fetch(`${baseUrl}/v1/models`, {
headers: authHeaders,
signal: AbortSignal.timeout(timeoutMs),
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = await res.json();
const ids: string[] = (body.data ?? []).map((m: { id: string }) => m.id);
console.log(`GET /v1/models -> ${ids.length ? ids.join(', ') : '(empty list)'}`);
if (!discoveredModel && ids.length) discoveredModel = ids[0];
} catch (err) {
console.error(`${TAG} GET /v1/models failed: ${(err as Error).message}`);
console.error(`${TAG} Is the server actually running at ${baseUrl}? Aborting.`);
process.exit(1);
}
if (!discoveredModel) {
console.error(`${TAG} No --model given and none discovered from /v1/models. Aborting.`);
process.exit(1);
}
// Live-tested against a real llama-swap server: a POST issued right after a GET on
// the same Node process reliably HANGS indefinitely (reproduced repeatedly — GET
// alone ~30ms, POST alone ~1-2s, GET-then-immediate-POST times out completely; a
// 2s pause between them fixed it every time). This looks like Node's fetch (undici)
// reusing a pooled keep-alive connection the server doesn't handle cleanly for a
// second request right behind a first. A short pause is the simplest portable fix
// (no extra deps, no need for undici's Agent/dispatcher API).
await new Promise((resolve) => setTimeout(resolve, 2000));
try {
const res = await fetch(`${baseUrl}/v1/chat/completions`, {
method: 'POST',
headers: { 'Content-Type': 'application/json', ...authHeaders },
body: JSON.stringify({
model: discoveredModel,
messages: [{ role: 'user', content: prompt }],
}),
signal: AbortSignal.timeout(timeoutMs),
});
if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
const body = await res.json();
const reply: string = body.choices?.[0]?.message?.content ?? '';
if (!reply.trim()) throw new Error('empty reply');
console.log(`POST /v1/chat/completions -> "${reply.trim().slice(0, 200)}"`);
console.log('Server baseline: PASS\n');
} catch (err) {
console.error(`${TAG} POST /v1/chat/completions failed: ${(err as Error).message}`);
console.error(`${TAG} Server responded to /v1/models but not to a chat request. Aborting.`);
process.exit(1);
}
return discoveredModel;
}
// --- per-harness run ----------------------------------------------------------
interface ChildResult {
code: number | null;
stdout: string;
stderr: string;
timedOut: boolean;
}
function runChild(bin: string, argv: string[], env: Record<string, string>, searchDirs: string[], timeoutMs: number) {
return new Promise<ChildResult>((resolve) => {
let stdout = '';
let stderr = '';
let settled = false;
const child = spawn(bin, argv, {
env: { ...process.env, ...env, PATH: pathWithExtraDirs(searchDirs) },
stdio: ['ignore', 'pipe', 'pipe'],
});
const timer = setTimeout(() => {
if (settled) return;
settled = true;
child.kill('SIGKILL');
resolve({ code: null, stdout, stderr, timedOut: true });
}, timeoutMs);
child.stdout.on('data', (d) => (stdout += d.toString()));
child.stderr.on('data', (d) => (stderr += d.toString()));
child.on('error', (err) => {
if (settled) return;
settled = true;
clearTimeout(timer);
resolve({ code: null, stdout, stderr: `${stderr}\n${err.message}`, timedOut: false });
});
child.on('close', (code) => {
if (settled) return;
settled = true;
clearTimeout(timer);
resolve({ code, stdout, stderr, timedOut: false });
});
});
}
interface HarnessResult {
id: string;
confidence: Confidence | 'unsupported' | 'no-one-shot-recipe';
status: 'PASS' | 'FAIL' | 'UNCONFIRMED' | 'SKIP' | 'LIST';
detail: string;
}
async function runHarness(
entry: CliEntry,
opts: Opts,
model: string,
endpoint: CustomModelEndpoint
): Promise<HarnessResult> {
const id = entry.id;
const injectionCap = entry.capabilities.customModelInjection;
// Dynamic: driven by the REGISTRY's own capability, not a hardcoded id check.
// A future CLI declared unsupported is skipped automatically, same as antigravity today.
if (injectionCap.kind === 'unsupported') {
return {
id,
confidence: 'unsupported',
status: 'SKIP',
detail: 'no known custom-model mechanism (registry: unsupported)',
};
}
const oneShot = ONE_SHOT[id];
if (!oneShot) {
return {
id,
confidence: 'no-one-shot-recipe',
status: 'SKIP',
detail:
'registry supports custom-model injection for this CLI, but this script has no ONE_SHOT invocation entry yet — add one to test it',
};
}
const binary = entry.discovery.binaries[0] ?? id;
const searchDirs = entry.discovery.searchDirs;
const resolved = resolveBinary(binary, searchDirs);
if (!resolved) {
return {
id,
confidence: oneShot.confidence,
status: 'SKIP',
detail: `binary "${binary}" not found on PATH or search dirs`,
};
}
// The REAL injection logic — same function the production route calls.
const injection = buildCustomModelInjection(entry, endpoint, model);
let env: Record<string, string> = {};
let tempDir: string | null = null;
if (injection.kind === 'env') {
env = injection.envOverrides;
} else if (injection.kind === 'configDir') {
tempDir = mkdtempSync(join(tmpdir(), `codeman-local-llm-test-${id}-`));
env = applyConfigDirInjection(tempDir, injection);
}
// injection.kind === 'unsupported' already handled via injectionCap above.
const argv = oneShot.argv(opts.prompt, model);
if (opts.list) {
const detail = `${binary} ${argv.join(' ')} | env: ${Object.keys(env).join(', ')}${tempDir ? ` | configDir: ${tempDir}` : ''}`;
if (tempDir && !opts.keepTemp) rmSync(tempDir, { recursive: true, force: true });
return { id, confidence: oneShot.confidence, status: 'LIST', detail };
}
const { code, stdout, stderr, timedOut } = await runChild(binary, argv, env, searchDirs, opts.timeout);
let detailSuffix = '';
if (tempDir && !opts.keepTemp) rmSync(tempDir, { recursive: true, force: true });
else if (tempDir) detailSuffix = ` [config kept at ${tempDir}]`;
if (timedOut) {
return {
id,
confidence: oneShot.confidence,
status: 'FAIL',
detail: `timed out after ${opts.timeout}ms. stderr: ${stderr.slice(-300)}${detailSuffix}`,
};
}
const reply = stdout.trim();
const matched = /hello/i.test(reply) && /world/i.test(reply);
const softStatus: HarnessResult['status'] = oneShot.confidence === 'verified' ? 'FAIL' : 'UNCONFIRMED';
if (code !== 0) {
return {
id,
confidence: oneShot.confidence,
status: softStatus,
detail: `exit ${code}. stderr: ${stderr.trim().slice(-300) || '(empty)'}${detailSuffix}`,
};
}
if (!reply) {
return { id, confidence: oneShot.confidence, status: softStatus, detail: `exit 0 but empty stdout${detailSuffix}` };
}
if (matched) {
return { id, confidence: oneShot.confidence, status: 'PASS', detail: `${reply.slice(0, 200)}${detailSuffix}` };
}
return {
id,
confidence: oneShot.confidence,
status: 'UNCONFIRMED',
detail: `reply didn't match heuristic, judge by eye: "${reply.slice(0, 300)}"${detailSuffix}`,
};
}
// --- main ---------------------------------------------------------------------
async function main(): Promise<void> {
const configDefaults = loadConfigFile();
const opts = parseArgs(process.argv.slice(2), configDefaults);
if (opts.help) {
printUsage();
process.exit(0);
}
// Dynamic: pulled from the live registry, not a hardcoded id list. `kind === 'agent'`
// excludes 'shell' (no model/endpoint concept). Antigravity stays in this list (it IS
// an enabled agent CLI) — it's the `unsupported` capability check in runHarness that
// skips it, not an exclusion here.
const allEntries = enabledClis().filter((e) => e.kind === 'agent');
const byId = new Map(allEntries.map((e) => [e.id, e]));
const ids = opts.only ?? [...byId.keys()];
const unknownIds = ids.filter((id) => !byId.has(id));
if (unknownIds.length) {
console.error(`${TAG} unknown harness id(s): ${unknownIds.join(', ')}`);
console.error(`${TAG} known ids (from the live CLI registry): ${[...byId.keys()].join(', ')}`);
process.exit(1);
}
const entries = ids.map((id) => byId.get(id)!);
// --probe-help never touches the network — no --base-url needed for it.
if (opts.probeHelp) {
for (const entry of entries) {
const binary = entry.discovery.binaries[0] ?? entry.id;
const resolved = resolveBinary(binary, entry.discovery.searchDirs);
console.log(`\n=== ${entry.id} (${binary}) ===`);
if (!resolved) {
console.log('(not found on PATH or search dirs)');
continue;
}
printHelp(binary, entry.discovery.searchDirs);
}
process.exit(0);
}
if (!opts.baseUrl) {
console.error(`${TAG} --base-url is required (pass it, or set "baseUrl" in ${CONFIG_PATH}).`);
console.error(`${TAG} See ${CONFIG_EXAMPLE_PATH} for the config file shape.\n`);
printUsage();
process.exit(1);
}
opts.baseUrl = opts.baseUrl.replace(/\/+$/, '');
const endpoint: CustomModelEndpoint = {
id: 'standalone-test',
label: 'standalone test',
baseUrl: opts.baseUrl,
apiKey: opts.apiKey,
};
// --list is a pure dry run: never touch the network, even if --model was given.
let model: string;
if (opts.list) {
model = opts.model ?? 'local-model';
console.log(`\n=== Step 0 skipped (--list never hits the network; using placeholder "${model}") ===\n`);
} else {
model = await baselineCheck(opts.baseUrl, opts.apiKey, opts.authStyle, opts.model, opts.prompt, opts.timeout);
}
console.log(`=== Testing ${entries.length} harness(es) ===`);
const results: HarnessResult[] = [];
for (const entry of entries) {
process.stdout.write(`\n--- ${entry.id} ---\n`);
const result = await runHarness(entry, opts, model, endpoint);
results.push(result);
console.log(`${result.status}: ${result.detail}`);
}
console.log('\n=== Summary ===');
const width = Math.max(...results.map((r) => r.id.length)) + 2;
for (const r of results) {
console.log(`${r.id.padEnd(width)} [${r.confidence.padEnd(20)}] ${r.status.padEnd(11)} ${r.detail.slice(0, 100)}`);
}
const hardFail = results.some((r) => r.status === 'FAIL' && r.confidence === 'verified');
if (hardFail) {
console.error(
`\n${TAG} at least one VERIFIED harness FAILed — that's a real regression, not just an unconfirmed guess.`
);
process.exit(1);
}
process.exit(0);
}
main().catch((err) => {
console.error(`${TAG} unexpected error:`, err);
process.exit(1);
});
+1 -1
View File
@@ -333,7 +333,7 @@ const capabilitiesSchema = z
kind: z.literal('configDir'),
dirEnvVar: envName,
fileName: z.string().min(1).max(80),
template: z.enum(['codex-toml', 'pi-models-json', 'omp-models-yml']),
template: z.enum(['codex-toml', 'pi-models-json', 'omp-models-yml', 'grok-toml']),
})
.strict(),
z.object({ kind: z.literal('unsupported') }).strict(),
+48 -28
View File
@@ -749,20 +749,30 @@ const PI: CliEntry = {
// just answer "yes" to, so omitting --approve is not itself a clamp — MATERIALIZE
// approveProjectTrust:false so buildPiCommand emits --no-approve outright.
privilegedParams: [{ param: 'approveProjectTrust', clampTo: false, materializeWhenAbsent: true }],
// Web-researched, unverified. pi's models.json hot-reloads, but this feature always
// restarts the CLI on switch for consistency with the other 8 harnesses. Written to an
// isolated PI_CONFIG_DIR so the user's real ~/.pi/agent/models.json is never touched.
// CORRECTED after live-testing: `PI_CONFIG_DIR` does NOT exist anywhere in pi's own
// bundled source (grepped the installed package directly) — it does nothing for pi
// itself, despite being a real Codeman env var that OTHER things (omp) read. The
// confirmed working redirect is `HOME` itself: pi hardcodes `~/.pi/agent/models.json`
// with no dedicated override, so redirecting the CHILD PROCESS's HOME is what
// actually relocates it (verified: a model written under an isolated HOME's
// `.pi/agent/models.json` shows up in `pi --list-models` and answers a real prompt
// against a real llama-swap server; PI_CONFIG_DIR alone left it silently unable to
// see any provider). ⚠️ This is a bigger blast radius than a dedicated config-dir
// var: it also redirects pi's real sessions/auth/extensions for the DURATION of a
// custom-model session, not just its provider config — document this trade-off
// wherever this capability is surfaced.
customModelInjection: {
kind: 'configDir',
dirEnvVar: 'PI_CONFIG_DIR',
fileName: 'agent/models.json',
dirEnvVar: 'HOME',
fileName: '.pi/agent/models.json',
template: 'pi-models-json',
},
// PI_CONFIG_DIR already matches the PI_ allowedPrefix above, so it was ALREADY
// reachable via plain envOverrides before this feature existed — and pi executes
// repo-local .pi/extensions TypeScript (see the External CLI modes note in CLAUDE.md),
// so redirecting this dir is a code-execution surface, not just a config swap.
privilegedEnvKeys: ['PI_CONFIG_DIR'],
// HOME is not `PI_`-prefixed, so unlike the old (wrong) PI_CONFIG_DIR guess this was
// never reachable via the generic envOverrides allowlist at all — listed here anyway,
// matching the documented pattern for every other CLI's dir-redirect var, since a
// redirected HOME is at least as sensitive as CODEX_HOME/GROK_HOME (pi executes
// repo-local .pi/extensions TypeScript — see the External CLI modes note in CLAUDE.md).
privilegedEnvKeys: ['HOME'],
},
overlays: {
credStore: {
@@ -858,16 +868,24 @@ const GROK: CliEntry = {
// already its safe interactive ask-mode, so the multi-user clamp only needs to force an
// EXPLICITLY-SENT bypass flag back off — nothing is materialized when config is absent.
privilegedParams: [{ param: 'alwaysApprove', clampTo: false }],
// Web-researched, unverified.
// CORRECTED after live-testing against a real grok binary: the original `env` kind
// (GROK_BASE_URL/GROK_MODEL/XAI_API_KEY) produced "Not signed in" — those env vars
// are NOT grok's real custom-endpoint mechanism. The real one (verified against
// xAI's own docs) is a `[model.<name>]` block in a config.toml under GROK_HOME,
// the same configDir shape as codex/pi/omp. `api_backend = "chat_completions"` is
// explicitly supported (unlike codex, which dropped it) — grok CAN talk to a plain
// OpenAI Chat-Completions server directly.
customModelInjection: {
kind: 'env',
baseUrlVar: 'GROK_BASE_URL',
apiKeyVar: 'XAI_API_KEY',
modelVars: ['GROK_MODEL'],
kind: 'configDir',
dirEnvVar: 'GROK_HOME',
fileName: 'config.toml',
template: 'grok-toml',
},
// All three already match the GROK_/XAI_ allowedPrefixes above, so they were ALREADY
// reachable via plain envOverrides before this feature existed.
privilegedEnvKeys: ['GROK_BASE_URL', 'XAI_API_KEY', 'GROK_MODEL'],
// GROK_HOME already matches the GROK_ allowedPrefix above, so it was ALREADY
// reachable via plain envOverrides before this feature existed — same reasoning
// as CODEX_HOME: a redirected config dir can restate policy the argv-level
// `alwaysApprove` clamp above cannot see.
privilegedEnvKeys: ['GROK_HOME'],
},
overlays: {
// ~/.grok also holds sessions/, memory/, downloads/ (the ~160MB binary), completions/,
@@ -1133,18 +1151,20 @@ const OMP: CliEntry = {
// Where omp resolves its auth from. No known concrete exfiltration path today (omp
// forwards no operator-held key into a pane), but a non-granted owner redirecting where
// a shared multi-tenant deployment resolves auth is not something to allow silently.
// PI_CONFIG_DIR added for custom-model-injection.ts's omp recipe, which reuses pi's
// dir-redirect mechanism (see the customModelInjection comment below) — already
// reachable via the PI_ allowedPrefix (pi's own entry), so this closes the same
// pre-existing gap for an omp session that PI's own entry closes for a pi session.
privilegedEnvKeys: ['OMP_AUTH_BROKER_URL', 'OMP_AUTH_BROKER_TOKEN', 'PI_CONFIG_DIR'],
// Web-researched, unverified. omp's ~/.omp tree is itself relocatable via PI_CONFIG_DIR
// (see the DeepSeek/OMP note in CLAUDE.md), so this reuses that same redirect rather
// than inventing an OMP-specific dir env var.
// HOME added for custom-model-injection.ts's omp recipe (see below). Unlike pi,
// PI_CONFIG_DIR genuinely IS one of the env vars omp reads (per the DeepSeek/OMP
// note in CLAUDE.md) — but live-testing this feature found it did NOT relocate
// omp's model config the way expected, while redirecting HOME itself (like pi)
// worked immediately (verified end-to-end: a real "hello world" reply came back).
privilegedEnvKeys: ['OMP_AUTH_BROKER_URL', 'OMP_AUTH_BROKER_TOKEN', 'HOME'],
// Verified end-to-end against a real llama-swap server (live-tested, not just
// researched — a real "hello world" reply came back). Same HOME-redirect mechanism
// as pi (see its customModelInjection comment for the full reasoning) — omp hardcodes
// `~/.omp/agent/models.yml` with no dedicated config-dir override either.
customModelInjection: {
kind: 'configDir',
dirEnvVar: 'PI_CONFIG_DIR',
fileName: 'agent/models.yml',
dirEnvVar: 'HOME',
fileName: '.omp/agent/models.yml',
template: 'omp-models-yml',
},
},
+13 -4
View File
@@ -452,9 +452,18 @@ export interface CliCapabilities {
* carried in one env var (opencode's `OPENCODE_CONFIG_CONTENT`).
* `configDir`: a generated config file under an isolated, dir-redirect-env-
* pointed directory so the user's real CLI config is never touched
* (codex's `CODEX_HOME`/`config.toml`, pi/omp's `PI_CONFIG_DIR`).
* `unsupported`: no known mechanism (antigravity) — the toolbar entry
* stays disabled for this CLI.
* (codex's `CODEX_HOME`/`config.toml`, pi/omp's `PI_CONFIG_DIR`, grok's
* `GROK_HOME`/`config.toml`). `unsupported`: no known mechanism
* (antigravity) — the toolbar entry stays disabled for this CLI.
*
* ⚠️ grok was ORIGINALLY declared as `env` kind (`GROK_BASE_URL`/
* `GROK_MODEL`/`XAI_API_KEY`) — that recipe was WRONG, not just unverified:
* live-tested against a real grok binary, it produced "Not signed in",
* because those env vars are not grok's real custom-endpoint mechanism at
* all. The real one is a `[model.<name>]` block in a `config.toml` under
* `GROK_HOME` (verified against xAI's own docs), same shape as codex/pi/
* omp — this is why the confidence table in deployment_plan.md exists:
* "researched" web docs can still be plausible-sounding and wrong.
*
* Every env var name this introduces that can redirect a session's
* traffic MUST also appear in `privilegedEnvKeys` above, exactly like
@@ -469,7 +478,7 @@ export interface CliCapabilities {
kind: 'configDir';
dirEnvVar: string;
fileName: string;
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml';
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' | 'grok-toml';
}
| { kind: 'unsupported' };
}
+42
View File
@@ -0,0 +1,42 @@
/**
* @fileoverview The one IO wrapper around `custom-model-injection.ts`'s pure
* `ConfigDirInjection` output — deliberately split out so that file, the
* discovery routes, and `scripts/test-local-llm-harnesses.ts` (via tsx) can
* all share EXACTLY one "write these files, merge this env" implementation.
* Before this existed, the route and the standalone script each carried
* their own copy of this logic, which is exactly the kind of drift the CLI
* registry's "declare once, consume everywhere" design exists to prevent —
* see deployment_plan.md and the "dynamic to support cli-registry changes"
* requirement it was written against.
*/
import { mkdirSync, writeFileSync, rmSync } from 'node:fs';
import { join, dirname } from 'node:path';
import type { ConfigDirInjection } from './custom-model-injection.js';
/**
* Writes a `ConfigDirInjection`'s files under `baseDir` and returns the full
* envOverrides object a caller should merge into the session/process env
* (the dir-redirect var plus any `extraEnv` the config file references by
* name). Never touches anything outside `baseDir` — the caller is
* responsible for choosing an isolated directory (never the user's real
* `~/.codex`, `~/.pi`, etc.).
*/
export function applyConfigDirInjection(baseDir: string, injection: ConfigDirInjection): Record<string, string> {
for (const file of injection.files) {
const filePath = join(baseDir, file.relPath);
mkdirSync(dirname(filePath), { recursive: true });
writeFileSync(filePath, file.content, 'utf8');
}
return { [injection.dirEnvVar]: baseDir, ...injection.extraEnv };
}
/** Best-effort recursive removal of a previously-written configDir. Never throws. */
export function removeConfigDir(dir: string | undefined): void {
if (!dir) return;
try {
rmSync(dir, { recursive: true, force: true });
} catch {
// best-effort cleanup only
}
}
+46 -5
View File
@@ -14,7 +14,7 @@
* llama-swap server (a real "hello world" reply came back). `codex`'s
* config.toml STRUCTURE is now verified (an earlier `[model].default` table
* shape was rejected by a real codex binary with "invalid type: map,
* expected a string" — caught by `scripts/test-local-llm-harnesses.mjs`),
* expected a string" — caught by `scripts/test-local-llm-harnesses.ts`),
* but `wire_api = "responses"` is the only value codex still accepts
* (support for `"chat"` was dropped in Feb 2026), and a plain OpenAI
* Chat-Completions server (llama.cpp, llama-swap, most local setups) does
@@ -134,8 +134,12 @@ function renderConfigContent(
const CODEX_API_KEY_ENV_VAR = 'CODEMAN_CUSTOM_MODEL_API_KEY';
/** The `[model.<name>]` block name grok's config.toml uses for the injected model — also
* what `-m <name>` in the standalone script's ONE_SHOT argv must reference to select it. */
export const GROK_CUSTOM_MODEL_NAME = 'codeman-custom';
function renderConfigFile(
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml',
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' | 'grok-toml',
endpoint: CustomModelEndpoint,
modelId: string,
apiKey: string
@@ -145,7 +149,7 @@ function renderConfigFile(
case 'codex-toml': {
// Verified against real codex (>= Feb 2026): `model` is a top-level STRING, never
// a `[model].default` table — codex rejects that with "invalid type: map, expected
// a string" (caught by scripts/test-local-llm-harnesses.mjs against a real llama-swap
// a string" (caught by scripts/test-local-llm-harnesses.ts against a real llama-swap
// server). The API key is NEVER a literal TOML field: codex's schema only supports
// `env_key`, the NAME of an env var it reads the credential from at runtime, so the
// actual value must ride along as an extra env var, never embedded in the file.
@@ -169,11 +173,23 @@ function renderConfigFile(
return { content, extraEnv: { [CODEX_API_KEY_ENV_VAR]: apiKey } };
}
case 'pi-models-json':
// Verified against pi's OWN bundled docs (models.md): `models` is an ARRAY of
// `{id: "..."}` objects, NOT an object keyed by model id — the earlier shape here
// silently loaded zero models ("No models available"), confirmed live. `authHeader:
// true` is required too: pi does not automatically send `Authorization: Bearer
// <apiKey>` just because `apiKey` is set (per the same doc) — without it, a real
// (non-llama.cpp) endpoint that actually checks the key would reject every request.
return {
content: JSON.stringify(
{
providers: {
custom: { baseUrl, apiKey, api: 'openai-completions', models: { [modelId]: {} } },
custom: {
baseUrl,
apiKey,
api: 'openai-completions',
authHeader: true,
models: [{ id: modelId }],
},
},
},
null,
@@ -181,8 +197,33 @@ function renderConfigFile(
),
};
case 'omp-models-yml':
// Mirrors the pi-models-json fix above (omp shares pi's config lineage per
// CLAUDE.md — it reads several of pi's own env vars): a flat list of bare model
// name strings under `models` is UNCONFIRMED against real omp docs (none are
// bundled with the binary) — this now matches pi's `{id: "..."}` object-list
// shape and adds `authHeader: true` on the same reasoning, but has not itself
// been live-tested the way pi's fix was. Verify before raising its confidence.
return {
content: `providers:\n custom:\n baseUrl: ${quoted(baseUrl)}\n apiKey: ${quoted(apiKey)}\n models:\n - ${quoted(modelId)}\n`,
content: `providers:\n custom:\n baseUrl: ${quoted(baseUrl)}\n apiKey: ${quoted(apiKey)}\n api: openai-completions\n authHeader: true\n models:\n - id: ${quoted(modelId)}\n`,
};
case 'grok-toml': {
// Verified against xAI's own docs (docs.x.ai/build/settings/reference): a
// `[model.<name>]` block, NOT plain env vars — an earlier `env`-kind recipe for
// grok was wrong, not just unverified (see the customModelInjection doc comment
// in cli-registry/types.ts). `api_backend = "chat_completions"` is explicitly
// supported (unlike codex, which dropped it after Feb 2026), so this one CAN
// talk to a plain OpenAI-compatible server directly. `env_key` reuses grok's own
// documented fallback var name (XAI_API_KEY) rather than inventing a new one.
const content = [
`[model.${GROK_CUSTOM_MODEL_NAME}]`,
`model = ${quoted(modelId)}`,
`base_url = ${quoted(baseUrl)}`,
`name = "Custom Endpoint"`,
`env_key = "XAI_API_KEY"`,
`api_backend = "chat_completions"`,
'',
].join('\n');
return { content, extraEnv: { XAI_API_KEY: apiKey } };
}
}
}
+6 -13
View File
@@ -8,7 +8,7 @@ import { FastifyInstance, type FastifyReply } from 'fastify';
import { z } from 'zod';
import { join, dirname, extname, basename } from 'node:path';
import { homedir } from 'node:os';
import { existsSync, statSync, mkdirSync, writeFileSync, rmSync } from 'node:fs';
import { existsSync, statSync, mkdirSync, writeFileSync } from 'node:fs';
import { execFile } from 'node:child_process';
import fs from 'node:fs/promises';
import { randomBytes } from 'node:crypto';
@@ -55,6 +55,7 @@ import {
} from '../schemas.js';
import { readCustomModelHosts } from '../../custom-model-hosts.js';
import { buildCustomModelInjection } from '../../custom-model-injection.js';
import { applyConfigDirInjection, removeConfigDir } from '../../custom-model-injection-apply.js';
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
import { TabLayoutValidationError } from '../../tab-layout.js';
import {
@@ -1184,7 +1185,7 @@ export function registerSessionRoutes(
if ('clear' in body) {
const previousConfigDir = session.setCustomModel(undefined);
if (previousConfigDir) rmSync(previousConfigDir, { recursive: true, force: true });
removeConfigDir(previousConfigDir);
const restarted = await session.restartCli();
persistAndBroadcastSession(ctx, session);
return { customModel: session.customModel, restarted };
@@ -1216,16 +1217,8 @@ export function registerSessionRoutes(
} else if (injection.kind === 'configDir') {
// Isolated per-session dir — never the user's real CLI config path.
configDir = join(dataPath('custom-model-configs'), session.id);
for (const file of injection.files) {
const filePath = join(configDir, file.relPath);
mkdirSync(dirname(filePath), { recursive: true });
writeFileSync(filePath, file.content, 'utf8');
}
// extraEnv: vars the written config file REFERENCES by name (codex's `env_key`
// convention) rather than embedding a literal value — must ride alongside
// dirEnvVar or the config points at a credential that was never actually set.
envOverrides = { [injection.dirEnvVar]: configDir, ...injection.extraEnv };
envKeys = [injection.dirEnvVar, ...Object.keys(injection.extraEnv ?? {})];
envOverrides = applyConfigDirInjection(configDir, injection);
envKeys = Object.keys(envOverrides);
} else {
// 'unsupported' is already handled above; this keeps the switch exhaustive.
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
@@ -1238,7 +1231,7 @@ export function registerSessionRoutes(
// Clean up the OLD config dir on disk, unless the new one happens to reuse the same
// path (same session, configDir kind again) — never delete the dir we just wrote.
if (previousConfigDir && previousConfigDir !== configDir) {
rmSync(previousConfigDir, { recursive: true, force: true });
removeConfigDir(previousConfigDir);
}
const restarted = await session.restartCli();
+17 -13
View File
@@ -12,7 +12,7 @@
* proves "if the CLI honors its documented env/config contract, it will hit
* the right endpoint with the right model." It does NOT prove the real CLI
* binary actually reads that env var / config file the way its docs say —
* that's still the job of `scripts/test-local-llm-harnesses.mjs` against a
* that's still the job of `scripts/test-local-llm-harnesses.ts` against a
* real endpoint and real binaries. This suite catches regressions in
* Codeman's own injection logic; it cannot catch a CLI changing its env-var
* name in a future release.
@@ -142,16 +142,17 @@ describe('custom-model-injection contract (mock server)', () => {
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
});
// gemini/grok/deepseek's `env` kind passes the base URL through UNCHANGED (unlike
// opencode/codex/pi/omp, which build a structured config and explicitly append /v1) —
// matching Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the SDK
// appends the path itself. Whether each of these THREE CLIs' own OpenAI-compatible
// gemini/deepseek's `env` kind passes the base URL through UNCHANGED (unlike
// opencode/codex/pi/omp/grok, which build a structured config and explicitly append
// /v1) — matching Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the
// SDK appends the path itself. Whether each of these TWO CLIs' own OpenAI-compatible
// client expects the var to already include /v1 (the common OpenAI-SDK convention) or
// appends it itself is genuinely CLI-specific and UNVERIFIED (see the confidence table
// in deployment_plan.md) — these tests model the common OpenAI-SDK convention (base_url
// ends in /v1) since that's the more likely behavior for an OpenAI-compatible client,
// but that assumption should be corrected here the moment it's checked against a real
// binary.
// binary. (grok WAS in this group too, until live-testing showed the whole `env` recipe
// was wrong for it — see its own test below.)
it('gemini: GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY reach the mock', async () => {
const injection = buildCustomModelInjection(entryOrThrow('gemini'), endpointFor(mock), 'qwen3');
@@ -168,15 +169,18 @@ describe('custom-model-injection contract (mock server)', () => {
expect((mock.requests[0].body as { model: string }).model).toBe('qwen3');
});
it('grok: GROK_BASE_URL/XAI_API_KEY reach the mock', async () => {
it('grok: config.toml [model.<name>] block base_url/env_key + extraEnv reach the mock over /v1/chat/completions', async () => {
const injection = buildCustomModelInjection(entryOrThrow('grok'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'env') throw new Error('unreachable');
if (injection.kind !== 'configDir') throw new Error('unreachable');
const toml = injection.files[0].content;
const baseUrl = /base_url = "([^"]+)"/.exec(toml)?.[1];
const model = /^model = "([^"]+)"/m.exec(toml)?.[1];
expect(baseUrl).toBe(`${mock.baseUrl}/v1`);
expect(model).toBe('qwen3');
expect(toml).toContain('api_backend = "chat_completions"');
expect(injection.extraEnv).toEqual({ XAI_API_KEY: 'contract-test-key' });
await callOpenAiCompat(
`${injection.envOverrides.GROK_BASE_URL}/v1`,
injection.envOverrides.XAI_API_KEY,
injection.envOverrides.GROK_MODEL
);
await callOpenAiCompat(baseUrl!, injection.extraEnv!.XAI_API_KEY, model!);
expect(mock.requests[0].path).toBe('/v1/chat/completions');
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
+22 -13
View File
@@ -91,21 +91,25 @@ describe('buildCustomModelInjection', () => {
expect(result.files[0].content).toContain('model = "weird\\"model"');
});
it('pi: configDir writes models.json under agent/', () => {
it('pi: configDir writes .pi/agent/models.json, redirected via HOME (verified live — PI_CONFIG_DIR does nothing for pi)', () => {
const result = buildCustomModelInjection(entryOrThrow('pi'), endpoint, 'qwen3');
if (result.kind !== 'configDir') throw new Error('unreachable');
expect(result.dirEnvVar).toBe('PI_CONFIG_DIR');
expect(result.files[0].relPath).toBe('agent/models.json');
expect(result.dirEnvVar).toBe('HOME');
expect(result.files[0].relPath).toBe('.pi/agent/models.json');
const parsed = JSON.parse(result.files[0].content);
expect(parsed.providers.custom.baseUrl).toBe('http://192.168.1.50:8080/v1');
expect(parsed.providers.custom.authHeader).toBe(true);
expect(parsed.providers.custom.models).toEqual([{ id: 'qwen3' }]); // array, NOT keyed by id
});
it('omp: configDir writes models.yml under agent/, redirected via PI_CONFIG_DIR', () => {
it('omp: configDir writes .omp/agent/models.yml, redirected via HOME (verified live end-to-end)', () => {
const result = buildCustomModelInjection(entryOrThrow('omp'), endpoint, 'qwen3');
if (result.kind !== 'configDir') throw new Error('unreachable');
expect(result.dirEnvVar).toBe('PI_CONFIG_DIR');
expect(result.files[0].relPath).toBe('agent/models.yml');
expect(result.dirEnvVar).toBe('HOME');
expect(result.files[0].relPath).toBe('.omp/agent/models.yml');
expect(result.files[0].content).toContain('baseUrl: "http://192.168.1.50:8080/v1"');
expect(result.files[0].content).toContain('authHeader: true');
expect(result.files[0].content).toContain('- id: "qwen3"');
});
it('gemini: env kind sets GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY/GEMINI_MODEL', () => {
@@ -118,14 +122,19 @@ describe('buildCustomModelInjection', () => {
});
});
it('grok: env kind sets GROK_BASE_URL/XAI_API_KEY/GROK_MODEL', () => {
it('grok: configDir writes a config.toml [model.<name>] block, key rides as extraEnv (XAI_API_KEY)', () => {
const result = buildCustomModelInjection(entryOrThrow('grok'), endpoint, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides).toEqual({
GROK_BASE_URL: 'http://192.168.1.50:8080',
XAI_API_KEY: 'my-key',
GROK_MODEL: 'qwen3',
});
expect(result.kind).toBe('configDir');
if (result.kind !== 'configDir') throw new Error('unreachable');
expect(result.dirEnvVar).toBe('GROK_HOME');
expect(result.files).toHaveLength(1);
expect(result.files[0].relPath).toBe('config.toml');
expect(result.files[0].content).toContain('model = "qwen3"');
expect(result.files[0].content).toContain('base_url = "http://192.168.1.50:8080/v1"');
expect(result.files[0].content).toContain('api_backend = "chat_completions"');
expect(result.files[0].content).toContain('env_key = "XAI_API_KEY"');
expect(result.files[0].content).not.toContain('api_key ='); // never a literal TOML field
expect(result.extraEnv).toEqual({ XAI_API_KEY: 'my-key' });
});
it('deepseek: env kind sets base URL/key only, no model var', () => {