mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
41416566aa
commit
61779745aa
+49
-34
@@ -69,11 +69,11 @@ config blob for opencode, a TOML file for Codex, etc. Devvyn gave the
|
||||
starting recipes for those three; the rest (Gemini, Pi, Grok, DeepSeek, OMP,
|
||||
Antigravity) were researched for this plan and are flagged by confidence
|
||||
below. A real end-to-end pass against Devvyn's own llama-swap server
|
||||
(`scripts/test-local-llm-harnesses.mjs`, inside a `codeman/agent:llm-test`
|
||||
(`scripts/test-local-llm-harnesses.ts`, inside a `codeman/agent:llm-test`
|
||||
Docker image with all 9 CLIs installed) then confirmed **claude and
|
||||
opencode work end-to-end**, corrected a real Codex config.toml schema bug
|
||||
the given recipe had (see the Codex row below), and surfaced that Codex's
|
||||
*protocol* — not just its config shape — does not work against a plain
|
||||
_protocol_ — not just its config shape — does not work against a plain
|
||||
OpenAI-Chat-Completions server like llama.cpp/llama-swap at all. Confidence
|
||||
below reflects what was actually observed, not just what was planned.
|
||||
|
||||
@@ -104,17 +104,17 @@ declared capability, never an `if (mode === 'claude')` branch.
|
||||
|
||||
## Per-CLI injection recipes (confidence-ranked)
|
||||
|
||||
| CLI | Mechanism | Confidence |
|
||||
|---|---|---|
|
||||
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
|
||||
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
|
||||
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
|
||||
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` (or `GOOGLE_VERTEX_BASE_URL`) + `GEMINI_API_KEY`; CLI needs a restart to pick them up (matches our restart-on-switch design). Model selection via `--model`/`GEMINI_MODEL`-style override — verify exact var name against the installed `gemini-cli` version before shipping | Web-researched, unverified |
|
||||
| `pi` | Config file `~/.pi/agent/models.json` (hot-reloadable) with a custom provider block: `baseUrl`, `apiKey`, `api:"openai-completions"`. Redirect via `PI_CONFIG_DIR` (already allowlisted per CLAUDE.md) pointed at an isolated dir containing just this file, rather than overwriting the user's real one | Web-researched, unverified |
|
||||
| `grok` | Env vars `GROK_BASE_URL`, `XAI_API_KEY` (dummy ok for local; a real key for most cloud endpoints), `GROK_MODEL`. All three already fit inside the existing `XAI_*`/CLI-specific allowlist shape | Web-researched, unverified |
|
||||
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts:879-913`, already in `privilegedEnvKeys`). Model selection is murkier — CLAUDE.md notes dsh model is "a profile composition entry," not a flag/env var, so redirecting the endpoint is solid but forcing a specific model name may not fully work; document as best-effort and verify against a real profile | Web-researched, unverified, partial |
|
||||
| `omp` | Config file `~/.omp/agent/models.yml`-equivalent with a custom provider `baseUrl`. CLAUDE.md notes omp's config tree is itself relocatable via `PI_CONFIG_DIR` — reuse the same isolated-dir-redirect approach as `pi` | Web-researched, unverified |
|
||||
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
|
||||
| CLI | Mechanism | Confidence |
|
||||
| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
|
||||
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
|
||||
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
|
||||
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
|
||||
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
|
||||
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
|
||||
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`, already in `privilegedEnvKeys`). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working |
|
||||
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
|
||||
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
|
||||
|
||||
Everything web-researched-but-unverified gets implemented but must be
|
||||
smoke-tested against real installs of those CLIs before being called done —
|
||||
@@ -139,8 +139,13 @@ union on each `CliEntry.capabilities`:
|
||||
type CustomModelInjection =
|
||||
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[] }
|
||||
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json' }
|
||||
| { kind: 'configDir'; dirEnvVar: string; fileName: string; template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' }
|
||||
| { kind: 'unsupported' }
|
||||
| {
|
||||
kind: 'configDir';
|
||||
dirEnvVar: string;
|
||||
fileName: string;
|
||||
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml';
|
||||
}
|
||||
| { kind: 'unsupported' };
|
||||
```
|
||||
|
||||
Declared per stock.ts entry per the table above. A pure function in a new
|
||||
@@ -241,7 +246,7 @@ dir-redirects, plus the already-privileged `DEEPSEEK_BASE_URL` — must be
|
||||
added to each CLI's `capabilities.privilegedEnvKeys` so
|
||||
`clampEnvOverridesForOwner()` strips them for a non-granted multi-user
|
||||
owner, exactly the precedent already documented for `DEEPSEEK_BASE_URL`/
|
||||
`OMP_AUTH_BROKER_URL`. This matters *more*, not less, now that endpoints can
|
||||
`OMP_AUTH_BROKER_URL`. This matters _more_, not less, now that endpoints can
|
||||
be cloud URLs: redirecting a non-granted user's session to an attacker's
|
||||
cloud endpoint is a credential-exfiltration path, not just a mischief
|
||||
redirect to a LAN box. Endpoint CRUD itself stays admin-only in multi-user
|
||||
@@ -259,14 +264,14 @@ mode, same as remote/docker hosts.
|
||||
- `src/web/public/index.html`, `settings-ui.js`, `session-ui.js`, `styles.css` — settings group, toolbar button/menu, badge, accent CSS
|
||||
- `src/web/sse-events.ts` + `constants.js` — if a dedicated SSE event is warranted for the badge (or just ride existing session-update broadcasts)
|
||||
- `test/fixtures/mock-openai-server.ts` (new) + `test/custom-model-injection-contract.test.ts` (new) — see Mock-server validation below
|
||||
- `scripts/test-local-llm-harnesses.mjs` (already added, this branch) — the standalone real-CLI-and-real-endpoint smoke test; despite the filename (kept for continuity with when it was written) it already supports any `--base-url`, local or cloud
|
||||
- `scripts/test-local-llm-harnesses.ts` (already added, this branch; run via `npx tsx`) — the standalone real-CLI-and-real-endpoint smoke test, supporting any `--base-url` (local or cloud). Dynamic: derives its harness list and every env var/config it injects from the live CLI registry + `buildCustomModelInjection()` rather than a second hand-maintained copy — only the one-shot invocation flags (`ONE_SHOT` table) are CLI-specific info the registry doesn't model and stay hand-maintained
|
||||
- `docs/custom-model-endpoints.md` (new) + a CLAUDE.md pointer bullet under External CLI modes / envOverrides
|
||||
|
||||
## Mock-server validation strategy (CI-runnable, no real CLI binaries needed)
|
||||
|
||||
Spawning nine real CLI binaries in CI isn't realistic, and neither Devvyn's
|
||||
llama.cpp box nor a real cloud subscription can be a CI dependency. So the
|
||||
injection *logic* gets a tier of automated coverage that sits between the
|
||||
injection _logic_ gets a tier of automated coverage that sits between the
|
||||
pure unit tests and the live manual checks in Verification:
|
||||
|
||||
1. **`test/fixtures/mock-openai-server.ts`** — a small in-process HTTP
|
||||
@@ -306,16 +311,21 @@ pure unit tests and the live manual checks in Verification:
|
||||
per-session dir rather than the user's real config path.
|
||||
|
||||
3. **Explicit, stated limitation** (goes in the test file's `@fileoverview`
|
||||
and in this doc, not left implicit): this proves *"if the CLI honors its
|
||||
and in this doc, not left implicit): this proves _"if the CLI honors its
|
||||
documented env/config contract, it will hit the right endpoint with the
|
||||
right model."* It does **not** prove the real CLI binary actually reads
|
||||
right model."_ It does **not** prove the real CLI binary actually reads
|
||||
that env var / config file the way its docs say — that's still the job
|
||||
of the live manual checks in Verification step 4-5 below, and is exactly
|
||||
why the confidence table above stays "unverified" for six of the nine
|
||||
CLIs until someone runs those binaries for real. The mock-server suite
|
||||
catches regressions in Codeman's own logic; it cannot catch a CLI
|
||||
changing its env-var name in a future release, or a real cloud endpoint
|
||||
behaving differently from the mock.
|
||||
why the confidence table above did not stop at "researched" — every CLI
|
||||
except antigravity (no mechanism at all) has since been run against a
|
||||
real llama-swap server via `scripts/test-local-llm-harnesses.ts`:
|
||||
claude/opencode/pi/grok/omp are confirmed PASS end-to-end, codex is
|
||||
confirmed FAIL for a real documented protocol reason (Responses-API-only
|
||||
since Feb 2026), and gemini/deepseek are confirmed reaching the server
|
||||
but failing for reasons not yet root-caused (see their table rows). The
|
||||
mock-server suite catches regressions in Codeman's own logic; it cannot
|
||||
catch a CLI changing its env-var name in a future release, or a real
|
||||
cloud endpoint behaving differently from a local llama.cpp box.
|
||||
|
||||
## Verification
|
||||
|
||||
@@ -326,15 +336,20 @@ pure unit tests and the live manual checks in Verification:
|
||||
3. Route tests (`app.inject`) for the new CRUD + discover-models endpoint
|
||||
(mock `fetch` for `/v1/models`), and for the multi-user clamp on the new
|
||||
privileged keys (mirror `test/routes/external-cli-bypass-clamp.test.ts`).
|
||||
4. **Standalone real-binary smoke test**: `scripts/test-local-llm-harnesses.mjs`
|
||||
(already written on this branch) exercises every harness against a real
|
||||
`--base-url` — local or cloud — outside of Codeman's UI entirely. Run it
|
||||
against Devvyn's llama.cpp server first (`claude`/`opencode`/`codex`
|
||||
should PASS, since those recipes are verified; the rest report
|
||||
UNCONFIRMED/SKIP until their guessed flags are corrected via
|
||||
`--probe-help`), then again against a real cloud endpoint (e.g. an Azure
|
||||
AI Foundry deployment) once one is available, to prove the `authStyle`/
|
||||
deployment-name handling holds up outside llama.cpp.
|
||||
4. **Standalone real-binary smoke test**: `scripts/test-local-llm-harnesses.ts`
|
||||
exercises every harness the CLI registry declares `customModelInjection`
|
||||
support for against a real `--base-url` — local or cloud — outside of
|
||||
Codeman's UI entirely, and is DYNAMIC (reads `enabledClis()` + calls the
|
||||
real `buildCustomModelInjection()`, so a future registry change is picked
|
||||
up automatically with zero edits to the script). Already run to
|
||||
completion against Devvyn's llama-swap server (`http://10.10.11.241:8080`,
|
||||
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
|
||||
claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected**
|
||||
(Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED**
|
||||
(reach the server, fail for undiagnosed reasons — see their table rows),
|
||||
antigravity **SKIP** (no mechanism). Re-run this against a real cloud
|
||||
endpoint (e.g. an Azure AI Foundry deployment) once one is available, to
|
||||
prove the `authStyle`/deployment-name handling holds up outside llama.cpp.
|
||||
5. Once the full feature (not just the standalone script) is built: add an
|
||||
endpoint via the real UI, hit discover-models, confirm the returned model
|
||||
list, pick Claude + the model on a real session, confirm via
|
||||
|
||||
Reference in New Issue
Block a user