From 61779745aa78e8dba055b24569157bfe472bfcce Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Sat, 5 Sep 2026 22:14:22 +0800 Subject: [PATCH] test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI registry (enabledClis()) and call the real production buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a second hand-maintained copy of every CLI's env/config shape. A future registry change (new CLI, edited env var, fixed config template) is now picked up automatically with zero edits to this script; only the one-shot invocation flags (info the registry genuinely doesn't model) stay in a small hand-maintained ONE_SHOT table, and a registry CLI with no entry there reports UNKNOWN rather than being silently skipped. Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/ removeConfigDir) so the production route and this script share one implementation instead of two. Full end-to-end run against a real llama-swap server, inside a codeman/agent:llm-test Docker image with all 9 CLI binaries: - claude, opencode, pi, grok, omp: PASS, real "hello world" replies - codex: confirmed FAIL for a real protocol reason, not a bug — it only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement - gemini: confirmed FAIL, unresolved after real investigation — an undocumented GATEWAY AuthType gemini-cli selects once GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override tried - deepseek: reaches the server (env vars are read) but gets a consistent HTTP_404; root cause not identified, documented as best-effort/unknown - antigravity: SKIP, no known mechanism (unchanged) Two real bugs found and fixed along the way (grok, pi/omp registry entries in stock.ts): grok's original recipe (env vars) was flat-out wrong, not just unverified — the real mechanism is a config.toml [model.] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR does nothing for either (grepped pi's entire bundled source — the string appears nowhere); the real redirect is the child process's own HOME, and both need `models` as an array of {id} objects, not an object keyed by id (silently loaded zero models otherwise). deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md updated with the final confidence table reflecting all of the above. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3 --- CLAUDE.md | 2 +- PR.md | 111 ++- deployment_plan.md | 83 ++- docs/custom-model-endpoints.md | 33 +- scripts/test-local-llm-harnesses.mjs | 631 ----------------- scripts/test-local-llm-harnesses.ts | 699 +++++++++++++++++++ src/config/cli-registry/schema.ts | 2 +- src/config/cli-registry/stock.ts | 76 +- src/config/cli-registry/types.ts | 17 +- src/custom-model-injection-apply.ts | 42 ++ src/custom-model-injection.ts | 51 +- src/web/routes/session-routes.ts | 19 +- test/custom-model-injection-contract.test.ts | 30 +- test/custom-model-injection.test.ts | 35 +- 14 files changed, 1061 insertions(+), 770 deletions(-) delete mode 100644 scripts/test-local-llm-harnesses.mjs create mode 100644 scripts/test-local-llm-harnesses.ts create mode 100644 src/custom-model-injection-apply.ts diff --git a/CLAUDE.md b/CLAUDE.md index fd87d132..c35302c2 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -225,7 +225,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. The session version worked and was still wrong in use: it put a terminal tab on screen next to the web tab the user actually asked for, every single time, and nothing about a long-lived HTTP server needs to be a tab. ⚠️ What a session gave for free now has to be paid for explicitly, and every piece is load-bearing: **exactly one** server (a second click REUSES it rather than racing it for a port, which two sessions structurally could not do), **restarted when the browser authority changes** (`--trusted-host` fences dsh's `/api` against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two, so whoever asks last wins: the asker is by definition the origin about to load the page), **killed on shutdown** (`stopDeepSeekWeb()` in the server teardown, because the child is detached so its whole plugin tree can be signalled at once, which also means it would OUTLIVE Codeman and hold its port against the next start), and **failures returned to the caller**, since with no tab there is nowhere for a stack trace to land. ⚠️ The port search starts at dsh's own default 3080 and walks 40, never fixed: that default is precisely the port most likely to be taken already by the user's own `dsh web`, and hardcoding it killed this feature with EADDRINUSE once. Free-port detection BINDS rather than connects (a connect probe cannot tell "free" from "listening but not answering yet"), so it is racy by nature and the caller still waits for the server to really answer before reporting success. ⚠️ Both `POST` and `DELETE` sit at the **same privilege bar as the profile installer** (`canUsernameRunPrivilegedCommands`) even though the action reads as "open a page": booting a dsh profile executes the plugin code in it, and the server is a single shared instance, so stopping it in multi-user mode takes it out from under other users' tabs. -**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `deployment_plan.md`): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` (default `'both'`) sends BOTH `Authorization: Bearer` and `api-key` headers on discovery since cloud gateways (Azure) and local servers (llama.cpp) disagree on the convention and there is no way to know in advance which one a given endpoint wants. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/grok/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi` config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `PI_CONFIG_DIR`, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. Confidence is per-CLI: claude/opencode/codex are hand-verified against a real llama.cpp/llama-swap server; gemini/pi/grok/deepseek/omp have their ONE-SHOT INVOCATION flags confirmed against real installed binaries' `--help` output, but their custom-endpoint env/config conventions remain unverified — see the confidence table in `deployment_plan.md`. The standalone `scripts/test-local-llm-harnesses.mjs` (reads a gitignored `scripts/local-llm-test.config.json`, template `.example.json` tracked) smoke-tests real CLI binaries against a real endpoint outside the web UI entirely, independent of tmux/sessions. +**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `deployment_plan.md`): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct but codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement — a confirmed protocol gap, not a bug; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek reaches the server but gets a consistent `HTTP_404` (root cause not identified); antigravity has no known mechanism at all. See the confidence table in `deployment_plan.md` for the full detail on each. **Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w-` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. ⚠️ **Closing has the mirror-image race and one owner**: `closeSession()` reads `wasActive` BEFORE its `await` and announces the delete via `_closingSessions`, while `_onSessionDeleted` skips the active-session handoff for an id in that set. Both used to read `activeSessionId` after the fact, so the `session_deleted` broadcast for your own delete could null it first and closing the tab you were on landed on the welcome screen instead of the next session, on the same build, depending on timing. The fallback also picks the first order entry that is still in `sessions` (a dead id can linger in `sessionOrder`, same reason Alt+N indexes a live-filtered list). A delete from ANOTHER client still shows the welcome screen, which is the honest answer when what you were looking at was taken away. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization) diff --git a/PR.md b/PR.md index 03e5ebd6..ebfe3fc9 100644 --- a/PR.md +++ b/PR.md @@ -50,7 +50,7 @@ just "add the gateway once, everything behind it shows up." ## Why The maintainer pays for a Claude Code subscription but also runs a capable -local model. Every harness Codeman drives already *has* its own mechanism +local model. Every harness Codeman drives already _has_ its own mechanism for pointing at a custom endpoint (env vars for Claude, a JSON config blob for opencode, a TOML file for Codex, etc.) — Codeman just never exposed a UI for it. Full motivation, the per-CLI recipe table, and the on-prem @@ -77,7 +77,7 @@ hardware use cases are written up in **[`deployment_plan.md`](deployment_plan.md web tabs use. - **`src/web/schemas.ts`** — `customModelEndpointsEnabled` (synced, default OFF) + the endpoint payload schema. -- **`scripts/test-local-llm-harnesses.mjs`** — standalone smoke-test script +- **`scripts/test-local-llm-harnesses.ts`** — standalone smoke-test script that spawns each real CLI binary one-shot against a real endpoint and checks it can answer "hello world," independent of the web UI. Reads defaults from a gitignored `scripts/local-llm-test.config.json` (see the @@ -130,11 +130,13 @@ format + tests green), ⬜ = not started. until chunk 6 lands) + a CLAUDE.md pointer bullet Also done outside the chunk list: the standalone -`scripts/test-local-llm-harnesses.mjs` smoke-test script + its gitignored -config file, the on-prem-hardware use-case writeup in `deployment_plan.md` -(DGX Spark, Strix Halo, Qwen5090), and a `codeman/agent:llm-test` Docker -image (all 9 CLI binaries, built from `docker/agent.Dockerfile`) for the -real end-to-end test against a live llama-swap server. +`scripts/test-local-llm-harnesses.ts` smoke-test script (now dynamic — +reads the live CLI registry rather than a hand-maintained harness list) + +its gitignored config file, the on-prem-hardware use-case writeup in +`deployment_plan.md` (DGX Spark, Strix Halo, Qwen5090), a +`codeman/agent:llm-test` Docker image (all 9 CLI binaries, built from +`docker/agent.Dockerfile`), and a **completed real end-to-end run of all 9 +harnesses** against a live llama-swap server — see Testing below. ## Testing performed so far @@ -145,14 +147,67 @@ test/custom-model-injection-contract.test.ts test/routes/custom-model-routes.tes test/routes/session-custom-model.test.ts test/routes/external-cli-bypass-clamp.test.ts` — 245+ tests passing, including the existing multi-user clamp suite (no regressions from the `privilegedEnvKeys` additions) -- `node --check scripts/test-local-llm-harnesses.mjs` + manual `--help` run +- Refactored `scripts/test-local-llm-harnesses.ts` (now `npx tsx`-run, was + plain `.mjs`) to import `enabledClis()` and `buildCustomModelInjection()` + directly from source instead of keeping a second hand-maintained copy of + every CLI's env/config shape — a registry change now needs zero edits to + the test script. Extracted the config-dir-write logic shared with the + production route into `custom-model-injection-apply.ts` so both places + call exactly one implementation. - **Real end-to-end run against the maintainer's live llama-swap server** (`http://10.10.11.241:8080`), inside `codeman/agent:llm-test` (all 9 CLI binaries, built via `docker/agent.Dockerfile`), against the smallest available model (`qwen3.5-0.8b-ud-q8_k_xl`, 1.1GB — picked by parsing the - server's own reported model sizes). Real findings, not simulated: + server's own reported model sizes). **Full 9-harness result: claude, + opencode, pi, grok, omp all PASS with a genuine "hello world" reply + round-tripped through the real endpoint; codex FAILs for a confirmed + protocol reason (not a bug — see below); gemini and deepseek reach the + server but fail for reasons not yet root-caused; antigravity SKIPs (no + known mechanism); all correctly classified by the now-dynamic + `scripts/test-local-llm-harnesses.ts`, which reads the live CLI registry + rather than a hand-maintained harness list.** Real findings, not + simulated: - **opencode: PASS.** Genuinely round-tripped a "hello world" reply through the real endpoint. + - **pi: PASS, after two real bugs found and fixed.** `PI_CONFIG_DIR` does + nothing for pi at all (grepped pi's entire bundled JS source — the + string appears nowhere); the real redirect is the child process's own + `HOME`, since pi hardcodes `~/.pi/agent/models.json` with no dedicated + override. Separately, pi's `models` field must be an **array** of + `{id}` objects, not an object keyed by id (confirmed against pi's own + bundled `docs/models.md`) — the object shape silently loaded zero + models. Also needs an explicit `--model custom/` on invocation. + - **grok: PASS, after the original recipe turned out to be flat-out + wrong**, not just unverified — the env-var recipe in this table's first + draft (`GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) produced "Not signed + in" against a real binary. Researched xAI's actual docs and corrected + to a `config.toml` with a `[model.]` block redirected via + `GROK_HOME`, with the key riding as an `env_key`-named env var — then + confirmed working end-to-end. + - **omp: PASS**, after the same two fixes as pi (array-shaped `models`, + `HOME`-redirect instead of `PI_CONFIG_DIR`) plus `--model custom/`. + Unverified against omp's own official docs (none are bundled in the + install), but empirically confirmed working live. + - **gemini: confirmed broken, unresolved after real investigation.** + Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an + undocumented `AuthType.GATEWAY` path with validation requirements a + live run never satisfies (`Invalid auth method selected`, regardless of + key format). Tried and ruled out: a Google-format dummy key, + `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` + override, and a hand-written `settings.json`. `--skip-trust` is a real, + separate fix for a different symptom (an untrusted-folder check + silently overriding `--approval-mode yolo`) and is kept, but does not + touch this auth failure. Left as an open, documented gap rather than + claimed as working. + - **deepseek: confirmed reaching the server, still failing, unresolved.** + A real run returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` + consistently — the env vars are read (the request reaches the network + rather than failing locally), but the root cause was not identified in + the time available. By analogy with codex's Responses-API gap, `dsh` + may expect DeepSeek's own API response shape rather than a generic + OpenAI-compatible one, but this was not confirmed by reading dsh's own + bundled source the way the pi/grok questions were resolved. Documented + as best-effort/unknown, matching its pre-existing lowest confidence tag. - **codex: real bug found and fixed.** The recipe's TOML shape (`[model].default`) was rejected by a real codex binary ("invalid type: map, expected a string") — codex wants a top-level `model` @@ -193,7 +248,7 @@ test/routes/session-custom-model.test.ts test/routes/external-cli-bypass-clamp.t is a tooling-correctness fix (affects the script's own baseline check), not a claim about how any CLI's own HTTP client behaves. - **Also found and fixed**: an earlier design sent BOTH `Authorization: - Bearer` and `api-key` auth header conventions on every discovery/ +Bearer` and `api-key` auth header conventions on every discovery/ baseline request, on the theory that an unused header is harmless. Live-tested against the real server, sending both reliably HUNG the request (reproduced 3×: either header alone ~500-600ms, both together @@ -206,13 +261,33 @@ test/routes/session-custom-model.test.ts test/routes/external-cli-bypass-clamp.t ## Not yet done / open questions for review -- Six of nine per-CLI recipes (Gemini, Pi, Grok, DeepSeek, OMP) are - **web-researched, not verified** against real binaries — see the - confidence table in `deployment_plan.md`. Antigravity has no known - mechanism at all and stays unsupported. -- Chunk 5's session-restart design needs a careful look before - implementation: switching a session's endpoint restarts its CLI process - in place (confirmed acceptable with the maintainer — these harnesses - read endpoint config at process start, not per-turn). +- **Chunk 6 (frontend)** — settings group, toolbar picker, tab badge — is + still entirely unbuilt; the feature is currently HTTP-API-only (see + `docs/custom-model-endpoints.md`). +- **Gemini is confirmed broken end-to-end** (`Invalid auth method + selected`, traced to an undocumented `GATEWAY` AuthType gemini-cli + selects once `GOOGLE_GEMINI_BASE_URL` is set) — needs upstream + investigation before it can be called supported. Documented in full in + `deployment_plan.md`'s confidence table rather than silently shipped as + working. +- **DeepSeek is confirmed reaching the server but failing** with a + consistent `HTTP_404`, root cause not identified — documented as + best-effort/unknown, same as its pre-existing lowest confidence tag. +- **Codex cannot work against a plain OpenAI-Chat-Completions server** + (llama.cpp/llama-swap/Ollama/vLLM's default) — it only speaks the + Responses API since Feb 2026. This is an external protocol + incompatibility, not something this PR can fix; codex support is real + only against a Responses-API-compatible endpoint. +- Antigravity has no known mechanism at all and stays unsupported. +- Chunk 5's session-restart design needs a careful look before merge: + switching a session's endpoint restarts its CLI process in place + (confirmed acceptable with the maintainer — these harnesses read + endpoint config at process start, not per-turn). Whether an INTERACTIVE + claude session with a custom model hits the same async-title-generation + hang the standalone script worked around with `--bare` (vs. just a + harmless background warning) is untested and should be checked before + calling claude's chunk 5 support done — `--bare` itself must never be + applied to a real interactive session, since it disables hooks Codeman + depends on. 🤖 Generated with [Claude Code](https://claude.com/claude-code) diff --git a/deployment_plan.md b/deployment_plan.md index 2d44c90e..e00e6daf 100644 --- a/deployment_plan.md +++ b/deployment_plan.md @@ -69,11 +69,11 @@ config blob for opencode, a TOML file for Codex, etc. Devvyn gave the starting recipes for those three; the rest (Gemini, Pi, Grok, DeepSeek, OMP, Antigravity) were researched for this plan and are flagged by confidence below. A real end-to-end pass against Devvyn's own llama-swap server -(`scripts/test-local-llm-harnesses.mjs`, inside a `codeman/agent:llm-test` +(`scripts/test-local-llm-harnesses.ts`, inside a `codeman/agent:llm-test` Docker image with all 9 CLIs installed) then confirmed **claude and opencode work end-to-end**, corrected a real Codex config.toml schema bug the given recipe had (see the Codex row below), and surfaced that Codex's -*protocol* — not just its config shape — does not work against a plain +_protocol_ — not just its config shape — does not work against a plain OpenAI-Chat-Completions server like llama.cpp/llama-swap at all. Confidence below reflects what was actually observed, not just what was planned. @@ -104,17 +104,17 @@ declared capability, never an `if (mode === 'claude')` branch. ## Per-CLI injection recipes (confidence-ranked) -| CLI | Mechanism | Confidence | -|---|---|---| -| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude | -| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"":{}}}},"model":"custom/"}` | **Verified by user** | -| `codex` | TOML `config.toml`: top-level `model = ""` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box | -| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` (or `GOOGLE_VERTEX_BASE_URL`) + `GEMINI_API_KEY`; CLI needs a restart to pick them up (matches our restart-on-switch design). Model selection via `--model`/`GEMINI_MODEL`-style override — verify exact var name against the installed `gemini-cli` version before shipping | Web-researched, unverified | -| `pi` | Config file `~/.pi/agent/models.json` (hot-reloadable) with a custom provider block: `baseUrl`, `apiKey`, `api:"openai-completions"`. Redirect via `PI_CONFIG_DIR` (already allowlisted per CLAUDE.md) pointed at an isolated dir containing just this file, rather than overwriting the user's real one | Web-researched, unverified | -| `grok` | Env vars `GROK_BASE_URL`, `XAI_API_KEY` (dummy ok for local; a real key for most cloud endpoints), `GROK_MODEL`. All three already fit inside the existing `XAI_*`/CLI-specific allowlist shape | Web-researched, unverified | -| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts:879-913`, already in `privilegedEnvKeys`). Model selection is murkier — CLAUDE.md notes dsh model is "a profile composition entry," not a flag/env var, so redirecting the endpoint is solid but forcing a specific model name may not fully work; document as best-effort and verify against a real profile | Web-researched, unverified, partial | -| `omp` | Config file `~/.omp/agent/models.yml`-equivalent with a custom provider `baseUrl`. CLAUDE.md notes omp's config tree is itself relocatable via `PI_CONFIG_DIR` — reuse the same isolated-dir-redirect approach as `pi` | Web-researched, unverified | -| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism | +| CLI | Mechanism | Confidence | +| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude | +| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"":{}}}},"model":"custom/"}` | **Verified by user** | +| `codex` | TOML `config.toml`: top-level `model = ""` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box | +| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done | +| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" | +| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly | +| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`, already in `privilegedEnvKeys`). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working | +| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live | +| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism | Everything web-researched-but-unverified gets implemented but must be smoke-tested against real installs of those CLIs before being called done — @@ -139,8 +139,13 @@ union on each `CliEntry.capabilities`: type CustomModelInjection = | { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[] } | { kind: 'configContentEnv'; envVar: string; template: 'opencode-json' } - | { kind: 'configDir'; dirEnvVar: string; fileName: string; template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' } - | { kind: 'unsupported' } + | { + kind: 'configDir'; + dirEnvVar: string; + fileName: string; + template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml'; + } + | { kind: 'unsupported' }; ``` Declared per stock.ts entry per the table above. A pure function in a new @@ -241,7 +246,7 @@ dir-redirects, plus the already-privileged `DEEPSEEK_BASE_URL` — must be added to each CLI's `capabilities.privilegedEnvKeys` so `clampEnvOverridesForOwner()` strips them for a non-granted multi-user owner, exactly the precedent already documented for `DEEPSEEK_BASE_URL`/ -`OMP_AUTH_BROKER_URL`. This matters *more*, not less, now that endpoints can +`OMP_AUTH_BROKER_URL`. This matters _more_, not less, now that endpoints can be cloud URLs: redirecting a non-granted user's session to an attacker's cloud endpoint is a credential-exfiltration path, not just a mischief redirect to a LAN box. Endpoint CRUD itself stays admin-only in multi-user @@ -259,14 +264,14 @@ mode, same as remote/docker hosts. - `src/web/public/index.html`, `settings-ui.js`, `session-ui.js`, `styles.css` — settings group, toolbar button/menu, badge, accent CSS - `src/web/sse-events.ts` + `constants.js` — if a dedicated SSE event is warranted for the badge (or just ride existing session-update broadcasts) - `test/fixtures/mock-openai-server.ts` (new) + `test/custom-model-injection-contract.test.ts` (new) — see Mock-server validation below -- `scripts/test-local-llm-harnesses.mjs` (already added, this branch) — the standalone real-CLI-and-real-endpoint smoke test; despite the filename (kept for continuity with when it was written) it already supports any `--base-url`, local or cloud +- `scripts/test-local-llm-harnesses.ts` (already added, this branch; run via `npx tsx`) — the standalone real-CLI-and-real-endpoint smoke test, supporting any `--base-url` (local or cloud). Dynamic: derives its harness list and every env var/config it injects from the live CLI registry + `buildCustomModelInjection()` rather than a second hand-maintained copy — only the one-shot invocation flags (`ONE_SHOT` table) are CLI-specific info the registry doesn't model and stay hand-maintained - `docs/custom-model-endpoints.md` (new) + a CLAUDE.md pointer bullet under External CLI modes / envOverrides ## Mock-server validation strategy (CI-runnable, no real CLI binaries needed) Spawning nine real CLI binaries in CI isn't realistic, and neither Devvyn's llama.cpp box nor a real cloud subscription can be a CI dependency. So the -injection *logic* gets a tier of automated coverage that sits between the +injection _logic_ gets a tier of automated coverage that sits between the pure unit tests and the live manual checks in Verification: 1. **`test/fixtures/mock-openai-server.ts`** — a small in-process HTTP @@ -306,16 +311,21 @@ pure unit tests and the live manual checks in Verification: per-session dir rather than the user's real config path. 3. **Explicit, stated limitation** (goes in the test file's `@fileoverview` - and in this doc, not left implicit): this proves *"if the CLI honors its + and in this doc, not left implicit): this proves _"if the CLI honors its documented env/config contract, it will hit the right endpoint with the - right model."* It does **not** prove the real CLI binary actually reads + right model."_ It does **not** prove the real CLI binary actually reads that env var / config file the way its docs say — that's still the job of the live manual checks in Verification step 4-5 below, and is exactly - why the confidence table above stays "unverified" for six of the nine - CLIs until someone runs those binaries for real. The mock-server suite - catches regressions in Codeman's own logic; it cannot catch a CLI - changing its env-var name in a future release, or a real cloud endpoint - behaving differently from the mock. + why the confidence table above did not stop at "researched" — every CLI + except antigravity (no mechanism at all) has since been run against a + real llama-swap server via `scripts/test-local-llm-harnesses.ts`: + claude/opencode/pi/grok/omp are confirmed PASS end-to-end, codex is + confirmed FAIL for a real documented protocol reason (Responses-API-only + since Feb 2026), and gemini/deepseek are confirmed reaching the server + but failing for reasons not yet root-caused (see their table rows). The + mock-server suite catches regressions in Codeman's own logic; it cannot + catch a CLI changing its env-var name in a future release, or a real + cloud endpoint behaving differently from a local llama.cpp box. ## Verification @@ -326,15 +336,20 @@ pure unit tests and the live manual checks in Verification: 3. Route tests (`app.inject`) for the new CRUD + discover-models endpoint (mock `fetch` for `/v1/models`), and for the multi-user clamp on the new privileged keys (mirror `test/routes/external-cli-bypass-clamp.test.ts`). -4. **Standalone real-binary smoke test**: `scripts/test-local-llm-harnesses.mjs` - (already written on this branch) exercises every harness against a real - `--base-url` — local or cloud — outside of Codeman's UI entirely. Run it - against Devvyn's llama.cpp server first (`claude`/`opencode`/`codex` - should PASS, since those recipes are verified; the rest report - UNCONFIRMED/SKIP until their guessed flags are corrected via - `--probe-help`), then again against a real cloud endpoint (e.g. an Azure - AI Foundry deployment) once one is available, to prove the `authStyle`/ - deployment-name handling holds up outside llama.cpp. +4. **Standalone real-binary smoke test**: `scripts/test-local-llm-harnesses.ts` + exercises every harness the CLI registry declares `customModelInjection` + support for against a real `--base-url` — local or cloud — outside of + Codeman's UI entirely, and is DYNAMIC (reads `enabledClis()` + calls the + real `buildCustomModelInjection()`, so a future registry change is picked + up automatically with zero edits to the script). Already run to + completion against Devvyn's llama-swap server (`http://10.10.11.241:8080`, + inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries): + claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected** + (Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED** + (reach the server, fail for undiagnosed reasons — see their table rows), + antigravity **SKIP** (no mechanism). Re-run this against a real cloud + endpoint (e.g. an Azure AI Foundry deployment) once one is available, to + prove the `authStyle`/deployment-name handling holds up outside llama.cpp. 5. Once the full feature (not just the standalone script) is built: add an endpoint via the real UI, hit discover-models, confirm the returned model list, pick Claude + the model on a real session, confirm via diff --git a/docs/custom-model-endpoints.md b/docs/custom-model-endpoints.md index ba5290f4..f6584367 100644 --- a/docs/custom-model-endpoints.md +++ b/docs/custom-model-endpoints.md @@ -81,15 +81,30 @@ pointed at. ## Confidence per harness -Only Claude, opencode, and Codex have been verified against a real -llama.cpp server by hand. Gemini, Pi, Grok, DeepSeek, and OMP's recipes are -correct on their one-shot invocation flags (confirmed against real -installed binaries' own `--help` output) but their env-var/config -conventions for a _custom_ endpoint are still web-researched, not verified -end-to-end — see the confidence table in `deployment_plan.md` before relying -on one of those five in production. `scripts/test-local-llm-harnesses.mjs` -is the standalone script used to check a harness against a real endpoint -outside the web UI entirely; see its own `--help` for usage. +Every harness except Antigravity has now been run end-to-end against a real +llama-swap server via `scripts/test-local-llm-harnesses.ts` (a dynamic +script that reads the live CLI registry, so a registry change is picked up +automatically). Results: + +- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply + came back through the endpoint. +- **Codex** — the config is structurally correct, but Codex only speaks the + Responses API since Feb 2026, which llama.cpp/llama-swap don't implement. + This is a real protocol incompatibility, not a bug here; Codex support + needs a Responses-API-compatible endpoint. +- **Gemini** — fails with `Invalid auth method selected`, traced to an + undocumented `GATEWAY` auth path gemini-cli selects once + `GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation + (several auth workarounds were tried and ruled out); do not rely on + Gemini support yet. +- **DeepSeek** — the request reaches the server (env vars are read) but + gets a consistent `HTTP_404`. Root cause not identified; best-effort only. +- **Antigravity** — no known custom-endpoint mechanism at all; unsupported. + +See the confidence table in `deployment_plan.md` for the full detail behind +each result. `scripts/test-local-llm-harnesses.ts` is the standalone script +used to check a harness against a real endpoint outside the web UI +entirely; see its own `--help` for usage. ## Security note diff --git a/scripts/test-local-llm-harnesses.mjs b/scripts/test-local-llm-harnesses.mjs deleted file mode 100644 index 6e43d5ee..00000000 --- a/scripts/test-local-llm-harnesses.mjs +++ /dev/null @@ -1,631 +0,0 @@ -#!/usr/bin/env node -/** - * Standalone smoke-test for pointing each Codeman-supported harness CLI at a - * custom OpenAI-compatible endpoint — local (llama.cpp, Ollama, vLLM, ...) or - * cloud (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a - * self-hosted gateway, ...). Anything that answers GET /v1/models and POST - * /v1/chat/completions in the standard shape qualifies; --base-url is not - * assumed to be a LAN address. - * - * This is intentionally OUTSIDE the npm test suite and outside Codeman's own - * session/tmux machinery: it spawns each real CLI binary directly, one-shot, - * with the env vars / config files that CLI's own docs say redirect it to a - * custom endpoint, and checks it can answer "hello world". - * - * Cloud endpoints often differ from a bare llama.cpp box in two ways this - * script accounts for: (1) auth may be an `api-key` header (Azure's - * convention) rather than `Authorization: Bearer` — the baseline check in - * Step 0 sends both, since an extra header is harmless to servers that - * ignore it; each CLI's OWN auth convention (set via its env vars/config, - * not this script) still needs to match what that endpoint expects. (2) a - * cloud endpoint's "model" may actually be a deployment name distinct from - * the model family (Azure AI Foundry deployments) — always pass --model - * explicitly for those rather than relying on GET /v1/models discovery. - * - * IMPORTANT CONFIDENCE NOTE: only claude/opencode/codex recipes are verified - * (Devvyn confirmed them by hand). gemini/pi/grok/deepseek/omp are best - * guesses from public docs, not verified against this repo or against real - * binaries. antigravity has no known CLI/env mechanism at all and is always - * skipped. Read a harness's UNCONFIRMED/FAIL output before trusting it — use - * --probe-help to read that binary's real --help and fix the guessed flag. - * - * Usage: - * node scripts/test-local-llm-harnesses.mjs --base-url http://192.168.1.50:8080 [options] - * node scripts/test-local-llm-harnesses.mjs --base-url https://.services.ai.azure.com/openai/v1 --model --api-key $AZURE_AI_KEY - * - * Options: - * --base-url Required. Root URL of the OpenAI-compatible endpoint (local or cloud). - * --model Model/deployment id to request. Default: first from GET /v1/models. - * --api-key API key to send. Default: local-dummy-key (fine for llama.cpp; required for most cloud endpoints). - * --auth-style