mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
41416566aa
commit
61779745aa
@@ -81,15 +81,30 @@ pointed at.
|
||||
|
||||
## Confidence per harness
|
||||
|
||||
Only Claude, opencode, and Codex have been verified against a real
|
||||
llama.cpp server by hand. Gemini, Pi, Grok, DeepSeek, and OMP's recipes are
|
||||
correct on their one-shot invocation flags (confirmed against real
|
||||
installed binaries' own `--help` output) but their env-var/config
|
||||
conventions for a _custom_ endpoint are still web-researched, not verified
|
||||
end-to-end — see the confidence table in `deployment_plan.md` before relying
|
||||
on one of those five in production. `scripts/test-local-llm-harnesses.mjs`
|
||||
is the standalone script used to check a harness against a real endpoint
|
||||
outside the web UI entirely; see its own `--help` for usage.
|
||||
Every harness except Antigravity has now been run end-to-end against a real
|
||||
llama-swap server via `scripts/test-local-llm-harnesses.ts` (a dynamic
|
||||
script that reads the live CLI registry, so a registry change is picked up
|
||||
automatically). Results:
|
||||
|
||||
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
|
||||
came back through the endpoint.
|
||||
- **Codex** — the config is structurally correct, but Codex only speaks the
|
||||
Responses API since Feb 2026, which llama.cpp/llama-swap don't implement.
|
||||
This is a real protocol incompatibility, not a bug here; Codex support
|
||||
needs a Responses-API-compatible endpoint.
|
||||
- **Gemini** — fails with `Invalid auth method selected`, traced to an
|
||||
undocumented `GATEWAY` auth path gemini-cli selects once
|
||||
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
|
||||
(several auth workarounds were tried and ruled out); do not rely on
|
||||
Gemini support yet.
|
||||
- **DeepSeek** — the request reaches the server (env vars are read) but
|
||||
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
|
||||
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
|
||||
|
||||
See the confidence table in `deployment_plan.md` for the full detail behind
|
||||
each result. `scripts/test-local-llm-harnesses.ts` is the standalone script
|
||||
used to check a harness against a real endpoint outside the web UI
|
||||
entirely; see its own `--help` for usage.
|
||||
|
||||
## Security note
|
||||
|
||||
|
||||
Reference in New Issue
Block a user