#!/usr/bin/env -S npx tsx /** * Standalone smoke-test for pointing each Codeman-supported harness CLI at a * custom OpenAI-compatible endpoint — local (llama.cpp, Ollama, vLLM, ...) or * cloud (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a * self-hosted gateway, ...). Anything that answers GET /v1/models and POST * /v1/chat/completions in the standard shape qualifies; --base-url is not * assumed to be a LAN address. * * This is intentionally OUTSIDE the npm test suite and outside Codeman's own * session/tmux machinery: it spawns each real CLI binary directly, one-shot, * with the env vars / config files that CLI's own docs say redirect it to a * custom endpoint, and checks it can answer "hello world". * * DYNAMIC BY DESIGN: this file imports the SAME `enabledClis()` registry and * `buildCustomModelInjection()` builder the production feature uses (see * ../src/config/cli-registry/, ../src/custom-model-injection.ts, * ../src/custom-model-injection-apply.ts) rather than keeping a second, * hand-maintained copy of each CLI's env vars/config shape. A registry * change (a new CLI, an edited env var name, a fixed config template) is * picked up here automatically with zero edits to this file. Only the * ONE-SHOT INVOCATION FLAGS (how to make each CLI answer one prompt and * exit — information the registry doesn't model at all, since it only knows * how to launch the interactive TUI) stay in the small ONE_SHOT table below; * a CLI newly added to the registry with no ONE_SHOT entry is reported * UNKNOWN rather than silently skipped or guessed at. * * Cloud endpoints often differ from a bare llama.cpp box in two ways this * script accounts for: (1) auth may be an `api-key` header (Azure's * convention) rather than `Authorization: Bearer` — see --auth-style below. * (2) a cloud endpoint's "model" may actually be a deployment name distinct * from the model family (Azure AI Foundry deployments) — always pass * --model explicitly for those rather than relying on GET /v1/models * discovery. * * IMPORTANT CONFIDENCE NOTE: claude and opencode are verified end-to-end * against a real llama-swap server. codex's config STRUCTURE is verified, * but it only speaks the Responses API (dropped Chat-Completions support * Feb 2026) — expect it to fail against a plain OpenAI-compatible server, * that's a real protocol gap, not a bug here. gemini/pi/grok/omp have their * ONE-SHOT INVOCATION flags confirmed against real installed binaries' * `--help` output, but their custom-endpoint env/config conventions remain * web-researched, unverified. deepseek (dsh) is a profile launcher with no * documented one-shot prompt flag at all — best-effort only. antigravity * has no known CLI/env/config mechanism (GUI-only per public docs) — its * registry entry declares `customModelInjection: { kind: 'unsupported' }`, * which this script picks up dynamically and always skips. * * Usage: * npx tsx scripts/test-local-llm-harnesses.ts --base-url http://192.168.1.50:8080 [options] * npx tsx scripts/test-local-llm-harnesses.ts --base-url https://.services.ai.azure.com/openai/v1 --model --api-key $AZURE_AI_KEY * * Options: * --base-url Required. Root URL of the OpenAI-compatible endpoint (local or cloud). * --model Model/deployment id to request. Default: first from GET /v1/models. * --api-key API key to send. Default: local-dummy-key (fine for llama.cpp; required for most cloud endpoints). * --auth-style