mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
41416566aa
commit
61779745aa
@@ -12,7 +12,7 @@
|
||||
* proves "if the CLI honors its documented env/config contract, it will hit
|
||||
* the right endpoint with the right model." It does NOT prove the real CLI
|
||||
* binary actually reads that env var / config file the way its docs say —
|
||||
* that's still the job of `scripts/test-local-llm-harnesses.mjs` against a
|
||||
* that's still the job of `scripts/test-local-llm-harnesses.ts` against a
|
||||
* real endpoint and real binaries. This suite catches regressions in
|
||||
* Codeman's own injection logic; it cannot catch a CLI changing its env-var
|
||||
* name in a future release.
|
||||
@@ -142,16 +142,17 @@ describe('custom-model-injection contract (mock server)', () => {
|
||||
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
|
||||
});
|
||||
|
||||
// gemini/grok/deepseek's `env` kind passes the base URL through UNCHANGED (unlike
|
||||
// opencode/codex/pi/omp, which build a structured config and explicitly append /v1) —
|
||||
// matching Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the SDK
|
||||
// appends the path itself. Whether each of these THREE CLIs' own OpenAI-compatible
|
||||
// gemini/deepseek's `env` kind passes the base URL through UNCHANGED (unlike
|
||||
// opencode/codex/pi/omp/grok, which build a structured config and explicitly append
|
||||
// /v1) — matching Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the
|
||||
// SDK appends the path itself. Whether each of these TWO CLIs' own OpenAI-compatible
|
||||
// client expects the var to already include /v1 (the common OpenAI-SDK convention) or
|
||||
// appends it itself is genuinely CLI-specific and UNVERIFIED (see the confidence table
|
||||
// in deployment_plan.md) — these tests model the common OpenAI-SDK convention (base_url
|
||||
// ends in /v1) since that's the more likely behavior for an OpenAI-compatible client,
|
||||
// but that assumption should be corrected here the moment it's checked against a real
|
||||
// binary.
|
||||
// binary. (grok WAS in this group too, until live-testing showed the whole `env` recipe
|
||||
// was wrong for it — see its own test below.)
|
||||
|
||||
it('gemini: GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY reach the mock', async () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('gemini'), endpointFor(mock), 'qwen3');
|
||||
@@ -168,15 +169,18 @@ describe('custom-model-injection contract (mock server)', () => {
|
||||
expect((mock.requests[0].body as { model: string }).model).toBe('qwen3');
|
||||
});
|
||||
|
||||
it('grok: GROK_BASE_URL/XAI_API_KEY reach the mock', async () => {
|
||||
it('grok: config.toml [model.<name>] block base_url/env_key + extraEnv reach the mock over /v1/chat/completions', async () => {
|
||||
const injection = buildCustomModelInjection(entryOrThrow('grok'), endpointFor(mock), 'qwen3');
|
||||
if (injection.kind !== 'env') throw new Error('unreachable');
|
||||
if (injection.kind !== 'configDir') throw new Error('unreachable');
|
||||
const toml = injection.files[0].content;
|
||||
const baseUrl = /base_url = "([^"]+)"/.exec(toml)?.[1];
|
||||
const model = /^model = "([^"]+)"/m.exec(toml)?.[1];
|
||||
expect(baseUrl).toBe(`${mock.baseUrl}/v1`);
|
||||
expect(model).toBe('qwen3');
|
||||
expect(toml).toContain('api_backend = "chat_completions"');
|
||||
expect(injection.extraEnv).toEqual({ XAI_API_KEY: 'contract-test-key' });
|
||||
|
||||
await callOpenAiCompat(
|
||||
`${injection.envOverrides.GROK_BASE_URL}/v1`,
|
||||
injection.envOverrides.XAI_API_KEY,
|
||||
injection.envOverrides.GROK_MODEL
|
||||
);
|
||||
await callOpenAiCompat(baseUrl!, injection.extraEnv!.XAI_API_KEY, model!);
|
||||
|
||||
expect(mock.requests[0].path).toBe('/v1/chat/completions');
|
||||
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
|
||||
|
||||
@@ -91,21 +91,25 @@ describe('buildCustomModelInjection', () => {
|
||||
expect(result.files[0].content).toContain('model = "weird\\"model"');
|
||||
});
|
||||
|
||||
it('pi: configDir writes models.json under agent/', () => {
|
||||
it('pi: configDir writes .pi/agent/models.json, redirected via HOME (verified live — PI_CONFIG_DIR does nothing for pi)', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('pi'), endpoint, 'qwen3');
|
||||
if (result.kind !== 'configDir') throw new Error('unreachable');
|
||||
expect(result.dirEnvVar).toBe('PI_CONFIG_DIR');
|
||||
expect(result.files[0].relPath).toBe('agent/models.json');
|
||||
expect(result.dirEnvVar).toBe('HOME');
|
||||
expect(result.files[0].relPath).toBe('.pi/agent/models.json');
|
||||
const parsed = JSON.parse(result.files[0].content);
|
||||
expect(parsed.providers.custom.baseUrl).toBe('http://192.168.1.50:8080/v1');
|
||||
expect(parsed.providers.custom.authHeader).toBe(true);
|
||||
expect(parsed.providers.custom.models).toEqual([{ id: 'qwen3' }]); // array, NOT keyed by id
|
||||
});
|
||||
|
||||
it('omp: configDir writes models.yml under agent/, redirected via PI_CONFIG_DIR', () => {
|
||||
it('omp: configDir writes .omp/agent/models.yml, redirected via HOME (verified live end-to-end)', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('omp'), endpoint, 'qwen3');
|
||||
if (result.kind !== 'configDir') throw new Error('unreachable');
|
||||
expect(result.dirEnvVar).toBe('PI_CONFIG_DIR');
|
||||
expect(result.files[0].relPath).toBe('agent/models.yml');
|
||||
expect(result.dirEnvVar).toBe('HOME');
|
||||
expect(result.files[0].relPath).toBe('.omp/agent/models.yml');
|
||||
expect(result.files[0].content).toContain('baseUrl: "http://192.168.1.50:8080/v1"');
|
||||
expect(result.files[0].content).toContain('authHeader: true');
|
||||
expect(result.files[0].content).toContain('- id: "qwen3"');
|
||||
});
|
||||
|
||||
it('gemini: env kind sets GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY/GEMINI_MODEL', () => {
|
||||
@@ -118,14 +122,19 @@ describe('buildCustomModelInjection', () => {
|
||||
});
|
||||
});
|
||||
|
||||
it('grok: env kind sets GROK_BASE_URL/XAI_API_KEY/GROK_MODEL', () => {
|
||||
it('grok: configDir writes a config.toml [model.<name>] block, key rides as extraEnv (XAI_API_KEY)', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('grok'), endpoint, 'qwen3');
|
||||
if (result.kind !== 'env') throw new Error('unreachable');
|
||||
expect(result.envOverrides).toEqual({
|
||||
GROK_BASE_URL: 'http://192.168.1.50:8080',
|
||||
XAI_API_KEY: 'my-key',
|
||||
GROK_MODEL: 'qwen3',
|
||||
});
|
||||
expect(result.kind).toBe('configDir');
|
||||
if (result.kind !== 'configDir') throw new Error('unreachable');
|
||||
expect(result.dirEnvVar).toBe('GROK_HOME');
|
||||
expect(result.files).toHaveLength(1);
|
||||
expect(result.files[0].relPath).toBe('config.toml');
|
||||
expect(result.files[0].content).toContain('model = "qwen3"');
|
||||
expect(result.files[0].content).toContain('base_url = "http://192.168.1.50:8080/v1"');
|
||||
expect(result.files[0].content).toContain('api_backend = "chat_completions"');
|
||||
expect(result.files[0].content).toContain('env_key = "XAI_API_KEY"');
|
||||
expect(result.files[0].content).not.toContain('api_key ='); // never a literal TOML field
|
||||
expect(result.extraEnv).toEqual({ XAI_API_KEY: 'my-key' });
|
||||
});
|
||||
|
||||
it('deepseek: env kind sets base URL/key only, no model var', () => {
|
||||
|
||||
Reference in New Issue
Block a user