fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)

DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.

Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
  published dependencies) into a scratch dir purely to read
  dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
  completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
  the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
  completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
  same endpoint. dsh's own error template ("DeepSeek API error (HTTP
  ${status})") reproduces the originally-reported
  "dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.

- New registry field `appendV1Suffix` (env kind only, deepseek's entry
  alone — claude/gemini must NOT get it, since claude was already
  confirmed working against the unmodified baseUrl). When set,
  buildCustomModelInjection runs endpoint.baseUrl through the same
  withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
  use, instead of writing it verbatim.

Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.

2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 13:41:19 +08:00
co-authored by Claude Sonnet 5
parent 8520925e76
commit e2034177c5
10 changed files with 114 additions and 45 deletions
+21 -17
View File
@@ -142,17 +142,19 @@ describe('custom-model-injection contract (mock server)', () => {
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
});
// gemini/deepseek's `env` kind passes the base URL through UNCHANGED (unlike
// opencode/codex/pi/omp/grok, which build a structured config and explicitly append
// /v1) — matching Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the
// SDK appends the path itself. Whether each of these TWO CLIs' own OpenAI-compatible
// client expects the var to already include /v1 (the common OpenAI-SDK convention) or
// appends it itself is genuinely CLI-specific and UNVERIFIED (see the confidence table
// in docs/custom-model-endpoints-plan.md) — these tests model the common OpenAI-SDK convention (base_url
// ends in /v1) since that's the more likely behavior for an OpenAI-compatible client,
// but that assumption should be corrected here the moment it's checked against a real
// binary. (grok WAS in this group too, until live-testing showed the whole `env` recipe
// was wrong for it — see its own test below.)
// gemini's `env` kind still passes the base URL through UNCHANGED (matching
// Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the SDK appends
// the path itself) — whether gemini-cli's own OpenAI-compatible-ish client expects the
// var to already include /v1 or appends it itself remains genuinely UNVERIFIED (it
// fails for an unrelated auth reason before this would even matter — see the
// confidence table in docs/custom-model-endpoints-plan.md); this test models the
// common OpenAI-SDK convention as the best guess, to be corrected the moment it's
// checked against a real client. deepseek WAS in this "passes through unchanged"
// group too, until reading `@deepseek-ai/dsh-llm-deepseek`'s own bundled source
// confirmed it builds its request URL as `${DEEPSEEK_BASE_URL}/chat/completions` with
// no `/v1` of its own — `appendV1Suffix` now fixes that (see its own test below),
// the same way grok's whole `env` recipe turned out to be wrong before live-testing
// corrected it to a `configDir` one.
it('gemini: GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY reach the mock', async () => {
const injection = buildCustomModelInjection(entryOrThrow('gemini'), endpointFor(mock), 'qwen3');
@@ -186,16 +188,18 @@ describe('custom-model-injection contract (mock server)', () => {
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
});
it('deepseek: DEEPSEEK_BASE_URL/DEEPSEEK_API_KEY reach the mock (base URL/key only, no model var)', async () => {
it('deepseek: DEEPSEEK_BASE_URL already carries the /v1 suffix dsh itself never adds, reaching the mock at the real path dsh requests', async () => {
// Confirmed by reading dsh's own bundled source: it fetches
// `${DEEPSEEK_BASE_URL}/chat/completions` verbatim, no /v1 insertion of its own — so
// this call (unlike gemini's above) passes DEEPSEEK_BASE_URL to callOpenAiCompat
// UNMODIFIED, exactly mirroring what the real harness does, rather than the test
// helping it along.
const injection = buildCustomModelInjection(entryOrThrow('deepseek'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'env') throw new Error('unreachable');
expect(Object.keys(injection.envOverrides).sort()).toEqual(['DEEPSEEK_API_KEY', 'DEEPSEEK_BASE_URL']);
expect(injection.envOverrides.DEEPSEEK_BASE_URL).toBe(`${mock.baseUrl}/v1`);
await callOpenAiCompat(
`${injection.envOverrides.DEEPSEEK_BASE_URL}/v1`,
injection.envOverrides.DEEPSEEK_API_KEY,
'qwen3'
);
await callOpenAiCompat(injection.envOverrides.DEEPSEEK_BASE_URL, injection.envOverrides.DEEPSEEK_API_KEY, 'qwen3');
expect(mock.requests[0].path).toBe('/v1/chat/completions');
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
+18 -2
View File
@@ -179,15 +179,31 @@ describe('buildCustomModelInjection', () => {
expect(result.extraEnv).toEqual({ XAI_API_KEY: 'my-key' });
});
it('deepseek: env kind sets base URL/key only, no model var', () => {
it('deepseek: env kind sets base URL (with a /v1 suffix appended) and key, no model var', () => {
// appendV1Suffix is REQUIRED here, not cosmetic: confirmed by reading dsh's own
// bundled source (@deepseek-ai/dsh-llm-deepseek) that it builds the request URL as
// `${DEEPSEEK_BASE_URL}/chat/completions` with no "/v1" of its own, while
// llama-swap/llama.cpp only serves "/v1/chat/completions" — without this, every
// request 404s (confirmed live; this is the fix for the originally-reported
// "dsh: HTTP_404: DeepSeek API error (HTTP 404)").
const result = buildCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides).toEqual({
DEEPSEEK_BASE_URL: 'http://192.168.1.50:8080',
DEEPSEEK_BASE_URL: 'http://192.168.1.50:8080/v1',
DEEPSEEK_API_KEY: 'my-key',
});
});
it('deepseek: appending the /v1 suffix is idempotent against a baseUrl that already ends in /v1', () => {
const result = buildCustomModelInjection(
entryOrThrow('deepseek'),
{ ...endpoint, baseUrl: 'http://192.168.1.50:8080/v1' },
'qwen3'
);
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides.DEEPSEEK_BASE_URL).toBe('http://192.168.1.50:8080/v1');
});
it('antigravity: unsupported', () => {
const result = buildCustomModelInjection(entryOrThrow('antigravity'), endpoint, 'qwen3');
expect(result).toEqual({ kind: 'unsupported' });