feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)

Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).

- Registry: capabilities.customModelInjection per CLI entry (env /
  configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
  model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
  custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
  (POST /api/sessions/:id/custom-model), reusing the existing
  respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
  CLI's privilegedEnvKeys, closing a pre-existing gap where several were
  already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
  against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
  every CLI's injected values through a real HTTP shape

Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
  a top-level model string + [model_providers.custom]); fixing it then
  surfaced a genuine, documented protocol incompatibility (Codex only
  speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
  don't implement)
- Claude Code's async session-title-generation call validates
  ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
  hangs the whole -p invocation on an unrecognized name; documented for
  chunk 6, worked around in the standalone script only (--bare is NOT
  safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
  and api-key headers) reliably hung a real server; removed the option
  entirely rather than just changing the default

Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
This commit is contained in:
Devvyn
2026-09-13 17:42:35 +08:00
co-authored by Claude Sonnet 5
parent a017e9a8e0
commit 41416566aa
25 changed files with 2849 additions and 5 deletions
+86 -1
View File
@@ -8,7 +8,7 @@ import { FastifyInstance, type FastifyReply } from 'fastify';
import { z } from 'zod';
import { join, dirname, extname, basename } from 'node:path';
import { homedir } from 'node:os';
import { existsSync, statSync, mkdirSync, writeFileSync } from 'node:fs';
import { existsSync, statSync, mkdirSync, writeFileSync, rmSync } from 'node:fs';
import { execFile } from 'node:child_process';
import fs from 'node:fs/promises';
import { randomBytes } from 'node:crypto';
@@ -51,7 +51,10 @@ import {
SessionOrderUpdateSchema,
SessionWaitQuerySchema,
SessionWaitOutputQuerySchema,
CustomModelSelectionSchema,
} from '../schemas.js';
import { readCustomModelHosts } from '../../custom-model-hosts.js';
import { buildCustomModelInjection } from '../../custom-model-injection.js';
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
import { TabLayoutValidationError } from '../../tab-layout.js';
import {
@@ -1161,6 +1164,88 @@ export function registerSessionRoutes(
return { color: session.color };
});
// ========== Custom Model Endpoint Profiles (deployment_plan.md) ==========
//
// Applies (or clears) a session's custom OpenAI-compatible endpoint selection and
// RESTARTS the pane's CLI process — these harnesses read endpoint config at process
// start, not per-turn, so a live hot-swap isn't possible (confirmed with the
// maintainer). Endpoints come from the admin-configured custom-model-hosts store
// (chunk 3's CRUD routes), never raw client-supplied env — that's what keeps this
// route safe to let any session owner call for their own session, unlike the
// generic envOverrides field the privilegedEnvKeys clamp exists to guard.
app.post('/api/sessions/:id/custom-model', async (req) => {
const { id } = req.params as { id: string };
const body = parseBody(CustomModelSelectionSchema, req.body, 'Invalid request body');
const session = findSessionOrFail(ctx, id, req);
if (session.isBusy()) {
return createErrorResponse(ApiErrorCode.SESSION_BUSY, 'Session is busy');
}
if ('clear' in body) {
const previousConfigDir = session.setCustomModel(undefined);
if (previousConfigDir) rmSync(previousConfigDir, { recursive: true, force: true });
const restarted = await session.restartCli();
persistAndBroadcastSession(ctx, session);
return { customModel: session.customModel, restarted };
}
const entry = getCli(session.mode);
if (!entry) {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, `No CLI registry entry for mode ${session.mode}`);
}
if (entry.capabilities.customModelInjection.kind === 'unsupported') {
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
}
const hosts = await readCustomModelHosts(getDataDir());
const endpoint = hosts.find((h) => h.id === body.endpointId);
if (!endpoint) {
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
}
const injection = buildCustomModelInjection(entry, endpoint, body.modelId);
let envOverrides: Record<string, string>;
let envKeys: string[];
let configDir: string | undefined;
if (injection.kind === 'env') {
envOverrides = injection.envOverrides;
envKeys = Object.keys(injection.envOverrides);
} else if (injection.kind === 'configDir') {
// Isolated per-session dir — never the user's real CLI config path.
configDir = join(dataPath('custom-model-configs'), session.id);
for (const file of injection.files) {
const filePath = join(configDir, file.relPath);
mkdirSync(dirname(filePath), { recursive: true });
writeFileSync(filePath, file.content, 'utf8');
}
// extraEnv: vars the written config file REFERENCES by name (codex's `env_key`
// convention) rather than embedding a literal value — must ride alongside
// dirEnvVar or the config points at a credential that was never actually set.
envOverrides = { [injection.dirEnvVar]: configDir, ...injection.extraEnv };
envKeys = [injection.dirEnvVar, ...Object.keys(injection.extraEnv ?? {})];
} else {
// 'unsupported' is already handled above; this keeps the switch exhaustive.
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
}
const previousConfigDir = session.setCustomModel(
{ endpointId: endpoint.id, modelId: body.modelId, label: endpoint.label, envKeys, configDir },
envOverrides
);
// Clean up the OLD config dir on disk, unless the new one happens to reuse the same
// path (same session, configDir kind again) — never delete the dir we just wrote.
if (previousConfigDir && previousConfigDir !== configDir) {
rmSync(previousConfigDir, { recursive: true, force: true });
}
const restarted = await session.restartCli();
persistAndBroadcastSession(ctx, session);
return { customModel: session.customModel, restarted };
});
// ========== Delete Session ==========
app.delete('/api/sessions/:id', async (req) => {