mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. Covers local hardware (llama.cpp, Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter). Off by default (customModelEndpointsEnabled, synced, default OFF). - Registry: capabilities.customModelInjection per CLI entry (env / configContentEnv / configDir / unsupported kinds) - Pure injection builder (custom-model-injection.ts) turning an endpoint + model id into the real env vars / config content per CLI - Endpoint store + CRUD routes (custom-model-hosts.ts, custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded - Session integration: Session.setCustomModel()/restartCli() (POST /api/sessions/:id/custom-model), reusing the existing respawn-pane -k primitive to restart the CLI process with new env - Multi-user hardening: every new redirect-capable env var added to its CLI's privilegedEnvKeys, closing a pre-existing gap where several were already reachable via the generic envOverrides field's prefix allowlist - Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries against a real endpoint outside the web UI, independent of tmux/sessions - Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying every CLI's injected values through a real HTTP shape Real end-to-end validation against a live llama-swap server (inside a codeman/agent:llm-test Docker image with all 9 CLI binaries) found and fixed three real bugs before they shipped: - Codex's config.toml schema was wrong ([model].default table instead of a top-level model string + [model_providers.custom]); fixing it then surfaced a genuine, documented protocol incompatibility (Codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement) - Claude Code's async session-title-generation call validates ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and hangs the whole -p invocation on an unrecognized name; documented for chunk 6, worked around in the standalone script only (--bare is NOT safe for a real interactive session, which needs hooks) - The discovery route's authStyle: 'both' option (send both Authorization and api-key headers) reliably hung a real server; removed the option entirely rather than just changing the default Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see PR.md and deployment_plan.md for the full chunk breakdown and confidence table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
a017e9a8e0
commit
41416566aa
@@ -0,0 +1,59 @@
|
||||
/**
|
||||
* @fileoverview Read/write-array store for user-configured custom OpenAI-compatible
|
||||
* model endpoints (local or cloud — deployment_plan.md). Same shape as
|
||||
* `remote-hosts.ts` / `webview-store.ts`: `~/.codeman/custom-model-hosts.json`
|
||||
* holding a plain array, read/written whole.
|
||||
*/
|
||||
|
||||
import { existsSync, mkdirSync } from 'node:fs';
|
||||
import fs from 'node:fs/promises';
|
||||
import { join } from 'node:path';
|
||||
|
||||
const CUSTOM_MODEL_HOSTS_FILE = 'custom-model-hosts.json';
|
||||
|
||||
export type CustomModelAuthStyle = 'bearer' | 'api-key';
|
||||
|
||||
export interface CustomModelHost {
|
||||
id: string;
|
||||
label: string;
|
||||
/** Root URL, local or cloud — e.g. "http://192.168.1.50:8080" or an Azure AI Foundry URL. */
|
||||
baseUrl: string;
|
||||
apiKey?: string;
|
||||
/**
|
||||
* Defaults to 'bearer' (the common `Authorization: Bearer` convention — matches
|
||||
* llama.cpp, OpenAI-compatible servers, and most gateways). Pick 'api-key' for
|
||||
* endpoints that specifically want the `api-key` header, e.g. Azure AI Foundry.
|
||||
*
|
||||
* ⚠️ There is deliberately NO 'both' option. An earlier design sent BOTH headers
|
||||
* on every discovery request on the theory that an unused header is harmless —
|
||||
* live-tested against a real llama-swap server, sending both reliably HUNG the
|
||||
* request indefinitely (reproduced 3× — Bearer alone: ~500ms, api-key alone:
|
||||
* ~600ms, both together: no response inside a 15s timeout). Whatever auth
|
||||
* middleware some servers run apparently does not handle two simultaneous
|
||||
* credential conventions gracefully, so "send everything and let the server
|
||||
* ignore what it doesn't need" is not a safe default — it can silently turn a
|
||||
* working endpoint into one that always times out.
|
||||
*/
|
||||
authStyle?: CustomModelAuthStyle;
|
||||
models?: string[];
|
||||
lastDiscoveredAt?: string;
|
||||
}
|
||||
|
||||
export function customModelHostsPath(configDir: string): string {
|
||||
return join(configDir, CUSTOM_MODEL_HOSTS_FILE);
|
||||
}
|
||||
|
||||
export async function readCustomModelHosts(configDir: string): Promise<CustomModelHost[]> {
|
||||
try {
|
||||
const raw = await fs.readFile(customModelHostsPath(configDir), 'utf-8');
|
||||
const parsed = JSON.parse(raw);
|
||||
return Array.isArray(parsed) ? (parsed as CustomModelHost[]) : [];
|
||||
} catch {
|
||||
return [];
|
||||
}
|
||||
}
|
||||
|
||||
export async function writeCustomModelHosts(configDir: string, hosts: CustomModelHost[]): Promise<void> {
|
||||
if (!existsSync(configDir)) mkdirSync(configDir, { recursive: true });
|
||||
await fs.writeFile(customModelHostsPath(configDir), JSON.stringify(hosts, null, 2));
|
||||
}
|
||||
Reference in New Issue
Block a user