mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-03 22:19:42 +02:00
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.
POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.
Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.
Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
128 lines
8.1 KiB
Markdown
128 lines
8.1 KiB
Markdown
# Custom Model Endpoints
|
|
|
|
Point a harness at your own OpenAI-compatible server instead of its native cloud backend, for
|
|
one session at a time. "Custom endpoint" covers **local** hardware (llama.cpp, Ollama, vLLM,
|
|
a home GPU rig, DGX Spark, Strix Halo) and **cloud** services (Azure AI Foundry's
|
|
OpenAI-compatible endpoint, OpenRouter, a company gateway) alike, anything answering
|
|
`GET /v1/models` and `POST /v1/chat/completions` in the standard shape.
|
|
|
|
**Off by default.** Turn it on in App Settings → Models → **Custom model endpoints**.
|
|
|
|
## Adding an endpoint
|
|
|
|
Still in App Settings → Models → Custom model endpoints:
|
|
|
|
1. **+ Add endpoint** — give it an id, a label, and the base URL (`http://192.168.1.50:8080`,
|
|
say). An API key is optional; most local servers don't check one.
|
|
2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it.
|
|
3. Pick a **default model** from what was discovered. This is the model the Run-menu entry
|
|
applies directly when only one model is discovered; with two or more, it's just the one
|
|
pre-marked in the picker dialog described below, not a silent default.
|
|
|
|
Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker
|
|
hosts — these are machine-level infra, not a per-user setting.
|
|
|
|
**Model lists refresh themselves.** Every saved endpoint is re-discovered automatically every
|
|
5 minutes in the background, so a model the server starts serving later — or stops serving —
|
|
shows up without another manual click of **Discover**. One endpoint being unreachable on a
|
|
given cycle (powered off, wrong network) never blocks the others from refreshing.
|
|
|
|
**Context length is picked up automatically where it can be, safely.** Against a
|
|
llama.cpp/llama-swap server, discovery also learns each *currently loaded* model's real
|
|
context window and applies it to the launched session (Claude Code today — see below), so
|
|
the harness stops assuming a large default window for a model name it doesn't recognise and
|
|
overflowing a much smaller real one. It's deliberately never probed for a model that isn't
|
|
already loaded, since asking a llama-swap server about an unloaded model can trigger an
|
|
actual, slow model swap as a side effect — a model just not currently loaded keeps whatever
|
|
context length an earlier cycle already learned for it instead.
|
|
|
|
## Running a session against one
|
|
|
|
With the setting on and at least one endpoint carrying a discovered model, the **Run**
|
|
dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a
|
|
custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a
|
|
session on that harness exactly the way its own entry would. It is a one-off "try this
|
|
endpoint" action, not a sticky mode — the plain **Run** button still means "this harness,
|
|
native cloud" afterward, and a fresh session never inherits whatever the last one was
|
|
pointed at.
|
|
|
|
**Which model it uses depends on how many the endpoint has discovered.** With exactly one,
|
|
the session launches straight away on that model — nothing to choose. With two or more, a
|
|
small dialog asks which one to use for this launch before starting the session; the
|
|
endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately
|
|
use a different one without changing the saved default.
|
|
|
|
**For opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP, picking an entry launches
|
|
straight onto the endpoint** — no restart, because the endpoint is applied before the
|
|
session's process ever starts. **Claude still restarts the harness's process in place** —
|
|
same tab, same conversation (`--resume`) — after a normal native launch, since that restart
|
|
is far less jarring for Claude than for the other seven, whose own TUI can fully
|
|
reinitialize on a restart. Either way, every supported harness reads its endpoint config at
|
|
process start, never per turn, so there is no live hot-swap while a turn is running.
|
|
|
|
Picking an entry that launches a **brand-new** Claude session waits (up to 20 seconds) for it to
|
|
finish its own startup before applying — a freshly started CLI reports itself as busy for its
|
|
boot sequence, and applying to a genuinely busy session is refused so a real, in-progress
|
|
turn is never interrupted out from under you. A session that is still busy after that wait
|
|
(a very slow-starting CLI, or one you started typing into right away) surfaces that refusal
|
|
as an ordinary error, which now stays on screen with a close button instead of vanishing
|
|
after a few seconds — read it, it names the actual reason rather than a generic failure.
|
|
|
|
Entries are hidden entirely for a session in a **remote (SSH) or Docker case** — support for
|
|
redirecting those hasn't landed yet, see below. The picker also only appears in the desktop
|
|
**Run** dropdown; the phone home screen builds its own run picker separately and does not
|
|
currently offer these entries.
|
|
|
|
**Claude Code specifically gets two extra fixes applied automatically:**
|
|
|
|
- Its discovered context length (see above) is passed through as
|
|
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`, so it doesn't send a full-size prompt against a much
|
|
smaller real local context and overflow it.
|
|
- Its session runs with an isolated `CLAUDE_CONFIG_DIR`, so the injected API key never sits
|
|
in the same directory as a stored claude.ai login — that combination is harmless for actual
|
|
requests (the API key wins) but the CLI still prints a "both claude.ai and
|
|
ANTHROPIC_API_KEY set" warning about it, which this avoids entirely. The isolated directory
|
|
keeps a link back to your real session history so the response viewer and similar features
|
|
still work for that session. That isolated directory starts with no prior approvals of its
|
|
own, so Codeman also pre-approves the injected key the same way answering Claude Code's own
|
|
"Detected a custom API key" prompt once would — without it, that prompt would otherwise
|
|
reappear on every single launch with nobody there to answer it.
|
|
|
|
## Which harnesses actually work
|
|
|
|
| Harness | Status |
|
|
| ------- | ------ |
|
|
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
|
|
| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. |
|
|
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
|
|
| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. |
|
|
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
|
|
|
|
Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry,
|
|
not a fixed list here, so this table can go stale before this page does — a greyed-out or
|
|
missing entry is the more current answer.
|
|
|
|
## What it does not do
|
|
|
|
- **No remote or Docker sessions yet.** Both restart their agent differently under the hood
|
|
(reattaching a durable tmux session rather than relaunching the process), so redirecting
|
|
them needs its own plumbing that hasn't been built.
|
|
- **No live hot-swap mid-conversation.** Applying a selection always restarts the process.
|
|
- **No button to un-point a session from the UI yet.** Clearing back to native cloud is an
|
|
HTTP call (`POST .../custom-model {"clear": true}`) or deleting the session; the settings
|
|
panel manages saved endpoints, not what a running session is currently pointed at.
|
|
- **Nothing is shared with your real cloud credentials.** The endpoint's own key, if any,
|
|
never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate,
|
|
explicit choice per session.
|
|
|
|
## Security
|
|
|
|
An endpoint's base URL can't point at a link-local or cloud-metadata address (both at save
|
|
time and against the address it actually resolves to), the same guard Web Tabs uses for
|
|
saved dashboards. Endpoint records and any per-session config files a harness needs are
|
|
written with owner-only permissions. See
|
|
[custom-model-endpoints-plan.md](https://github.com/Ark0N/Codeman/blob/master/docs/custom-model-endpoints-plan.md)
|
|
in the repository for the full design reasoning, including why this feature closed a
|
|
pre-existing gap in how session environment overrides were guarded rather than opening a new
|
|
one.
|