mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
The llama-swap conflict check on the apply/create routes only ever runs at THAT session's own launch/apply moment, and cannot see a swap caused by a DIFFERENT session's later, ordinary use. Confirmed live: a second Codex session picking a different model launched with no warning at all — nothing conflicted at that exact instant — yet it silently evicted the first session's model regardless (llama.cpp runs one model at a time). Reproduced and root-caused via direct API calls against a live test-picker instance rather than guessing. - detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups live sessions with a customModel by endpointId, checks each group's endpoint via GET /running once, and flags a session whose own modelId is no longer in the running list. Read-only, best-effort per endpoint like refreshAllCustomModelHosts's sibling sweep. - Notifies once per displacement via a caller-owned de-dupe Set: a session id is added when displaced, removed once its own model is loaded/ready again, so a later genuinely-new displacement can notify again. - New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS, 20s — much shorter than the 5-minute model-list refresh, since this is time-sensitive) broadcasts a new custom-model:swapped-out SSE event per displacement. De-dupe Set cleared per-session on session cleanup to avoid an unbounded leak. - Frontend: global toast (not tied to the displaced session's tab, since the point is warning before the user types into it) naming the session, its previous model, and what's currently loaded. Chose the "detect after the fact" scope (vs. checking before every message send, which would add a round-trip to every turn on every custom-model session) per explicit user decision after being presented the trade-off. 9 new tests for the detection logic (flag/clear/re-flag cycle, unreachable/deleted endpoints, non-llama-swap servers, multiple sessions on one endpoint). SSE registry bumped 158->159, parity test passing. Typecheck/lint/frontend-syntax clean; full suite shows no new regressions (9 more passing than baseline, matching the new tests; same pre-existing Windows-environment failures). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
175 lines
12 KiB
Markdown
175 lines
12 KiB
Markdown
# Custom Model Endpoints
|
||
|
||
Point a harness at your own OpenAI-compatible server instead of its native cloud backend, for
|
||
one session at a time. "Custom endpoint" covers **local** hardware (llama.cpp, Ollama, vLLM,
|
||
a home GPU rig, DGX Spark, Strix Halo) and **cloud** services (Azure AI Foundry's
|
||
OpenAI-compatible endpoint, OpenRouter, a company gateway) alike, anything answering
|
||
`GET /v1/models` and `POST /v1/chat/completions` in the standard shape.
|
||
|
||
**Off by default.** Turn it on in App Settings → Models → **Custom model endpoints**.
|
||
|
||
## Adding an endpoint
|
||
|
||
Still in App Settings → Models → Custom model endpoints:
|
||
|
||
1. **+ Add endpoint** — give it an id, a label, and the base URL (`http://192.168.1.50:8080`,
|
||
say). An API key is optional; most local servers don't check one.
|
||
2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it.
|
||
3. Pick a **default model** from what was discovered. This is the model the Run-menu entry
|
||
applies directly when only one model is discovered; with two or more, it's just the one
|
||
pre-marked in the picker dialog described below, not a silent default.
|
||
|
||
Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker
|
||
hosts — these are machine-level infra, not a per-user setting.
|
||
|
||
**Model lists refresh themselves.** Every saved endpoint is re-discovered automatically every
|
||
5 minutes in the background, so a model the server starts serving later — or stops serving —
|
||
shows up without another manual click of **Discover**. One endpoint being unreachable on a
|
||
given cycle (powered off, wrong network) never blocks the others from refreshing.
|
||
|
||
**Context length is picked up automatically where it can be, safely.** Against a
|
||
llama.cpp/llama-swap server, discovery also learns each _currently loaded_ model's real
|
||
context window and applies it to the launched session (Claude Code today — see below), so
|
||
the harness stops assuming a large default window for a model name it doesn't recognise and
|
||
overflowing a much smaller real one. It's deliberately never probed for a model that isn't
|
||
already loaded, since asking a llama-swap server about an unloaded model can trigger an
|
||
actual, slow model swap as a side effect — a model just not currently loaded keeps whatever
|
||
context length an earlier cycle already learned for it instead.
|
||
|
||
## Running a session against one
|
||
|
||
With the setting on and at least one endpoint carrying a discovered model, the **Run**
|
||
dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a
|
||
custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a
|
||
session on that harness exactly the way its own entry would. It is a one-off "try this
|
||
endpoint" action, not a sticky mode — the plain **Run** button still means "this harness,
|
||
native cloud" afterward, and a fresh session never inherits whatever the last one was
|
||
pointed at.
|
||
|
||
**Which model it uses depends on how many the endpoint has discovered.** With exactly one,
|
||
the session launches straight away on that model — nothing to choose. With two or more, a
|
||
small dialog asks which one to use for this launch before starting the session; the
|
||
endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately
|
||
use a different one without changing the saved default.
|
||
|
||
**For opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP, picking an entry launches
|
||
straight onto the endpoint** — no restart, because the endpoint is applied before the
|
||
session's process ever starts. **Claude still restarts the harness's process in place** —
|
||
same tab, same conversation (`--resume`) — after a normal native launch, since that restart
|
||
is far less jarring for Claude than for the other seven, whose own TUI can fully
|
||
reinitialize on a restart. Either way, every supported harness reads its endpoint config at
|
||
process start, never per turn, so there is no live hot-swap while a turn is running.
|
||
|
||
Picking an entry that launches a **brand-new** Claude session waits (up to 20 seconds) for it to
|
||
finish its own startup before applying — a freshly started CLI reports itself as busy for its
|
||
boot sequence, and applying to a genuinely busy session is refused so a real, in-progress
|
||
turn is never interrupted out from under you. A session that is still busy after that wait
|
||
(a very slow-starting CLI, or one you started typing into right away) surfaces that refusal
|
||
as an ordinary error, which now stays on screen with a close button instead of vanishing
|
||
after a few seconds — read it, it names the actual reason rather than a generic failure.
|
||
|
||
Entries are hidden entirely for a session in a **remote (SSH) or Docker case** — support for
|
||
redirecting those hasn't landed yet, see below. The picker also only appears in the desktop
|
||
**Run** dropdown; the phone home screen builds its own run picker separately and does not
|
||
currently offer these entries.
|
||
|
||
**Against llama-swap, applying a selection also starts the actual model load, rather than
|
||
waiting on your first prompt to do it.** llama-swap has no "switch model" button of its own
|
||
— the only thing that starts a swap is a real request naming the model, and confirmed live:
|
||
just applying a selection never reached llama-swap's own logs at all until something asked
|
||
it to load. Picking an entry now also sends the smallest real request that will trigger
|
||
that load, in the background, the moment the target model isn't already loaded and ready.
|
||
|
||
**The centred loading banner shows a live countdown, and a real timeout is an error, not a
|
||
shrug.** When it knows the model's discovered file size (its GB figure, when llama-swap
|
||
states one), it shows both a rough expected-time estimate and a live countdown against it —
|
||
e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1–3 min) on llama-swap — 47s remaining". If
|
||
the countdown reaches zero and the model still isn't ready, the banner turns into a sticky
|
||
error telling you to check the llama-swap server's own logs, and **the session that load was
|
||
for is closed automatically** — a console left open and pointed at a model that never
|
||
finished loading would just be confusing to leave sitting there.
|
||
|
||
**You'll also be told if a session's model gets swapped out from under it later, not just
|
||
at launch.** The conflict warning above only fires at the moment you launch or apply a
|
||
model — llama.cpp only runs one model at a time, so if a DIFFERENT session using the same
|
||
endpoint later triggers its own load, whatever was loaded before (including a session you
|
||
already had running) gets silently evicted, with no warning at that instant since nothing
|
||
conflicted when it was first set up. A background check (every 20 seconds) catches this
|
||
after the fact and shows a toast naming which session lost its model and what's loaded now
|
||
— so you know before typing into that session that it's about to reload (and, in turn,
|
||
evict whatever displaced it).
|
||
|
||
**Claude Code specifically gets three extra fixes applied automatically:**
|
||
|
||
- Its discovered context length (see above) is passed through as
|
||
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`, so it doesn't send a full-size prompt against a much
|
||
smaller real local context and overflow it.
|
||
- Its session runs with an isolated `CLAUDE_CONFIG_DIR`, so the injected API key never sits
|
||
in the same directory as a stored claude.ai login — that combination is harmless for actual
|
||
requests (the API key wins) but the CLI still prints a "both claude.ai and
|
||
ANTHROPIC_API_KEY set" warning about it, which this avoids entirely. The isolated directory
|
||
keeps a link back to your real session history so the response viewer and similar features
|
||
still work for that session. That isolated directory starts with no prior approvals of its
|
||
own, so Codeman also pre-approves the injected key the same way answering Claude Code's own
|
||
"Detected a custom API key" prompt once would — without it, that prompt would otherwise
|
||
reappear on every single launch with nobody there to answer it.
|
||
- **That same fresh isolated directory also looks like a brand-new Claude Code profile**, so
|
||
without this fix it replayed the WHOLE first-run sequence every single launch: the theme
|
||
picker, the security-notes screen, the "trust this folder?" dialog, and a one-time warning
|
||
about running with permissions bypassed — none of which a real, already-used profile shows
|
||
again. Codeman now pre-seeds that same "already been through this once" state (onboarding
|
||
completed, this session's own project marked trusted, the bypass-permissions warning
|
||
acknowledged) so a custom-model launch reaches the actual conversation exactly as fast as a
|
||
native cloud one does, instead of stopping at a wizard with nobody there to click through it.
|
||
|
||
**If a model's real context is too small for Claude Code to even get started, you get a
|
||
warning instead of a confusing failure.** Claude Code's own system prompt and tools take up
|
||
roughly 40K tokens on their own, before you've typed anything — a small local model with a
|
||
smaller real context than that fails outright on the very first message, no matter what
|
||
context size Codeman tells it to expect (raising the declared context only changes when
|
||
Claude Code trims _conversation history_, and there is none yet on message one). Picking
|
||
such a model now shows an in-app dialog naming the model, its discovered context and what's
|
||
needed, before anything launches or restarts, with the fix spelled out: reconfigure
|
||
llama-swap to give that model (or a smaller one) an explicit larger context instead of
|
||
relying on auto-fit (`--fit-ctx`), which sizes the context around fitting the biggest model
|
||
rather than the biggest context — for example adding `-c 65536` to that model's llama-swap
|
||
entry. "Launch anyway" is still there if you want to try regardless.
|
||
|
||
## Which harnesses actually work
|
||
|
||
| Harness | Status |
|
||
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
|
||
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
|
||
| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. |
|
||
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
|
||
| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. |
|
||
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
|
||
|
||
Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry,
|
||
not a fixed list here, so this table can go stale before this page does — a greyed-out or
|
||
missing entry is the more current answer.
|
||
|
||
## What it does not do
|
||
|
||
- **No remote or Docker sessions yet.** Both restart their agent differently under the hood
|
||
(reattaching a durable tmux session rather than relaunching the process), so redirecting
|
||
them needs its own plumbing that hasn't been built.
|
||
- **No live hot-swap mid-conversation.** Applying a selection always restarts the process.
|
||
- **No button to un-point a session from the UI yet.** Clearing back to native cloud is an
|
||
HTTP call (`POST .../custom-model {"clear": true}`) or deleting the session; the settings
|
||
panel manages saved endpoints, not what a running session is currently pointed at.
|
||
- **Nothing is shared with your real cloud credentials.** The endpoint's own key, if any,
|
||
never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate,
|
||
explicit choice per session.
|
||
|
||
## Security
|
||
|
||
An endpoint's base URL can't point at a link-local or cloud-metadata address (both at save
|
||
time and against the address it actually resolves to), the same guard Web Tabs uses for
|
||
saved dashboards. Endpoint records and any per-session config files a harness needs are
|
||
written with owner-only permissions. See
|
||
[custom-model-endpoints-plan.md](https://github.com/Ark0N/Codeman/blob/master/docs/custom-model-endpoints-plan.md)
|
||
in the repository for the full design reasoning, including why this feature closed a
|
||
pre-existing gap in how session environment overrides were guarded rather than opening a new
|
||
one.
|