docs(custom-model): bring CLAUDE.md and api-reference.md up to date

Full documentation review pass across the branch's 30 commits.
CLAUDE.md's Custom Model Endpoint Profiles entry hadn't been touched
since the initial backend+picker cut (3 early commits) despite 27
follow-up commits adding real behavior — it described restart-in-place
as universal (now claude-only; 7 other CLIs launch one-shot) and
claimed codex's Responses-API gap as a flat protocol break (now
re-verified as a more precise tool-calling gap). Corrected both and
added a new paragraph covering everything landed since: the llama-swap
conflict check, the after-the-fact swap-displacement sweep, the
/running-cmd-based context-length fix, the context-window floor
warning, skipFirstRunPrompts, the real-time /api/events-based log
status, and the countdown-to-Cancel-button change.

docs/api-reference.md's custom-model-endpoints section was missing the
running-status route, the requiresConfirmation/requiresContextWarning
response shapes, and POST /api/quick-start's customModel field
entirely (the primary launch path for 7 of 8 supported CLIs) — added
all three. Also fixed a real markdown bug in custom-model-endpoints.md:
an inline code span (`POST <baseUrl>/v1/chat/completions`) split across
a line break, which CommonMark renders with the line ending collapsed
to a space, so it displayed as ".../v1/chat/ completions" with a
spurious space inside the path.

Verified: origin/master and upstream/master are both already an
ancestor of this branch (identical at bd286bf5, no new commits since
this branch was cut) — nothing to merge, no conflicts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 13:18:04 +08:00
co-authored by Claude Sonnet 5
parent db9729e1fc
commit 8520925e76
4 changed files with 171 additions and 107 deletions
+3 -1
View File
File diff suppressed because one or more lines are too long
+90 -34
View File
@@ -67,7 +67,7 @@ The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
the HTTP status.
| `errorCode` | HTTP | Meaning |
|-------------|------|---------|
| ------------------ | ---- | --------------------------------------------------- |
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
| `UNAUTHORIZED` | 401 | Authentication required or failed |
| `NOT_FOUND` | 404 | Resource does not exist |
@@ -88,7 +88,7 @@ driving the API from a shell tool cannot practically hold a stream and parse
events inline.
| Call | Blocks until |
|------|--------------|
| --------------------------------------------- | -------------------------------------------------- |
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
@@ -141,7 +141,7 @@ contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
### Signals
| Signal | Source | Actually fires for |
|--------|--------|--------------------|
| --------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
@@ -156,12 +156,12 @@ can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit
once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no *further* `idle`, so it is the
a session parked on the trust dialog emits no _further_ `idle`, so it is the
startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died *inside* its tmux pane
covers a session that exited — including a worker that died _inside_ its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
@@ -184,7 +184,7 @@ blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
whatever the pane is doing, so it never emits a _transition_. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
@@ -219,7 +219,7 @@ with `from=buffer` keeps matching long after the dialog is gone. A worked versio
### `GET /api/v1/sessions/:id/wait`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| --------- | -------------------------------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
@@ -240,7 +240,7 @@ a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| --------- | ------------------------------- | -------- | ----------------------------------------------------------------------------------------------------------- |
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
@@ -297,7 +297,7 @@ hand-written query string decodes to a space.
Two optional fields on the existing endpoint:
| Field | Type | Notes |
|-------|------|-------|
| ------------- | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
@@ -330,16 +330,24 @@ All three nest the wait result under `data.wait`, so one client helper works aga
any of them:
```json
{ "success": true, "data": {
{
"success": true,
"data": {
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"status": "idle",
"limitPaused": false,
"wait": {
"signal": "stop", "until": ["stop", "idle", "exit"],
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
"waitedMs": 8421, "timeoutMs": 60000
"signal": "stop",
"until": ["stop", "idle", "exit"],
"timedOut": false,
"immediate": false,
"ended": false,
"aborted": false,
"waitedMs": 8421,
"timeoutMs": 60000
}
}
}
}}
```
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
@@ -354,7 +362,7 @@ client that reads `delivered === false` as "duplicate" silently treats a failed
as a success.
| Field | Type | Meaning |
|-------|------|---------|
| ---------------- | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
@@ -391,7 +399,7 @@ read the timeout as "the worker is wedged" and kill a session that was working f
### Errors
| `errorCode` | HTTP | When |
|-------------|------|------|
| --------------- | ---- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
@@ -548,24 +556,72 @@ user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md).
- `DELETE /api/v1/model-endpoints/:id` removes one.
- `POST /api/v1/model-endpoints/:id/discover-models` fetches the endpoint's
own `GET /v1/models` and stores the result as `models`, updating
`lastDiscoveredAt`. A `defaultModelId` that no longer appears in the fresh
list is dropped rather than carried forward invalid. Failures answer
`502 OPERATION_FAILED` with the underlying connection error, or a named
egress refusal if the resolved address turned out to be blocked. The same
refresh also runs automatically for every saved endpoint every 5 minutes
in the background (`refreshAllCustomModelHosts()`, `custom-model-routes.ts`,
started from `server.ts`), so there is no route for triggering "refresh
all" — one endpoint being unreachable on a cycle never blocks the others.
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId } |
{ clear: true }` applies (or clears) the session's selection and
**restarts the session's CLI process in place** — every supported harness
reads its endpoint config at process start, never per turn, so there is
no live hot-swap. A Claude session resumes its existing conversation
across the restart; pi/omp/grok additionally get a forced `--model`/`-m`
value, since for those three the config file alone does not select it.
`400 INVALID_INPUT` for a remote (SSH) or Docker session — both restart
their agent differently under the hood, and applying to one would report
success while changing nothing.
`lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own
response already reports loaded) `modelContextLengths` and `modelSizesGB`.
A `defaultModelId` that no longer appears in the fresh list is dropped
rather than carried forward invalid. Failures answer `502 OPERATION_FAILED`
with the underlying connection error, or a named egress refusal if the
resolved address turned out to be blocked. The same refresh also runs
automatically for every saved endpoint every 5 minutes in the background
(`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, started from
`server.ts`), so there is no route for triggering "refresh all" — one
endpoint being unreachable on a cycle never blocks the others.
- `GET /api/v1/model-endpoints/:id/running-status` -> `{ isLlamaSwap,
running: [{model, state, cmd?}], logLine? }`, read-only, no admin gate
(any session owner who could already point a session at this endpoint can
equally ask what it currently has loaded). `isLlamaSwap` is
feature-detected via the endpoint's own `GET /running` — a plain
llama.cpp/OpenAI-compatible server has none and always answers `false`.
`logLine`, present only when `isLlamaSwap` is true, is the most recent
REAL backend `llama-server` process log line (`load_model: ...`,
`llama_server: model loaded`, etc.), sourced from the endpoint's own
`GET /api/events` SSE stream and filtered to `source: "upstream"` frames
only (never llama-swap's own `source: "proxy"` request-access log) — one
connection is held open per endpoint and reused across every poller,
idle-closed after 30s of nobody asking. This is what the Run-menu
picker's loading banner polls once a second while a model is loading.
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId,
confirmed? } | { clear: true }` applies (or clears) the session's
selection and **restarts the session's CLI process in place** — every
supported harness reads its endpoint config at process start, never per
turn, so there is no live hot-swap. (`POST /api/v1/quick-start`'s own
`customModel: { endpointId, modelId, confirmed? }` field is the
no-restart equivalent for a session that doesn't exist yet — see below.)
A Claude session resumes its existing conversation across the restart;
pi/omp/grok additionally get a forced `--model`/`-m` value, since for
those three the config file alone does not select it. `400 INVALID_INPUT`
for a remote (SSH) or Docker session — both restart their agent
differently under the hood, and applying to one would report success
while changing nothing. Two more responses replace the normal
`{customModel, restarted}` shape, neither an error — both require
retrying the same call with `confirmed: true` to proceed anyway, and
neither restarts or creates anything on the first ask:
- `{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}` —
llama.cpp/llama-swap only runs one model at a time, and switching would
unload a model another **live session's own selection** is actively
using. Never returned for a plain (non-llama-swap) server, and never
just because a swap is needed at all — only when it would disrupt
someone else.
- `{requiresContextWarning: true, modelId, contextLength,
minSafeContextTokens}` — Claude Code's own fixed per-turn overhead
(system prompt + tool schemas) can exceed a small model's entire
discovered context on its own, before any conversation history exists
to compact, guaranteeing the very first message fails regardless of
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`. Gated on the CLI registry declaring a
`contextLengthVar` (claude only today), so it never fires for another
harness.
- `POST /api/v1/quick-start`'s `customModel: { endpointId, modelId,
confirmed? }` field (alongside its normal `caseName`/`mode`/etc. body)
computes the same injection **before** the session exists and launches
directly on the endpoint — no restart, because there was never a
native-backend boot to restart away from. Runs the identical checks as
the dedicated route above (`requiresConfirmation`/`requiresContextWarning`,
same shapes, same `confirmed: true` retry), and is refused the same way
for a remote or Docker case. This is what the Run-menu picker uses for
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; Claude still uses the
dedicated restart route above (its `--resume`-based restart is far less
jarring than a full relaunch, and folding it into the one-shot path is
separate work — see `docs/custom-model-endpoints-plan.md`).
## Voice dictation
+6 -2
View File
@@ -352,8 +352,12 @@ pure unit tests and the live manual checks in Verification:
up automatically with zero edits to the script). Already run to
completion against the author's llama-swap server (a LAN address,
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected**
(Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED**
claude/opencode/pi/grok/omp **PASS**, codex **partially works and still
isn't usable** (plain chat succeeds against a llama-swap deployment that
answers `/v1/responses`, but a real tool-call attempt comes back as
inert text rather than an executable `function_call` — see the
confidence table row for the full, re-verified picture), gemini/deepseek
**UNCONFIRMED**
(reach the server, fail for undiagnosed reasons — see their table rows),
antigravity **SKIP** (no mechanism). Re-run this against a real cloud
endpoint (e.g. an Azure AI Foundry deployment) once one is available, to
+9 -7
View File
@@ -12,11 +12,12 @@ recipe confidence table, and security reasoning:
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
> **Status**: fully wired end to end — registry capability, the injection
> engine, the endpoint store + discovery route, the session restart route,
> a settings-panel CRUD surface, and the Run-menu picker described below.
> Antigravity has no known custom-endpoint mechanism and is not supported.
> The HTTP API (examples below) still works directly and is what the picker
> itself calls under the hood.
> engine, the endpoint store + discovery route, both the restart-in-place
> apply route (Claude) and the one-shot quick-start launch path (every
> other supported harness), a settings-panel CRUD surface, and the Run-menu
> picker described below. Antigravity has no known custom-endpoint
> mechanism and is not supported. The HTTP API (examples below) still works
> directly and is what the picker itself calls under the hood.
## Turning it on
@@ -351,8 +352,9 @@ which can take anywhere from a few seconds to well over a minute:
naming the model, and confirmed live: applying a selection alone never
reached llama-swap at all (nothing in its own server logs), since nothing
had actually asked it to load anything yet. Both apply routes now also
send the smallest real request that will — `POST <baseUrl>/v1/chat/
completions` with `max_tokens: 1` and one throwaway message — whenever the
send the smallest real request that will —
`POST <baseUrl>/v1/chat/completions` with `max_tokens: 1` and one
throwaway message — whenever the
target model isn't already the one loaded and ready, fire-and-forget (its
response is never read; `GET /api/model-endpoints/:id/running-status`,
polled client-side, is what actually confirms readiness). The response