docs(custom-model): bring CLAUDE.md and api-reference.md up to date

Full documentation review pass across the branch's 30 commits.
CLAUDE.md's Custom Model Endpoint Profiles entry hadn't been touched
since the initial backend+picker cut (3 early commits) despite 27
follow-up commits adding real behavior — it described restart-in-place
as universal (now claude-only; 7 other CLIs launch one-shot) and
claimed codex's Responses-API gap as a flat protocol break (now
re-verified as a more precise tool-calling gap). Corrected both and
added a new paragraph covering everything landed since: the llama-swap
conflict check, the after-the-fact swap-displacement sweep, the
/running-cmd-based context-length fix, the context-window floor
warning, skipFirstRunPrompts, the real-time /api/events-based log
status, and the countdown-to-Cancel-button change.

docs/api-reference.md's custom-model-endpoints section was missing the
running-status route, the requiresConfirmation/requiresContextWarning
response shapes, and POST /api/quick-start's customModel field
entirely (the primary launch path for 7 of 8 supported CLIs) — added
all three. Also fixed a real markdown bug in custom-model-endpoints.md:
an inline code span (`POST <baseUrl>/v1/chat/completions`) split across
a line break, which CommonMark renders with the line ending collapsed
to a space, so it displayed as ".../v1/chat/ completions" with a
spurious space inside the path.

Verified: origin/master and upstream/master are both already an
ancestor of this branch (identical at bd286bf5, no new commits since
this branch was cut) — nothing to merge, no conflicts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 13:18:04 +08:00
co-authored by Claude Sonnet 5
parent db9729e1fc
commit 8520925e76
4 changed files with 171 additions and 107 deletions
+3 -1
View File
File diff suppressed because one or more lines are too long
+153 -97
View File
@@ -66,17 +66,17 @@ The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
`src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on `src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on
the HTTP status. the HTTP status.
| `errorCode` | HTTP | Meaning | | `errorCode` | HTTP | Meaning |
|-------------|------|---------| | ------------------ | ---- | --------------------------------------------------- |
| `INVALID_INPUT` | 400 | Malformed request / failed validation | | `INVALID_INPUT` | 400 | Malformed request / failed validation |
| `UNAUTHORIZED` | 401 | Authentication required or failed | | `UNAUTHORIZED` | 401 | Authentication required or failed |
| `NOT_FOUND` | 404 | Resource does not exist | | `NOT_FOUND` | 404 | Resource does not exist |
| `SESSION_BUSY` | 409 | Session is busy | | `SESSION_BUSY` | 409 | Session is busy |
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) | | `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
| `ALREADY_EXISTS` | 409 | Resource already exists | | `ALREADY_EXISTS` | 409 | Resource already exists |
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed | | `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
| `RATE_LIMITED` | 429 | Too many requests | | `RATE_LIMITED` | 429 | Too many requests |
| `INTERNAL_ERROR` | 500 | Unexpected server error | | `INTERNAL_ERROR` | 500 | Unexpected server error |
Adding a new error code is non-breaking; removing or renaming one is a major change. Adding a new error code is non-breaking; removing or renaming one is a major change.
@@ -87,10 +87,10 @@ exist because SSE is Codeman's only other "tell me when" channel, and an agent
driving the API from a shell tool cannot practically hold a stream and parse driving the API from a shell tool cannot practically hold a stream and parse
events inline. events inline.
| Call | Blocks until | | Call | Blocks until |
|------|--------------| | --------------------------------------------- | -------------------------------------------------- |
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires | | `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output | | `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires | | `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
`POST .../input` with `wait` is not the same as a `POST` followed by a separate `POST .../input` with `wait` is not the same as a `POST` followed by a separate
@@ -140,13 +140,13 @@ contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
### Signals ### Signals
| Signal | Source | Actually fires for | | Signal | Source | Actually fires for |
|--------|--------|--------------------| | --------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) | | `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) | | `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only | | `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below | | `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode | | `exit` | no process is behind the session | every mode |
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic `stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
fallback that can flap mid-turn when a spinner pauses. The default set when `until` fallback that can flap mid-turn when a spinner pauses. The default set when `until`
@@ -156,12 +156,12 @@ can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit
once the session is up: the default set's `idle` also resolves on a spinner pause, once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up) and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured: can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no *further* `idle`, so it is the a session parked on the trust dialog emits no _further_ `idle`, so it is the
startup transition, not the dialog, that produces the false success below. startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The ⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died *inside* its tmux pane covers a session that exited — including a worker that died _inside_ its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
@@ -184,7 +184,7 @@ blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell ⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever, session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait whatever the pane is doing, so it never emits a _transition_. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there: requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
@@ -218,11 +218,11 @@ with `from=buffer` keeps matching long after the dialog is gone. A worked versio
### `GET /api/v1/sessions/:id/wait` ### `GET /api/v1/sessions/:id/wait`
| Param | Type | Default | Notes | | Param | Type | Default | Notes |
|-------|------|---------|-------| | --------- | -------------------------------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback | | `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` | | `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time | | `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
```bash ```bash
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000" curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
@@ -239,12 +239,12 @@ a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output` ### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes | | Param | Type | Default | Notes |
|-------|------|---------|-------| | --------- | ------------------------------- | -------- | ----------------------------------------------------------------------------------------------------------- |
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found | | `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing | | `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking | | `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` | | `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a **Matching is literal, never a pattern.** A `regex` parameter is rejected with a
`400` rather than ignored, so a caller that assumed otherwise finds out immediately `400` rather than ignored, so a caller that assumed otherwise finds out immediately
@@ -296,10 +296,10 @@ hand-written query string decodes to a space.
Two optional fields on the existing endpoint: Two optional fields on the existing endpoint:
| Field | Type | Notes | | Field | Type | Notes |
|-------|------|-------| | ------------- | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" | | `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` | | `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
"absent" rather than failing validation. That is deliberate: `.optional()` would "absent" rather than failing validation. That is deliberate: `.optional()` would
@@ -330,16 +330,24 @@ All three nest the wait result under `data.wait`, so one client helper works aga
any of them: any of them:
```json ```json
{ "success": true, "data": { {
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464", "success": true,
"status": "idle", "data": {
"limitPaused": false, "sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"wait": { "status": "idle",
"signal": "stop", "until": ["stop", "idle", "exit"], "limitPaused": false,
"timedOut": false, "immediate": false, "ended": false, "aborted": false, "wait": {
"waitedMs": 8421, "timeoutMs": 60000 "signal": "stop",
"until": ["stop", "idle", "exit"],
"timedOut": false,
"immediate": false,
"ended": false,
"aborted": false,
"waitedMs": 8421,
"timeoutMs": 60000
}
} }
}} }
``` ```
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`, `POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
@@ -353,21 +361,21 @@ redelivery (harmless, the turn it refers to may be long over), while with
client that reads `delivered === false` as "duplicate" silently treats a failed send client that reads `delivered === false` as "duplicate" silently treats a failed send
as a success. as a success.
| Field | Type | Meaning | | Field | Type | Meaning |
|-------|------|---------| | ---------------- | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) | | `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) | | `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) | | `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) | | `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) | | `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` | | `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) | | `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met | | `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome | | `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting | | `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied | | `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand | | `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard | | `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
Read the outcome by discriminator, in this order: Read the outcome by discriminator, in this order:
@@ -390,12 +398,12 @@ read the timeout as "the worker is wedged" and kill a session that was working f
### Errors ### Errors
| `errorCode` | HTTP | When | | `errorCode` | HTTP | When |
|-------------|------|------| | --------------- | ---- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` | | `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own | | `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full | | `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem | | `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
The two capacity codes are deliberately different. A process-wide cap reported as The two capacity codes are deliberately different. A process-wide cap reported as
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The `SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
@@ -446,9 +454,9 @@ Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md).
- `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first, - `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first,
ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId, ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId,
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?, sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
toolSummary?, message?, cwd?, context?, options?: {n, label}[], toolSummary?, message?, cwd?, context?, options?: {n, label}[],
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame; acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
`options` is present only when the dialog's numbered choices parsed `options` is present only when the dialog's numbered choices parsed
confidently; `acknowledgedAt` marks an item a human has already looked at confidently; `acknowledgedAt` marks an item a human has already looked at
(see `/viewed` below) and tells clients not to re-arm its tab alert. Listing (see `/viewed` below) and tells clients not to re-arm its tab alert. Listing
@@ -466,7 +474,7 @@ Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md).
first, `422 OPERATION_FAILED` when the session refused input. first, `422 OPERATION_FAILED` when the session refused input.
- `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes. - `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes.
- `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId, - `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId,
acknowledged: itemId | null }`. Marks the session's pending **idle** item as acknowledged: itemId | null }`. Marks the session's pending **idle** item as
seen by a human (the web UI calls it when you open the session's tab): the seen by a human (the web UI calls it when you open the session's tab): the
item stays pending and answerable, but stops arming the yellow tab alert on item stays pending and answerable, but stops arming the yellow tab alert on
every client, including after a reload. Permission/question items are never every client, including after a reload. Permission/question items are never
@@ -491,7 +499,7 @@ user guide: [`readmymind.md`](readmymind.md).
- `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the - `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the
session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals, session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals,
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
50, each <= 500 chars). A case with nothing recorded answers an empty 50, each <= 500 chars). A case with nothing recorded answers an empty
profile with `updatedAt: 0`; nothing is persisted by reads. profile with `updatedAt: 0`; nothing is persisted by reads.
- `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict - `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict
@@ -528,12 +536,12 @@ user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md).
- `GET /api/v1/model-endpoints` -> `CustomModelHost[]`, an unwrapped bare - `GET /api/v1/model-endpoints` -> `CustomModelHost[]`, an unwrapped bare
array like every other list route (still riding the standard `{success, array like every other list route (still riding the standard `{success,
data}` envelope on the wire — unwrap it the same way). Answers `[]` for a data}` envelope on the wire — unwrap it the same way). Answers `[]` for a
non-admin in multi-user mode. `apiKey` is never returned; `apiKeySet: non-admin in multi-user mode. `apiKey` is never returned; `apiKeySet:
boolean` reports whether one is stored, so a client can render "unchanged boolean` reports whether one is stored, so a client can render "unchanged
if left blank" without ever holding the real value. if left blank" without ever holding the real value.
- `POST /api/v1/model-endpoints` with `{ id, label, baseUrl, apiKey?, - `POST /api/v1/model-endpoints` with `{ id, label, baseUrl, apiKey?,
authStyle?, defaultModelId? }` creates one. `id` must match authStyle?, defaultModelId? }` creates one. `id` must match
`^[a-zA-Z0-9_-]+$`; `authStyle` is `bearer` (default) or `api-key`, never `^[a-zA-Z0-9_-]+$`; `authStyle` is `bearer` (default) or `api-key`, never
both (a real server hung indefinitely when sent both headers on one both (a real server hung indefinitely when sent both headers on one
request); `baseUrl` must be `http(s)`, carry no embedded credentials, and request); `baseUrl` must be `http(s)`, carry no embedded credentials, and
@@ -548,24 +556,72 @@ user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md).
- `DELETE /api/v1/model-endpoints/:id` removes one. - `DELETE /api/v1/model-endpoints/:id` removes one.
- `POST /api/v1/model-endpoints/:id/discover-models` fetches the endpoint's - `POST /api/v1/model-endpoints/:id/discover-models` fetches the endpoint's
own `GET /v1/models` and stores the result as `models`, updating own `GET /v1/models` and stores the result as `models`, updating
`lastDiscoveredAt`. A `defaultModelId` that no longer appears in the fresh `lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own
list is dropped rather than carried forward invalid. Failures answer response already reports loaded) `modelContextLengths` and `modelSizesGB`.
`502 OPERATION_FAILED` with the underlying connection error, or a named A `defaultModelId` that no longer appears in the fresh list is dropped
egress refusal if the resolved address turned out to be blocked. The same rather than carried forward invalid. Failures answer `502 OPERATION_FAILED`
refresh also runs automatically for every saved endpoint every 5 minutes with the underlying connection error, or a named egress refusal if the
in the background (`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, resolved address turned out to be blocked. The same refresh also runs
started from `server.ts`), so there is no route for triggering "refresh automatically for every saved endpoint every 5 minutes in the background
all" — one endpoint being unreachable on a cycle never blocks the others. (`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, started from
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId } | `server.ts`), so there is no route for triggering "refresh all" — one
{ clear: true }` applies (or clears) the session's selection and endpoint being unreachable on a cycle never blocks the others.
**restarts the session's CLI process in place** — every supported harness - `GET /api/v1/model-endpoints/:id/running-status` -> `{ isLlamaSwap,
reads its endpoint config at process start, never per turn, so there is running: [{model, state, cmd?}], logLine? }`, read-only, no admin gate
no live hot-swap. A Claude session resumes its existing conversation (any session owner who could already point a session at this endpoint can
across the restart; pi/omp/grok additionally get a forced `--model`/`-m` equally ask what it currently has loaded). `isLlamaSwap` is
value, since for those three the config file alone does not select it. feature-detected via the endpoint's own `GET /running` — a plain
`400 INVALID_INPUT` for a remote (SSH) or Docker session — both restart llama.cpp/OpenAI-compatible server has none and always answers `false`.
their agent differently under the hood, and applying to one would report `logLine`, present only when `isLlamaSwap` is true, is the most recent
success while changing nothing. REAL backend `llama-server` process log line (`load_model: ...`,
`llama_server: model loaded`, etc.), sourced from the endpoint's own
`GET /api/events` SSE stream and filtered to `source: "upstream"` frames
only (never llama-swap's own `source: "proxy"` request-access log) — one
connection is held open per endpoint and reused across every poller,
idle-closed after 30s of nobody asking. This is what the Run-menu
picker's loading banner polls once a second while a model is loading.
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId,
confirmed? } | { clear: true }` applies (or clears) the session's
selection and **restarts the session's CLI process in place** — every
supported harness reads its endpoint config at process start, never per
turn, so there is no live hot-swap. (`POST /api/v1/quick-start`'s own
`customModel: { endpointId, modelId, confirmed? }` field is the
no-restart equivalent for a session that doesn't exist yet — see below.)
A Claude session resumes its existing conversation across the restart;
pi/omp/grok additionally get a forced `--model`/`-m` value, since for
those three the config file alone does not select it. `400 INVALID_INPUT`
for a remote (SSH) or Docker session — both restart their agent
differently under the hood, and applying to one would report success
while changing nothing. Two more responses replace the normal
`{customModel, restarted}` shape, neither an error — both require
retrying the same call with `confirmed: true` to proceed anyway, and
neither restarts or creates anything on the first ask:
- `{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}` —
llama.cpp/llama-swap only runs one model at a time, and switching would
unload a model another **live session's own selection** is actively
using. Never returned for a plain (non-llama-swap) server, and never
just because a swap is needed at all — only when it would disrupt
someone else.
- `{requiresContextWarning: true, modelId, contextLength,
minSafeContextTokens}` — Claude Code's own fixed per-turn overhead
(system prompt + tool schemas) can exceed a small model's entire
discovered context on its own, before any conversation history exists
to compact, guaranteeing the very first message fails regardless of
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`. Gated on the CLI registry declaring a
`contextLengthVar` (claude only today), so it never fires for another
harness.
- `POST /api/v1/quick-start`'s `customModel: { endpointId, modelId,
confirmed? }` field (alongside its normal `caseName`/`mode`/etc. body)
computes the same injection **before** the session exists and launches
directly on the endpoint — no restart, because there was never a
native-backend boot to restart away from. Runs the identical checks as
the dedicated route above (`requiresConfirmation`/`requiresContextWarning`,
same shapes, same `confirmed: true` retry), and is refused the same way
for a remote or Docker case. This is what the Run-menu picker uses for
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; Claude still uses the
dedicated restart route above (its `--resume`-based restart is far less
jarring than a full relaunch, and folding it into the one-shot path is
separate work — see `docs/custom-model-endpoints-plan.md`).
## Voice dictation ## Voice dictation
@@ -575,7 +631,7 @@ same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synce
[`claude-voice-plan.md`](claude-voice-plan.md). [`claude-voice-plan.md`](claude-voice-plan.md).
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?, - `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
signed in to Claude Code on the server), `expired` (the access token elapsed; signed in to Claude Code on the server), `expired` (the access token elapsed;
running any Claude session refreshes it) or `malformed`. The OAuth token running any Claude session refreshes it) or `malformed`. The OAuth token
itself is never returned by this or any other endpoint. itself is never returned by this or any other endpoint.
+6 -2
View File
@@ -352,8 +352,12 @@ pure unit tests and the live manual checks in Verification:
up automatically with zero edits to the script). Already run to up automatically with zero edits to the script). Already run to
completion against the author's llama-swap server (a LAN address, completion against the author's llama-swap server (a LAN address,
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries): inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected** claude/opencode/pi/grok/omp **PASS**, codex **partially works and still
(Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED** isn't usable** (plain chat succeeds against a llama-swap deployment that
answers `/v1/responses`, but a real tool-call attempt comes back as
inert text rather than an executable `function_call` — see the
confidence table row for the full, re-verified picture), gemini/deepseek
**UNCONFIRMED**
(reach the server, fail for undiagnosed reasons — see their table rows), (reach the server, fail for undiagnosed reasons — see their table rows),
antigravity **SKIP** (no mechanism). Re-run this against a real cloud antigravity **SKIP** (no mechanism). Re-run this against a real cloud
endpoint (e.g. an Azure AI Foundry deployment) once one is available, to endpoint (e.g. an Azure AI Foundry deployment) once one is available, to
+9 -7
View File
@@ -12,11 +12,12 @@ recipe confidence table, and security reasoning:
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md). [`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
> **Status**: fully wired end to end — registry capability, the injection > **Status**: fully wired end to end — registry capability, the injection
> engine, the endpoint store + discovery route, the session restart route, > engine, the endpoint store + discovery route, both the restart-in-place
> a settings-panel CRUD surface, and the Run-menu picker described below. > apply route (Claude) and the one-shot quick-start launch path (every
> Antigravity has no known custom-endpoint mechanism and is not supported. > other supported harness), a settings-panel CRUD surface, and the Run-menu
> The HTTP API (examples below) still works directly and is what the picker > picker described below. Antigravity has no known custom-endpoint
> itself calls under the hood. > mechanism and is not supported. The HTTP API (examples below) still works
> directly and is what the picker itself calls under the hood.
## Turning it on ## Turning it on
@@ -351,8 +352,9 @@ which can take anywhere from a few seconds to well over a minute:
naming the model, and confirmed live: applying a selection alone never naming the model, and confirmed live: applying a selection alone never
reached llama-swap at all (nothing in its own server logs), since nothing reached llama-swap at all (nothing in its own server logs), since nothing
had actually asked it to load anything yet. Both apply routes now also had actually asked it to load anything yet. Both apply routes now also
send the smallest real request that will — `POST <baseUrl>/v1/chat/ send the smallest real request that will —
completions` with `max_tokens: 1` and one throwaway message — whenever the `POST <baseUrl>/v1/chat/completions` with `max_tokens: 1` and one
throwaway message — whenever the
target model isn't already the one loaded and ready, fire-and-forget (its target model isn't already the one loaded and ready, fire-and-forget (its
response is never read; `GET /api/model-endpoints/:id/running-status`, response is never read; `GET /api/model-endpoints/:id/running-status`,
polled client-side, is what actually confirms readiness). The response polled client-side, is what actually confirms readiness). The response