fix(custom-model): address second pre-merge review (Ark0N)

Blocker: .center-status-banner never actually disappears.

- Add `.center-status-banner[hidden] { display: none; }`, same trap as
  `.home-sessions[hidden]`: the author-level `display: flex` beat the
  UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
  card stayed laid out at `opacity: 0` with its text/cancel/close
  children still `pointer-events: auto` -- an invisible 442x67 click
  blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
  banner (10001) and the swap-confirm/context-warning modals (10010)
  in CLAUDE.md's Z-index layers list.

Stale wording pointed at the reverted sticky-toast default:

- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
  `.toast-message` comment in styles.css all still said "toasts
  default to sticky" after 1f32128c put the flat 3s default back.
  Reworded all three to describe the actual behaviour: one call site
  passes an explicit `duration: 0`.

Smaller items from the same review:

- docs/api-reference.md said discovery failures answer
  `502 OPERATION_FAILED`; OPERATION_FAILED is 422 per src/types/api.ts
  and the error-code table earlier in the same file.
- The periodic re-discovery sweep (server.ts) never read
  customModelEndpointsEnabled, so turning the feature off left
  Codeman polling every saved endpoint forever. Added
  readCustomModelEndpointsEnabled() (custom-model-routes.ts, same
  shape as readPlanUsageTelemetryEnabled) and gated the interval
  callback on it.
- Reverted the formatting-only Prettier pass docs/api-reference.md
  picked up (table padding, *x* to _x_, JSON re-indent) by re-merging
  the new Custom Model Endpoints section onto the pre-PR file, so the
  diff is reviewable. No prose content was lost -- verified by diffing
  the result against the pre-revert file (formatting-only) and against
  the merge-base file (only the new section added).
- docs/custom-model-endpoints.md now states that a custom-model Claude
  session's isolated CLAUDE_CONFIG_DIR loses the user's global
  settings.json, user-level skills/agents/commands, and MCP servers
  from ~/.claude.json -- only `projects` is symlinked back.

Design question left open in the review (does `confirmed: true` need
to be two flags so "launch anyway" on the context warning doesn't also
skip the llama-swap displacement warning): keeping the single flag, as
offered. The 20s displacement sweep still catches a resulting swap
after the fact, so it's a surprise rather than a silent failure, and
splitting it is real behavioural surface I have no way to verify live
in this environment.

`npm run test:browser` could not be run in this environment (no tmux,
no downloaded Playwright browser binary) -- none of its suite's files
touch code this fix changes, but it still needs a real pass before
merge, same as any frontend change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
This commit is contained in:
Devvyn
2026-09-18 21:45:50 +08:00
co-authored by Claude Sonnet 5
parent 1f32128ca9
commit 9982a1325f
9 changed files with 136 additions and 95 deletions
+1 -1
View File
@@ -7,7 +7,7 @@
Everything below was found and fixed against a **real llama-swap server**, not just unit tests:
- **Session-busy false refusal.** A freshly launched CLI reports itself `busy` for its own startup (spinner, workspace-trust check) well before the apply call would reach it, and the apply route correctly refuses to restart a session mid-turn — indistinguishable from a fresh boot. The picker now waits for the new session to go idle (bounded at 20s, never an error on timeout) before applying.
- **Errors and confirmations you can actually read.** Toasts now default to sticky with a close button (errors always were meant to stay, but a fixed 3s timer silently hid them); a failed apply's real server-side reason (not a generic message) reaches the toast.
- **Errors and confirmations you can actually read.** A failed apply's real server-side reason (not a generic message) reaches the toast, and that specific message stays on screen with a close button instead of vanishing on the usual 3s timer.
- **"Both claude.ai and ANTHROPIC_API_KEY set" warning.** A custom-model Claude session now runs with an isolated `CLAUDE_CONFIG_DIR` (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — `projects` is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on _every_ launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would.
- **Context-window overflow.** Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length and applies it as `CLAUDE_CODE_MAX_CONTEXT_TOKENS` — sourced primarily from llama-swap's own `GET /running`, whose `cmd` field carries the launch flags (`--fit-ctx`/`-c`/`--ctx-size`) actually in effect, since `GET /props`'s `n_ctx` was confirmed live to report the model's theoretical/trained maximum rather than the real `--fit-ctx`-shrunk runtime context (a 154112-vs-16384 discrepancy, caught only because the fixed value still overflowed) — `/props` is now a fallback for a plain llama.cpp server with no `/running` at all.
- **Context floor too small for Claude Code to even start.** Fixing the overflow above surfaced a second, unfixable-by-injection failure: Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on their own (confirmed live via an `in:0 out:0` failure on the very first message), which can exceed a small model's entire real context before any conversation history exists to trim — no `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes that, since it only governs when history gets compacted. Applying such a model now returns a warning (gated on the CLI registry declaring a `contextLengthVar`, so it's a no-op for every other harness) instead of launching straight into a guaranteed first-message failure, and the Run-menu picker shows it as an in-app dialog naming the model, its discovered context and the ~40K safe floor, with the actual fix spelled out: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on auto-fit, which optimizes for the biggest model that fits rather than the biggest context. "Launch anyway" is still one click away.
+2 -2
View File
File diff suppressed because one or more lines are too long
+77 -85
View File
@@ -66,17 +66,17 @@ The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
`src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on
the HTTP status.
| `errorCode` | HTTP | Meaning |
| ------------------ | ---- | --------------------------------------------------- |
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
| `UNAUTHORIZED` | 401 | Authentication required or failed |
| `NOT_FOUND` | 404 | Resource does not exist |
| `SESSION_BUSY` | 409 | Session is busy |
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
| `ALREADY_EXISTS` | 409 | Resource already exists |
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
| `RATE_LIMITED` | 429 | Too many requests |
| `INTERNAL_ERROR` | 500 | Unexpected server error |
| `errorCode` | HTTP | Meaning |
|-------------|------|---------|
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
| `UNAUTHORIZED` | 401 | Authentication required or failed |
| `NOT_FOUND` | 404 | Resource does not exist |
| `SESSION_BUSY` | 409 | Session is busy |
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
| `ALREADY_EXISTS` | 409 | Resource already exists |
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
| `RATE_LIMITED` | 429 | Too many requests |
| `INTERNAL_ERROR` | 500 | Unexpected server error |
Adding a new error code is non-breaking; removing or renaming one is a major change.
@@ -87,10 +87,10 @@ exist because SSE is Codeman's only other "tell me when" channel, and an agent
driving the API from a shell tool cannot practically hold a stream and parse
events inline.
| Call | Blocks until |
| --------------------------------------------- | -------------------------------------------------- |
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| Call | Blocks until |
|------|--------------|
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
@@ -140,13 +140,13 @@ contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
### Signals
| Signal | Source | Actually fires for |
| --------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode |
| Signal | Source | Actually fires for |
|--------|--------|--------------------|
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode |
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
@@ -156,12 +156,12 @@ can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit
once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no _further_ `idle`, so it is the
a session parked on the trust dialog emits no *further* `idle`, so it is the
startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died _inside_ its tmux pane
covers a session that exited — including a worker that died *inside* its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
@@ -184,7 +184,7 @@ blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a _transition_. Since send-and-wait
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
@@ -218,11 +218,11 @@ with `from=buffer` keeps matching long after the dialog is gone. A worked versio
### `GET /api/v1/sessions/:id/wait`
| Param | Type | Default | Notes |
| --------- | -------------------------------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
```bash
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
@@ -239,12 +239,12 @@ a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes |
| --------- | ------------------------------- | -------- | ----------------------------------------------------------------------------------------------------------- |
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
@@ -296,10 +296,10 @@ hand-written query string decodes to a space.
Two optional fields on the existing endpoint:
| Field | Type | Notes |
| ------------- | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
| Field | Type | Notes |
|-------|------|-------|
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
"absent" rather than failing validation. That is deliberate: `.optional()` would
@@ -330,24 +330,16 @@ All three nest the wait result under `data.wait`, so one client helper works aga
any of them:
```json
{
"success": true,
"data": {
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"status": "idle",
"limitPaused": false,
"wait": {
"signal": "stop",
"until": ["stop", "idle", "exit"],
"timedOut": false,
"immediate": false,
"ended": false,
"aborted": false,
"waitedMs": 8421,
"timeoutMs": 60000
}
{ "success": true, "data": {
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"status": "idle",
"limitPaused": false,
"wait": {
"signal": "stop", "until": ["stop", "idle", "exit"],
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
"waitedMs": 8421, "timeoutMs": 60000
}
}
}}
```
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
@@ -361,21 +353,21 @@ redelivery (harmless, the turn it refers to may be long over), while with
client that reads `delivered === false` as "duplicate" silently treats a failed send
as a success.
| Field | Type | Meaning |
| ---------------- | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
| Field | Type | Meaning |
|-------|------|---------|
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
Read the outcome by discriminator, in this order:
@@ -398,12 +390,12 @@ read the timeout as "the worker is wedged" and kill a session that was working f
### Errors
| `errorCode` | HTTP | When |
| --------------- | ---- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
| `errorCode` | HTTP | When |
|-------------|------|------|
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
The two capacity codes are deliberately different. A process-wide cap reported as
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
@@ -454,9 +446,9 @@ Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md).
- `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first,
ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId,
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
toolSummary?, message?, cwd?, context?, options?: {n, label}[],
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
toolSummary?, message?, cwd?, context?, options?: {n, label}[],
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
`options` is present only when the dialog's numbered choices parsed
confidently; `acknowledgedAt` marks an item a human has already looked at
(see `/viewed` below) and tells clients not to re-arm its tab alert. Listing
@@ -474,7 +466,7 @@ acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
first, `422 OPERATION_FAILED` when the session refused input.
- `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes.
- `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId,
acknowledged: itemId | null }`. Marks the session's pending **idle** item as
acknowledged: itemId | null }`. Marks the session's pending **idle** item as
seen by a human (the web UI calls it when you open the session's tab): the
item stays pending and answerable, but stops arming the yellow tab alert on
every client, including after a reload. Permission/question items are never
@@ -499,7 +491,7 @@ user guide: [`readmymind.md`](readmymind.md).
- `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the
session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals,
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
50, each <= 500 chars). A case with nothing recorded answers an empty
profile with `updatedAt: 0`; nothing is persisted by reads.
- `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict
@@ -559,7 +551,7 @@ authStyle?, defaultModelId? }` creates one. `id` must match
`lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own
response already reports loaded) `modelContextLengths` and `modelSizesGB`.
A `defaultModelId` that no longer appears in the fresh list is dropped
rather than carried forward invalid. Failures answer `502 OPERATION_FAILED`
rather than carried forward invalid. Failures answer `422 OPERATION_FAILED`
with the underlying connection error, or a named egress refusal if the
resolved address turned out to be blocked. The same refresh also runs
automatically for every saved endpoint every 5 minutes in the background
@@ -631,7 +623,7 @@ same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synce
[`claude-voice-plan.md`](claude-voice-plan.md).
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
signed in to Claude Code on the server), `expired` (the access token elapsed;
running any Claude session refreshes it) or `malformed`. The OAuth token
itself is never returned by this or any other endpoint.
+7 -1
View File
@@ -290,7 +290,13 @@ size`. No entry for the model in `modelContextLengths` means the var is
[`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied
automatically here. Best-effort: a platform that refuses the symlink keeps
the pre-existing blind-response-viewer side effect rather than failing the
whole custom-model apply over it.
whole custom-model apply over it. ⚠️ **This relocates the whole `.claude`
tree, not just transcripts**: a custom-model Claude session also loses the
user's global `settings.json`, user-level skills (the codeman agent skill
included), user-level agents and commands, and the MCP servers configured
in `~/.claude.json` — none of those are symlinked back, only `projects` is.
A fine trade for "point this session at my local llama.cpp," but worth
knowing before it surprises you mid-session.
**That isolated directory needed one more fix to actually be usable
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
+11 -2
View File
@@ -8526,8 +8526,8 @@ kbd {
.toast-message {
flex: 1;
/* Errors are sticky by default (showToast) precisely so a longer, specific
message survives to be read — let it wrap instead of clipping. */
/* A sticky toast (showToast's opts.duration: 0) can carry a longer, specific
message — let it wrap instead of clipping. */
white-space: pre-wrap;
word-break: break-word;
}
@@ -8588,6 +8588,15 @@ kbd {
transform: translate(-50%, -50%) scale(1);
}
/* `hidden` has to be re-asserted over the `display: flex` above, or `dismiss()`
setting `el.hidden = true` does nothing (same trap as `.home-sessions[hidden]`
below): the card stays laid out at `opacity: 0` with its text/cancel/close
children still `pointer-events: auto`, an invisible click-blocker dead centre
over the terminal until the page reloads. */
.center-status-banner[hidden] {
display: none;
}
.center-status-spinner {
flex-shrink: 0;
width: 18px;
+14 -1
View File
@@ -16,7 +16,7 @@
import type { FastifyInstance, FastifyRequest } from 'fastify';
import { ApiErrorCode, createErrorResponse, type ApiResponse } from '../../types.js';
import { isAdmin, parseBody } from '../route-helpers.js';
import { isAdmin, parseBody, readJsonConfig, SETTINGS_PATH } from '../route-helpers.js';
import { isMultiUserMode } from '../../config/multiuser.js';
import { getDataDir } from '../../config/instance.js';
import { isBlockedWebviewUrl } from '../webview-egress-policy.js';
@@ -565,6 +565,19 @@ function applyDiscoveredModels(host: CustomModelHost, result: DiscoveryResult):
};
}
/**
* `customModelEndpointsEnabled` defaults OFF (unlike `showPlanUsageLimits`'s
* absent-means-on in `readPlanUsageTelemetryEnabled`), so mirror the frontend's
* own gate (`session-ui.js`'s `!settings.customModelEndpointsEnabled`) rather
* than that reader's default. Exists so the periodic re-discovery sweep in
* server.ts can skip entirely while the feature is off, instead of polling
* every saved endpoint forever regardless of the setting.
*/
export async function readCustomModelEndpointsEnabled(): Promise<boolean> {
const settings = await readJsonConfig<Record<string, unknown>>(SETTINGS_PATH, 'settings.json', {});
return settings.customModelEndpointsEnabled === true;
}
/**
* Re-discovers every saved endpoint's models, best-effort. One endpoint being
* unreachable (powered off, wrong network) must not stop the others from
+1
View File
@@ -30,6 +30,7 @@ export { registerTabLayoutRoutes } from './tab-layout-routes.js';
export {
registerCustomModelRoutes,
refreshAllCustomModelHosts,
readCustomModelEndpointsEnabled,
detectCustomModelSwapDisplacements,
pruneIdleLlamaSwapLogTails,
type CustomModelSessionLike,
+13 -3
View File
@@ -191,6 +191,7 @@ import {
registerTabLayoutRoutes,
registerCustomModelRoutes,
refreshAllCustomModelHosts,
readCustomModelEndpointsEnabled,
detectCustomModelSwapDisplacements,
pruneIdleLlamaSwapLogTails,
tryWebviewRefererFallback,
@@ -2761,9 +2762,18 @@ export class WebServer extends EventEmitter {
if (!this.testMode) {
this.cleanup.setInterval(
() => {
refreshAllCustomModelHosts().catch((err) => {
console.error('[custom-model] periodic re-discovery failed:', getErrorMessage(err));
});
// Reads the setting fresh on every tick, same reasoning as
// readPlanUsageTelemetryEnabled() beside it: a live toggle takes effect
// on the very next cycle, not just at server boot, and turning the
// feature off actually stops the polling instead of only hiding the UI.
void readCustomModelEndpointsEnabled()
.then((enabled) => {
if (!enabled) return;
return refreshAllCustomModelHosts();
})
.catch((err) => {
console.error('[custom-model] periodic re-discovery failed:', getErrorMessage(err));
});
},
CUSTOM_MODEL_REDISCOVER_INTERVAL_MS,
{ description: 'custom model endpoint re-discovery' }
+10
View File
@@ -969,6 +969,16 @@ describe('Custom Model Endpoint Profiles: _showCenterStatus Cancel button (real
expect(win.document.querySelector('.center-status-close')).not.toBeNull();
expect(win.document.querySelector('.center-status-cancel')).toBeNull();
});
it('re-asserts [hidden] over the flex display, so dismiss() actually hides it', () => {
// .center-status-banner is display:flex, which defeats the `hidden` attribute —
// dismiss()'s only visibility lever — unless this rule exists: without it the card
// stays laid out at opacity:0 with its text/cancel/close children still
// pointer-events:auto, an invisible click-blocker dead centre over the terminal
// until the page reloads. Same trap as .home-sessions[hidden], see home-sessions.test.ts.
const css = readFileSync(new URL('../src/web/public/styles.css', import.meta.url), 'utf-8');
expect(css).toMatch(/\.center-status-banner\[hidden\]\s*\{\s*display:\s*none;/);
});
});
describe("Custom Model Endpoint Profiles: requiresContextWarning (this CLI's own overhead can exceed a small model's real context)", () => {