mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-09 16:59:43 +02:00
fix(custom-model): address second pre-merge review (Ark0N)
Blocker: .center-status-banner never actually disappears.
- Add `.center-status-banner[hidden] { display: none; }`, same trap as
`.home-sessions[hidden]`: the author-level `display: flex` beat the
UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
card stayed laid out at `opacity: 0` with its text/cancel/close
children still `pointer-events: auto` -- an invisible 442x67 click
blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
banner (10001) and the swap-confirm/context-warning modals (10010)
in CLAUDE.md's Z-index layers list.
Stale wording pointed at the reverted sticky-toast default:
- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
`.toast-message` comment in styles.css all still said "toasts
default to sticky" after 1f32128c put the flat 3s default back.
Reworded all three to describe the actual behaviour: one call site
passes an explicit `duration: 0`.
Smaller items from the same review:
- docs/api-reference.md said discovery failures answer
`502 OPERATION_FAILED`; OPERATION_FAILED is 422 per src/types/api.ts
and the error-code table earlier in the same file.
- The periodic re-discovery sweep (server.ts) never read
customModelEndpointsEnabled, so turning the feature off left
Codeman polling every saved endpoint forever. Added
readCustomModelEndpointsEnabled() (custom-model-routes.ts, same
shape as readPlanUsageTelemetryEnabled) and gated the interval
callback on it.
- Reverted the formatting-only Prettier pass docs/api-reference.md
picked up (table padding, *x* to _x_, JSON re-indent) by re-merging
the new Custom Model Endpoints section onto the pre-PR file, so the
diff is reviewable. No prose content was lost -- verified by diffing
the result against the pre-revert file (formatting-only) and against
the merge-base file (only the new section added).
- docs/custom-model-endpoints.md now states that a custom-model Claude
session's isolated CLAUDE_CONFIG_DIR loses the user's global
settings.json, user-level skills/agents/commands, and MCP servers
from ~/.claude.json -- only `projects` is symlinked back.
Design question left open in the review (does `confirmed: true` need
to be two flags so "launch anyway" on the context warning doesn't also
skip the llama-swap displacement warning): keeping the single flag, as
offered. The 20s displacement sweep still catches a resulting swap
after the fact, so it's a surprise rather than a silent failure, and
splitting it is real behavioural surface I have no way to verify live
in this environment.
`npm run test:browser` could not be run in this environment (no tmux,
no downloaded Playwright browser binary) -- none of its suite's files
touch code this fix changes, but it still needs a real pass before
merge, same as any frontend change.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
1f32128ca9
commit
9982a1325f
@@ -7,7 +7,7 @@
|
|||||||
Everything below was found and fixed against a **real llama-swap server**, not just unit tests:
|
Everything below was found and fixed against a **real llama-swap server**, not just unit tests:
|
||||||
|
|
||||||
- **Session-busy false refusal.** A freshly launched CLI reports itself `busy` for its own startup (spinner, workspace-trust check) well before the apply call would reach it, and the apply route correctly refuses to restart a session mid-turn — indistinguishable from a fresh boot. The picker now waits for the new session to go idle (bounded at 20s, never an error on timeout) before applying.
|
- **Session-busy false refusal.** A freshly launched CLI reports itself `busy` for its own startup (spinner, workspace-trust check) well before the apply call would reach it, and the apply route correctly refuses to restart a session mid-turn — indistinguishable from a fresh boot. The picker now waits for the new session to go idle (bounded at 20s, never an error on timeout) before applying.
|
||||||
- **Errors and confirmations you can actually read.** Toasts now default to sticky with a close button (errors always were meant to stay, but a fixed 3s timer silently hid them); a failed apply's real server-side reason (not a generic message) reaches the toast.
|
- **Errors and confirmations you can actually read.** A failed apply's real server-side reason (not a generic message) reaches the toast, and that specific message stays on screen with a close button instead of vanishing on the usual 3s timer.
|
||||||
- **"Both claude.ai and ANTHROPIC_API_KEY set" warning.** A custom-model Claude session now runs with an isolated `CLAUDE_CONFIG_DIR` (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — `projects` is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on _every_ launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would.
|
- **"Both claude.ai and ANTHROPIC_API_KEY set" warning.** A custom-model Claude session now runs with an isolated `CLAUDE_CONFIG_DIR` (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — `projects` is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on _every_ launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would.
|
||||||
- **Context-window overflow.** Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length and applies it as `CLAUDE_CODE_MAX_CONTEXT_TOKENS` — sourced primarily from llama-swap's own `GET /running`, whose `cmd` field carries the launch flags (`--fit-ctx`/`-c`/`--ctx-size`) actually in effect, since `GET /props`'s `n_ctx` was confirmed live to report the model's theoretical/trained maximum rather than the real `--fit-ctx`-shrunk runtime context (a 154112-vs-16384 discrepancy, caught only because the fixed value still overflowed) — `/props` is now a fallback for a plain llama.cpp server with no `/running` at all.
|
- **Context-window overflow.** Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length and applies it as `CLAUDE_CODE_MAX_CONTEXT_TOKENS` — sourced primarily from llama-swap's own `GET /running`, whose `cmd` field carries the launch flags (`--fit-ctx`/`-c`/`--ctx-size`) actually in effect, since `GET /props`'s `n_ctx` was confirmed live to report the model's theoretical/trained maximum rather than the real `--fit-ctx`-shrunk runtime context (a 154112-vs-16384 discrepancy, caught only because the fixed value still overflowed) — `/props` is now a fallback for a plain llama.cpp server with no `/running` at all.
|
||||||
- **Context floor too small for Claude Code to even start.** Fixing the overflow above surfaced a second, unfixable-by-injection failure: Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on their own (confirmed live via an `in:0 out:0` failure on the very first message), which can exceed a small model's entire real context before any conversation history exists to trim — no `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes that, since it only governs when history gets compacted. Applying such a model now returns a warning (gated on the CLI registry declaring a `contextLengthVar`, so it's a no-op for every other harness) instead of launching straight into a guaranteed first-message failure, and the Run-menu picker shows it as an in-app dialog naming the model, its discovered context and the ~40K safe floor, with the actual fix spelled out: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on auto-fit, which optimizes for the biggest model that fits rather than the biggest context. "Launch anyway" is still one click away.
|
- **Context floor too small for Claude Code to even start.** Fixing the overflow above surfaced a second, unfixable-by-injection failure: Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on their own (confirmed live via an `in:0 out:0` failure on the very first message), which can exceed a small model's entire real context before any conversation history exists to trim — no `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes that, since it only governs when history gets compacted. Applying such a model now returns a warning (gated on the CLI registry declaring a `contextLengthVar`, so it's a no-op for every other harness) instead of launching straight into a guaranteed first-message failure, and the Run-menu picker shows it as an in-app dialog naming the model, its discovered context and the ~40K safe floor, with the actual fix spelled out: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on auto-fit, which optimizes for the biggest model that fits rather than the biggest context. "Launch anyway" is still one click away.
|
||||||
|
|||||||
+77
-85
@@ -66,17 +66,17 @@ The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
|
|||||||
`src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on
|
`src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on
|
||||||
the HTTP status.
|
the HTTP status.
|
||||||
|
|
||||||
| `errorCode` | HTTP | Meaning |
|
| `errorCode` | HTTP | Meaning |
|
||||||
| ------------------ | ---- | --------------------------------------------------- |
|
|-------------|------|---------|
|
||||||
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
|
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
|
||||||
| `UNAUTHORIZED` | 401 | Authentication required or failed |
|
| `UNAUTHORIZED` | 401 | Authentication required or failed |
|
||||||
| `NOT_FOUND` | 404 | Resource does not exist |
|
| `NOT_FOUND` | 404 | Resource does not exist |
|
||||||
| `SESSION_BUSY` | 409 | Session is busy |
|
| `SESSION_BUSY` | 409 | Session is busy |
|
||||||
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
|
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
|
||||||
| `ALREADY_EXISTS` | 409 | Resource already exists |
|
| `ALREADY_EXISTS` | 409 | Resource already exists |
|
||||||
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
|
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
|
||||||
| `RATE_LIMITED` | 429 | Too many requests |
|
| `RATE_LIMITED` | 429 | Too many requests |
|
||||||
| `INTERNAL_ERROR` | 500 | Unexpected server error |
|
| `INTERNAL_ERROR` | 500 | Unexpected server error |
|
||||||
|
|
||||||
Adding a new error code is non-breaking; removing or renaming one is a major change.
|
Adding a new error code is non-breaking; removing or renaming one is a major change.
|
||||||
|
|
||||||
@@ -87,10 +87,10 @@ exist because SSE is Codeman's only other "tell me when" channel, and an agent
|
|||||||
driving the API from a shell tool cannot practically hold a stream and parse
|
driving the API from a shell tool cannot practically hold a stream and parse
|
||||||
events inline.
|
events inline.
|
||||||
|
|
||||||
| Call | Blocks until |
|
| Call | Blocks until |
|
||||||
| --------------------------------------------- | -------------------------------------------------- |
|
|------|--------------|
|
||||||
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
|
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
|
||||||
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
|
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
|
||||||
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
|
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
|
||||||
|
|
||||||
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
|
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
|
||||||
@@ -140,13 +140,13 @@ contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
|
|||||||
|
|
||||||
### Signals
|
### Signals
|
||||||
|
|
||||||
| Signal | Source | Actually fires for |
|
| Signal | Source | Actually fires for |
|
||||||
| --------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|--------|--------|--------------------|
|
||||||
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
|
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
|
||||||
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
|
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
|
||||||
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
|
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
|
||||||
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
|
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
|
||||||
| `exit` | no process is behind the session | every mode |
|
| `exit` | no process is behind the session | every mode |
|
||||||
|
|
||||||
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
|
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
|
||||||
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
|
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
|
||||||
@@ -156,12 +156,12 @@ can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit
|
|||||||
once the session is up: the default set's `idle` also resolves on a spinner pause,
|
once the session is up: the default set's `idle` also resolves on a spinner pause,
|
||||||
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
|
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
|
||||||
can land inside your first wait window and report a turn that never ran. Measured:
|
can land inside your first wait window and report a turn that never ran. Measured:
|
||||||
a session parked on the trust dialog emits no _further_ `idle`, so it is the
|
a session parked on the trust dialog emits no *further* `idle`, so it is the
|
||||||
startup transition, not the dialog, that produces the false success below.
|
startup transition, not the dialog, that produces the false success below.
|
||||||
|
|
||||||
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
|
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
|
||||||
server answers from `pid === null` plus a mux-layer pane-death probe, and that
|
server answers from `pid === null` plus a mux-layer pane-death probe, and that
|
||||||
covers a session that exited — including a worker that died _inside_ its tmux pane
|
covers a session that exited — including a worker that died *inside* its tmux pane
|
||||||
while the local attach client (and therefore `pid`) lives on — one that was
|
while the local attach client (and therefore `pid`) lives on — one that was
|
||||||
detached, and one that was **created but never started**. So the first wait
|
detached, and one that was **created but never started**. So the first wait
|
||||||
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
|
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
|
||||||
@@ -184,7 +184,7 @@ blocked, and polling `blocked` alone will sit at its timeout.
|
|||||||
|
|
||||||
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
|
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
|
||||||
session emits its one `idle` at startup and then stays `status: "idle"` forever,
|
session emits its one `idle` at startup and then stays `status: "idle"` forever,
|
||||||
whatever the pane is doing, so it never emits a _transition_. Since send-and-wait
|
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
|
||||||
requires a transition (and so does `fresh=1`), both can only time out there:
|
requires a transition (and so does `fresh=1`), both can only time out there:
|
||||||
a documented default `wait` on a shell worker running `sleep 4` times out at the
|
a documented default `wait` on a shell worker running `sleep 4` times out at the
|
||||||
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
|
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
|
||||||
@@ -218,11 +218,11 @@ with `from=buffer` keeps matching long after the dialog is gone. A worked versio
|
|||||||
|
|
||||||
### `GET /api/v1/sessions/:id/wait`
|
### `GET /api/v1/sessions/:id/wait`
|
||||||
|
|
||||||
| Param | Type | Default | Notes |
|
| Param | Type | Default | Notes |
|
||||||
| --------- | -------------------------------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|-------|------|---------|-------|
|
||||||
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
|
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
|
||||||
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
|
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
|
||||||
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
|
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
|
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
|
||||||
@@ -239,12 +239,12 @@ a plain signal wait, so check the endpoint path before blaming the parameters.
|
|||||||
|
|
||||||
### `GET /api/v1/sessions/:id/wait-output`
|
### `GET /api/v1/sessions/:id/wait-output`
|
||||||
|
|
||||||
| Param | Type | Default | Notes |
|
| Param | Type | Default | Notes |
|
||||||
| --------- | ------------------------------- | -------- | ----------------------------------------------------------------------------------------------------------- |
|
|-------|------|---------|-------|
|
||||||
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
|
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
|
||||||
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
|
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
|
||||||
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
|
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
|
||||||
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
|
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
|
||||||
|
|
||||||
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
|
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
|
||||||
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
|
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
|
||||||
@@ -296,10 +296,10 @@ hand-written query string decodes to a space.
|
|||||||
|
|
||||||
Two optional fields on the existing endpoint:
|
Two optional fields on the existing endpoint:
|
||||||
|
|
||||||
| Field | Type | Notes |
|
| Field | Type | Notes |
|
||||||
| ------------- | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|-------|------|-------|
|
||||||
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
|
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
|
||||||
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
|
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
|
||||||
|
|
||||||
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
|
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
|
||||||
"absent" rather than failing validation. That is deliberate: `.optional()` would
|
"absent" rather than failing validation. That is deliberate: `.optional()` would
|
||||||
@@ -330,24 +330,16 @@ All three nest the wait result under `data.wait`, so one client helper works aga
|
|||||||
any of them:
|
any of them:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{ "success": true, "data": {
|
||||||
"success": true,
|
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
|
||||||
"data": {
|
"status": "idle",
|
||||||
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
|
"limitPaused": false,
|
||||||
"status": "idle",
|
"wait": {
|
||||||
"limitPaused": false,
|
"signal": "stop", "until": ["stop", "idle", "exit"],
|
||||||
"wait": {
|
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
|
||||||
"signal": "stop",
|
"waitedMs": 8421, "timeoutMs": 60000
|
||||||
"until": ["stop", "idle", "exit"],
|
|
||||||
"timedOut": false,
|
|
||||||
"immediate": false,
|
|
||||||
"ended": false,
|
|
||||||
"aborted": false,
|
|
||||||
"waitedMs": 8421,
|
|
||||||
"timeoutMs": 60000
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
}
|
}}
|
||||||
```
|
```
|
||||||
|
|
||||||
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
|
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
|
||||||
@@ -361,21 +353,21 @@ redelivery (harmless, the turn it refers to may be long over), while with
|
|||||||
client that reads `delivered === false` as "duplicate" silently treats a failed send
|
client that reads `delivered === false` as "duplicate" silently treats a failed send
|
||||||
as a success.
|
as a success.
|
||||||
|
|
||||||
| Field | Type | Meaning |
|
| Field | Type | Meaning |
|
||||||
| ---------------- | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|-------|------|---------|
|
||||||
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
|
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
|
||||||
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
|
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
|
||||||
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
|
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
|
||||||
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
|
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
|
||||||
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
|
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
|
||||||
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
|
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
|
||||||
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
|
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
|
||||||
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
|
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
|
||||||
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
|
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
|
||||||
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
|
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
|
||||||
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
|
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
|
||||||
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
|
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
|
||||||
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
|
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
|
||||||
|
|
||||||
Read the outcome by discriminator, in this order:
|
Read the outcome by discriminator, in this order:
|
||||||
|
|
||||||
@@ -398,12 +390,12 @@ read the timeout as "the worker is wedged" and kill a session that was working f
|
|||||||
|
|
||||||
### Errors
|
### Errors
|
||||||
|
|
||||||
| `errorCode` | HTTP | When |
|
| `errorCode` | HTTP | When |
|
||||||
| --------------- | ---- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|-------------|------|------|
|
||||||
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
|
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
|
||||||
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
|
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
|
||||||
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
|
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
|
||||||
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
|
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
|
||||||
|
|
||||||
The two capacity codes are deliberately different. A process-wide cap reported as
|
The two capacity codes are deliberately different. A process-wide cap reported as
|
||||||
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
|
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
|
||||||
@@ -454,9 +446,9 @@ Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md).
|
|||||||
|
|
||||||
- `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first,
|
- `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first,
|
||||||
ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId,
|
ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId,
|
||||||
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
|
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
|
||||||
toolSummary?, message?, cwd?, context?, options?: {n, label}[],
|
toolSummary?, message?, cwd?, context?, options?: {n, label}[],
|
||||||
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
|
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
|
||||||
`options` is present only when the dialog's numbered choices parsed
|
`options` is present only when the dialog's numbered choices parsed
|
||||||
confidently; `acknowledgedAt` marks an item a human has already looked at
|
confidently; `acknowledgedAt` marks an item a human has already looked at
|
||||||
(see `/viewed` below) and tells clients not to re-arm its tab alert. Listing
|
(see `/viewed` below) and tells clients not to re-arm its tab alert. Listing
|
||||||
@@ -474,7 +466,7 @@ acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
|
|||||||
first, `422 OPERATION_FAILED` when the session refused input.
|
first, `422 OPERATION_FAILED` when the session refused input.
|
||||||
- `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes.
|
- `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes.
|
||||||
- `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId,
|
- `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId,
|
||||||
acknowledged: itemId | null }`. Marks the session's pending **idle** item as
|
acknowledged: itemId | null }`. Marks the session's pending **idle** item as
|
||||||
seen by a human (the web UI calls it when you open the session's tab): the
|
seen by a human (the web UI calls it when you open the session's tab): the
|
||||||
item stays pending and answerable, but stops arming the yellow tab alert on
|
item stays pending and answerable, but stops arming the yellow tab alert on
|
||||||
every client, including after a reload. Permission/question items are never
|
every client, including after a reload. Permission/question items are never
|
||||||
@@ -499,7 +491,7 @@ user guide: [`readmymind.md`](readmymind.md).
|
|||||||
|
|
||||||
- `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the
|
- `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the
|
||||||
session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals,
|
session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals,
|
||||||
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
|
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
|
||||||
50, each <= 500 chars). A case with nothing recorded answers an empty
|
50, each <= 500 chars). A case with nothing recorded answers an empty
|
||||||
profile with `updatedAt: 0`; nothing is persisted by reads.
|
profile with `updatedAt: 0`; nothing is persisted by reads.
|
||||||
- `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict
|
- `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict
|
||||||
@@ -559,7 +551,7 @@ authStyle?, defaultModelId? }` creates one. `id` must match
|
|||||||
`lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own
|
`lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own
|
||||||
response already reports loaded) `modelContextLengths` and `modelSizesGB`.
|
response already reports loaded) `modelContextLengths` and `modelSizesGB`.
|
||||||
A `defaultModelId` that no longer appears in the fresh list is dropped
|
A `defaultModelId` that no longer appears in the fresh list is dropped
|
||||||
rather than carried forward invalid. Failures answer `502 OPERATION_FAILED`
|
rather than carried forward invalid. Failures answer `422 OPERATION_FAILED`
|
||||||
with the underlying connection error, or a named egress refusal if the
|
with the underlying connection error, or a named egress refusal if the
|
||||||
resolved address turned out to be blocked. The same refresh also runs
|
resolved address turned out to be blocked. The same refresh also runs
|
||||||
automatically for every saved endpoint every 5 minutes in the background
|
automatically for every saved endpoint every 5 minutes in the background
|
||||||
@@ -631,7 +623,7 @@ same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synce
|
|||||||
[`claude-voice-plan.md`](claude-voice-plan.md).
|
[`claude-voice-plan.md`](claude-voice-plan.md).
|
||||||
|
|
||||||
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
|
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
|
||||||
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
|
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
|
||||||
signed in to Claude Code on the server), `expired` (the access token elapsed;
|
signed in to Claude Code on the server), `expired` (the access token elapsed;
|
||||||
running any Claude session refreshes it) or `malformed`. The OAuth token
|
running any Claude session refreshes it) or `malformed`. The OAuth token
|
||||||
itself is never returned by this or any other endpoint.
|
itself is never returned by this or any other endpoint.
|
||||||
|
|||||||
@@ -290,7 +290,13 @@ size`. No entry for the model in `modelContextLengths` means the var is
|
|||||||
[`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied
|
[`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied
|
||||||
automatically here. Best-effort: a platform that refuses the symlink keeps
|
automatically here. Best-effort: a platform that refuses the symlink keeps
|
||||||
the pre-existing blind-response-viewer side effect rather than failing the
|
the pre-existing blind-response-viewer side effect rather than failing the
|
||||||
whole custom-model apply over it.
|
whole custom-model apply over it. ⚠️ **This relocates the whole `.claude`
|
||||||
|
tree, not just transcripts**: a custom-model Claude session also loses the
|
||||||
|
user's global `settings.json`, user-level skills (the codeman agent skill
|
||||||
|
included), user-level agents and commands, and the MCP servers configured
|
||||||
|
in `~/.claude.json` — none of those are symlinked back, only `projects` is.
|
||||||
|
A fine trade for "point this session at my local llama.cpp," but worth
|
||||||
|
knowing before it surprises you mid-session.
|
||||||
|
|
||||||
**That isolated directory needed one more fix to actually be usable
|
**That isolated directory needed one more fix to actually be usable
|
||||||
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
|
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
|
||||||
|
|||||||
@@ -8526,8 +8526,8 @@ kbd {
|
|||||||
|
|
||||||
.toast-message {
|
.toast-message {
|
||||||
flex: 1;
|
flex: 1;
|
||||||
/* Errors are sticky by default (showToast) precisely so a longer, specific
|
/* A sticky toast (showToast's opts.duration: 0) can carry a longer, specific
|
||||||
message survives to be read — let it wrap instead of clipping. */
|
message — let it wrap instead of clipping. */
|
||||||
white-space: pre-wrap;
|
white-space: pre-wrap;
|
||||||
word-break: break-word;
|
word-break: break-word;
|
||||||
}
|
}
|
||||||
@@ -8588,6 +8588,15 @@ kbd {
|
|||||||
transform: translate(-50%, -50%) scale(1);
|
transform: translate(-50%, -50%) scale(1);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* `hidden` has to be re-asserted over the `display: flex` above, or `dismiss()`
|
||||||
|
setting `el.hidden = true` does nothing (same trap as `.home-sessions[hidden]`
|
||||||
|
below): the card stays laid out at `opacity: 0` with its text/cancel/close
|
||||||
|
children still `pointer-events: auto`, an invisible click-blocker dead centre
|
||||||
|
over the terminal until the page reloads. */
|
||||||
|
.center-status-banner[hidden] {
|
||||||
|
display: none;
|
||||||
|
}
|
||||||
|
|
||||||
.center-status-spinner {
|
.center-status-spinner {
|
||||||
flex-shrink: 0;
|
flex-shrink: 0;
|
||||||
width: 18px;
|
width: 18px;
|
||||||
|
|||||||
@@ -16,7 +16,7 @@
|
|||||||
|
|
||||||
import type { FastifyInstance, FastifyRequest } from 'fastify';
|
import type { FastifyInstance, FastifyRequest } from 'fastify';
|
||||||
import { ApiErrorCode, createErrorResponse, type ApiResponse } from '../../types.js';
|
import { ApiErrorCode, createErrorResponse, type ApiResponse } from '../../types.js';
|
||||||
import { isAdmin, parseBody } from '../route-helpers.js';
|
import { isAdmin, parseBody, readJsonConfig, SETTINGS_PATH } from '../route-helpers.js';
|
||||||
import { isMultiUserMode } from '../../config/multiuser.js';
|
import { isMultiUserMode } from '../../config/multiuser.js';
|
||||||
import { getDataDir } from '../../config/instance.js';
|
import { getDataDir } from '../../config/instance.js';
|
||||||
import { isBlockedWebviewUrl } from '../webview-egress-policy.js';
|
import { isBlockedWebviewUrl } from '../webview-egress-policy.js';
|
||||||
@@ -565,6 +565,19 @@ function applyDiscoveredModels(host: CustomModelHost, result: DiscoveryResult):
|
|||||||
};
|
};
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* `customModelEndpointsEnabled` defaults OFF (unlike `showPlanUsageLimits`'s
|
||||||
|
* absent-means-on in `readPlanUsageTelemetryEnabled`), so mirror the frontend's
|
||||||
|
* own gate (`session-ui.js`'s `!settings.customModelEndpointsEnabled`) rather
|
||||||
|
* than that reader's default. Exists so the periodic re-discovery sweep in
|
||||||
|
* server.ts can skip entirely while the feature is off, instead of polling
|
||||||
|
* every saved endpoint forever regardless of the setting.
|
||||||
|
*/
|
||||||
|
export async function readCustomModelEndpointsEnabled(): Promise<boolean> {
|
||||||
|
const settings = await readJsonConfig<Record<string, unknown>>(SETTINGS_PATH, 'settings.json', {});
|
||||||
|
return settings.customModelEndpointsEnabled === true;
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Re-discovers every saved endpoint's models, best-effort. One endpoint being
|
* Re-discovers every saved endpoint's models, best-effort. One endpoint being
|
||||||
* unreachable (powered off, wrong network) must not stop the others from
|
* unreachable (powered off, wrong network) must not stop the others from
|
||||||
|
|||||||
@@ -30,6 +30,7 @@ export { registerTabLayoutRoutes } from './tab-layout-routes.js';
|
|||||||
export {
|
export {
|
||||||
registerCustomModelRoutes,
|
registerCustomModelRoutes,
|
||||||
refreshAllCustomModelHosts,
|
refreshAllCustomModelHosts,
|
||||||
|
readCustomModelEndpointsEnabled,
|
||||||
detectCustomModelSwapDisplacements,
|
detectCustomModelSwapDisplacements,
|
||||||
pruneIdleLlamaSwapLogTails,
|
pruneIdleLlamaSwapLogTails,
|
||||||
type CustomModelSessionLike,
|
type CustomModelSessionLike,
|
||||||
|
|||||||
+13
-3
@@ -191,6 +191,7 @@ import {
|
|||||||
registerTabLayoutRoutes,
|
registerTabLayoutRoutes,
|
||||||
registerCustomModelRoutes,
|
registerCustomModelRoutes,
|
||||||
refreshAllCustomModelHosts,
|
refreshAllCustomModelHosts,
|
||||||
|
readCustomModelEndpointsEnabled,
|
||||||
detectCustomModelSwapDisplacements,
|
detectCustomModelSwapDisplacements,
|
||||||
pruneIdleLlamaSwapLogTails,
|
pruneIdleLlamaSwapLogTails,
|
||||||
tryWebviewRefererFallback,
|
tryWebviewRefererFallback,
|
||||||
@@ -2761,9 +2762,18 @@ export class WebServer extends EventEmitter {
|
|||||||
if (!this.testMode) {
|
if (!this.testMode) {
|
||||||
this.cleanup.setInterval(
|
this.cleanup.setInterval(
|
||||||
() => {
|
() => {
|
||||||
refreshAllCustomModelHosts().catch((err) => {
|
// Reads the setting fresh on every tick, same reasoning as
|
||||||
console.error('[custom-model] periodic re-discovery failed:', getErrorMessage(err));
|
// readPlanUsageTelemetryEnabled() beside it: a live toggle takes effect
|
||||||
});
|
// on the very next cycle, not just at server boot, and turning the
|
||||||
|
// feature off actually stops the polling instead of only hiding the UI.
|
||||||
|
void readCustomModelEndpointsEnabled()
|
||||||
|
.then((enabled) => {
|
||||||
|
if (!enabled) return;
|
||||||
|
return refreshAllCustomModelHosts();
|
||||||
|
})
|
||||||
|
.catch((err) => {
|
||||||
|
console.error('[custom-model] periodic re-discovery failed:', getErrorMessage(err));
|
||||||
|
});
|
||||||
},
|
},
|
||||||
CUSTOM_MODEL_REDISCOVER_INTERVAL_MS,
|
CUSTOM_MODEL_REDISCOVER_INTERVAL_MS,
|
||||||
{ description: 'custom model endpoint re-discovery' }
|
{ description: 'custom model endpoint re-discovery' }
|
||||||
|
|||||||
@@ -969,6 +969,16 @@ describe('Custom Model Endpoint Profiles: _showCenterStatus Cancel button (real
|
|||||||
expect(win.document.querySelector('.center-status-close')).not.toBeNull();
|
expect(win.document.querySelector('.center-status-close')).not.toBeNull();
|
||||||
expect(win.document.querySelector('.center-status-cancel')).toBeNull();
|
expect(win.document.querySelector('.center-status-cancel')).toBeNull();
|
||||||
});
|
});
|
||||||
|
|
||||||
|
it('re-asserts [hidden] over the flex display, so dismiss() actually hides it', () => {
|
||||||
|
// .center-status-banner is display:flex, which defeats the `hidden` attribute —
|
||||||
|
// dismiss()'s only visibility lever — unless this rule exists: without it the card
|
||||||
|
// stays laid out at opacity:0 with its text/cancel/close children still
|
||||||
|
// pointer-events:auto, an invisible click-blocker dead centre over the terminal
|
||||||
|
// until the page reloads. Same trap as .home-sessions[hidden], see home-sessions.test.ts.
|
||||||
|
const css = readFileSync(new URL('../src/web/public/styles.css', import.meta.url), 'utf-8');
|
||||||
|
expect(css).toMatch(/\.center-status-banner\[hidden\]\s*\{\s*display:\s*none;/);
|
||||||
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
describe("Custom Model Endpoint Profiles: requiresContextWarning (this CLI's own overhead can exceed a small model's real context)", () => {
|
describe("Custom Model Endpoint Profiles: requiresContextWarning (this CLI's own overhead can exceed a small model's real context)", () => {
|
||||||
|
|||||||
Reference in New Issue
Block a user