Files
Codeman/.changeset/run-menu-custom-model-picker.md
T
Codeman maintainer 4205f6930f fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine.

**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.

**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.

**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.

**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.

The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:58:33 +02:00

3.0 KiB

aicodeman
aicodeman
minor

feat(custom-model): pick a custom endpoint straight from the Run menu

#393 landed the backend for custom model endpoints and left it reachable only over the HTTP API. This is the rest of it. Turn on Custom model endpoints in App Settings, save an endpoint, and the Run dropdown grows a Custom Endpoints section built live off the CLI registry, one entry per harness that can actually redirect plus each endpoint you saved. Pick one and it launches that harness pointed at your server, asking which model first when the endpoint has more than one. Endpoints re-discover themselves every five minutes, and one unreachable endpoint never blocks the others. App Settings gains full add, edit and delete for endpoints.

Seven of the harnesses (opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP) now launch directly onto the endpoint with no restart at all, where before you watched a native boot followed immediately by a second one. Claude still launches and then restarts in place, which its own resume makes far less jarring.

Most of this release's work went into things that only show up against a real server, and each was found that way rather than in tests: a freshly launched CLI reporting itself busy for its own startup and getting refused; Claude Code assuming a large context window for a model it does not recognise and silently overflowing a small one; a model whose real context is below what Claude Code's own system prompt costs, which no setting can fix and which now warns before launching into a certain failure; and the big one, llama.cpp running exactly one model at a time, so applying a selection can unload the model another session is using. That last case now asks first, tells you which session it affects, and keeps a "loading model" notice on screen for the whole swap window, so a prompt sent mid-swap reads as loading rather than as an answer from whatever was loaded a moment ago. A background sweep also catches the reverse: your session's model being evicted later by somebody else's ordinary use.

Two things worth knowing if you drive this over the HTTP API or run multi-user. The two questions an apply can ask (the model's context window is too small, and loading it will unload the model another session is using) are now answered by separate confirmedContext and confirmedSwap fields rather than one confirmed. They shared a flag until now, and since the context check runs first, confirming that one silently agreed to evict another session's model as well. The old confirmed still means both. And CLAUDE_CONFIG_DIR is now admin-only in multi-user mode: it joined claude's privileged env keys, so a non-granted owner can no longer set it through envOverrides, and an already-persisted one is dropped on reboot-restore, which returns that session to the default Claude account rather than the per-client one it was pointed at. Single-user installs are unaffected.

Remote SSH and Docker sessions are refused for now, since their restart reattaches a durable tmux rather than relaunching the agent.