fix(custom-model): wait for a freshly launched session to go idle before applying

Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.

Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.

Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 10:38:37 +08:00
co-authored by Claude Sonnet 5
parent 409a6e65f9
commit 5c25a52f95
5 changed files with 68 additions and 6 deletions
+1 -1
View File
File diff suppressed because one or more lines are too long
+15 -5
View File
@@ -105,11 +105,21 @@ the point of asking is letting one launch deliberately differ from the
saved default, not just confirming it. Whichever way the model was decided,
the launch itself runs a single session on that harness exactly the way its
own Run-menu entry would (same case creation, env overrides, everything),
then immediately applies the endpoint and model to it via the route below.
It is a one-off "try this endpoint" action, not a sticky mode: the plain
Run button still means "this harness, native cloud" afterward. Entries are
hidden entirely for a remote or Docker active case, since the apply route
refuses both (see the next section).
then **waits for the new session to go idle** (`GET .../wait?until=idle`,
bounded at 20s — a normal 200 either way, never an error, per the wait
endpoint's own contract) before applying the endpoint and model to it via
the route below. That wait exists because a freshly launched CLI reports
itself as `busy` for its own startup (a boot spinner, a workspace-trust
check) well before the apply call would otherwise reach it, and the apply
route correctly refuses to restart a session mid-turn — a fresh boot looks
exactly like one from the outside. A session still busy after the wait
reaches the apply call anyway and gets that route's own honest
`SESSION_BUSY` error, now visible as a sticky toast with a close button
rather than a generic message that vanished in three seconds. It is a
one-off "try this endpoint" action, not a sticky mode: the plain Run button
still means "this harness, native cloud" afterward. Entries are hidden
entirely for a remote or Docker active case, since the apply route refuses
both (see the next section).
## Applying a model to a session
+8
View File
@@ -48,6 +48,14 @@ conversation where the harness supports resuming one, fresh environment. That re
necessary, not incidental: every supported harness reads its endpoint config at process
start, never per turn, so there is no live hot-swap while a turn is running.
Picking an entry that launches a **brand-new** session waits (up to 20 seconds) for it to
finish its own startup before applying — a freshly started CLI reports itself as busy for its
boot sequence, and applying to a genuinely busy session is refused so a real, in-progress
turn is never interrupted out from under you. A session that is still busy after that wait
(a very slow-starting CLI, or one you started typing into right away) surfaces that refusal
as an ordinary error, which now stays on screen with a close button instead of vanishing
after a few seconds — read it, it names the actual reason rather than a generic failure.
Entries are hidden entirely for a session in a **remote (SSH) or Docker case** — support for
redirecting those hasn't landed yet, see below. The picker also only appears in the desktop
**Run** dropdown; the phone home screen builds its own run picker separately and does not
+13
View File
@@ -721,6 +721,19 @@ Object.assign(CodemanApp.prototype, {
const sessionId = this.activeSessionId;
if (!sessionId || sessionId === before) return;
// A freshly launched CLI reports its OWN startup as 'busy' (spinner, the
// workspace-trust check, whatever else it does before its first prompt) —
// measured landing well before this line reliably reaches it — and the
// apply route's isBusy() guard correctly refuses to restart a session
// mid-turn, "mid-turn" included, which this fresh boot looks exactly
// like from the outside. Give it a bounded chance to settle first rather
// than raising a false "Session is busy" on every single launch. Per the
// wait contract a timeout here is a normal 200, never an error — a
// session still busy after 20s just reaches the apply call below and
// gets the route's own honest, now-visible SESSION_BUSY error instead of
// this guessing about it.
await this._apiJson(`/api/sessions/${sessionId}/wait?until=idle&timeout=20000`);
// _apiJson() (used everywhere else in this file) unwraps a success body to
// its `data`, but on failure it swallows the response entirely and returns
// null — exactly the `error` text a caller needs to tell "the endpoint is
+31
View File
@@ -340,6 +340,37 @@ describe('Custom Model Endpoint Profiles: applying a picked entry', () => {
expect(calls[0].body).toEqual({ endpointId: 'llama-box', modelId: 'qwen3' });
});
it('waits for the freshly launched session to go idle before applying, so its own boot activity is never mistaken for a busy turn', async () => {
// Measured live: a just-launched CLI reports 'busy' for its own startup
// (spinner, workspace-trust check) well before the apply call could
// otherwise reach it, and the apply route's isBusy() guard correctly
// refuses to restart a session mid-turn — which a fresh boot looks
// exactly like from the outside. This pins the fix: wait for idle FIRST.
const { app } = bootApp({});
app.activeSessionId = 'old-session';
app.run = async () => {
app.activeSessionId = 'new-session';
};
const calls: string[] = [];
app._apiJson = async (path: string) => {
calls.push(path);
if (path === '/api/model-endpoints') return [];
return null; // the wait call's return value is unused — a timeout is a normal 200
};
app._api = async (path: string) => {
calls.push(path);
return { ok: true, json: async () => ({ success: true, data: {} }) };
};
await app.runCustomModelEntry('claude', 'llama-box', 'qwen3');
const waitIndex = calls.findIndex((p) => p.includes('/wait?'));
const applyIndex = calls.findIndex((p) => p.endsWith('/custom-model'));
expect(waitIndex).toBeGreaterThanOrEqual(0);
expect(calls[waitIndex]).toBe('/api/sessions/new-session/wait?until=idle&timeout=20000');
expect(applyIndex).toBeGreaterThan(waitIndex);
});
it('surfaces the real server error in the toast on a failed apply, rather than a generic message', async () => {
const { app } = bootApp({});
app.activeSessionId = 'old-session';