mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 20:49:41 +02:00
master
1029
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bb8ada7e5f |
fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does. 1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred it from the name: a session the user renamed by hand to something shaped like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the next prompt overwrote their name. The route persists right after, so the loss went to disk. `restoreMuxSessions()` already passes it. 2. The already-live sets were snapshotted once before a loop that awaits a real `startInteractive()` per entry, so by the tenth entry the snapshot was tens of seconds old and a conversation resumed by hand from the Resume list in that window was invisible to it: two panes on one transcript, the exact thing the check exists to prevent. Both sets are now read per iteration, and the late case is spent rather than re-offered for the same reason the batch case is. 3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path. The stamp predates the reboot and the pane is new, so honouring it meant one click had every restored session type `continue` into itself about a minute later, unattended, against the route header's own promise that a restored session comes back idle and disarmed. The setting stays ENABLED, so it re-arms on the next real limit message. A Codeman restart still re-arms from the stamp, because the limit footer will not reprint on its own; the new option exists only to tell the two paths apart. 4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()` and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession` performs that it was missing. Cosmetic, but a run left open reads as still going in the away digest. 5. A restored claude session gets `seedAgentSessionPreamble()` like both create paths, so the agent skill's bootstrap stays a two-line loader. 6. The heuristic's container comment was wrong in one direction and quiet about the real gap: after a genuine host reboot a containerized Codeman sees the host's short uptime and the banner does appear. What it cannot see is a container-only restart, which is where this would help most. 7. The banner is hidden in a solo window, which shows one session and has no tab strip to put restored ones in. Also reverts 17 of the 18 hunks in docs/api-reference.md, which were Prettier reformatting of prose the PR does not otherwise touch (docs/ is outside the format glob), keeping only the Reboot restore section and repairing the two continuation lines that reformat de-indented; renumbers reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which loads after it; and gives the feature its CLAUDE.md entry plus a route test for the multi-user workspace-forbidden branch, the only new rule that had nothing behind it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ea5323d990 |
test(input): pin the batched commit-plus-Enter ordering #441 fixes
The unit harness proves WHICH candidate gets forwarded; the ordering is the half that shipped the bug, and only a real xterm shows it. The new browser case dispatches the character's keydown, its composed insertText and Enter's keydown in ONE page task, the shape an Android soft keyboard delivers through a single InputConnection transaction, and asserts what reaches the send path. Verified in both directions on this machine: with the drain in place the wire is `o\r`; with the drain removed (master's behaviour) it is `\r` and the character is gone entirely, because by the time the zero-delay timer runs xterm has emitted the `\r` and bumped the canonical counter past the candidate's snapshot, so the candidate stands down. The other four cases pass in both states. CLAUDE.md now names the decision point, what it costs (a keydown decides with less evidence than the timer did) and why that is safe for Enter, and says that the pin lives in a suite the CI gate does not run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f32c4f60d5 |
Merge pull request #442 from irisitymichaelgrundberg/feat/restore-sessions-after-reboot
feat(sessions): offer to rebuild the sessions a host reboot destroyed |
||
|
|
9a503872d9 |
Merge pull request #441 from shenlvkang-collab/fix/android-last-char
fix(input): deliver a recovered keystroke before the Enter that submits it |
||
|
|
9d7b29d899 |
Merge pull request #429 from opticon454/chore/cli-catalog-followups
chore(cli-registry): clean up dead code and stale claims left after #380 |
||
|
|
ff8dc92187 |
Merge pull request #424 from Ark0N/fix/terminal-history-anchor-after-parse
fix(terminal): restore the history anchor after xterm parses, not before |
||
|
|
62ceb4e87b |
fix(sessions): correct what the missing-pid rule actually recognises
A second real reboot disproved the mechanism the previous commit was built on. Typing `/exit` does not persist `pid: null`, and the session was restored anyway. The pid a session record carries is its `tmux attach-session` process, not the agent. `/exit` ends the CLI inside the pane, `remain-on-exit` keeps the pane, and the attach process stays alive throughout — so Codeman's PTY never exits, no exit handler runs, and the record keeps both its pid and `status: 'idle'`. The lifecycle log for the session that came back shows created, started, stale_cleaned and recovered, with no exit event at all, which is the proof: Codeman never learned the agent was gone. So nothing durable distinguishes an exited agent from a session that was idle when the power went, and this pass restores both. Ark0N/Codeman#446 is about making Codeman notice the dead pane; contrary to what the previous commit's message claimed, this genuinely does wait on that. Until a record can say the agent is gone, the user dismisses or closes those sessions. The rule itself is kept, because a record with no attach process does describe a session that never started or whose pane died outright, and refusing it is right. Only its documentation was wrong. The module header, the branch comment and the test names now say what it recognises instead of claiming the case it cannot see. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5108a24bf0 |
fix(sessions): never restore a session whose agent was already exited
Found by a real reboot, which is the first thing to catch it. Typing `/exit` ends the CLI process and leaves the session record behind, and the process-exit handler persists `pid: null` with `status: 'idle'` before anything else runs. By status alone that is indistinguishable from a session sitting idle when the power went, so the boot pass offered those sessions back and a click spawned the agents the user had deliberately closed — the exact case the eligibility rule exists to exclude. The absent pid is what tells the two apart, and the plan step now refuses a record without one, under its own `not-running` reason so the boot log says why. On a healthy board every running session carries a pid; a record with none describes an agent that is already gone. Deliberately the conservative direction. A session that somehow persisted no pid while genuinely running is not offered, and its conversation stays reachable from the Resume list, which is where every session would be without this feature. The opposite error spawns processes nobody asked for. Ark0N/Codeman#446 covers the dead panes those exits leave behind, but this does not wait on it: the rule belongs here whether or not the record's shape changes later. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
18ab2ab595 |
docs(sessions): correct what a failed rebuild is actually likely to be
Ran the feature against a real server for the first time, on an isolated instance, and two claims in the code turned out to be wrong. A rebuild that fails after the session is registered was documented as commonly caused by a CLI binary missing from a freshly booted machine's PATH. It is not: the resolver finds its binary by absolute path, so PATH never enters into it, and a server started without claude on PATH restored every session normally. Nor does an un-enterable workspace fail — tmux falls back to another directory and the pane comes up there. Neither obvious cause throws, so the discard path is defended rather than expected, and the comments now say that instead of naming a cause that cannot happen. The four review rounds that shaped this path all reasoned about a trigger none of them could test. The path itself is still worth having, since a mux failure would reach it, but its comments should not claim a likelihood the machine disagrees with. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e2034177c5 |
fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.
Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
published dependencies) into a scratch dir purely to read
dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
same endpoint. dsh's own error template ("DeepSeek API error (HTTP
${status})") reproduces the originally-reported
"dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.
- New registry field `appendV1Suffix` (env kind only, deepseek's entry
alone — claude/gemini must NOT get it, since claude was already
confirmed working against the unmodified baseUrl). When set,
buildCustomModelInjection runs endpoint.baseUrl through the same
withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
use, instead of writing it verbatim.
Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.
2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
db9729e1fc |
feat(custom-model): remove loading-banner countdown, add manual Cancel
Replaces the size-scaled expected-time estimate + matching auto-timeout with a generic hardware/model-size disclaimer and a user-driven Cancel button, per explicit request. Real load time depends on hardware this feature has no way to know (VRAM, storage speed, GPU contention), so the old estimate/timeout was a guess dressed up as a fact — worse, one that could kill a genuinely slow load partway through on slower hardware. - _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline entirely — polls indefinitely until ready or cancelled, no automatic give-up. Message is now "Loading <model> (<size>) on <endpoint> — this can take a while depending on your hardware and the model size.", with the real llama.cpp log line still on its own second line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/ _formatRemaining (dead code once the countdown is gone) — _lookupModelSizeGB is kept, the GB figure still shows. - _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real "Cancel" button (distinct from the error-type "×" close button, since Cancel has a real consequence) that calls it on click. Caller owns what cancelling actually means, same split as the swap-confirm modal's promise-resolving buttons. - Cancelling dismisses the banner, shows an info toast (not an error — this was deliberate), and closes the session, mirroring what the old timeout used to do automatically but now on the user's own call. - New .center-status-cancel CSS (bordered pill button, distinct from the plain "×" close glyph). Test changes: removed the now-invalid timeout-auto-close/estimate tests, added cancel-flow tests (dismiss/toast-type/session-close, never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests for the new Cancel button (bootAppWithRealCenterStatus, evaluating panels-ui.js instead of stubbing _showCenterStatus, since this button is worth verifying for real rather than just through the stub every other test in the file uses). Typecheck/lint/frontend-syntax clean; full suite shows no new regressions. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
2d3fc65758 |
feat(custom-model): show real-time llama.cpp backend status in the loading banner
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.
- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
(custom-model-routes.ts): one persistent GET /api/events (SSE)
connection held open per endpoint, parsing logData frames and
keeping the latest source:"upstream" (backend llama-server) line —
filtering out llama-swap's own source:"proxy" request-access lines.
Idle-closed after 30s of no polling, same 20s sweep as the existing
swap-displacement check.
- running-status route now returns logLine alongside the existing
isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
("llama.cpp: <line>", bootlog timestamp/level/component prefix
stripped for display) that stays on the last real thing llama.cpp
said rather than clearing to blank between polls.
⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).
12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
5ddc028a2f |
feat(custom-model): detect and notify when a session's model gets swapped out later
The llama-swap conflict check on the apply/create routes only ever runs at THAT session's own launch/apply moment, and cannot see a swap caused by a DIFFERENT session's later, ordinary use. Confirmed live: a second Codex session picking a different model launched with no warning at all — nothing conflicted at that exact instant — yet it silently evicted the first session's model regardless (llama.cpp runs one model at a time). Reproduced and root-caused via direct API calls against a live test-picker instance rather than guessing. - detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups live sessions with a customModel by endpointId, checks each group's endpoint via GET /running once, and flags a session whose own modelId is no longer in the running list. Read-only, best-effort per endpoint like refreshAllCustomModelHosts's sibling sweep. - Notifies once per displacement via a caller-owned de-dupe Set: a session id is added when displaced, removed once its own model is loaded/ready again, so a later genuinely-new displacement can notify again. - New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS, 20s — much shorter than the 5-minute model-list refresh, since this is time-sensitive) broadcasts a new custom-model:swapped-out SSE event per displacement. De-dupe Set cleared per-session on session cleanup to avoid an unbounded leak. - Frontend: global toast (not tied to the displaced session's tab, since the point is warning before the user types into it) naming the session, its previous model, and what's currently loaded. Chose the "detect after the fact" scope (vs. checking before every message send, which would add a round-trip to every turn on every custom-model session) per explicit user decision after being presented the trade-off. 9 new tests for the detection logic (flag/clear/re-flag cycle, unreachable/deleted endpoints, non-llama-swap servers, multiple sessions on one endpoint). SSE registry bumped 158->159, parity test passing. Typecheck/lint/frontend-syntax clean; full suite shows no new regressions (9 more passing than baseline, matching the new tests; same pre-existing Windows-environment failures). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
211b872335 |
feat(custom-model): skip Claude Code's first-run wizard on custom-model launches
A fresh, isolated CLAUDE_CONFIG_DIR (used to keep an injected API key away from a stored claude.ai OAuth login) looks like a brand-new Claude Code profile to the CLI, so it replays its ENTIRE first-run sequence on every single launch: the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running with --dangerously-skip-permissions) a one-time bypass-permissions warning — confirmed live, none of which a real, already-onboarded profile shows again. - New registry-declared env-kind field `skipFirstRunPrompts` (alongside apiKeyTrustFile, which it reuses) — claude's entry only, carried through buildCustomModelInjection (pure) into applyCustomModelInjection (IO). - seedFirstRunOnboardingState(): merges hasCompletedOnboarding: true and this session's own projects[workingDir].hasTrustDialogAccepted: true into the same <configDir>/.claude.json the API-key trust file already writes to — other projects and other fields on this session's own entry are left untouched. - seedSkipBypassPermissionsPrompt(): merges skipDangerousModePermissionPrompt: true into <configDir>/settings.json, a separate file, same corrupt-tolerant merge behavior. - applyCustomModelInjection() gains an optional workingDir parameter, threaded from session.workingDir (dedicated apply route) / resolvedCasePath (quick-start route) — boot recovery omits it (a dialog already answered once needs no re-seed on the same, persisted isolated directory). Tests added at the pure-builder, IO-wrapper (including merge-preserves- other-fields and corrupt-file-tolerance cases), and existing directory- listing assertions updated for the new settings.json file. Typecheck/ lint/format clean; full suite shows no new regressions (baseline pre-existing Windows-environment failures unchanged, 8 more passing tests than before — the ones added here). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
b45a96358e |
feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.
- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
registry entry declares contextLengthVar (currently only claude) and
the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
(40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
quick-start customModel path) check this before the swap-conflict
check and before launching/restarting anything, returning
{requiresContextWarning, modelId, contextLength, minSafeContextTokens}
— skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
_resolveContextWarningConfirm (session-ui.js), wired into both
_quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
(the path Claude actually uses) ahead of the swap-confirmation check.
Explains the fix in-modal: give the model an explicit larger -c/
--ctx-size in llama-swap instead of relying on --fit-ctx, which
optimizes for the biggest model that fits rather than the biggest
context.
Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
29984c639d |
fix(remote): stop the flush losing a chunk, and reset the host form's wake fields
Own review pass over the PR:
- `_flush` took the chunk out of the buffer only AFTER awaiting the write. Input
arriving during that await is enqueued (`waking` is still set, so it takes the
buffer path), and the 4 KB cap then drops the OLDEST chunk — which is the one
already on its way to the pane. The `shift()` that followed removed the NEXT
chunk instead, so the drop-oldest bookkeeping silently lost a chunk that was
never written, while the log line blamed the one that was. The chunk is now
removed before the await and re-inserted at the FRONT on a failed write, so the
order of the queue behind it is preserved. Regression test: a chunk enqueued
during the first write of a full buffer must still reach the pane (red against
the old order).
- `showCreateCaseModal()` reset the remote-host form fields but not the two new
wake inputs, so one host's MAC/command carried over into the next host that
form saved.
- The banner's pre-poll `wakeConfigured` labelled a command-only host as 'mac'.
Nothing reads the distinction, but the field is documented as which path is
configured, so it says the truth until the first poll corrects it.
- Stale `resolveRemote` comment ("only for sessions that have no usable target of
their own"): after the host config became authoritative in both directions it is
consulted on the TTL regardless.
|
||
|
|
7b947fa3f1 |
fix(remote): close the wake-state leaks and the dishonest wake budget
Review follow-up on the wake-on-LAN PR (five findings, all of them about the state the feature keeps and the budgets it inherits): - Wake state is dropped by `WebServer.cleanupSession` instead of the two delete routes, so it now goes with the session on EVERY cleanup path (cron, admin, scheduled-run teardown, error paths) instead of surviving with up to 4 KB of the user's buffered keystrokes. `registerSessionRoutes` returns the registry so the server can own its lifetime without the wake-capable code living in `server.ts`; the wiring guard is updated to allow that and gains a second assertion that `server.ts` calls nothing but `drop`/`stop` on it. - `_effectiveRemote` returns before `_state`, so a LOCAL session no longer gets a wake-state entry — the input gate runs on every keystroke, so that entry used to be allocated for every session the user types in. - An input chunk larger than the 4 KB cap is dropped OUTRIGHT instead of being head-trimmed and then written as a fragment: one paste is one `input` value and was never typed character by character, so its tail is a partial command the user never sent. The drop is logged. - The manual wake button passes `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s) like the create/attach paths, instead of inheriting the 90 s session default that the dashboard's reverse proxy cuts off at 60 s. - `RemoteWakeRegistry.stop()` aborts in-flight readiness polls (abortable sleep) and refuses new wakes, and `WebServer.stop()` calls it, so a restart during a wake no longer waits the poll out. - The banner/toast wording keys off a new `queuedInput` flag on the two SSE events, which is true only when the server actually holds bytes: browser keystrokes travel over the WebSocket, which never passes through the registry, so the wake BUTTON must not promise queued input. The failed-wake path also stops pattern-matching the error message (it re-asks the reachability route) and the WoL dialog says "admin-only" instead of "host not found" for a non-admin in multi-user mode. |
||
|
|
39976041e0 |
fix(sessions): let a dismiss reach the entries a restore is holding
Fourth review of the reboot-restore branch, and the third to find a defect in the previous round's fix. This one is the same shape as its predecessor: a counter keyed on one thing, compared against a set keyed on another. The generation counter was indexed by the entry's owner, while the in-flight set holds the caller doing the restoring. Those are the same person exactly when a user restores their own sessions, which is every case the tests covered. The route deliberately supports the other case: an admin may spend another user's entries. So when an admin restored Bob's sessions and Bob dismissed the banner, nothing matched, the entries came back, and a plan Bob had explicitly dismissed was re-armed for another twenty-four hours. Rather than reconcile the two key spaces, the counter is gone. `take()` now parks the entries it hands out, remembering which caller is spending them, and they stay parked until that restore ends. A dismiss filters the parked entries by `canAccess(entry.owner)` — the same predicate it already applies to the plan — so it reaches them wherever they are. `releaseFlight()` puts back only what is still parked. Expiry and a fresh boot plan unpark everything, for the same reason. There is one key space now, the entry's owner, and the spender is only ever used to tell two concurrent flights apart. That removes `generations`, `snapshotGenerations()`, `bump()`, `bumpAll()` and the argument threaded through the route. The discard grew the teardown it still lacked. A rebuild can fail after startInteractive() resolved, and a restored workspace still carries Codeman's hooks, so the CLI can post a hook event within milliseconds; the transcript watcher that starts from it, the attachment registry, the wait registry and the approvals inbox all outlive the listeners and would meet the retry, which reuses the session id by design. Its steps also run in reverse order now, so no live listener can reach a tracker that has already stopped, and the mux kill has its own guard, because stop() kills the pane in its last block after destroying four trackers. Tests. The run-summary test named an interval and asserted a map entry, so dropping stop() left it green; it now spies on stop(). Nothing pinned that before-spawn must precede setupSessionListeners, which reads the flag that phase restores, so swapping the two lines was silent; the ordering test now includes the listener setup. The retry assertion was a tautology and now asserts a different refs object. Both strengthened tests were verified by reverting their fix. Two new tests cover the admin-restores-another-owner cases this round was about. The server in the discard test is built once and stopped, since its constructor registers handlers on module-level watchers, and the workspace is removed through safeRmHomeTree. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
71ed7b127c |
fix(sessions): make the discard a real inverse of the construction
Third review of the reboot-restore branch. The narrow discard the previous commit introduced avoided everything cleanupSession() did wrongly, and in dropping so much of it also dropped four things it had to keep. The worst broke the retry the whole design rests on. setupSessionListeners() returns early while sessionListenerRefs still holds the session id, and the discard never cleared that entry. So the advertised flow — a rebuild fails because the agent binary is missing, the user fixes their PATH and clicks again — reused the same id, wired no listeners at all, and produced a tab that never showed output, never updated its status and never persisted. That is worse than the leak the discard was added to prevent. Three more registrations leaked with it: a RunSummaryTracker and its interval, an image watcher on the workspace, and the Ralph fix-plan watcher. The discard now undoes each registration setupSessionListeners() makes, in its order, and the per-session custom-model config directory, which holds the endpoint's API key literally and which nothing else would ever remove. The image-watcher flag was restored after the code that reads it, so a session came back reporting the feature as on with nothing watching. It moves to the before-spawn phase, and that phase now runs before the listeners rather than after them. The generation counter that lets a mid-restore dismiss win was global while clear() is ownership-scoped, so one user's dismiss discarded another user's unspent entries, permanently, because nothing rebuilds an in-memory plan. It is now per owner. Bumping only the owners of entries the dismiss removed was not enough either: take() has already emptied the plan by then, so a dismiss landing mid-restore saw nothing of that owner's to remove and invalidated nothing. The owners that matter are those with a restore in flight, filtered by what the dismissing user may access, and that is what clear() now bumps. Plan expiry bumps too, so a restore straddling the 24-hour boundary cannot hand entries back and give an expired plan another full day. Tests. discardPartiallyBuiltSession had no test at all: the only implementation any test ran was the mock's one-line stub, which is why every defect above was invisible. test/discard-partially-built-session.ts drives the real WebServer, and the retry assertion fails if the listener refs are left behind — verified by reverting the fix. The dismiss-race test drove the registry by hand, so deleting the route's generation argument left it green; it now goes through the route, and two further tests cover the multi-user cases. The mock context has now gone stale twice, because route tests pass it as `ctx as never` and tsconfig.json includes only src, so nothing ever compares it to the ports. A type-level guard is therefore inert — I wrote one and confirmed it never fires. test/mocks/mock-route-context-completeness.ts compares the mock's keys against WebServer.createRouteContext() at runtime instead, and names what is missing. Also: the API reference now says workspace-forbidden is judged against the owner's grant, the banner's module header no longer claims Restore always dismisses it, and the detail span gets the same min-width: 0 the phone rule already needed. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fa52753e8b |
fix(sessions): undo a failed rebuild without deleting the user's data
A second review of the previous commit found that its own repair for the session leak introduced three defects, all from reaching for cleanupSession() to undo a half-built session. That function is the user-initiated delete, not an undo. It banked the session's historical token and cost totals into the lifetime figures, and a reboot never runs cleanup, so those totals had never been counted before; every failed rebuild added them again. It saw the pin that had just been restored and demoted the record to `stopped`, which this pass reads as the durable marker of a deliberate kill, so a pinned session whose rebuild failed became permanently unrestorable. And it recursively removed `.claude-images` from the working directory, which belongs to the workspace rather than to the session, so a failed rebuild destroyed the pasted images of any other live session in that repo. discardPartiallyBuiltSession() now undoes only what the construction did: the map entry, the tab-layout slot, the listeners and any pane the launch created before throwing. The persisted record, the lifetime totals, the Ralph state and the workspace's files are left alone. Re-applying the persisted state also splits in two, which removes the first two defects at the root rather than only at the call site. The half that shapes the pane, the custom-model environment and the nice priority, still runs before the spawn. The half that is the session's own history now runs after it, so a session whose pane never started carries no totals and no pin for anything downstream to misread. The rest of that review. The multi-user workspace confinement re-check read the requesting user's grant, and returns true for an admin, so the case its own comment described was the one it missed; it now resolves the entry owner's grant through isWorkingDirAllowedForUsername, the way cron does. A forbidden workspace goes back on offer, matching both the registry's stated contract and the API reference. The client re-reads the plan after a restore instead of blanking the banner, so entries the server put back stay reachable, and a 409 now says a restore is already running rather than reporting a failure. A dismiss arriving mid-restore wins, through a generation counter the route carries across its take. The re-application also restores the tab colour, the image-watcher flag and the original pinnedAt, via a new Session.restorePin that does not re-stamp the pin time. The phone breakpoint gains min-width: 0, without which a nowrap flex item never shrinks and the buttons still overflow, and it folds into the existing phone block. Ralph's loop configuration still does not survive a restore, because toState() reads it off a live tracker and there is no way to keep it without arming the loop. The method now says so rather than leaving it implied. Tests. The capacity test could not fail on the property it existed for: it filled the board past the cap before the loop, so a single pre-loop check would have passed it. It now leaves one seat, so only a per-iteration check restores exactly one entry. New tests cover the ordering around the spawn, a throw before the loop returning the whole plan and releasing the flight, the dismiss-during-restore race, and that the failure path calls the narrow discard rather than the delete. The shared mock context gains the port method it was missing, which is what made the first run of these tests fail for the wrong reason. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
993710263d |
fix(custom-model): stop trusting /props's n_ctx, parse the real context size from /running's cmd
Root cause of the context-overflow regression reported live: "API Error: 400 request (36437 tokens) exceeds the available context size (16384 tokens)". Discovery had stored modelContextLengths.qwen3.8-27b-ud-q4_k_xl = 154112, so CLAUDE_CODE_MAX_CONTEXT_TOKENS told Claude Code it had a huge window and it never compacted - but the real llama-swap server was launched with --fit-ctx 16384 (confirmed against /running's own cmd field) and refused the request right at that real limit. /props?model=<id>'s n_ctx (the field discovery read) is confirmed live to be unreliable for a --fit-ctx-launched backend: it reported 154112 for the same model /running says was launched with --fit-ctx 16384 - appears to report the model's theoretical/trained maximum context, not the runtime- configured one. discoverModels() now parses the REAL configured size straight out of llama-swap's own launch command instead (parseCtxFromCmd(), reading /running's cmd field - --fit-ctx first, then the plain llama.cpp -c/ --ctx-size a hand-written command might use), and only falls back to the old /props probe when cmd states no recognizable flag at all. One /running call now covers every loaded model's context length in a single request, same as it already did for the swap-conflict check and the load trigger. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
7bbe408e44 |
feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout (updated every poll, so every second by default) instead of a static "this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1-3 min) on llama-swap - 47s remaining". If the countdown reaches zero and the model still isn't ready, this is now treated as a real failure rather than a "keep waiting" shrug: - The banner turns into a sticky error (_showCenterStatus gains a `type` option - 'error' drops the spinner and adds a close button, since nothing is "in progress" anymore and a sticky message needs a way to dismiss it), naming the llama-swap server's own logs as where to look for detail. - The session that load was for is closed automatically (closeSession) - requested explicitly: a console left open and pointed at a model that never finished loading is worse than no console at all. Both apply paths now thread the new session's id through to _watchLlamaSwapLoading for this (new required 3rd parameter, after endpointId/modelId). _watchLlamaSwapGeneration's existing stale-call guard extends naturally to this: a superseded call's own eventual timeout recognises it no longer owns the banner and neither shows the error nor closes a session that may by then belong to a different, newer launch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
55dae31530 |
feat(custom-model): estimate model load time from its discovered size
Discovery now also parses a GB figure out of an auto-discovered model's own description (llama-swap writes "Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike context length this needs no /props probe (the figure is right there in /v1/models) so it is populated for every model regardless of loaded state. A hand-configured profile's own description has no such figure and correctly gets no entry. The loading banner (_watchLlamaSwapLoading) now looks this up and, when known, shows it plus a rough estimate from a small size->time matrix (_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) - "Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on llama-swap... this can take a while" - and uses that same estimate's own bracket to scale the banner's default give-up timeout for a very large model, instead of a flat 5 minutes for everything. Explicitly labelled as an UNMEASURED, typical-hardware estimate in every relevant comment - this is not benchmarked against any real endpoint's actual storage/GPU, just a reasonable expectation-setter. A model with no discoverable size (a hand-configured profile) gets no size/estimate shown at all, matching the "never a guess" convention modelContextLengths already established. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
0af233c96c |
fix(custom-model): poll llama-swap readiness every 1s, check immediately, extend the cap
Reported: the "Loading..." banner stayed up past 2 minutes even though llama-swap itself had already finished loading the model. Three fixes: 1. pollIntervalMs default 3000ms -> 1000ms (as asked). 2. The loop now checks readiness IMMEDIATELY on entry rather than sleeping a full interval first - a model that's already ready (a fast load, or a re-apply onto one already loaded) shouldn't sit on "Loading..." at all. 3. maxWaitMs default 120000ms (2 min) -> 300000ms (5 min): a large (20GB+) model reading from disk can genuinely take longer than 2 minutes, which would have looked identical to the reported symptom - "still stuck past the point it should have resolved" - except it would have actually flipped to a "still waiting" warning toast at the 2-minute mark rather than staying on "Loading" indefinitely, so this alone doesn't explain what was reported, but is a real, separate improvement worth making. Also fixes a real, separate bug this surfaced while reasoning through the report: _showCenterStatus's banner is ONE shared, reused DOM node. A second call to _watchLlamaSwapLoading (e.g. switching models again before the first switch's loop had finished) would take over that shared banner, but the FIRST loop was still running and would eventually dismiss or overwrite it once ITS OWN deadline or readiness check resolved - clobbering whatever the second, current loop had put there. A generation counter (_watchLlamaSwapGeneration) now lets each call recognise when it no longer owns the banner and stop touching it silently, rather than only the last call to actually start ever safely reading or writing it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
fbede5cd2a |
fix(sessions): act on the dual review of the reboot-restore route
Fifteen findings from two independent reviews of #442, three of them blocking. Every one is addressed here. The three blockers all sat in the restore route. A rebuild that threw after addSession left a registered session with no pane behind it, visible on the board, holding a layout slot and written to state.json, with its plan entry already spent; the catch now cleans the session up and puts the entry back. The loop checked neither the global nor the per-user session cap, so one click could take a board past a documented limit; capacity is now re-checked per iteration, because the loop is itself creating the sessions it counts. Worst of the three, a rebuilt session carried none of the state its constructor has no parameter for and then persisted itself over the record that held it, zeroing token and cost totals and dropping the pin. The pin matters most: pruning keeps a record only while it is pinned, so discarding it handed the record to the next stale sweep. A new reapplyPersistedSessionState() on the session port restores the pin, the token totals, auto-compact, auto-clear, auto-resume, nice priority, the flicker filter and the custom-model selection, and it runs before both startInteractive and the first persist. The rest, in the order they bite a user. Every rebuild failure was reported as workspace-missing, so the banner told users their repo was gone when the agent had simply failed to start; there are now distinct reasons, and the toast names each one. The client read restored and skipped off the outer response object rather than through the uniform envelope, so every count came back zero and neither toast ever fired. A board left open across the reboot never learned an offer existed, because the banner was seeded only on the page-load path; it now re-reads on every SSE init. The workspace check was existence-only, skipping the multi-user confinement that the create route applies, so a withdrawn grant would not be noticed. The banner had no phone breakpoint while its text was nowrap and its buttons could not shrink. Smaller: a missing workspace is now re-offered rather than dropped, while an already-open conversation is dropped rather than re-offered forever; a throw anywhere in the route returns the unspent entries instead of discarding the plan; the single flight is keyed by owner, since take() already stops two callers receiving one entry; the env clamp's header no longer claims a protection it cannot provide on this path today, and names the check that does bite; the three endpoints are documented in docs/api-reference.md; and the module header now says that os.uptime() reads the host's clock, so the feature is effectively off inside a container. The review also explained why the tests missed all of this: they proved the construction claim through their own copy of the construction rather than through the route, and the route tests used workspaces that did not exist, so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts mocks the Session module to drive the route's real path, and covers the cleanup, the reason reported, the re-application ordering, the broadcast and the caps. The mock route context gains the port method and the mux call the route needs. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a7f74f374f |
fix(remote): keep the wake banner hidden after switching to a local session
refreshHostWakeBanner clears _hostWake before calling _hostWakeTick, so the clear branch's `if (this._hostWake)` guard skipped the repaint: once the banner had appeared for an unreachable remote session it stayed up on every chat (local ones included) until a reload, and the 30s ticker never cleared it either. Render unconditionally in that branch — _renderHostWakeBanner is idempotent with a null state. Reproduced in a real browser (Puppeteer, mobile viewport): state went null but banner.hidden stayed false. Regression test added in test/host-wake-banner.test.ts (red before, green after). |
||
|
|
0929694012 |
fix(custom-model): actually trigger the llama-swap load, not just watch for it
Root cause of "it doesn't look like llama-swap is actually switching the model" (confirmed live: no load_model line in llama-swap's own logs after applying a selection). llama-swap has no "switch model" admin endpoint - the ONLY thing that starts a swap is a real inference request naming the model. Every previous fix (the conflict check, the loading banner) assumed a swap would start on its own; nothing ever actually asked llama-swap to load anything until the launched CLI's first real prompt did, which could be much later than "applying the selection" implied. Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest real request that will start a load - POST <baseUrl>/v1/chat/completions, max_tokens: 1, one throwaway message - fire-and-forget (never awaited by the caller; the frontend's own running-status polling is what actually confirms readiness). Wired into both apply paths (the dedicated restart route and the one-shot quick-start route), fired whenever the target model isn't already the one loaded and ready - a broader condition than the existing swapNeeded (which only gates the "this will evict another session's model" confirmation ask and deliberately stays narrow to that). modelSwapInProgress in both routes' responses now reflects this same broader condition too, so the frontend's loading banner actually correlates with a real in-flight load rather than only firing when something else happened to be loaded already. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
01b32ee6cd |
fix(custom-model): move the switching/loading status to a centred banner
The "Claude started - switching to <endpoint>..." and "Loading <model> on <endpoint>... this can take a while" messages lived in the top-right toast corner along with everything else, easy to miss given they can each sit on screen for well over a minute (a real llama-swap model load). Adds _showCenterStatus() (panels-ui.js): a single, reused, screen-centred banner with a spinner, non-blocking (no backdrop, pointer-events: none on the wrapper) so it never gets in the way of using the app while it's up. Both call sites (_runCustomModelEntryViaRestart's switching message, _watchLlamaSwapLoading's loading message) now use it instead of showToast. Every OTHER status in these two flows - the llama-swap conflict warning already moved to its own modal, apply failures, cancellation, and _watchLlamaSwapLoading's own final "ready"/"still waiting" outcome - stays exactly where it was, in the corner. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
2936ba6e3d |
fix(custom-model): replace the native confirm() popup with an in-app modal
The llama-swap "this will unload it for session X" warning used a native browser confirm() popup, which looks out of place next to the rest of the app's own modals. Adds #customModelSwapConfirmModal (index.html) with Cancel/Switch-anyway buttons, styled to match the app. _confirmModelSwap(message) shows it and returns a promise that resolves true/false the same way confirm() would; _resolveModelSwapConfirm(proceed) (wired to both buttons and the backdrop click) settles it. Both llama-swap conflict call sites (_quickStartWithCustomModelConfirm for the one-shot launch path, _runCustomModelEntryViaRestart for Claude's restart path) now await this instead of calling confirm() directly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
83033b4299 |
fix(custom-model): show a status toast during Claude's native-boot-then-restart window
Claude stays on the launch-then-restart path (see runCustomModelEntry's own comment for why), but with nothing on screen during that window, a native boot that briefly talks to the cloud model read as "the endpoint didn't apply" rather than "the switch hasn't happened yet". A sticky "Claude started - switching to <endpoint>..." toast now covers the whole window from the native launch through the apply call, updated in place (never stacked) as the outcome resolves: dismissed on cancel or failure (replaced by the existing cancellation/error toast), handed off to _watchLlamaSwapLoading's own sticky toast when a model swap is in progress, or updated to the existing "Pointed at ... - restarting" message and auto-dismissed after 3s on a plain success. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
f865f74a0f |
feat(custom-model): launch directly on the endpoint, no restart, for 7 of 8 CLIs
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.
POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.
Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.
Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
bcebc81fcd |
feat(custom-model): detect llama-swap model conflicts before switching
Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.
1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
server) via its own GET /running, which plain llama.cpp has no concept of
at all. New GET /api/model-endpoints/:id/running-status route exposes this
read-only, for the frontend's polling loop below.
2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
what llama-swap currently has loaded. If it differs from the requested
model AND another live session's own customModel selection is actively
using that loaded model, the apply is refused with a
{requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
instead of silently switching. A "confirmed: true" field on the retry
skips the check. Switching with nothing else affected proceeds
immediately, no confirmation asked, only ever when there is something to
warn about.
3. The frontend (runCustomModelEntry) shows a native confirm() naming the
affected session(s) and the model they'd lose, matching this codebase's
existing convention for this class of decision (delete case, kill
session, etc.) rather than a new modal. On a successful apply the response
also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
poll shows a sticky "Loading <model>..." toast via the new running-status
route until llama-swap reports the target model ready (bounded at 2
minutes), so a prompt sent mid-swap reads as "loading", never as silence
or an answer from whatever was loaded a moment before.
Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
da933d70be |
feat(sessions): offer to rebuild the sessions a host reboot destroyed
A host reboot takes the tmux server down with it, so every pane dies, reconciliation finds nothing to attach to, and the board comes up empty. Picking yesterday's work back up meant finding each conversation in history and resuming it by hand, one at a time. The boot pass now works out what the reboot killed and leaves it on offer. It runs inside restoreMuxSessions(), in the window where reconciliation has reported the dead sessions and cleanupStaleSessions() has not pruned their records yet, which is the only place the records can still be read. The board shows a banner, and nothing is created until the user clicks it. A click rather than an automatic restore is what makes the reboot heuristic acceptable. The heuristic cannot tell a reboot from a crash that took tmux down inside the same window, so it decides whether to ASK, never whether to act: a wrong yes costs a line of text the user dismisses instead of N CLI processes nobody asked for. Four things are re-checked when the click arrives rather than trusted from boot, because hours can pass and the board moves on. The owner's privilege grant re-resolves through the env clamp. The workspace must still be on disk. A conversation the user already resumed by hand from the Resume list is skipped, since two panes running --resume on one conversation would fight over the same transcript. Entries leave the plan synchronously before the first await, and the route is single-flighted, so a double-click or two devices cannot both reach the same entry. A restored session comes back attached, idle and disarmed. Respawn controllers and Ralph loops are deliberately not re-armed: a machine that just came up is the worst moment to turn an autonomous run loose. Its workspace hooks are installed by the restore route itself, because the boot-time sweep sits behind a gate that is false after a reboot and has finished long before the click; without them a session goes silently blind, with no stop or idle events for respawn, no Approvals Inbox item and no red tab on a blocking dialog. Stats collection starts the same way. The pane is new, so the conversation continues and the terminal scrollback does not. The banner says so rather than letting an empty pane read as a broken restore. The plan lives in memory only. A server restart drops it, which costs the convenience this adds and never the conversation: the conversation is the transcript under ~/.claude/projects, which the Welcome screen's Resume list and the Session Manager already read, so a dropped plan returns the user to resuming by hand. clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the question it answers is about session privilege rather than about HTTP and it now has a caller outside the route layer. Its test hook stays re-exported from session-routes.ts. Claude sessions only for this pass. The other CLIs name their thread in their own config object, which this does not thread through yet. Remote and docker sessions are skipped on purpose, because both need another host or a container to be up and a freshly booted machine cannot promise either. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
25f22b9839 |
test(custom-model): update session-custom-model route test for CLAUDE_CONFIG_DIR isolation
Fixes the CI failure on the last two commits: this route test asserted an exact envKeys list for a claude-mode apply that predates the CLAUDE_CONFIG_DIR isolation fix, so it failed on the new CLAUDE_CONFIG_DIR entry it correctly started appending. Updates the expected list and adds assertions for the isolated config dir path and the pre-seeded .claude.json trust-approval file, matching the behavior added in the two prior commits rather than just tolerating it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
97464bfa27 |
fix(custom-model): pre-approve the injected API key in the isolated Claude config dir
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.
Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
0e8b1981af |
fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker: 1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same config directory and warns about it (confirmed cosmetic - the API key wins for actual requests, verified via a real session's own API Usage Billing line). A custom-model claude session now gets an isolated CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty, no files written into it) so there is nothing to conflict with. projects is symlinked (junction on Windows) back into the real config dir so the response viewer, subagent windows and Read My Mind keep working for that session, best-effort. 2. Context-window overflow. Claude Code assumes a large default context window for a model id it doesn't recognise and never compacts, so a custom endpoint's real, much smaller context (verified live: a 400 exceeding a 16384-token llama-swap model with a stock ~33.7K-token system prompt) silently overflows. Discovery now also learns each model's real context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx), but ONLY for a model llama-swap's own /v1/models response already marks status.value === 'loaded' - never an unloaded one, since llama-swap treats ?model= as a routing hint and probing an unloaded model risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery. A server with no status field at all gets no enrichment rather than a guess; a model not probed this round keeps its previously-learned value until it disappears from the list entirely. Stored per model (CustomModelHost.modelContextLengths) and applied via a new contextLengthVar registry field, set to CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude. Both new fields live on the existing env-kind customModelInjection capability shape, declared only on claude's registry entry - every other CLI's injection is unaffected (pinned by test). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
5c25a52f95 |
fix(custom-model): wait for a freshly launched session to go idle before applying
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.
Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.
Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
409a6e65f9 |
fix(custom-model,toast): surface the real apply error, and make error toasts sticky with a close button
Two related fixes, both needed to actually diagnose 'Session started on the native backend — could not apply the custom endpoint' reports from live testing: 1. runCustomModelEntry()'s apply call went through _apiJson(), which unwraps a success body but SWALLOWS a failure response entirely and returns null — discarding the one thing (error, errorCode) that would tell 'endpoint unreachable' apart from 'not a discovered model', 'remote/Docker session', or a dozen other real causes the apply route already reports distinctly. Switched to _api() so the actual response body is read on failure too, and the toast now includes the real message. 2. showToast() defaulted every toast, error or not, to a 3s auto-dismiss with no way to read it again — exactly what made the above generic message impossible to act on even before the fix above. Error toasts now default to sticky (duration: 0, no auto-dismiss) unless a caller opts into a duration, and every toast — sticky or not — gets an explicit close (x) button, since a sticky toast with no way to dismiss it would just accumulate across repeated failures. Tests: custom-model-run-menu-ui.test.ts's two apply tests updated for the _api() switch (their mocks previously stubbed _apiJson, which the apply call no longer goes through), plus a new test pinning that the real server error string reaches the toast on a failure. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
a1c35da0d8 |
fix(input): deliver a recovered keystroke before the Enter that submits it
Every message typed on an Android phone lost its last character. An Android soft keyboard commits the last typed character and sends the Enter key in ONE InputConnection transaction, so the committed-text `input` event and the Enter keydown are both processed before any zero-delay timer runs. The orphaned-input recovery from #388 resolved its candidate only on such a timer, and that lost the character twice over: * ORDER — xterm emits `\r` synchronously from the Enter keydown, and the local-echo composer submits `pendingText` right there. The recovered character arrived one macrotask too late to be part of the prompt. * LOSS — that same `\r` bumps the canonical counter, so by the time the candidate resolved, `canonicalCount > snapshot` read as "xterm spoke for this keystroke" and stood the recovery down. The character was not merely late, it was dropped. Drain pending candidates synchronously at the next keydown instead, from xterm's custom key handler, which runs before xterm processes that key. The counter then still holds the value it had while the candidate's own keystroke was current, so the stand-down decision is made against the right keystroke, and the recovered byte reaches the composer ahead of whatever the new key emits. The timer stays as the fallback for a keystroke with no key after it. Physical keyboards are unaffected: there the timer has already resolved the candidate long before the next key arrives. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5a9ff07f57 |
feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real llama.cpp server: 1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to apply the endpoint's defaultModelId (or the first discovered model) silently. Now, via the new selectCustomModelEntry() (session-ui.js): - exactly one discovered model launches straight away, same as before - two or more open a new #customModelPickModal listing every discovered model; defaultModelId (if set) is marked but never auto-chosen, since the point of asking is letting ONE launch deliberately differ from the saved default, not just confirming it The endpoint is re-fetched at click time rather than trusting anything cached from the dropdown's own render, since the model list can have changed (the sweep below, or a settings-panel edit) since it opened. runCustomModelEntry() itself — the actual launch, routed through run() for the in-flight lock, snapshot-guarded against applying to the wrong session — is unchanged; it now just always receives an explicit model id from one of these two paths instead of computing one itself. 2. Periodic re-discovery. Every saved endpoint's models now refresh automatically every 5 minutes in the background (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way as the Codex plan-usage poll it sits beside — this.cleanup.setInterval, off under testMode), so a model the server starts or stops serving shows up without another manual "Discover" click. The manual POST .../discover-models route and the new refreshAllCustomModelHosts() sweep (custom-model-routes.ts) now share one pure merge step (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId that no longer appears) rather than two copies that could drift. The sweep is best-effort per host — one endpoint being unreachable on a cycle never blocks the others — and re-reads the store before each host's write, keyed by id, so a concurrent edit or delete from the settings panel always wins over a sweep that started before it. Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated file for the sweep (kept separate from custom-model-routes.test.ts because that file's data dir is shared across every test in it — one temp HOME per FILE, not per test — which would make a sweep-touches-every-host assertion meaningless there). test/custom-model-run-menu-ui.test.ts gained a new describe block driving the real picker modal through JSDOM: single-model bypass, multi-model dialog with the default marked-not-chosen, picking a row closes the modal and launches with that exact model, the endpoint re-fetch, and the two "vanished by click time" toast paths. Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md, docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated — the last of these also caught up two sentences that had gone stale after the draft-review fixes landed (the picker routes through run() now, not a raw run*() call). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
60e1bd52f7 |
fix(custom-model): act on the draft review — unparseable onclick, unwrapped envelope, wrong-session apply, missing lock, no tests
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.
Blockers:
1. Every generated inline onclick was unparseable. JSON.stringify's own
double quotes terminated the double-quoted HTML attribute at the first
one, leaving btn.onclick null on every picker entry and every Discover/
Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
argument, the same idiom deleteCase's onclick already uses four lines
away in session-ui.js. This also closes the live-HTML-injection route
through modelId (server-controlled, from the endpoint's own /v1/models
reply): with quoting intact, a `>` inside it can no longer terminate the
<button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
like every other /api route (server.ts's preSerialization hook applies
to arrays too), so Array.isArray(hosts) was always false in production
and the picker/settings panel silently saw nothing. Both call sites now
go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
returns normally without ever changing activeSessionId, so the apply
step used to silently re-point and restart whatever session the user was
already looking at. runCustomModelEntry() now snapshots activeSessionId
before the launch and requires it to have actually changed.
Majors:
4. Routes the launch through run() itself via a temporary _runMode swap
(never persisted — setRunMode() would sync it to the server) instead of
a parallel hardcoded dispatch table, so a custom-model launch now holds
the same _runInFlight lock every other Run click gets. This also
resolves the "hardcoded runners map contradicts the PR's own design"
minor: dispatch is run()'s own, so a CLI whose customModelInjection
recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
parses markup this module generated itself) for exactly the DOM-level
facts the review said needed no Playwright and no tmux: a generated
button's onclick genuinely compiles and fires, a dangerous modelId never
produces a live element, the envelope unwrap works, the session-changed
guard holds, run() actually gets called (proving the in-flight lock
engages), and _runMode is restored afterward. Confirmed against the
pre-fix code first (reproduces btn.onclick === null exactly) so this
isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
render-index-html.test.ts for the other fixes below.
Minors:
- Generated entries now filter through isCliAvailable(), matching
_refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
(applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
and to settings-modal open) instead of always rendering; the endpoint GET
no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
redactApiKey() replaces the field with a computed apiKeySet: boolean, and
a PUT with no apiKey now keeps the stored one server-side
(applyStoredApiKey()) instead of the client resending a value it was
never given. New tests cover both directions (kept vs. replaced) by
observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
(_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
event, since the real role can resolve after settings were first opened)
— endpoint writes were already admin-only server-side, but the button
used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
(CLAUDE.md already records that exact literal turning the settings
preview into a grey slab on light skins), .run-mode-custom-models gets
the same gap: 2px .run-mode-menu's own flex gap only applies one level
up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
</script> (CliEntry.label is user-clis.json-settable, unlike
__codemanCliAvailable's booleans-only payload) via a new exported
escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
docs/api-reference.md; left the "no zh-CN for the new Models-section
group" minor unaddressed only insofar as the wider Models section (task
routing, thinking effort, etc.) has never had zh-CN coverage either —
everything this PR itself introduces (labels, hints, button text, the
Run-menu's "Custom Endpoints" header) IS translated in i18n.js.
Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
fed6582d3e |
fix(test): strip the custom-model Run-menu picker's injected script too
CI on PR #430 failed test/server-index-title.test.ts's byte-identity check: renderIndexHtml now injects a second unconditional <script> before </head> (window.__codemanCustomModelClis, added alongside the existing __codemanCliAvailable one), and the test only knew to strip the older one before comparing the rendered HTML against the raw template. Strip both. Unlike __codemanCliAvailable (an object, historically injected only where something resolved), the new one is a plain array injected unconditionally, possibly empty, so it needs stripping on every machine, not just one with CLIs installed. Verified the two replace() calls compose correctly against the exact strings server.ts actually produces (simulated in isolation; this box has no tmux, so the real WebServer-backed test file cannot run here at all -- same environment gap noted throughout this PR's review). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
25fae9ad10 |
feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment: "generate those entries from the saved profiles rather than a fixed duplicate per harness, and put it in a follow-up PR so this one stays the backend... The Run-menu picker is yours if you want it." Adds the frontend surface the backend has been waiting on: - Run menu: a "Custom Endpoints" section lists one entry per (harness that supports customModelInjection, saved endpoint) pair, e.g. "Claude Code (llama.cpp)". The harness list comes from window.__codemanCustomModelClis, injected at page render straight off the CLI registry's own capabilities (never a hardcoded id list in the frontend), so a CLI whose injection recipe lands later appears with no frontend change. Picking an entry runs that harness's own existing run*() function unmodified (case creation, env overrides, everything, forced to a single instance) and then applies the endpoint's default model to the session it creates via the existing POST /api/sessions/:id/custom-model route. Entries are hidden for a remote/docker active case, since that route already refuses both. - Settings: App Settings -> Models gets a "Custom model endpoints" group wiring up the customModelEndpointsEnabled toggle (declared since #393, read by nothing until now) plus CRUD against the existing /api/model-endpoints routes: list, add/edit (inline form), delete, discover models. - Backend: CustomModelHost gains an optional defaultModelId, the model the picker applies with no further choice per endpoint (one generated menu entry per CLI+endpoint pair, not per CLI+endpoint+model). The route refuses a value that isn't one of the endpoint's own discovered models, and a fresh discovery drops a default that no longer appears rather than carrying an invalid one forward. Docs: docs/custom-model-endpoints.md describes the new picker and settings panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the "backend-only" status note and documents the picker's generation mechanism. Tests: four new route tests cover defaultModelId validation, acceptance, and the drop/keep behaviour across a re-discovery; a new render-index-html test pins the __codemanCustomModelClis injection (present, agent CLIs supporting the capability, antigravity and shell excluded) and its solo-window skip. No browser test was added for the Run-menu picker itself or the settings CRUD panel (this box has no tmux, so the live server used by test:browser/test:mobile could not be exercised here) -- worth a Playwright pass before merge, same as any other frontend PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
4a30f510e6 |
fix(remote): let the host config turn wake-on-LAN OFF for a live session too
Found by driving the real UI: with a MAC configured in remote-hosts.json, removing it (here: to reach the "Configure WoL" dialog) changed nothing for a running session — _effectiveRemote short-circuited on the session's own snapshot whenever that snapshot HAD a target, so the resolver was only ever consulted in the one direction where the feature was missing. The documented "host config is authoritative" promise therefore failed in the direction a user can actually observe, and a wake target could live on invisibly after being deleted from the config. The resolver is now consulted on the TTL regardless, and wins for the wake fields in both directions. Also adds a route test for the browser's real input shape: one POST per keystroke, all buffered during a wake, replayed IN ORDER. |
||
|
|
d0a5a583cd |
feat(remote): wake a sleeping host when a session is created or attached
Pressing Run on a remote case whose host was asleep failed with `could not verify tmux on remote host 192.168.50.137: …` — an ssh error that blames tmux for a machine that is merely suspended. The only wake paths were typed input on an established session and the banner's Wake button, so OPENING a session (the moment the user actually decides to use that host) had none. `RemoteWakeRegistry.ensureHostAwake()` reuses the existing probe/wake/readiness machinery for a host that has no session yet, and is wired into the two user-initiated create paths: `POST /api/quick-start` for a remote case (before the tmux prereq probe, which is what surfaced the misleading error) and `POST /api/sessions` with `attachRemoteSession`. A host without a wake target is not even probed, so its behavior and latency are byte-identical. The wake is blocking — the caller gets the session or an error — but bounded by REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS (40 s) instead of the 90 s session default, because the dashboard sits behind a reverse proxy whose default `proxy_read_timeout` is 60 s: a longer wait would be cut off at the proxy while the session was still being created. The budget has to cover the whole request (40 s wake + 1.5 s probe + the tmux probe's own 15 s = 56.5 s worst case), which is why it is 40 s and not 45. A timeout now says the host did not come back, and an unreachable host without a wake target says so instead of pointing at tmux. The wiring is deliberately in the HTTP ROUTE, never in the shared session service: `cron-service.ts` builds sessions there with nobody waiting on the answer, and a wake on that path would power the host on for every schedule — the timer-driven re-wake invariant #1 exists to prevent. Both halves are asserted (importers of `remote-wake`, and `ensureHostAwake` having exactly one caller file), so a future caller has to come through the guard test. A rejection from the wake IO is caught too: a broken target must fail the wake, not the route. `remote:hostWaking`/`remote:hostWakeFailed` now carry `forNewSession` for the session-less case, where "input is queued" would be untrue; the toast then reads "the session starts when it is back". Live wake numbers are unchanged (this reuses the measured ~12 s S3 path); the route behavior is covered by new tests in session-routes.test.ts with an injected registry, so no test opens a real socket or ssh. |
||
|
|
8dfc965d13 |
fix(remote): stop the wake handlers shadowing each other; enforce the input cap
Two findings from a final review pass over the wake-on-LAN feature. `_onRemoteHostWaking` / `_onRemoteHostWakeFailed` were defined in BOTH `panels-ui.js` (toasts) and `host-wake-ui.js` (banner). Both files mix into `CodemanApp.prototype` and `host-wake-ui.js` loads later, so the panels-ui copies were silently shadowed: the toast never fired, and a wake started for a BACKGROUND session (input on a non-active tab) produced no notification at all, since the banner handler only acts on the active session. The handlers now live only in `host-wake-ui.js`, show the toast unconditionally, and update the banner when the woken session is the active one. `appendBoundedPending` dropped only WHOLE chunks, so a single input value over the cap (one large paste is one `input` value, up to the 100 KB input schema) was kept in full: "bounded at 4 KB" held per chunk, not per session, and nothing was logged. The surviving chunk's head is now trimmed too, code-point aware so a multi-byte character is never split into a replacement char. Adds the guard that would have caught the first one: every SSE dispatch handler must be defined in exactly ONE frontend module. The existing test only asserts a handler EXISTS somewhere, which two modules both satisfy while one is shadowed. |
||
|
|
c9515b1d4c |
fix(terminal): keep the output a pane capture could not contain
Live terminal events are queued while a buffer load runs, and the load discards that queue when it ends. That is right when the loaded buffer is the server's accumulated byte history. The route appends to that history right up to the moment it serializes the response, so a queued event already appears in it and replaying it would duplicate output, most visibly Ink's cursor-up redraws. A tmux pane capture is a photograph, current only as of the instant `capture-pane` ran. Output printed afterwards was queued and then dropped, and nothing scheduled a re-fetch to recover it: `_onSessionNeedsRefresh` is wired only to the 128KB overflow path. The CLI's next partial redraw then landed on a frame the terminal never received. How much went missing depended on which capture the route served. A `?full=1` load returns the capture alone, with no history in front of it, so it lost everything from the capture to the end of the chunked write. A `?tail=` load returns history, a clear, and then the capture, and the route reads that history after the capture, so it lost everything from the response to the end of that write. The chunked write dominates either way. An agent CLI hides the loss on its next full redraw; a shell session does not, because its output is linear and nothing repaints it. Queue entries now carry their arrival time, and `_finishBufferLoad` takes a `since` cutoff, so a capture load replays exactly the tail that arrived after the response headers. The earlier events stay dropped, because a payload that carries history does hold those. All four paths that fetch a terminal buffer and write it now decide this the same way, through one `_bufferLoadFinishOpts` helper, so they cannot drift apart: `selectSession`, `_onSessionNeedsRefresh`, `_onSessionClearTerminal` and `_maybeRefetchFullHistory`. The second of those is the one that stings. It exists to restore output the client already dropped once under backpressure, and it was dropping more output while performing that recovery. The cache-hit write inside `selectSession` stays on discard deliberately: it runs before the fetch, so its queue holds only events the capture that follows already contains. Two further things had to change for that tail to still exist when the load ends, and a browser test is what found both. `chunkedTerminalWrite` is what ends the load for every non-empty buffer, so the flush policy travels to its own finish calls; the call in `selectSession` runs only when the write was skipped. `_beginBufferLoad` no longer empties the queue when one load re-enters it, which it does on every write, because that reset discarded the whole fetch window before anything could replay it. The response already distinguishes the sources. `source` reads `mux-visible` or `mux-full-history` for a capture and `history` for the byte stream. Follows #395, #396 and #397, which fixed the ways the replayed frame itself could disagree with the terminal. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5b920cb43d |
feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream and named tabs wrong five ways (every prompt, every write path, a bare Esc eating the next prompt's first character, pasted newlines as Enter, any CSI clearing the draft) and replaced the whole name, which dropped the case from the tab and reset the w<n> counter. This lands the feature with each of those closed: - First prompt means the first: applyAutoName() flips a placeholder to `auto` whether or not the string changed. nameSource is now the tri-state placeholder | auto | manual; the name setter is the only manual path. - Only user-originated input counts: write()/writeViaMux() take SessionWriteOptions.fromUser, set by the browser WS path and POST /input only, so Ralph, respawn, cron, approvals and the trust-dialog keys can never name a tab. A startMode 'shell' CLI never feeds the tracker (a capability, not an id check); the send-key route feeds trackUserInput() because its line feed bypasses the session. - Prefix form `w3-case: title`: parseSessionPrefix() already renders it as the title with the prefix in the tooltip and the next-session counter still matches it. Composed within MAX_SESSION_NAME_LENGTH. - Tracker rules per key: bare Esc resolves at chunk end; mouse/focus reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R taint the draft so Enter submits nothing rather than a fragment; bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space; the draft keeps its head past 8192 code points; an escape past 64 bytes is abandoned. - Title: slash commands by shape (a path is a prompt), `!` escapes refused, first sentence only past 8 code points ("e.g." is not a title), 72 code points on a word boundary. - Synced `autoNameSessions` setting, default OFF (the prompt reaches mux-sessions.json, session:updated and /api/search), App Settings -> Appearance -> Tabs, read fresh per prompt after the eligibility check. Tests: test/session-auto-name.test.ts (tracker, title, composition, ownership, emit gating), the wiring test (once, prefix, setting off, manual protected), test/routes/session-name-routes.test.ts (PUT /name flips to manual and persists). Verified live on an isolated instance: API and browser-typed prompts name the tab, a second prompt does not, shells and renamed tabs are untouched, nameSource survives a restart. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
c4322513d9 | Merge pull request #376 from shenlvkang-collab/feat/auto-session-names-upstream | ||
|
|
1380b023e2 |
fix(remote): make the wake banner's poller page-wide and independent of tab switches
Reported as 'the tab shows no banner' while the host was verifiably unreachable: the banner only started polling from selectSession, which RETURNS EARLY for the tab you are already on (so a page loaded with the remote tab active never polled), and a long-lived tab keeps running the JS it loaded — the feature was invisible to anyone who did not switch tabs after the deploy. The poller is now page-wide: one interval (created on init and on the first session switch), re-targeted whenever the active session changes, plus a visibilitychange wake-up. It no longer depends on any single selection path running. Also adds test/sse-dispatch-table.test.ts: a static guard that every [SSE_EVENTS.X, '_onFoo'] entry names an event constants.js defines AND a handler some module defines. Both halves fail silently (a typo'd constant is an undefined table key; a renamed handler just never runs), which is exactly how a new banner can never appear with no error anywhere. |