mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
ecb95b5d6774d1d1b1b9eee8e4c205b64bcf115b
1644
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ecb95b5d67 |
Merge pull request #459 from opticon454/feature/run-menu-picker-currently-loaded-model
feat(custom-model): promote the currently-loaded/last-used model in the Run-menu picker |
||
|
|
88e5b7b200 |
fix(custom-model): address Ark0N's PR review — client-side probe timeout, defer "last used" past confirmation, docs, zh-CN
Four things from the maintainer's review on PR #459, all fixed: 1. Bound _getCustomModelCurrentlyLoaded's probe client-side (~800ms via Promise.race, on top of — never instead of — the route's own 5s server-side timeout). Without it, an asleep/firewalled endpoint behind a saved model list left the picker completely invisible for up to 5s after the Run menu had already closed, with no spinner or toast. `timeoutMs` is an optional param (default 800, real callers never pass it) so a test can drive it in milliseconds, same pattern as `_watchLlamaSwapLoading`'s own `pollIntervalMs` — this code runs in a JSDOM window's own realm, whose setTimeout vi.useFakeTimers() cannot patch. 2. "Last used" is now written only once a launch actually applies, never on the mere click. It moved out of runCustomModelEntry (unconditional) and into each path's own success point: _quickStartWithCustomModelConfirm after the final post succeeds, and _runCustomModelEntryViaRestart right after the apply's success check. A context-window-warning decline means this exact model cannot work with this CLI at all, so the old unconditional write would promote, next time the picker opened, the one model guaranteed to fail again. 3. Documented the promotion/tag precedence and the new codeman:customModelLastUsed:<mode>:<endpointId> localStorage key in both CLAUDE.md's Custom Model Endpoint Profiles section and docs/custom-model-endpoints.md's Run-menu picker section. 4. Added zh-CN entries for "Currently loaded" and "Last used" in i18n.js, next to this modal's existing "Choose a model"/"Custom Endpoints" pair. New tests: the client-side timeout (endpoint that never answers, one that answers within the bound, and a rejected-after-timeout probe settling quietly), and "last used" recording on success vs. NOT recording on either confirmation's decline, for both the restart and one-shot paths. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
73607663fd |
fix(custom-model): guard the picker's async open against a slower, superseded probe
Code review (high effort) on the previous commit found a real race: making _openCustomModelPickModal async (it now awaits the currently-loaded-model probe before rendering) meant a second, faster call for a different endpoint could render first, only for the first call's slower probe to resolve afterwards and overwrite the modal with the wrong endpoint's model list — while _pendingCustomModelPick (set synchronously, before either await) still named the second, correct endpoint. Picking a model in that state would launch/apply the wrong model on the wrong endpoint. Fixed with the same mutable-generation-counter guard _watchLlamaSwapLoading already uses for an identical async-superseded-by- newer-call shape: every DOM write, including _pendingCustomModelPick itself, is deferred until after the awaited probe, and a call that finds its generation already superseded bails out untouched instead of clobbering whatever a newer call already rendered. Added a regression test driving two overlapping opens with a controlled promise so the earlier, slower probe resolves after the later, faster one renders, asserting the late response is a no-op. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
458ca578e7 |
feat(custom-model): promote the currently-loaded/last-used model in the Run-menu picker
Custom Model Endpoint Profiles' "which model" picker (session-ui.js's _openCustomModelPickModal) always listed models in their raw discovery order, so on a host with several downloaded GGUFs the user had to remember (or eyeball the "Default" tag) which one llama-swap actually had hot before picking — the whole point of the picker being fast is undone if it makes you think first. The picker now promotes exactly one model to the top of the list: - If llama-swap reports a model from this host's own list `ready` right now (via the existing GET /api/model-endpoints/:id/running-status route), that model is promoted and tagged "Currently loaded" — it's what a launch attaches to with zero wait. - Otherwise, the last model actually launched on this exact (harness, endpoint) pair is promoted and tagged "Last used", read from a new per-device localStorage key (codeman:customModelLastUsed:<mode>:<endpointId>), written by runCustomModelEntry on every launch attempt regardless of outcome. - A plain (non-llama-swap) OpenAI-compatible server, an unreachable endpoint, or a loaded-but-not-yet-ready model never promotes anything — the rest of the list keeps its discovery order. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
2df9355367 |
fix(cli-registry): address PR B2 review — fix two test guards, drop unused catalogue
Two required fixes from Ark0N's review of #458: 1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on <file>::<line>::<expression>. A single inserted line anywhere above an entry shifted every subsequent line number, so all 21 entries went stale simultaneously and the same 21 branches were reported as "new" — on a file six other open PRs also touch. Dropped the line number from the key (<file>::<expression>, matching the backend guard's own design), which collapses 21 line-keyed entries to 11 or-collapse where the same expression recurs at multiple call sites in the same file. 2. test/run-mode-ui.test.ts's terminal-ownership guard scanned method bodies via `^ {2}async (run[A-Za-z]*)\(\) \{$`, which matched the 8 one-line run<Mode>() wrappers PR B2 introduced but not _runCliMode(mode), where the real logic (and the actual risk the guard exists to catch) now lives. Fixed the regex to `^ {2}async (_?run[A-Za-z]*)\(\w*\) \{$` and added _runCliMode to the sanity list. Same-class fix in test/opencode-resize.test.ts, which had the identical blind spot via runOpenCode.toString(). Both reproduced live before fixing (inserted the same comment line; added this.terminal.clear() to _runCliMode) to confirm the bug, then confirmed the fix catches it and the suite stays green otherwise. Also resolves Open Question 2 by dropping window.__codemanCliCatalog entirely: nothing consumed it, and a registry DECLARED_FOR_LATER field costs nothing until read while an unconsumed script tag on every page render is a different trade. Reverts Phase 1 cleanly — server.ts's injection, shortBadge back in types.ts's DECLARED_FOR_LATER list and the pinned guard test, and the three associated render-index-html.test.ts / server-index-title.test.ts assertions. Full gate: 405 files / 7717 tests / 0 failures (net unchanged), typecheck/ lint/format clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
cd64b0a3f7 |
feat(cli-registry): drive the run-menu frontend from the CLI catalogue (PR B2)
PR #380 (PR B) held back the frontend half of the CLI registry refactor, explicitly deferring window.__codemanCliCatalog and making session-ui.js / mobile-overview.js catalogue-driven as "PR B2". - Inject window.__codemanCliCatalog in renderIndexHtml(), following the existing __codemanCustomModelClis pattern (escapeScriptJson-guarded, resolved per-request). Reading CliEntry.shortBadge here is what makes it genuinely read, so it drops out of types.ts's DECLARED_FOR_LATER list. - Consolidate session-ui.js's 8 near-duplicate run<Mode>() launch functions (opencode/codex/gemini/antigravity/pi/omp/grok/deepseek) into one shared _runCliMode() plus a local RUN_MODE_LAUNCH config table. The 8 method names stay as thin wrappers (index.html calls them by name; tests assert on the name). Also collapses a duplicated 8-way isAltMode/isExternalCli OR-chain (same expression, copy-pasted twice in openSessionOptions) into one EXTERNAL_CLI_MODES check. - Add test/frontend-cli-no-id-branching.test.ts, a guard scoped to session-ui.js/mobile-overview.js only (not the rest of src/web/public/, which stays explicitly out of scope per CLAUDE.md), mirroring the backend's own no-id-branching guard. mobile-overview.js and the wiring of accent/echo/wheelForward/ keyboardAccessory were investigated and deliberately left alone: the first is already a single, tested, gated table (not duplicated logic); the second set belongs to terminal-ui.js/keyboard-accessory.js/styles.css, files outside this PR's mandate. Verified on a tmux-capable devbox (this sandbox has no tmux): full CI gate at 405 files / 7717 tests / 0 failures, typecheck clean, 94 targeted tests covering exact per-CLI wire-body shapes unmodified and passing, and a live anti-vacuity check on the new guard (injected a real branch, confirmed it fails, reverted, confirmed green). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
4205f6930f |
fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine. **The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext` and `confirmedSwap` changed the wire field without moving three assertions that check it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts` (the swap modal and the context modal, each of which already receives exactly the right per-question flag). Moved, with the titles. **Worse, my own tests for the split never ran.** The four cases in `session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`, which was declared inside a sibling `describe`, so they threw a ReferenceError during setup. The split would have shipped with no passing server-side coverage while the gate reported the failure as four broken tests rather than as four tests that were never written. `mockRunning` is hoisted to the outer describe. **The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so the other eight modes fell back to claude's `❯`. That is also starship's default shell prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line `❯ npm run build` sits on screen for as long as the command runs, the verifier reads it as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N prompt, `read -p`, an installer or a pager, where it takes the default. The module's own fileoverview already stated the rule this broke. Now `?? ''`, which `promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI that actually declares a composer. **My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing three numbered rules. The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told an integrator to retry with `confirmed: true` for both questions, which is precisely the thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md` and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR` multi-user consequence, both user-visible and both previously absent, and #454's gained the one exception to its own claim: a Custom Endpoints launch ignores the Instance count stepper and always starts one session. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9af12afb57 |
docs(custom-model): make the docs match the code, and trim the changeset
More from the review of
|
||
|
|
fe3bd0074c |
fix(custom-model): split the two confirmation questions, and seed the API key the way claude reads it
Two findings from the review of
|
||
|
|
3b55957d79 |
fix(custom-model): merge-time fixes for the Run-menu picker
Conflict resolution against the five PRs that landed while this was in review, plus the items left for merge on the thread. The real one was `session-ui.js`. #454 refactored all eight non-Claude `run*()` functions to funnel through one `_launchQuickStartInstances()` helper that does the POST itself, while this PR replaced that same POST in each of them with `_quickStartWithCustomModelConfirm()`. Resolved in the helper rather than seven times over: the helper now goes through the confirm path, and each body builder carries the `customModel` spread. `runAntigravity` deliberately does NOT, since antigravity's `customModelInjection` is `unsupported`; parity with this PR's own per-mode choices is asserted rather than assumed. That merge creates a question neither feature had alone: the confirm dialog now runs inside a loop that can launch up to 20 instances. Both questions it can ask (context window too small, and loading this will unload the model another session is using) are decisions about the ENDPOINT, and every instance in a batch targets the same one, so the answer is taken once and carried to the rest. Without that a 20-instance launch asks the same question 20 times. Also: `sse-events.ts` is 161 constants (master added two for remote wake, this adds one, verified by counting rather than by arithmetic), `server.ts` keeps both new SSE prefixes, the two comments pointing at code that no longer exists are corrected, and CLAUDE.md's SSE and route counts move to 161 / ~236 / custom-model (6). `pumpLlamaSwapLogTail`'s unparsed remainder is now capped at 64 KiB. It only shrank at a `\n\n` frame boundary, so a backend that streams without one would grow it for the life of a deliberately indefinite connection. NOT changed, deliberately: the context warning and the swap-conflict warning still share one `confirmed` flag with the context check first, so confirming "launch anyway" on a too-small context also skips the "this unloads it for another session" ask. That is the author's documented choice and the reviewer's own note calls it minor. Both fixes are worse to make here than to defer: separate flags are new wire surface landed unreviewed during a release, and reordering the checks adds a network round trip to a path that currently short-circuits. Raised as a follow-up instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1a99b5836c | Merge pull request #430 from opticon454/custom-model-run-menu | ||
|
|
035bfbc2fe |
fix(remote): merge-time fixes for Wake-on-LAN
The MAC-count limit lived in two places that disagreed. RemoteHostSchema.wakeMac's 128-character cap admits seven comma-separated MACs while parseMacList takes at most four, all-or-nothing, so a five-MAC value validated, was written to remote-hosts.json, and then resolved to NO wake target: POST /api/sessions/:id/wake answered "No wake-on-LAN target configured for this host" and the banner offered "Configure WoL" for a host the user had just configured. MAX_WAKE_MACS now lives in src/config/remote-wake-limits.ts and both sides refine against it. Its own module because src/remote-wake.ts is import-fenced to session-routes.ts and server.ts (the wiring guard that stops a watcher waking a host), and because schemas.ts must not drag dgram/net/child_process into every request-validating module. The documented 40 s request budget also omitted the wake's own cost. A `command` target is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS and runs BEFORE the readiness poll, so a slow one pushed a wakeCommand host's worst case to ~68 s, past the 60 s proxy_read_timeout the budget exists to stay under. _wakeAndWait now subtracts the wake's measured elapsed time from the readiness budget, floored at one poll interval so a wake that ate the whole budget still gets one probe. A magic packet is effectively instant and is unaffected, which is why live testing never saw it. Also: the two new endpoints are documented in docs/api-reference.md with the import fence stated as the rule it is, CLAUDE.md's frontend module count moves to 34, and the release changesets carry the Thanks section. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4c705094f7 |
fix(terminal): ship the copy clean as a trailing trim, without the shared dedent
#451 cleaned two things on copy. The trailing trim is right and every native terminal does it. The shared leading-indent strip is this project's own rule, and it is dropped here rather than shipped. Measured against the shipped transform over 401,445 three-row windows across 1,010 tracked files in this repo, it fired on 73% of them: 92% inside a YAML workflow, 76% over `git log` output, 48% in a TypeScript source. No width threshold separates a margin from content because they are the same widths, a live Claude Code pane's own margins measuring 2 and 5 columns while the most common non-TUI shared run is 4. The failure modes are not symmetric either: a wrong trailing trim costs nothing, while a wrong dedent silently deletes information that was on the screen, with nothing in the clipboard to hint at it, on git log bodies, on indented code read out of cat (semantic in Python), on git diff context rows where the leading space is the marker, and on stack traces. It also could not be made self-consistent cheaply. Whether the first row joined the measurement depended on the mousedown COLUMN, which the user never sees, so one block of three rows produced three different clipboard results; and the flag read getSelectionPosition().start, which is xterm's mousedown anchor and is never normalised, so dragging UP through a block read it off the bottom row. The PR's test stub hardcoded a downward drag, so its suite could not express that case. The transform, the wiring, the tests, the invariants, CLAUDE.md, the wiki page and the changeset all move together. The test block now pins the ABSENCE as a contract, with the git log, Python and git diff cases as its examples, so this is not re-derived later. If it is ever revisited, the one qualification that measured clean is painted trailing padding: zero false positives over all 401,445 windows. Also from the review: the comments and invariant rule justifying the padding-only clear described the pre-change code (the Ctrl+C gate reads the CLEANED selection now, so such a selection falls through to the PTY on its own and the clear is feedback rather than protection), the new 'Nothing to copy' toast gained its zh-CN entry, and the invariants paragraph no longer repeats its own opening sentence. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c376534a50 |
fix(run,terminal): merge-time fixes for the Instance count stepper and capture geometry
#454: the behaviour the PR adds had no test, so a regression test drives runGrok() at tabCount 3 and asserts three quick-start POSTs with sequential w<n>-<case> names (verified to fail against master's session-ui.js). Each caller now reads the count BEFORE its opening banner and announces it there, the way runClaude() already did, so a launch no longer prints two headers and a launch with another session already active still says how many are starting. runClaude() calls the shared _readTabCount() instead of its own copy of the 1..20 clamp, and that helper optional-chains the element read, since hoisting it above each caller's try block would otherwise let a missing #tabCount throw where the launch-error path cannot report it. #435: sizeMovedUnderLoad derived from data.source alone. `mux-visible` is not sufficient: a failed display-message cursor query makes capturePaneBuffer skip the snapshot repaint and return the raw capture, which the route still labels mux-visible, so a size that moved during such a load bought a full forced reload to repair a frame that was never positioned. It now tests Number.isFinite(data.captureRows) like its two siblings. Plus the invariants and CLAUDE.md lines promised on #435: a visible capture reports its geometry and omits it when nothing was positioned, the comparison runs on mux-visible only, and the replay is capped at one attempt and latches per session when it cannot converge. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2c3ccdf030 |
Merge pull request #439
feat(remote): wake a sleeping host (Wake-on-LAN) from input, banner and native magic packet |
||
|
|
475436242c |
Merge pull request #455
fix(input): make sure a prompt sent through the API actually leaves the composer |
||
|
|
613b774bf1 |
Merge pull request #451
fix(terminal): trim the padding and shared indent out of a copied selection |
||
|
|
60c9af0599 |
Merge pull request #435
fix(terminal): replay a pane capture at the geometry it was taken at |
||
|
|
5fc391a47c |
fix(custom-model): address fourth pre-merge review + merge upstream master (Ark0N)
Merged upstream/master (22 commits: reboot-restore recovery feature,
terminal keycode229 recovery work, install.sh/CLI-catalog generator
changes, CHANGELOG/version bump to 1.30.0) into this branch. No
conflicts; git auto-merged every overlapping file (CLAUDE.md,
docs/api-reference.md, app.js, index.html, styles.css, routes/index.ts,
session-routes.ts, schemas.ts, server.ts).
Two required fixes from the latest review:
1. privilegedEnvKeys widening (stock.ts) changes behaviour outside this
feature. The reviewer decided to keep both CLAUDE_CODE_MAX_CONTEXT_TOKENS
and CLAUDE_CONFIG_DIR listed (types.ts's rule that every traffic-
redirecting var this feature introduces must appear there stays
literally true), and asked for the real consequences documented
instead of hidden:
- Corrected session-env-clamp.ts's fileoverview, which stated the
opposite of what the code now does (reboot-restore's clamp call
used to be able to strip nothing for claude; it now strips a
persisted CLAUDE_CONFIG_DIR for a non-granted owner).
- Corrected the rationale comments in stock.ts: privilegedEnvKeys
has exactly one consumer (ownerClampedEnvKeys, feeding the
generic envOverrides clamp on create/quick-start/reboot-restore),
not the custom-model routes.
- Added a CLAUDE.md line to the CLAUDE_CONFIG_DIR gotcha covering
the admin-only-in-multi-user-mode and reboot-restore-strips-it
consequences.
- Added a "Claude multi-user clamp" test next to the existing
DeepSeek/OMP ones, pinning the new stripping behaviour.
2. GET .../running-status (custom-model-routes.ts) no longer passes
the raw llama-swap `cmd` field (the literal launch line, which can
carry model paths and --api-key) to the browser -- the frontend
only ever reads model/state, cmd exists solely for server-side
parseCtxFromCmd() during discovery. Added a test asserting the
response never contains cmd or a planted secret.
Also regenerated config/clis.stock.json and install.sh's catalogue
block (npm run generate:cli-catalog) to clear drift introduced by the
upstream merge, since it was failing the sync check.
Left to the reviewer, as they said they'd take at merge: the two
"comments pointing at removed code" cleanups, the two stale CLAUDE.md
counts, and the small items list (mode==='claude' frontend branch,
isCliAvailable() unknown-id gap, shared confirmed flag ordering,
one-shot cancel toast severity, pumpLlamaSwapLogTail buffer cap).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
|
||
|
|
5bb489addb |
fix(remote): authorize the attach wake first; tell the caller what happened to its bytes
Review round 3 on #439. - The attachRemoteSession branch of POST /api/sessions ran `ensureHostAwake` before the multi-user gates, so a non-admin could have any configured host's `wakeCommand` spawned (or a packet broadcast) and the request held for the wake budget, then be refused for the workingDir. The admin gate now comes first, before the host is even looked up; remote hosts are admin-only infrastructure everywhere else. Route test: wake spy empty, 403. - The non-wait input route answers `{buffered:true}` when the registry took the chunk and `{buffered:true, dropped:true}` when it was over the cap and is gone (`RemoteInputOutcome` gains 'dropped'); additive to the bare `{}`. - The send-and-wait path answers OPERATION_FAILED when the host never comes back, like create and attach, instead of writing into the stalled pane and reporting delivered:true plus a timeout. - The flush writes with `fromUser: true`, so a first prompt buffered through a wake can still name the tab. Docs: api-reference (input route), remote-sessions.md (two invariants), CLAUDE.md key pattern. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG |
||
|
|
19ffe9b7a8 |
fix(input): make sure a prompt sent through the API actually leaves the composer
Claude Code 2.1.277 takes typed text the moment its composer paints but ignores Enter for the first 30 to 50 seconds after it (measured 2026-09-19 through the input route: an Enter at 28 s stranded the prompt, one at 51 s submitted it). The text+Enter pair `sendInput` sends 50 ms apart therefore left every programmatic prompt sitting unsent, and every waiter burned its timeout on a turn that never started. Server: `SubmitVerifier` (session-submit-verifier.ts), armed from `writeViaMux` for every mux write that carried a carriage return, reads the pane on a 2 s to 60 s schedule and re-sends Enter only while the last composer line (the CLI's own prompt glyph) still holds the head of what was sent. An empty composer, other text, or no composer line at all ends it; a newer write replaces the schedule. Skill: `sendwait` gets the same loop (`_composer_text`, no-break space stripped by its bytes for BSD sed) for servers that predate this, and the preamble version moves to 1.30.1 so seeded agents pick up the fresh copy. SKILL.md's heredoc and the plugin mirror are regenerated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
95dc6fe944 |
fix(terminal): remember a geometry replay that did not converge
`resizeRetry` caps the recursion inside one select and says nothing about the next one, so a pane this browser cannot size reported the same mismatch on every select and bought the same failed repair each time: two fetches per tab switch for the life of the page, measured as a running count of 2, 4, 6 across three selects. That is the case this branch describes as happening every time rather than occasionally, a phone whose resize `Session.resize` declines while a desktop claim is live, and it is not the only one — any pane Codeman cannot size lands there, including one a second tmux client is also holding. Each wasted pass costs another `capture-pane`, which is `execSync` and blocks the server's event loop, plus a reset and chunked rewrite, a discarded snapshot and cache entry, and a dropped and reopened WebSocket. `_geometryRetryUseless` mirrors the existing `_fullHistoryRepullUseless`: a retry pass whose frame still does not fit adds the session, geometry that fits removes it, and the replay gate consults it. The proof has to come from a retry pass rather than a first one, because the retry ran at the size that stuck and the pane ignored it. Clearing on a fitting frame is what stops a pane that becomes sizeable again, once the desktop tab closes or its claim goes idle, from staying permanently unrepaired. The race case never reaches the latch, since it converges on its first attempt. The new browser case walks all of that: three selects reading 2, 3, 4 instead of 2, 4, 6, then a fitting frame, then a mismatch diagnosed afresh. Without the gate it fails on the second switch with `expected 4 to be 3`. Rebased onto master, which has moved to 1.30.0 and taken #436. The one conflict was `config/test-suites.ts`, where both branches appended a glob to `BROWSER_TEST_GLOBS`; both are kept. Everything else merged clean, #436's own changes to the same buffer-load path included. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
383f834704 |
fix(terminal): flush unsent local echo before the geometry replay
On a touch device the characters the user has typed live only in the local-echo overlay until Enter; they have never reached the PTY. The replay re-enters `selectSession` with `forceReload` on the session that is still active, and that branch nulled `activeSessionId` before `_cleanupPreviousSession` ran. The flush there is guarded on a session it can still see, so it was skipped, and the unconditional `_localEchoOverlay.clear()` that follows took the characters with it. Measured in chromium against the previous head: typing into the overlay and then making the call the replay makes left `pendingText` empty with nothing crossing into the delivery layer on either transport. The flush moves into `_flushLocalEchoTo(sessionId)`, called from both `_cleanupPreviousSession` and the `forceReload` branch before it nulls the id. The session is a parameter because the two callers mean different ones: cleanup flushes to the tab being left, the branch to the tab being reloaded. This was reachable before this branch, through the one gesture that already takes the `forceReload` path on an active session. What is new is that nothing the user does triggers it. The replay fires on its own the moment a tab switch finishes, which is exactly when someone typing into a still-loading terminal has text in the overlay, and on a phone beside an active desktop tab that is every tab switch. A seventh browser case pins it: it forces the overlay on, since headless chromium reports no touch support and the case would otherwise pass vacuously, asserts the typed characters really are sitting unsent, then triggers the replay and asserts they reached the session. Without the fix it fails with nothing delivered at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e0d4477edc |
fix(terminal): keep the geometry replay to the pass that can converge
Three follow-ups to the source gate, each one measured rather than reasoned. A pane already drawing at the size the client just requested is left alone. The replay runs at `dimsAfterLoad`, so it can only change what is on screen if the pane was drawing at some other size; when the reported geometry already IS that size, the second pass captures the identical frame and pays a full reload to do it, including a visible re-flash, a dropped and reopened WebSocket and a deleted xterm snapshot. That equality is the signature of a clamp rather than a race: `getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` does not, so a terminal narrower than 40 columns or shorter than 10 rows reports a pane permanently bigger than itself and replayed on every tab switch without ever converging. A race never produces the equality, since its premise is that the pane was still at the size it was asked to leave. The declined-resize case does not produce it either, so that one still costs the single capped attempt and needs the pane-ownership question this does not touch. The full-history re-arm is unreachable and now says so. A pass that consumed the flag sent `full=1`, and the route answers `full=1` with `mux-full-history` or `history`, never `mux-visible`, so the source gate already rules out every such pass. The line stays for the invariant, but its comment no longer reads as if a page load retries, and the suite pins that it does not. The response no longer reports geometry for a body that carries no capture. The full-history path writes `capturedGeometry` from the cursor query and then returns '' for a pane holding nothing visible, which drops the source to `history` with the geometry already recorded: a `full=1` request whose capture reported 100x50 and returned nothing answered `source: "history"` with both fields set. Nothing acted on it, because the client ignores geometry on any other source, but the field said a frame had been drawn at a size when none had. The browser stub now derives `source` from the request the way the route does, rather than answering `full=1` with `mux-visible`, which the route cannot produce. Each case reaches a visible-frame response the way production does, by not being the first select of the page. Three cases pin the new behaviour and each fails without its guard: the clamp case sees two fetches instead of one, the scope case and the full-history case both see a replay the gate forbids, and the width case sees one fetch instead of two. The changeset now describes the change from 1.29.x rather than the difference between the two commits on this branch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5cfb98fb8b |
fix(terminal): compare capture geometry only on a visible-frame response
Only a visible-frame capture positions its rows absolutely, so only that frame can be damaged by a terminal of the wrong size. A `full=1` body is linear scrollback closed by a relative cursor move, which is relative precisely so the browser's row count need not match the pane's, and a `history` body is the byte stream, which carries no row alignment to protect. The geometry comparison ran on all three, so it fired most often on the one response it cannot help: `_fullHistoryLoaded` is empty on the first select of every non-shell session per page, and a session whose pane a desktop tab holds too tall to ever fit then paid a second whole-scrollback capture, reset and replay on every page load and every first tab switch. `framePositionsRowsAbsolutely` gates both the captured-geometry comparison and `sizeMovedUnderLoad`. A size that moved under a byte-stream or scrollback replay is healed by xterm's own reflow plus the SIGWINCH the trailing `sendResize` already sends. A pane WIDER than the terminal damages the same frame a second way, so `captureCols` is now compared rather than only logged. `formatPaneSnapshot` paints each row out to the pane's own width, so a narrower browser wraps every painted row, and the wrap on the last one scrolls the whole frame up by a row. The terminal response no longer falls back to `session.ptyCols`/`ptyRows` when the capture reported no geometry. The cursor query is what produces the absolute addressing in the first place, so a capture that lost it returned a raw frame that was never positioned, and a byte-history response was never positioned either. Naming the session's own PTY size there described a frame that does not exist and invited a repair for damage that is not present. `_ptyCols` is also written only by `resize()` while the PTY is spawned at the size queried from tmux, so it can be wrong on its own terms. Both fields are now absent instead, and the `Session` getters added for that fallback go with it. Two browser cases cover the new behaviour and each fails without its fix: a `mux-full-history` response with both dimensions mismatched asserts one fetch (two without the gate), and a `mux-visible` response wider than the terminal but short enough to fit asserts two (one without the width comparison). Corrects a claim in the comment above `capturedGeometry` in tmux-manager.ts. Both replay paths do not address rows absolutely; the full-history one ends in a relative move, which is the whole reason the gate is right. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3edf9aae2f |
fix(terminal): replay a pane capture at the geometry it was taken at
A visible-frame capture repaints each row at an absolute position, counting up to the pane's height. A terminal shorter than that clamps every address past its own height onto its last line. The overflow rows then overwrite one another, and the rows underneath are lost. Replaying a real 50-row capture into a 30-row terminal rendered 28 lines of a 45-line command and drew the frame twice. Nothing in the response said what height the frame was built for, so the client could not detect this. A capture now reports the geometry it was really taken at through `capturedGeometry` on `PaneCaptureOptions`, and the terminal response carries it as `captureCols` and `captureRows`. When the captured pane is taller than the terminal, or the size that produced the capture did not survive the load, `selectSession` replays once at the size that stuck. `resizeRetry` caps that at one attempt, so two competing fits cannot trade replays forever. The retry re-arms the full-history flag only when the pass that ran had consumed it. A tab switch takes the bounded tail, so its retry takes the tail too: clearing the flag unconditionally would upgrade that switch into a fresh scrollback capture the user never asked for, which the route's own comments put at tens of megabytes. What this repairs is a capture that won a race against the resize meant to precede it. It does not repair a capture whose pane was too tall because `Session.resize` declined the resize outright, which it does for a small viewport while a desktop viewport's size claim is live. The retry re-sends the same declined resize and captures the same pane, and `resizeRetry` then stops it. Repairing that means changing who owns the pane size, which is a policy question this does not touch. The reported geometry still helps there, because the client can see the mismatch at all rather than being blind to it. Follows #395, #396 and #397, which fixed the other ways the replayed frame and the terminal could disagree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
358aef16e3 |
fix(run): make the Instance count stepper work for every non-Claude mode
runOpenCode(), runCodex(), runGemini(), runAntigravity(), runPi(), runOmp(), runGrok(), and runDeepSeek() all ignored the "Instance count" stepper next to the Run button and hardcoded a single quick-start call — bumping the counter to 2 or 3 while on any of these modes silently launched exactly one session, with no error. Only runClaude() ever read it. Extract the shared launch-N-sessions-and-select-the-first loop into _launchQuickStartInstances(), reused by all eight modes, and _readTabCount() for the shared clamp-and-parse. Each mode still builds its own quick-start body (config differs per CLI), just via a closure passed to the shared loop instead of a single inline fetch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1040f6c489 |
fix(remote): a proxied host is reachability-unknown; scope remote: SSE per session
Review round 2 on #439. 1. The bare TCP probe connects to host:port, which a host behind a jump host or SOCKS proxy does not answer even while ssh works. Acting on that verdict drew a permanent banner over a healthy session, replaced a real "needs tmux" error with "not reachable" in quick-start, and - with a wake target - buffered every HTTP input for the life of the session, since the readiness poll could never succeed. `WakeableRemote` now carries `jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such a host into reachability-UNKNOWN: input is delivered, `checkReachable` / `checkHostReachable` answer `null` (never `false`), `ensureHostAwake` returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate fires on `=== false` only, and `GET …/reachability` reports `reachable: null, probeable: false` so the banner has nothing to key on. A wake target can still be fired for it, blind: no readiness poll, no reattach, no toast - the response says only whether the packet went out. 2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake has no session yet, so the registry names the requesting user (`ensureHostAwake({ requestedBy })` -> `username` in the payload) and `deriveSseHint` routes on it; with neither it fails closed to admins. Single-user mode is unaffected. Smaller, from the same review: - A flush write that fails now drops the remaining buffer (logged) instead of retaining it: the wake still resolved and marked the host reachable, so the retained chunk waited for the NEXT wake and was replayed hours later, after everything typed since. Same policy as the oversized paste. - The banner polls on tab activation (a user action) and on its 30 s timer only for a host with a wake target; a timer connecting to a host Codeman cannot wake is the traffic invariant #2 rejects keepalives for. A proxied host is never polled. - `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP socket refuse under VITEST, as remote-files.ts does. The guard caught a leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the probe but still polled readiness with the real one, so the shutdown test had been connecting to a production address. The poll now uses the injected probe. - docs/remote-sessions.md is additions only again (the reformatting is gone); the architecture-invariants overlap resolved itself in the merge. Live, against a throwaway instance with a non-routable ghost host: proxied -> no probe, no wake, the genuine ssh error after 10 s; direct (control) -> probe, magic packet, "did not come back" after the 40 s budget. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG |
||
|
|
e271a65e79 |
Merge origin/master into feat/remote-host-wake
Resolves CLAUDE.md count tables (route counts recounted on the merged tree: 235 handlers, sessions 37) and keeps both the host-wake and the reboot-restore banner in index.html. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG |
||
|
|
3cdb4bf42e |
docs(terminal): the merge-time notes promised on #436
The four edits the review said would be folded in at merge, none of them code: the changeset becomes one user-facing paragraph, since it is what CHANGELOG.md and the release notes print; the `_bufferLoadFinishOpts` comment now names the second contributor to the duplicate window (`captureActivePaneBuffer` is `execSync`, so anything painted into the pane before the server read it is in the capture and is broadcast after the reply) and says why a `history` payload keeps the pre-existing discard when its exposure is the same; the `_finishBufferLoad` doc block moves from above `_beginBufferLoad` onto the function it documents; and the test file's header describes both rules the file now pins instead of only COD-144. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
492f8d8ddf |
Merge pull request #436 from irisitymichaelgrundberg/fix/replay-output-that-arrived-after-the-capture
fix(terminal): keep the output a pane capture could not contain |
||
|
|
56209e7829 | Merge remote-tracking branch 'upstream/master' into feature/run-menu-custom-model-picker | ||
|
|
afb6754453 |
fix(custom-model): address third pre-merge review (Ark0N)
Blocker 1: the loading banner hides itself ~200ms after it reopens. - _showCenterStatus reuses one shared DOM node; dismiss() scheduled el.hidden = true 200ms later with nothing to cancel it. On the Claude path, switchingToast.dismiss() is followed by one same- origin request (5-30ms locally) before _watchLlamaSwapLoading opens the new banner -- well inside that window -- so the stale timer fired against the shared node and hid the fresh banner, leaving the whole model-load wait with no progress text, no log line and no reachable Cancel button. - Fixed by parking the pending timeout on the element and clearing it at the top of _showCenterStatus. Added a regression test that reproduces the exact repro (open, dismiss, reopen 20ms later, advance past 200ms) alongside the existing Cancel-button DOM tests; confirmed it fails without the fix and passes with it. Blocker 2: the swap-conflict warning named other users' sessions. - Both affectedSessions scans (POST .../custom-model and quick-start) walked the whole session map with no ownership filter, so in multi- user mode a non-admin pointing their own session at a shared endpoint learned another user's session name and id -- which with autoNameSessions on is that user's own prompt. - The swap is still blocked pending confirmation regardless of ownership (a foreign session is just as real a disruption); only which ones get NAMED back to the caller is scoped, via the already-imported canAccessOwned. Added a two-owner test to test/routes/session-custom-model.test.ts covering both the foreign-owner (blocked, not named) and same-owner (named) cases. Smaller ride-along fixes: - server.ts boot recovery now passes contextLength into applyCustomModelInjection, so CLAUDE_CODE_MAX_CONTEXT_TOKENS is correctly rebuilt into _envOverrides after a restart instead of surviving only because tmux retains the old setenv. - pumpLlamaSwapLogTail's finally now deletes by IDENTITY, not just by key, so an aborted pump finishing after a newer entry was created for the same endpoint can no longer delete that newer entry and orphan its connection. - docs/custom-model-endpoints.md now notes that clearing a custom model removes injected keys by name, including CLAUDE_CONFIG_DIR -- so a session that also had CLAUDE_CONFIG_DIR set via envOverrides (the per-client-account case) silently falls back to the default account on clear. Left for later, as flagged in the review itself: the quick-start case-scaffolding/cancel ordering (real behavioural reordering across a large handler, too risky to make without a live re-test), and retiring runCustomModelEntry's mode === 'claude' branch behind a launchStrategy registry field (explicitly deferred by the reviewer to "the next one"). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R |
||
|
|
f9edb33d15 |
fix(terminal): trim the padding and shared indent out of a copied selection
xterm hands back whole screen rows and trims only the cells that were never written to, so the real spaces a full-screen TUI paints across the unused part of a row count as content and reach the clipboard. Measured against Claude Code in a 282-column pane, single lines arrived carrying 138 trailing spaces, and every line carried the two-space transcript indent as well. Windows Terminal, iTerm2 and GNOME Terminal all trim that for you, decideAutoCopy already calls a wall of spaces "never what the gesture meant", and _selectTouchSelectionLine already treats those cells as padding — the mouse and keyboard paths never had the same rule. CodemanCopySelection.clean lives in constants.js beside decideAutoCopy, its pure sibling. It drops the trailing run from each line, and removes the leading run only where every selected row shares one. A selection of a single row keeps its run, because one row shares nothing with anything and stripping it would silently reindent one line of `git log` body text or one line out of `less`. A drag that began inside a row keeps its partial first line untouched and out of the measurement, which otherwise pins the shared run to zero and leaves every following row indented. Every pass over a line is a scan rather than a regex. `/[ \t]+(\r?)$/` is quadratic on a line whose spaces are followed by a non-space character, which is what right-aligned or centred TUI content looks like: measured over 50 000 rows with a 280-column run it took 2.9s, against 1.3ms for the scan, and a 2 000-column run took 16s. The scan is also the faster of the two on an ordinary padded row. cleanedTerminalSelection in terminal-ui.js is the half that needs the live terminal. It returns a COLUMN selection untouched: Alt+drag makes one, and a rectangle's rows lining up is the point of the gesture, so both halves of the clean would destroy it. xterm exposes the mode nowhere public, so the check reads terminal._core._selectionService, the way this file already reads terminal._core for cell dimensions, and cleans normally if a future xterm renames the field. A test pins that assumption against the library rather than against a stub repeating the literal. The Ctrl+C chord decides on the cleaned selection, not the raw one. A drag across the blank part of a row selects real padding spaces, so the raw text is truthy, and testing it would spend that press on a copy of nothing and make the user press again to interrupt. A padding-only selection is now dropped and the press falls through to the PTY, while Ctrl+Shift+C still never falls through. copyTerminalSelection gates on trim() for the same reason, since a multi-row drag across padding cleans to line breaks alone and a bare newline pasted into a chat composer submits it. All four of the main terminal's copy paths go through it: the Ctrl+C chord, right-click, the phone selection button and Auto Copy. The browser's own Edit menu copy, a disabled copy shortcut and the subagent windows still copy raw rows, as they did before, and the invariants doc now says so rather than claiming every copy is cleaned. Auto Copy resolves its own toggle before it reads the selection, since it is off by default and a selection can run to the 50 000-row scrollback ceiling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3730bc7df5 |
docs(terminal): correct what selectSession does with the viewport
The JSDoc on `_syncStickyScrollBaseline` said `selectSession` deliberately ends
at the bottom, so the baseline the replay samples is already true there. It
does not. `selectSession` calls `scrollToBottom()` after the write and then
ends at `scrollToLastNonEmptyLine()` (app.js:6512), which targets
`lastNonEmptyLine - rows + 2` and therefore parks ABOVE `baseY` whenever the
replayed frame keeps trailing blank rows — which a full capture does on
purpose, since no transform that can delete a line may run over one.
Its baseline really is a stale true. What covers it is the sticky snap itself:
since
|
||
|
|
75a028e825 |
fix(terminal): re-take the sticky-scroll baseline after a replay
A capture load now replays its queued tail, and that replay runs through `batchTerminalWrite`, which samples `_wasAtBottomBeforeWrite` before it queues. It runs inside `chunkedTerminalWrite`, before that promise resolves, with the terminal freshly reset and rewritten — so the sample is always true. The caller then restored the reader's position and the next `flushPendingWrites` scrolled straight back to the bottom off the latched flag, undoing it. The only thing in the way was `_hasRecentUserScrollUp()`, a 1500ms window a server-triggered refresh is usually past. `_syncStickyScrollBaseline()` re-takes the flag from wherever the viewport now sits, and the two paths that restore a position call it right after doing so: `_onSessionNeedsRefresh` and `_maybeRefetchFullHistory`. Those are the paths #259 and #205 exist for, and they are also where a non-empty queue is most likely, since a needsRefresh fires when output is flooding. Re-taking rather than suppressing the sampling: suppressing leaves whatever stale value the flag held from before the load, which on the full-history re-pull has no reason to be false. `selectSession` and `_onSessionClearTerminal` deliberately end at the bottom, so the sampled true is already the truth there and they do not call it. `_bufferLoadFinishOpts` gains the coverage the CI gate can see: both mux sources flush, `history` does not, and a payload naming no source does not. Its only coverage was the browser suite, which CI does not run. The JSDoc and the changeset now record the one duplicate window this cutoff cannot close. The server appends output to the byte buffer in the same tick it emits, but broadcasts on a batch timer — 8ms over WebSocket, 16 to 50ms over SSE — so a batch pending when `capture-pane` ran leaves the server after the reply and is replayed although the capture holds it. It is one batch interval wide against a recovery window spanning the whole chunked write, and closing it means flushing that batch server side before the capture. The second browser test asserts its session was created, so a failed create fails it instead of passing with zero hits. docs/architecture-invariants.md no longer claims the replay leaves the queued-event discard window alone. That clause now describes what decides how a load ends, the baseline rule, the batch window, and the three covering tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9982a1325f |
fix(custom-model): address second pre-merge review (Ark0N)
Blocker: .center-status-banner never actually disappears.
- Add `.center-status-banner[hidden] { display: none; }`, same trap as
`.home-sessions[hidden]`: the author-level `display: flex` beat the
UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
card stayed laid out at `opacity: 0` with its text/cancel/close
children still `pointer-events: auto` -- an invisible 442x67 click
blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
banner (10001) and the swap-confirm/context-warning modals (10010)
in CLAUDE.md's Z-index layers list.
Stale wording pointed at the reverted sticky-toast default:
- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
`.toast-message` comment in styles.css all still said "toasts
default to sticky" after
|
||
|
|
1f32128ca9 |
fix(custom-model): address PR #430 pre-merge review (Ark0N)
Four blockers from the 2026-09-18 review: - PUT /api/model-endpoints/:id now merges modelContextLengths/ modelSizesGB back in from the stored record instead of trusting the editor's body, so renaming an endpoint or changing its default model no longer silently drops the context-window floor check and CLAUDE_CODE_MAX_CONTEXT_TOKENS injection. - custom-model:swapped-out is now session-scoped (added to SESSION_PREFIXES) instead of broadcasting to every connected client. - The quick-start custom-model path now hands setCustomModel() only the endpoint's own injected env vars, not the full merged set, matching the restart-in-place path — the full set put CLAUDE_CODE_EFFORT_LEVEL back after the Session constructor had already stripped it. - The quick-start launchModel override for pi/grok/omp is now applied generically via the registry's legacyConfigField, mirroring Session._withCustomModelLaunchModel, instead of three hardcoded mode === '<id>' branches a future CLI's injection recipe would miss. Also scopes the sticky-toast default (item 5): reverted the blanket "all error toasts are sticky" default, which had no container cap or eviction, back to a flat 3s; the one message that needs a moment to read (a failed custom-model apply) now passes an explicit duration: 0 at its own call site. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R |
||
|
|
bb8ada7e5f |
fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does. 1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred it from the name: a session the user renamed by hand to something shaped like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the next prompt overwrote their name. The route persists right after, so the loss went to disk. `restoreMuxSessions()` already passes it. 2. The already-live sets were snapshotted once before a loop that awaits a real `startInteractive()` per entry, so by the tenth entry the snapshot was tens of seconds old and a conversation resumed by hand from the Resume list in that window was invisible to it: two panes on one transcript, the exact thing the check exists to prevent. Both sets are now read per iteration, and the late case is spent rather than re-offered for the same reason the batch case is. 3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path. The stamp predates the reboot and the pane is new, so honouring it meant one click had every restored session type `continue` into itself about a minute later, unattended, against the route header's own promise that a restored session comes back idle and disarmed. The setting stays ENABLED, so it re-arms on the next real limit message. A Codeman restart still re-arms from the stamp, because the limit footer will not reprint on its own; the new option exists only to tell the two paths apart. 4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()` and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession` performs that it was missing. Cosmetic, but a run left open reads as still going in the away digest. 5. A restored claude session gets `seedAgentSessionPreamble()` like both create paths, so the agent skill's bootstrap stays a two-line loader. 6. The heuristic's container comment was wrong in one direction and quiet about the real gap: after a genuine host reboot a containerized Codeman sees the host's short uptime and the banner does appear. What it cannot see is a container-only restart, which is where this would help most. 7. The banner is hidden in a solo window, which shows one session and has no tab strip to put restored ones in. Also reverts 17 of the 18 hunks in docs/api-reference.md, which were Prettier reformatting of prose the PR does not otherwise touch (docs/ is outside the format glob), keeping only the Reboot restore section and repairing the two continuation lines that reformat de-indented; renumbers reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which loads after it; and gives the feature its CLAUDE.md entry plus a route test for the multi-user workspace-forbidden branch, the only new rule that had nothing behind it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f32c4f60d5 |
Merge pull request #442 from irisitymichaelgrundberg/feat/restore-sessions-after-reboot
feat(sessions): offer to rebuild the sessions a host reboot destroyed |
||
|
|
9a503872d9 |
Merge pull request #441 from shenlvkang-collab/fix/android-last-char
fix(input): deliver a recovered keystroke before the Enter that submits it |
||
|
|
ff8dc92187 |
Merge pull request #424 from Ark0N/fix/terminal-history-anchor-after-parse
fix(terminal): restore the history anchor after xterm parses, not before |
||
|
|
62ceb4e87b |
fix(sessions): correct what the missing-pid rule actually recognises
A second real reboot disproved the mechanism the previous commit was built on. Typing `/exit` does not persist `pid: null`, and the session was restored anyway. The pid a session record carries is its `tmux attach-session` process, not the agent. `/exit` ends the CLI inside the pane, `remain-on-exit` keeps the pane, and the attach process stays alive throughout — so Codeman's PTY never exits, no exit handler runs, and the record keeps both its pid and `status: 'idle'`. The lifecycle log for the session that came back shows created, started, stale_cleaned and recovered, with no exit event at all, which is the proof: Codeman never learned the agent was gone. So nothing durable distinguishes an exited agent from a session that was idle when the power went, and this pass restores both. Ark0N/Codeman#446 is about making Codeman notice the dead pane; contrary to what the previous commit's message claimed, this genuinely does wait on that. Until a record can say the agent is gone, the user dismisses or closes those sessions. The rule itself is kept, because a record with no attach process does describe a session that never started or whose pane died outright, and refusing it is right. Only its documentation was wrong. The module header, the branch comment and the test names now say what it recognises instead of claiming the case it cannot see. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5108a24bf0 |
fix(sessions): never restore a session whose agent was already exited
Found by a real reboot, which is the first thing to catch it. Typing `/exit` ends the CLI process and leaves the session record behind, and the process-exit handler persists `pid: null` with `status: 'idle'` before anything else runs. By status alone that is indistinguishable from a session sitting idle when the power went, so the boot pass offered those sessions back and a click spawned the agents the user had deliberately closed — the exact case the eligibility rule exists to exclude. The absent pid is what tells the two apart, and the plan step now refuses a record without one, under its own `not-running` reason so the boot log says why. On a healthy board every running session carries a pid; a record with none describes an agent that is already gone. Deliberately the conservative direction. A session that somehow persisted no pid while genuinely running is not offered, and its conversation stays reachable from the Resume list, which is where every session would be without this feature. The opposite error spawns processes nobody asked for. Ark0N/Codeman#446 covers the dead panes those exits leave behind, but this does not wait on it: the rule belongs here whether or not the record's shape changes later. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5f55f9cb65 |
fix(sessions): never let the reboot-restore plan fail recovery
The plan build runs inside the try that decides whether restoreMuxSessions() succeeded, so a throw would be caught there, report restoration as failed, and block the stale cleanup and layout reconciliation that follow. An optional convenience would then break the recovery it exists to help. It is guarded on its own now: the correct way for this to fail is an offer nobody gets. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
18ab2ab595 |
docs(sessions): correct what a failed rebuild is actually likely to be
Ran the feature against a real server for the first time, on an isolated instance, and two claims in the code turned out to be wrong. A rebuild that fails after the session is registered was documented as commonly caused by a CLI binary missing from a freshly booted machine's PATH. It is not: the resolver finds its binary by absolute path, so PATH never enters into it, and a server started without claude on PATH restored every session normally. Nor does an un-enterable workspace fail — tmux falls back to another directory and the pane comes up there. Neither obvious cause throws, so the discard path is defended rather than expected, and the comments now say that instead of naming a cause that cannot happen. The four review rounds that shaped this path all reasoned about a trigger none of them could test. The path itself is still worth having, since a mux failure would reach it, but its comments should not claim a likelihood the machine disagrees with. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e2034177c5 |
fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.
Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
published dependencies) into a scratch dir purely to read
dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
same endpoint. dsh's own error template ("DeepSeek API error (HTTP
${status})") reproduces the originally-reported
"dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.
- New registry field `appendV1Suffix` (env kind only, deepseek's entry
alone — claude/gemini must NOT get it, since claude was already
confirmed working against the unmodified baseUrl). When set,
buildCustomModelInjection runs endpoint.baseUrl through the same
withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
use, instead of writing it verbatim.
Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.
2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
db9729e1fc |
feat(custom-model): remove loading-banner countdown, add manual Cancel
Replaces the size-scaled expected-time estimate + matching auto-timeout with a generic hardware/model-size disclaimer and a user-driven Cancel button, per explicit request. Real load time depends on hardware this feature has no way to know (VRAM, storage speed, GPU contention), so the old estimate/timeout was a guess dressed up as a fact — worse, one that could kill a genuinely slow load partway through on slower hardware. - _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline entirely — polls indefinitely until ready or cancelled, no automatic give-up. Message is now "Loading <model> (<size>) on <endpoint> — this can take a while depending on your hardware and the model size.", with the real llama.cpp log line still on its own second line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/ _formatRemaining (dead code once the countdown is gone) — _lookupModelSizeGB is kept, the GB figure still shows. - _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real "Cancel" button (distinct from the error-type "×" close button, since Cancel has a real consequence) that calls it on click. Caller owns what cancelling actually means, same split as the swap-confirm modal's promise-resolving buttons. - Cancelling dismisses the banner, shows an info toast (not an error — this was deliberate), and closes the session, mirroring what the old timeout used to do automatically but now on the user's own call. - New .center-status-cancel CSS (bordered pill button, distinct from the plain "×" close glyph). Test changes: removed the now-invalid timeout-auto-close/estimate tests, added cancel-flow tests (dismiss/toast-type/session-close, never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests for the new Cancel button (bootAppWithRealCenterStatus, evaluating panels-ui.js instead of stubbing _showCenterStatus, since this button is worth verifying for real rather than just through the stub every other test in the file uses). Typecheck/lint/frontend-syntax clean; full suite shows no new regressions. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
2d3fc65758 |
feat(custom-model): show real-time llama.cpp backend status in the loading banner
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.
- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
(custom-model-routes.ts): one persistent GET /api/events (SSE)
connection held open per endpoint, parsing logData frames and
keeping the latest source:"upstream" (backend llama-server) line —
filtering out llama-swap's own source:"proxy" request-access lines.
Idle-closed after 30s of no polling, same 20s sweep as the existing
swap-displacement check.
- running-status route now returns logLine alongside the existing
isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
("llama.cpp: <line>", bootlog timestamp/level/component prefix
stripped for display) that stays on the last real thing llama.cpp
said rather than clearing to blank between polls.
⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).
12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
5ddc028a2f |
feat(custom-model): detect and notify when a session's model gets swapped out later
The llama-swap conflict check on the apply/create routes only ever runs at THAT session's own launch/apply moment, and cannot see a swap caused by a DIFFERENT session's later, ordinary use. Confirmed live: a second Codex session picking a different model launched with no warning at all — nothing conflicted at that exact instant — yet it silently evicted the first session's model regardless (llama.cpp runs one model at a time). Reproduced and root-caused via direct API calls against a live test-picker instance rather than guessing. - detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups live sessions with a customModel by endpointId, checks each group's endpoint via GET /running once, and flags a session whose own modelId is no longer in the running list. Read-only, best-effort per endpoint like refreshAllCustomModelHosts's sibling sweep. - Notifies once per displacement via a caller-owned de-dupe Set: a session id is added when displaced, removed once its own model is loaded/ready again, so a later genuinely-new displacement can notify again. - New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS, 20s — much shorter than the 5-minute model-list refresh, since this is time-sensitive) broadcasts a new custom-model:swapped-out SSE event per displacement. De-dupe Set cleared per-session on session cleanup to avoid an unbounded leak. - Frontend: global toast (not tied to the displaced session's tab, since the point is warning before the user types into it) naming the session, its previous model, and what's currently loaded. Chose the "detect after the fact" scope (vs. checking before every message send, which would add a round-trip to every turn on every custom-model session) per explicit user decision after being presented the trade-off. 9 new tests for the detection logic (flag/clear/re-flag cycle, unreachable/deleted endpoints, non-llama-swap servers, multiple sessions on one endpoint). SSE registry bumped 158->159, parity test passing. Typecheck/lint/frontend-syntax clean; full suite shows no new regressions (9 more passing than baseline, matching the new tests; same pre-existing Windows-environment failures). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |