mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
e587d8459039666fefde966bbc810a8ce1dc9e34
895
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e587d84590 |
fix(terminal): make the failed-load notice fit the narrowest terminal
A third pass in a real browser, at the widths this app actually renders at. The notice a failed history load writes into the blanked pane was one 70-character sentence. At 430px that exactly filled the line; at 320px it wrapped and left a lone '.' on a line of its own. The floor this app will render at is 40 columns — reachable today by raising the font on a phone — so the notice is three lines now, none over 25 columns, one fact each: what failed, that the session is still alive, and what to do. It says RELOAD rather than "reopen the tab" because `selectSession` early-returns when the session is already active, so clicking the tab you are already on retries nothing. The earlier wording named no next step at all, which left a mostly-empty terminal and no way out of it. CLAUDE.md no longer cites "758px reachable to the right" as evidence: that figure is a property of the test content, not of the fix, and the file's value is that a reader can trust a claim without re-deriving it. What is pinned instead is the invariant that survives any content — the full pane width is reachable, and removing the class returns scrollLeft to 0, so a resolved mismatch cannot leave the pane parked off-screen. Verified at 430, 360 and 320px against the shipped bundle, with the test asserting the built asset carries the copy so an edit that never reached the build fails rather than passing on the source's wording. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
abf1d1f1ca |
fix(terminal): the PTY and the browser terminal must never disagree about size
Issue #464, "text gets muffled sometimes, in both TUI default and fullscreen". The screenshot is not a dropped frame or a frozen renderer — it is arithmetic. Claude Code's TUI wraps its frame at the width the PTY reported and erases the previous frame by walking the cursor up the rows it believes that frame took. A browser terminal of a different width makes each logical line occupy more physical rows than Ink counted, so `eraseLines(n)` clears too few and the new frame paints over rows nothing erased: doubled lines, and short tool summaries sitting inside longer prose rows with the prose's tail still visible. Reproduced against this repo's own xterm before changing anything — a 120-column PTY against a 62-column terminal renders every wrapped line twice. `test/ terminal-pty-geometry.test.ts` pins that, and pins the clean render at matching widths beside it, so the assertion cannot be satisfied by code that fixes nothing. Four ways the two drifted apart, none of them observable from either end: 1. `fitAddon.fit()` resizes xterm to `proposeDimensions()` RAW while every server-facing path reported those floored at 40x10. Measured in Chrome at 430px: font size 44 proposed 13 columns, the server was told 40, and xterm stayed at 13. Three call sites each did their own fit-then-floor, and two re-read the proposal after the fit — `_shrinkPaddingToFit()` runs exactly there, so the container had moved. 2. `throttledResize` (keyboard up) and `sendResize` (session detached into its own window) reflowed locally and withheld only the SIGWINCH. That is the one combination that cannot be right: a reflow nothing is rendering for buys nothing and costs correctness. Both now withhold everything, and the keyboard's settle timer still sends the one resize that stops the PTY going stale. 3. `setFontSize`/`setFontFamily`/`setFontWeight` move the cell size — a geometry change — and told the server nothing at all, so raising the font on a phone left the CLI wrapping at the old column count. 4. `Session.resize` DECLINES a small-viewport request while a desktop connection holds an active sizing claim, and said nothing, because resize was write-only. `syncTerminalGeometry()` is now the one function that may change the terminal's size: it fits, floors and applies as a single step, so the numbers xterm holds are the numbers the server is told. A test sweeps every module for a bare `fit()` on the main terminal, and finds exactly one — the owner's own. For (4) the client cannot win, so it is told the truth instead: both transports answer a resize with `session.ptyCols`/`ptyRows` (`{"t":"zc"}` on the socket, the body of the resize POST) and `_onPtyGeometryReport` adopts them. A terminal that keeps a shape the PTY refused does not render "too narrow", it renders garbled. Adopting can leave the pane wider than the screen and the container is `overflow: hidden`, so `.pty-oversized` grants horizontal reach for exactly as long as the mismatch lasts: correct-and-reachable beats correct-and-clipped beats garbled. That rule sets both overflow axes and its own `touch-action` because mobile.css loads later and sets `.terminal-container { overflow: visible; touch-action: none }` — a bare `overflow-x` would leave overflow-y computing to `auto` and hand the browser a vertical scroll container the terminal's touch handler knows nothing about. Verified in Chrome at 430px against a live server, with a desktop client holding the claim: the phone adopts 198x43, gets `overflow-x: auto` / `overflow-y: hidden` / `touch-action: pan-x`, 758px of reach to the right, and keeps its own vertical scrolling. The pre-fix build was measured in the same harness for the control. Two things this deliberately does not do. It does not change who owns the pane size — the desktop still wins, and `_startMobileResizeRetry` still takes it back once that goes idle. And `throttledResize` still holds the PTY's shape for the whole keyboard animation rather than sending a SIGWINCH per step; that decision predates this and was not re-tested here. Also in this commit, Ark0N's third-pass review items on #431: - The response viewer's byte-buffer fallback and `_onSessionClearTerminal` both used the no-param `/terminal` form, capped only by `terminalBufferMaxBytes` (32MB) — the largest body the frontend asks for anywhere. One carried no deadline at all and the other got the 15s tail budget. Both now take the full-history budget. - A `?full=1` capture that outruns its deadline falls back to the bounded tail. The pane is blanked before that fetch, so an abort used to leave a black rectangle, discard the queued live output and never reach `_connectWs`. A failed load now still opens the socket, says one dim line where the content would have been, and clears the tab's spinner — which nothing did, so a failed select left `aria-busy="true"` set forever. - `_wsOutputGapSession` is cleared at the repaint that settles it, not in a `finally` that also ran on the catch. A reconcile that threw, or hit the new deadline — the flaky link the marker exists for — dropped the gap with nothing to retry it. `ws.onopen` no longer clears it up front either. - The replay-clear invariant is pinned in the gate, which is the drift this PR exists to fix: `_resetTerminalForReplay` must be a queued write and nothing else, and no module may blank the terminal with a `clear()+reset()` pair. - `DIAG_ENTRY_MAX_CHARS` replaces the hardcoded 300, bound through a local first: `CodemanDiag?.x` still throws a ReferenceError when the identifier was never declared, and that is the one function in the app that must not throw. - panels-ui's two kill-all clears route through the same helper, and the xterm-version guard's comment says "resolved lockfile version" rather than "dependency RANGE", which is what it has pinned since the last round. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
abd39318e6 |
fix(terminal): deadline must cover the body, precache must ignore the cache-bust query
Review fixes. Two of these are defects in the previous commit.
1. The fetch deadline only covered time-to-headers. `await fetch()` settles on
response headers, so clearing the abort timer in a finally around it left the
body — the multi-megabyte `?full=1` capture the deadline exists for —
completely unbounded; it only ever bounded a server that accepts a connection
and never replies. Measured against a server that sends headers immediately
and stalls the body 4s under a 1s deadline: fetch resolved at 30ms, timer
cleared there, body completed at 4026ms unaborted. Now the body is read
inside `_fetchTerminalCapture`, which returns {json, headers, headersAt} —
headers because two callers read server-timing, headersAt because those same
callers measure header-vs-body time and can no longer observe that moment.
`_terminalCaptureInflight` is scoped the same way, so a body still streaming
counts toward a capture starting beside it. Same test now aborts at 1005ms.
2. The precache could never be hit, and the previous commit made that expensive
rather than free. `renderIndexHtml` runs `cacheBustAssets`, which appends
`?v=<mtime>` to every same-origin .js/.css reference INCLUDING content-hashed
names — confirmed against a running instance:
`vendor/xterm-zerolag-input.6fee72f2.js?v=1789402869101`. `caches.match` is
query-sensitive, so entries keyed on the bare hashed path were unreachable;
deriving the list from the manifest turned cheap 404s into ~1.3MB downloaded
at every install that nothing could read back, once per deploy now that
CACHE_NAME rotates. The fallback match takes `{ ignoreSearch: true }`, which
also lets runtime-cached entries survive an mtime change.
3. `_wsOutputGapSession` was only cleared in ws.onopen, so paths that already
repaint the buffer left it set and the socket replayed everything a second
time. `selectSession` loads the buffer and only THEN calls `_connectWs`, so
neither the _isLoadingBuffer nor the _terminalRefreshOwner guard applied.
`_markTerminalBufferReconciled()` is now called from _onSessionNeedsRefresh's
finally, from selectSession after its load, and from _cleanupSessionData.
The scope claim was also wrong and is corrected in the comment: when the
network drops, SSE drops with it and handleInit's keepTerminal branch already
reconciles. The genuinely uncovered case is the WS dying while SSE stays up,
where _onSSETerminal discards SSE terminal frames until _wsReady flips in
onclose — up to the ping+pong window of output nothing writes.
4. CLAUDE.md said "all of them measured rather than reasoned", which the PR's
own "not verified" section contradicted. Split explicitly: the replay race is
measured, the watchdog mechanism is verified against xterm 6.0.0 under jsdom
(field path resolves, a forced stale handle makes refreshRows a no-op, the
kick schedules a fresh frame), and the iOS rAF-discard premise is reasoned
and still wants a device. Adds the two missing entries — the WebSocket
reconcile and the sw.js/build.mjs "keep these in sync or the build throws"
contract.
Also: test/xterm-private-api.test.ts pins the RESOLVED lockfile version instead
of the declared `^6.0.0` range, which was the wrong assertion in both directions
— a real upgrade to 6.4.0 can rename a private field while resolving inside the
range, and an innocuous range edit failed while changing nothing installed. And
test/sw-precache-manifest.test.ts now parses HASHABLE out of scripts/build.mjs
rather than hand-copying it, which was the same drift this PR exists to fix; the
parse is guarded against silently matching nothing.
The deadline fix has a behavioural test against a real socket plus a source
guard asserting `await res.json()` precedes the finally — verified to fail when
the helper is reverted to the old shape, so it is not vacuous.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
c0422c4e21 |
feat(terminal): renderer watchdog, atomic replay clear, fetch deadlines, reconnect recovery
Four ways the terminal can silently stop being correct — in each case the
buffer keeps updating, nothing throws, and the only recourse is a reload.
1. Renderer freeze after backgrounding. iOS DISCARDS scheduled rAF callbacks
when a PWA backgrounds, and xterm's RenderDebouncer only clears its
`_animationFrame` handle from inside that callback — so one drop leaves it
permanently set and every later refresh() early-returns. Parsing is
decoupled from rendering, so bytes keep filling the buffer correctly while
nothing paints. Codeman has exactly ONE xterm for the whole page load, so a
single backgrounding wedges it until a reload. Adds a 2s liveness poll and
`_kickRenderer()`, which does what the dropped `_innerRefresh` would have.
2. Replay clears raced live output. xterm's write() is async-queued while
reset() is synchronous and, per upstream, "does not clear input buffers and
does not reset the parser" — so bytes queued before a reset are parsed after
it and fuse into the snapshot. Verified against the real xterm 6 here:
write('p8'); reset(); write('rmissions') renders "p8rmissions". The main
path was already safe via a queued erase; the needsRefresh and clearTerminal
paths were not. All three now share one queued `\x1bc` (RIS), which unlike
3J/H/2J also resets modes, charsets, scroll regions and SGR state.
3. Output lost on WebSocket reconnect. Input frames carry seq+cid and are
delivered exactly once; output frames carry nothing. ws.onopen re-sends dims
and flushes queued input, and needsRefresh only fires on external-CLI
startup and SSE backpressure drain — never on reconnect. Output produced
while offline was simply absent afterwards. Interim fix: reaching onclose
means the drop was unintentional, so the session is marked and the next open
reconciles from the server buffer. Sequencing output is the follow-up.
4. Terminal captures had no deadline. No AbortController anywhere in the
frontend, including `?full=1`, which the code itself calls "unbounded-ish
work: at the default history limit it can be megabytes". Adds a budget that
scales with full-vs-tail and with captures in flight, degrading to a plain
fetch where AbortController is missing.
Also: the service-worker precache was dead — the build content-hashes assets
but sw.js listed pre-hash names, so 15 of 23 entries 404'd (verified against a
running instance) and cache.add().catch() hid it. Offline still worked via
runtime caching, but CACHE_NAME was a constant so activate's cleanup never
deleted anything and every past release's assets accumulated. Both are now
derived from the build manifest. Crash-trail entries are flattened and capped,
since they are joined with \n into one value and one call site interpolates a
server-controlled WS close reason.
The watchdog reads xterm privates — there is no public API. Every access is
optional-chained so a shape change degrades to a no-op. `_renderService` only
exists after open(), which needs a real DOM, so the gate cannot assert the
field path; test/xterm-private-api.test.ts pins the dependency range instead.
Tests: 23 new (terminal-resilience, sw-precache-manifest, xterm-private-api),
all pure/static so they run in the gate, which excludes the mobile suite. One
static source guard in history-truncation-notice updated for the renamed call;
the behaviour it pins is unchanged.
Not verified: no browser available, so no runtime reproduction of the freeze
and no real-device test of the reconnect path. Both warrant a device pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
9466acfc1a |
chore: version packages (#461)
* chore: version packages * chore: sync CLAUDE.md version to 1.32.0 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Codeman maintainer <noreply@anthropic.com> |
||
|
|
d8e85285c9 |
fix(mobile): merge-time fixes for the prompt composer (#444)
- styles.css: restate the composer overlay's own bottom gutter after the fold rules (the generic .paste-overlay longhand erased it: 0px flat, hinge strip replacing it folded) and subtract the fold strip from the dialog's max-height - test/foldable-layout.test.ts: simulate the cascade for .paste-overlay.prompt-composer-overlay (fails without the CSS fix); pin the palette anchor by name instead of ELEMENTS.at(-1) - keyboard-accessory.js: guard the app global in refreshForActiveSession() like the rest of the file - keyboard-accessory.js: a whitespace-only draft is empty (Send no longer submits blank lines); the text still goes out untrimmed - keyboard-accessory.js: derive _composerMaxLength and the frame refusal from one 64 KiB frame limit minus both bracketed-paste markers so they cannot drift - keyboard-accessory.js: translate the textarea placeholder and label at build time, since the DOM translator skips <textarea> subtrees - i18n.js: zh-CN entries for the composer dialog copy - docs/wiki/Mobile-Guide.md: describe the Compose key instead of a clipboard key - CLAUDE.md: a "Mobile prompt composer" paragraph after the accessory bar one - test/mobile-prompt-composer.test.ts: pin the whitespace rule and the derived budget Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit f6725ba52da17b0bdbee8be3b5011e7cae514f69) |
||
|
|
0f955327b2 |
fix(cli-registry): merge-time fixes for the run-menu consolidation (#458)
- test/opencode-resize.test.ts: retarget the launcher guard at the real code (this.selectSession(firstSessionId), any this.activeSessionId assignment) with an anti-vacuity check; the old strings existed nowhere, so it could never fail - session-ui.js: restore as comments the two invariants the merged bodies lost (deepseek leaves statusReporting unset, i.e. ON; no effort field for external CLIs, it is Claude-specific) - docs/cli-registry.md: move the frontend-guard paragraph below the two backend-guard paragraphs so they keep their antecedent, and note the widened comparison shape - test/frontend-cli-no-id-branching.test.ts: the comparison shape accepts any left-hand identifier (const m = this._runMode; m === 'codex' was invisible), normalized to `mode`; the two `m !== 'shell'` display filters are allowlisted and the remaining blind spots documented - test/run-mode-dispatch.test.ts: table-driven pin of run() dispatch (claude to runClaude, each RUN_MODE_LAUNCH id to _runCliMode(id), shell to runShell, unknown to runClaude, lock held and released) - CLAUDE.md: name the second CI-gated guard next to the backend one - server.ts: every </head> injection passes a replacer function; a clis.json label containing $' re-injected the rest of the document past escapeScriptJson (two render tests pin it, proven failing on the string form) - _isAltCliMode(): no reference anywhere in the tree, nothing to fix Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 1ea363ff808a62861559bc141e724b163cc1c56e) |
||
|
|
d3f2ec0220 |
fix(custom-model): merge-time fixes for the promoted-model picker (#459)
The maintainer's promised follow-ups to opticon454's picker promotion, applied on the landing branch after the merge ( |
||
|
|
dcf9437308 |
Merge pull request #460 from Ark0N/feat/installer-v2
feat(install): three questions up front, an unattended build, and a URL you can scan |
||
|
|
9a2e14a93a |
Merge pull request #453 from timkjr/feat/split-pane-sessions
feat: split-pane sessions — view two live terminals side by side |
||
|
|
72d437ab63 |
fix(install): fold in both reviews of #460
The two reviews on the PR (DeepSeek Harness, then Claude) found one class of
bug twice and a list of smaller ones; all of them land here, each pinned in
test/install-sh-invariants.test.ts and, where it is bash logic, driven in the
bash:3.2 CI step as well.
The Start line the done screen prints is now composed in one place
(start_command_hint) from every non-default value, the same five the exec
branch exports through export_bind_env, so "do not start" under a sub-path or
a custom port no longer prints a bare `codeman web`. The --lan / --tailscale /
env preset paths read ${CODEMAN_PASSWORD:-$EXISTING_PASSWORD}: a flag re-run on
a unit that carried a password used to rewrite it without the password and
with the unauthenticated ack. --password and --port flip RECONFIGURE so they
reach the unit instead of taking the quiet update path, and `install.sh name`
re-syncs the unit's base URL after the mapping is re-added.
Also: the sudo keepalive is ended before the exec into the foreground server
(exec skips the EXIT trap, and the loop keys on $$); Ctrl+C in the HTTPS-toggle
poll is trapped for the poll only and skips Tailscale for the run instead of
killing the installer; uninstall asks before removing a LaunchDaemon this
installer never wrote; a foreign daemon gets a launchctl kickstart hint and the
done screen stops claiming the new build is running; the preflight summary
reads the Tailscale state with a line grep when node is not installed yet; the
LAN security notice uses the configured port; a bare re-run ends on the done
screen; a build failure after a rename names the install.sh tailscale
recovery; TS_JOINED_HERE (written, never read) is gone; the plan doc and
architecture-invariants say what the code does. A minor changeset is included.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
1ba0684438 |
docs(install): describe installer v2 and the Tailscale naming options
README, the Installation / Remote-Access / Running-As-A-Service wiki pages, docs/security-architecture.md and CLAUDE.md describe the three-question flow, the flags, the subcommands, the sub-path answer for an occupied :443 and why the rename is opt-in. docs/installer-v2-plan.md is the design and the verification record (what was measured, what still needs a fresh machine); docs/tailscale-installer-plan.md points at it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
0b3e086334 |
fix(split-pane): address Ark0N's fourth pass — PTY-less picker exclusion, hollow chord test, remaining key gates
- buildSplitPickerSessions() now excludes any session with pid === null (exited CLI, tripped PTY-exit breaker, a restore that never re-attached). Pane B has no equivalent of selectSession()'s auto re-attach POST, so a split opened onto one had nothing reading its tmux pane: no terminal events ever arrived and Session.write() silently dropped every keystroke with no ack either way, while the socket itself reported healthy. - Fixed the hollow chord regression test: the synthetic keydowns carried no keyCode, which is what xterm's evaluateKeyboardEvent switches on to produce a data frame at all, so the assertion held regardless of whether the gate fired. Adding real keyCodes surfaced a second, real bug in the Alt+B case: the event bubbles to app.js's own document-level shortcut dispatcher, which really toggles the sidebar and resets the layout attribute the gate reads before Pane B's own (later, non-capture) handler ever sees it — fixed by driving the app's real settings cache instead of only the DOM attribute. - Ported the two remaining primary-pane gates with real consequences: Ctrl+Z (SIGTSTP) is swallowed for every non-shell session, matching terminal-ui.js's reasoning (an Ink/TUI agent loop stops dead with no visible output otherwise), and Shift/Ctrl+Enter now POSTs to /api/sessions/:id/send-key for THIS pane's own session instead of letting xterm send a bare \r, which used to submit an incomplete prompt instead of inserting a newline. Smart-copy Ctrl+C is re-implemented against Pane B's own terminal (copying app.copyTerminalSelection() would have copied Pane A's selection instead). - Updated docs/architecture-invariants.md and docs/split-pane-sessions-plan.md to match, and added CLAUDE.md's missing .split-picker-menu z-index entry. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3152ec801d |
docs(split-pane): short CLAUDE.md rule, stale module count, wiki entries, shortcut-handler caveat
CLAUDE.md previously only mentioned split-pane in the load-order list, with nothing in the Architecture/frontend prose the way every other feature gets, and its own module count was one stale (34, should have been bumped to 35 when terminal-split.js was added). Add a short pointer-style paragraph next to the other terminal features, fix the count. docs/wiki/The-Dashboard.md's header button table and docs/wiki/Settings-Reference.md's header chips list are the two user-facing surfaces that never mention Split at all; added both, plus a note that the feature is desktop-only regardless of the setting. docs/split-pane-sessions-plan.md: recorded the one design note that isn't a code change — the global capture-phase shortcut handler always resolves against Pane A, so Ctrl+L/Ctrl+W typed into Pane B affects the other session. Not fixed for v1, same reasoning as the rest of the "deliberately plainer" section. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
97a1238c85 |
feat(split-pane): add SplitTerminalPane class for Pane B
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
88e5b7b200 |
fix(custom-model): address Ark0N's PR review — client-side probe timeout, defer "last used" past confirmation, docs, zh-CN
Four things from the maintainer's review on PR #459, all fixed: 1. Bound _getCustomModelCurrentlyLoaded's probe client-side (~800ms via Promise.race, on top of — never instead of — the route's own 5s server-side timeout). Without it, an asleep/firewalled endpoint behind a saved model list left the picker completely invisible for up to 5s after the Run menu had already closed, with no spinner or toast. `timeoutMs` is an optional param (default 800, real callers never pass it) so a test can drive it in milliseconds, same pattern as `_watchLlamaSwapLoading`'s own `pollIntervalMs` — this code runs in a JSDOM window's own realm, whose setTimeout vi.useFakeTimers() cannot patch. 2. "Last used" is now written only once a launch actually applies, never on the mere click. It moved out of runCustomModelEntry (unconditional) and into each path's own success point: _quickStartWithCustomModelConfirm after the final post succeeds, and _runCustomModelEntryViaRestart right after the apply's success check. A context-window-warning decline means this exact model cannot work with this CLI at all, so the old unconditional write would promote, next time the picker opened, the one model guaranteed to fail again. 3. Documented the promotion/tag precedence and the new codeman:customModelLastUsed:<mode>:<endpointId> localStorage key in both CLAUDE.md's Custom Model Endpoint Profiles section and docs/custom-model-endpoints.md's Run-menu picker section. 4. Added zh-CN entries for "Currently loaded" and "Last used" in i18n.js, next to this modal's existing "Choose a model"/"Custom Endpoints" pair. New tests: the client-side timeout (endpoint that never answers, one that answers within the bound, and a rejected-after-timeout probe settling quietly), and "last used" recording on success vs. NOT recording on either confirmation's decline, for both the restart and one-shot paths. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
51b4a1b758 |
chore: version packages (1.31.0)
* chore: version packages * chore: sync CLAUDE.md version to 1.31.0 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Codeman maintainer <noreply@anthropic.com> |
||
|
|
4205f6930f |
fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine. **The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext` and `confirmedSwap` changed the wire field without moving three assertions that check it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts` (the swap modal and the context modal, each of which already receives exactly the right per-question flag). Moved, with the titles. **Worse, my own tests for the split never ran.** The four cases in `session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`, which was declared inside a sibling `describe`, so they threw a ReferenceError during setup. The split would have shipped with no passing server-side coverage while the gate reported the failure as four broken tests rather than as four tests that were never written. `mockRunning` is hoisted to the outer describe. **The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so the other eight modes fell back to claude's `❯`. That is also starship's default shell prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line `❯ npm run build` sits on screen for as long as the command runs, the verifier reads it as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N prompt, `read -p`, an installer or a pager, where it takes the default. The module's own fileoverview already stated the rule this broke. Now `?? ''`, which `promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI that actually declares a composer. **My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing three numbered rules. The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told an integrator to retry with `confirmed: true` for both questions, which is precisely the thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md` and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR` multi-user consequence, both user-visible and both previously absent, and #454's gained the one exception to its own claim: a Custom Endpoints launch ignores the Instance count stepper and always starts one session. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9af12afb57 |
docs(custom-model): make the docs match the code, and trim the changeset
More from the review of
|
||
|
|
3b55957d79 |
fix(custom-model): merge-time fixes for the Run-menu picker
Conflict resolution against the five PRs that landed while this was in review, plus the items left for merge on the thread. The real one was `session-ui.js`. #454 refactored all eight non-Claude `run*()` functions to funnel through one `_launchQuickStartInstances()` helper that does the POST itself, while this PR replaced that same POST in each of them with `_quickStartWithCustomModelConfirm()`. Resolved in the helper rather than seven times over: the helper now goes through the confirm path, and each body builder carries the `customModel` spread. `runAntigravity` deliberately does NOT, since antigravity's `customModelInjection` is `unsupported`; parity with this PR's own per-mode choices is asserted rather than assumed. That merge creates a question neither feature had alone: the confirm dialog now runs inside a loop that can launch up to 20 instances. Both questions it can ask (context window too small, and loading this will unload the model another session is using) are decisions about the ENDPOINT, and every instance in a batch targets the same one, so the answer is taken once and carried to the rest. Without that a 20-instance launch asks the same question 20 times. Also: `sse-events.ts` is 161 constants (master added two for remote wake, this adds one, verified by counting rather than by arithmetic), `server.ts` keeps both new SSE prefixes, the two comments pointing at code that no longer exists are corrected, and CLAUDE.md's SSE and route counts move to 161 / ~236 / custom-model (6). `pumpLlamaSwapLogTail`'s unparsed remainder is now capped at 64 KiB. It only shrank at a `\n\n` frame boundary, so a backend that streams without one would grow it for the life of a deliberately indefinite connection. NOT changed, deliberately: the context warning and the swap-conflict warning still share one `confirmed` flag with the context check first, so confirming "launch anyway" on a too-small context also skips the "this unloads it for another session" ask. That is the author's documented choice and the reviewer's own note calls it minor. Both fixes are worse to make here than to defer: separate flags are new wire surface landed unreviewed during a release, and reordering the checks adds a network round trip to a path that currently short-circuits. Raised as a follow-up instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1a99b5836c | Merge pull request #430 from opticon454/custom-model-run-menu | ||
|
|
035bfbc2fe |
fix(remote): merge-time fixes for Wake-on-LAN
The MAC-count limit lived in two places that disagreed. RemoteHostSchema.wakeMac's 128-character cap admits seven comma-separated MACs while parseMacList takes at most four, all-or-nothing, so a five-MAC value validated, was written to remote-hosts.json, and then resolved to NO wake target: POST /api/sessions/:id/wake answered "No wake-on-LAN target configured for this host" and the banner offered "Configure WoL" for a host the user had just configured. MAX_WAKE_MACS now lives in src/config/remote-wake-limits.ts and both sides refine against it. Its own module because src/remote-wake.ts is import-fenced to session-routes.ts and server.ts (the wiring guard that stops a watcher waking a host), and because schemas.ts must not drag dgram/net/child_process into every request-validating module. The documented 40 s request budget also omitted the wake's own cost. A `command` target is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS and runs BEFORE the readiness poll, so a slow one pushed a wakeCommand host's worst case to ~68 s, past the 60 s proxy_read_timeout the budget exists to stay under. _wakeAndWait now subtracts the wake's measured elapsed time from the readiness budget, floored at one poll interval so a wake that ate the whole budget still gets one probe. A magic packet is effectively instant and is unaffected, which is why live testing never saw it. Also: the two new endpoints are documented in docs/api-reference.md with the import fence stated as the rule it is, CLAUDE.md's frontend module count moves to 34, and the release changesets carry the Thanks section. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4c705094f7 |
fix(terminal): ship the copy clean as a trailing trim, without the shared dedent
#451 cleaned two things on copy. The trailing trim is right and every native terminal does it. The shared leading-indent strip is this project's own rule, and it is dropped here rather than shipped. Measured against the shipped transform over 401,445 three-row windows across 1,010 tracked files in this repo, it fired on 73% of them: 92% inside a YAML workflow, 76% over `git log` output, 48% in a TypeScript source. No width threshold separates a margin from content because they are the same widths, a live Claude Code pane's own margins measuring 2 and 5 columns while the most common non-TUI shared run is 4. The failure modes are not symmetric either: a wrong trailing trim costs nothing, while a wrong dedent silently deletes information that was on the screen, with nothing in the clipboard to hint at it, on git log bodies, on indented code read out of cat (semantic in Python), on git diff context rows where the leading space is the marker, and on stack traces. It also could not be made self-consistent cheaply. Whether the first row joined the measurement depended on the mousedown COLUMN, which the user never sees, so one block of three rows produced three different clipboard results; and the flag read getSelectionPosition().start, which is xterm's mousedown anchor and is never normalised, so dragging UP through a block read it off the bottom row. The PR's test stub hardcoded a downward drag, so its suite could not express that case. The transform, the wiring, the tests, the invariants, CLAUDE.md, the wiki page and the changeset all move together. The test block now pins the ABSENCE as a contract, with the git log, Python and git diff cases as its examples, so this is not re-derived later. If it is ever revisited, the one qualification that measured clean is painted trailing padding: zero false positives over all 401,445 windows. Also from the review: the comments and invariant rule justifying the padding-only clear described the pre-change code (the Ctrl+C gate reads the CLEANED selection now, so such a selection falls through to the PTY on its own and the clear is feedback rather than protection), the new 'Nothing to copy' toast gained its zh-CN entry, and the invariants paragraph no longer repeats its own opening sentence. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c376534a50 |
fix(run,terminal): merge-time fixes for the Instance count stepper and capture geometry
#454: the behaviour the PR adds had no test, so a regression test drives runGrok() at tabCount 3 and asserts three quick-start POSTs with sequential w<n>-<case> names (verified to fail against master's session-ui.js). Each caller now reads the count BEFORE its opening banner and announces it there, the way runClaude() already did, so a launch no longer prints two headers and a launch with another session already active still says how many are starting. runClaude() calls the shared _readTabCount() instead of its own copy of the 1..20 clamp, and that helper optional-chains the element read, since hoisting it above each caller's try block would otherwise let a missing #tabCount throw where the launch-error path cannot report it. #435: sizeMovedUnderLoad derived from data.source alone. `mux-visible` is not sufficient: a failed display-message cursor query makes capturePaneBuffer skip the snapshot repaint and return the raw capture, which the route still labels mux-visible, so a size that moved during such a load bought a full forced reload to repair a frame that was never positioned. It now tests Number.isFinite(data.captureRows) like its two siblings. Plus the invariants and CLAUDE.md lines promised on #435: a visible capture reports its geometry and omits it when nothing was positioned, the comparison runs on mux-visible only, and the replay is capped at one attempt and latches per session when it cannot converge. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2c3ccdf030 |
Merge pull request #439
feat(remote): wake a sleeping host (Wake-on-LAN) from input, banner and native magic packet |
||
|
|
5fc391a47c |
fix(custom-model): address fourth pre-merge review + merge upstream master (Ark0N)
Merged upstream/master (22 commits: reboot-restore recovery feature,
terminal keycode229 recovery work, install.sh/CLI-catalog generator
changes, CHANGELOG/version bump to 1.30.0) into this branch. No
conflicts; git auto-merged every overlapping file (CLAUDE.md,
docs/api-reference.md, app.js, index.html, styles.css, routes/index.ts,
session-routes.ts, schemas.ts, server.ts).
Two required fixes from the latest review:
1. privilegedEnvKeys widening (stock.ts) changes behaviour outside this
feature. The reviewer decided to keep both CLAUDE_CODE_MAX_CONTEXT_TOKENS
and CLAUDE_CONFIG_DIR listed (types.ts's rule that every traffic-
redirecting var this feature introduces must appear there stays
literally true), and asked for the real consequences documented
instead of hidden:
- Corrected session-env-clamp.ts's fileoverview, which stated the
opposite of what the code now does (reboot-restore's clamp call
used to be able to strip nothing for claude; it now strips a
persisted CLAUDE_CONFIG_DIR for a non-granted owner).
- Corrected the rationale comments in stock.ts: privilegedEnvKeys
has exactly one consumer (ownerClampedEnvKeys, feeding the
generic envOverrides clamp on create/quick-start/reboot-restore),
not the custom-model routes.
- Added a CLAUDE.md line to the CLAUDE_CONFIG_DIR gotcha covering
the admin-only-in-multi-user-mode and reboot-restore-strips-it
consequences.
- Added a "Claude multi-user clamp" test next to the existing
DeepSeek/OMP ones, pinning the new stripping behaviour.
2. GET .../running-status (custom-model-routes.ts) no longer passes
the raw llama-swap `cmd` field (the literal launch line, which can
carry model paths and --api-key) to the browser -- the frontend
only ever reads model/state, cmd exists solely for server-side
parseCtxFromCmd() during discovery. Added a test asserting the
response never contains cmd or a planted secret.
Also regenerated config/clis.stock.json and install.sh's catalogue
block (npm run generate:cli-catalog) to clear drift introduced by the
upstream merge, since it was failing the sync check.
Left to the reviewer, as they said they'd take at merge: the two
"comments pointing at removed code" cleanups, the two stale CLAUDE.md
counts, and the small items list (mode==='claude' frontend branch,
isCliAvailable() unknown-id gap, shared confirmed flag ordering,
one-shot cancel toast severity, pumpLlamaSwapLogTail buffer cap).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
|
||
|
|
5bb489addb |
fix(remote): authorize the attach wake first; tell the caller what happened to its bytes
Review round 3 on #439. - The attachRemoteSession branch of POST /api/sessions ran `ensureHostAwake` before the multi-user gates, so a non-admin could have any configured host's `wakeCommand` spawned (or a packet broadcast) and the request held for the wake budget, then be refused for the workingDir. The admin gate now comes first, before the host is even looked up; remote hosts are admin-only infrastructure everywhere else. Route test: wake spy empty, 403. - The non-wait input route answers `{buffered:true}` when the registry took the chunk and `{buffered:true, dropped:true}` when it was over the cap and is gone (`RemoteInputOutcome` gains 'dropped'); additive to the bare `{}`. - The send-and-wait path answers OPERATION_FAILED when the host never comes back, like create and attach, instead of writing into the stalled pane and reporting delivered:true plus a timeout. - The flush writes with `fromUser: true`, so a first prompt buffered through a wake can still name the tab. Docs: api-reference (input route), remote-sessions.md (two invariants), CLAUDE.md key pattern. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG |
||
|
|
19ffe9b7a8 |
fix(input): make sure a prompt sent through the API actually leaves the composer
Claude Code 2.1.277 takes typed text the moment its composer paints but ignores Enter for the first 30 to 50 seconds after it (measured 2026-09-19 through the input route: an Enter at 28 s stranded the prompt, one at 51 s submitted it). The text+Enter pair `sendInput` sends 50 ms apart therefore left every programmatic prompt sitting unsent, and every waiter burned its timeout on a turn that never started. Server: `SubmitVerifier` (session-submit-verifier.ts), armed from `writeViaMux` for every mux write that carried a carriage return, reads the pane on a 2 s to 60 s schedule and re-sends Enter only while the last composer line (the CLI's own prompt glyph) still holds the head of what was sent. An empty composer, other text, or no composer line at all ends it; a newer write replaces the schedule. Skill: `sendwait` gets the same loop (`_composer_text`, no-break space stripped by its bytes for BSD sed) for servers that predate this, and the preamble version moves to 1.30.1 so seeded agents pick up the fresh copy. SKILL.md's heredoc and the plugin mirror are regenerated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1040f6c489 |
fix(remote): a proxied host is reachability-unknown; scope remote: SSE per session
Review round 2 on #439. 1. The bare TCP probe connects to host:port, which a host behind a jump host or SOCKS proxy does not answer even while ssh works. Acting on that verdict drew a permanent banner over a healthy session, replaced a real "needs tmux" error with "not reachable" in quick-start, and - with a wake target - buffered every HTTP input for the life of the session, since the readiness poll could never succeed. `WakeableRemote` now carries `jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such a host into reachability-UNKNOWN: input is delivered, `checkReachable` / `checkHostReachable` answer `null` (never `false`), `ensureHostAwake` returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate fires on `=== false` only, and `GET …/reachability` reports `reachable: null, probeable: false` so the banner has nothing to key on. A wake target can still be fired for it, blind: no readiness poll, no reattach, no toast - the response says only whether the packet went out. 2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake has no session yet, so the registry names the requesting user (`ensureHostAwake({ requestedBy })` -> `username` in the payload) and `deriveSseHint` routes on it; with neither it fails closed to admins. Single-user mode is unaffected. Smaller, from the same review: - A flush write that fails now drops the remaining buffer (logged) instead of retaining it: the wake still resolved and marked the host reachable, so the retained chunk waited for the NEXT wake and was replayed hours later, after everything typed since. Same policy as the oversized paste. - The banner polls on tab activation (a user action) and on its 30 s timer only for a host with a wake target; a timer connecting to a host Codeman cannot wake is the traffic invariant #2 rejects keepalives for. A proxied host is never polled. - `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP socket refuse under VITEST, as remote-files.ts does. The guard caught a leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the probe but still polled readiness with the real one, so the shutdown test had been connecting to a production address. The poll now uses the injected probe. - docs/remote-sessions.md is additions only again (the reformatting is gone); the architecture-invariants overlap resolved itself in the merge. Live, against a throwaway instance with a non-routable ghost host: proxied -> no probe, no wake, the genuine ssh error after 10 s; direct (control) -> probe, magic packet, "did not come back" after the 40 s budget. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG |
||
|
|
e271a65e79 |
Merge origin/master into feat/remote-host-wake
Resolves CLAUDE.md count tables (route counts recounted on the merged tree: 235 handlers, sessions 37) and keeps both the host-wake and the reboot-restore banner in index.html. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG |
||
|
|
56209e7829 | Merge remote-tracking branch 'upstream/master' into feature/run-menu-custom-model-picker | ||
|
|
9982a1325f |
fix(custom-model): address second pre-merge review (Ark0N)
Blocker: .center-status-banner never actually disappears.
- Add `.center-status-banner[hidden] { display: none; }`, same trap as
`.home-sessions[hidden]`: the author-level `display: flex` beat the
UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
card stayed laid out at `opacity: 0` with its text/cancel/close
children still `pointer-events: auto` -- an invisible 442x67 click
blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
banner (10001) and the swap-confirm/context-warning modals (10010)
in CLAUDE.md's Z-index layers list.
Stale wording pointed at the reverted sticky-toast default:
- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
`.toast-message` comment in styles.css all still said "toasts
default to sticky" after
|
||
|
|
20fc7b3c3d |
chore: version packages (#447)
* chore: version packages * chore: sync the CLAUDE.md version line to 1.30.0 The changesets bot does not touch this line, and pushing it to master after merging the version PR starts a second Release run that has raced the first before. Riding the bot's own branch keeps it to one push. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Codeman maintainer <noreply@anthropic.com> |
||
|
|
bb8ada7e5f |
fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does. 1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred it from the name: a session the user renamed by hand to something shaped like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the next prompt overwrote their name. The route persists right after, so the loss went to disk. `restoreMuxSessions()` already passes it. 2. The already-live sets were snapshotted once before a loop that awaits a real `startInteractive()` per entry, so by the tenth entry the snapshot was tens of seconds old and a conversation resumed by hand from the Resume list in that window was invisible to it: two panes on one transcript, the exact thing the check exists to prevent. Both sets are now read per iteration, and the late case is spent rather than re-offered for the same reason the batch case is. 3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path. The stamp predates the reboot and the pane is new, so honouring it meant one click had every restored session type `continue` into itself about a minute later, unattended, against the route header's own promise that a restored session comes back idle and disarmed. The setting stays ENABLED, so it re-arms on the next real limit message. A Codeman restart still re-arms from the stamp, because the limit footer will not reprint on its own; the new option exists only to tell the two paths apart. 4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()` and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession` performs that it was missing. Cosmetic, but a run left open reads as still going in the away digest. 5. A restored claude session gets `seedAgentSessionPreamble()` like both create paths, so the agent skill's bootstrap stays a two-line loader. 6. The heuristic's container comment was wrong in one direction and quiet about the real gap: after a genuine host reboot a containerized Codeman sees the host's short uptime and the banner does appear. What it cannot see is a container-only restart, which is where this would help most. 7. The banner is hidden in a solo window, which shows one session and has no tab strip to put restored ones in. Also reverts 17 of the 18 hunks in docs/api-reference.md, which were Prettier reformatting of prose the PR does not otherwise touch (docs/ is outside the format glob), keeping only the Reboot restore section and repairing the two continuation lines that reformat de-indented; renumbers reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which loads after it; and gives the feature its CLAUDE.md entry plus a route test for the multi-user workspace-forbidden branch, the only new rule that had nothing behind it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ea5323d990 |
test(input): pin the batched commit-plus-Enter ordering #441 fixes
The unit harness proves WHICH candidate gets forwarded; the ordering is the half that shipped the bug, and only a real xterm shows it. The new browser case dispatches the character's keydown, its composed insertText and Enter's keydown in ONE page task, the shape an Android soft keyboard delivers through a single InputConnection transaction, and asserts what reaches the send path. Verified in both directions on this machine: with the drain in place the wire is `o\r`; with the drain removed (master's behaviour) it is `\r` and the character is gone entirely, because by the time the zero-delay timer runs xterm has emitted the `\r` and bumped the canonical counter past the candidate's snapshot, so the candidate stands down. The other four cases pass in both states. CLAUDE.md now names the decision point, what it costs (a keydown decides with less evidence than the timer did) and why that is safe for Enter, and says that the pin lives in a suite the CI gate does not run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1f61d21298 |
docs: correct six stale counts and claims in CLAUDE.md
Each of these was measurable and wrong: the CI note listed 5 excluded Playwright tests where config/test-suites.ts has 9, never mentioned the packages/xterm-zerolag-input run that follows the gate, and never mentioned wiki-sync.yml at all; the format glob note omitted that lint covers only src/**/*.ts; app.js is ~6.9K lines, not ~6.7K, and voice-pcm-worklet.js is fetched from JS rather than sitting in the load order; src/config/ holds 23 files plus the cli-registry/ subdir, not 21, and nothing said that the repo-root config/ is a different directory; the route count is ~232 with cases at 34, not ~228 with cases at 30. Also adds the pointer to docs/wiki/ as the user-facing manual, which the header describes every other doc surface but not that one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e2034177c5 |
fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.
Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
published dependencies) into a scratch dir purely to read
dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
same endpoint. dsh's own error template ("DeepSeek API error (HTTP
${status})") reproduces the originally-reported
"dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.
- New registry field `appendV1Suffix` (env kind only, deepseek's entry
alone — claude/gemini must NOT get it, since claude was already
confirmed working against the unmodified baseUrl). When set,
buildCustomModelInjection runs endpoint.baseUrl through the same
withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
use, instead of writing it verbatim.
Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.
2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
8520925e76 |
docs(custom-model): bring CLAUDE.md and api-reference.md up to date
Full documentation review pass across the branch's 30 commits.
CLAUDE.md's Custom Model Endpoint Profiles entry hadn't been touched
since the initial backend+picker cut (3 early commits) despite 27
follow-up commits adding real behavior — it described restart-in-place
as universal (now claude-only; 7 other CLIs launch one-shot) and
claimed codex's Responses-API gap as a flat protocol break (now
re-verified as a more precise tool-calling gap). Corrected both and
added a new paragraph covering everything landed since: the llama-swap
conflict check, the after-the-fact swap-displacement sweep, the
/running-cmd-based context-length fix, the context-window floor
warning, skipFirstRunPrompts, the real-time /api/events-based log
status, and the countdown-to-Cancel-button change.
docs/api-reference.md's custom-model-endpoints section was missing the
running-status route, the requiresConfirmation/requiresContextWarning
response shapes, and POST /api/quick-start's customModel field
entirely (the primary launch path for 7 of 8 supported CLIs) — added
all three. Also fixed a real markdown bug in custom-model-endpoints.md:
an inline code span (`POST <baseUrl>/v1/chat/completions`) split across
a line break, which CommonMark renders with the line ending collapsed
to a space, so it displayed as ".../v1/chat/ completions" with a
spurious space inside the path.
Verified: origin/master and upstream/master are both already an
ancestor of this branch (identical at
|
||
|
|
acb8d4b0aa |
docs(remote): correct what the wake PR moved
- `host-wake-ui.js` joins the documented load order (12.2) and gets its `@dependency`/`@loadorder` tags; the frontend module count is 33, not 32. - `remote-wake` is not "(pure)" — the module uses `dgram`/`net`/`child_process`. - SSE counts: 160 constants, and the category is "Remote auto-reconnect / wake (5)"; the route table's per-file counts are refreshed (sessions 37, cases 34). - The CLAUDE.md wake rule now names the create/attach wake, the 40 s request budget, the whole-chunk paste drop, the registry's lifetime (drop on cleanup, stop on shutdown) and the deliberately non-wake-aware WebSocket keystroke path — that paragraph is what the next person reads. - Reverted the eight lines of unrelated Prettier markdown churn in `docs/architecture-invariants.md` (docs/ is not in the format glob, so it was an editor): only the new wake paragraph remains in the diff. |
||
|
|
5c25a52f95 |
fix(custom-model): wait for a freshly launched session to go idle before applying
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.
Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.
Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
5a9ff07f57 |
feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real llama.cpp server: 1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to apply the endpoint's defaultModelId (or the first discovered model) silently. Now, via the new selectCustomModelEntry() (session-ui.js): - exactly one discovered model launches straight away, same as before - two or more open a new #customModelPickModal listing every discovered model; defaultModelId (if set) is marked but never auto-chosen, since the point of asking is letting ONE launch deliberately differ from the saved default, not just confirming it The endpoint is re-fetched at click time rather than trusting anything cached from the dropdown's own render, since the model list can have changed (the sweep below, or a settings-panel edit) since it opened. runCustomModelEntry() itself — the actual launch, routed through run() for the in-flight lock, snapshot-guarded against applying to the wrong session — is unchanged; it now just always receives an explicit model id from one of these two paths instead of computing one itself. 2. Periodic re-discovery. Every saved endpoint's models now refresh automatically every 5 minutes in the background (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way as the Codex plan-usage poll it sits beside — this.cleanup.setInterval, off under testMode), so a model the server starts or stops serving shows up without another manual "Discover" click. The manual POST .../discover-models route and the new refreshAllCustomModelHosts() sweep (custom-model-routes.ts) now share one pure merge step (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId that no longer appears) rather than two copies that could drift. The sweep is best-effort per host — one endpoint being unreachable on a cycle never blocks the others — and re-reads the store before each host's write, keyed by id, so a concurrent edit or delete from the settings panel always wins over a sweep that started before it. Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated file for the sweep (kept separate from custom-model-routes.test.ts because that file's data dir is shared across every test in it — one temp HOME per FILE, not per test — which would make a sweep-touches-every-host assertion meaningless there). test/custom-model-run-menu-ui.test.ts gained a new describe block driving the real picker modal through JSDOM: single-model bypass, multi-model dialog with the default marked-not-chosen, picking a row closes the modal and launches with that exact model, the endpoint re-fetch, and the two "vanished by click time" toast paths. Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md, docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated — the last of these also caught up two sentences that had gone stale after the draft-review fixes landed (the picker routes through run() now, not a raw run*() call). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
25fae9ad10 |
feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment: "generate those entries from the saved profiles rather than a fixed duplicate per harness, and put it in a follow-up PR so this one stays the backend... The Run-menu picker is yours if you want it." Adds the frontend surface the backend has been waiting on: - Run menu: a "Custom Endpoints" section lists one entry per (harness that supports customModelInjection, saved endpoint) pair, e.g. "Claude Code (llama.cpp)". The harness list comes from window.__codemanCustomModelClis, injected at page render straight off the CLI registry's own capabilities (never a hardcoded id list in the frontend), so a CLI whose injection recipe lands later appears with no frontend change. Picking an entry runs that harness's own existing run*() function unmodified (case creation, env overrides, everything, forced to a single instance) and then applies the endpoint's default model to the session it creates via the existing POST /api/sessions/:id/custom-model route. Entries are hidden for a remote/docker active case, since that route already refuses both. - Settings: App Settings -> Models gets a "Custom model endpoints" group wiring up the customModelEndpointsEnabled toggle (declared since #393, read by nothing until now) plus CRUD against the existing /api/model-endpoints routes: list, add/edit (inline form), delete, discover models. - Backend: CustomModelHost gains an optional defaultModelId, the model the picker applies with no further choice per endpoint (one generated menu entry per CLI+endpoint pair, not per CLI+endpoint+model). The route refuses a value that isn't one of the endpoint's own discovered models, and a fresh discovery drops a default that no longer appears rather than carrying an invalid one forward. Docs: docs/custom-model-endpoints.md describes the new picker and settings panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the "backend-only" status note and documents the picker's generation mechanism. Tests: four new route tests cover defaultModelId validation, acceptance, and the drop/keep behaviour across a re-discovery; a new render-index-html test pins the __codemanCustomModelClis injection (present, agent CLIs supporting the capability, antigravity and shell excluded) and its solo-window skip. No browser test was added for the Run-menu picker itself or the settings CRUD panel (this box has no tmux, so the live server used by test:browser/test:mobile could not be exercised here) -- worth a Playwright pass before merge, same as any other frontend PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
3248f35081 |
chore: version packages (#437)
* chore: version packages * chore: sync the CLAUDE.md version line to 1.29.1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Codeman maintainer <noreply@anthropic.com> |
||
|
|
5b920cb43d |
feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream and named tabs wrong five ways (every prompt, every write path, a bare Esc eating the next prompt's first character, pasted newlines as Enter, any CSI clearing the draft) and replaced the whole name, which dropped the case from the tab and reset the w<n> counter. This lands the feature with each of those closed: - First prompt means the first: applyAutoName() flips a placeholder to `auto` whether or not the string changed. nameSource is now the tri-state placeholder | auto | manual; the name setter is the only manual path. - Only user-originated input counts: write()/writeViaMux() take SessionWriteOptions.fromUser, set by the browser WS path and POST /input only, so Ralph, respawn, cron, approvals and the trust-dialog keys can never name a tab. A startMode 'shell' CLI never feeds the tracker (a capability, not an id check); the send-key route feeds trackUserInput() because its line feed bypasses the session. - Prefix form `w3-case: title`: parseSessionPrefix() already renders it as the title with the prefix in the tooltip and the next-session counter still matches it. Composed within MAX_SESSION_NAME_LENGTH. - Tracker rules per key: bare Esc resolves at chunk end; mouse/focus reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R taint the draft so Enter submits nothing rather than a fragment; bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space; the draft keeps its head past 8192 code points; an escape past 64 bytes is abandoned. - Title: slash commands by shape (a path is a prompt), `!` escapes refused, first sentence only past 8 code points ("e.g." is not a title), 72 code points on a word boundary. - Synced `autoNameSessions` setting, default OFF (the prompt reaches mux-sessions.json, session:updated and /api/search), App Settings -> Appearance -> Tabs, read fresh per prompt after the eligibility check. Tests: test/session-auto-name.test.ts (tracker, title, composition, ownership, emit gating), the wiring test (once, prefix, setting off, manual protected), test/routes/session-name-routes.test.ts (PUT /name flips to manual and persists). Verified live on an isolated instance: API and browser-typed prompts name the tab, a second prompt does not, shells and renamed tabs are untouched, nameSource survives a restart. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
8b5a13435a |
feat(remote): host-unreachable banner, manual wake, and native MAC wake-on-LAN
The reactive wake (typing into a session whose host slept) left the state invisible: nothing told the user the machine was asleep, and with no wake target configured there was nothing to do about it. Adds: - RemoteHost.wakeMac (comma-separated) - Codeman builds and broadcasts the magic packet itself (UDP port 9), so the common case needs no external script. The existing wakeCommand stays as the explicit override. - GET /api/sessions/:id/reachability - probes (throttled, cached, and it never wakes) and reports HOW the host can be woken, or that nothing is configured. - POST /api/sessions/:id/wake - wakes, waits, reattaches the pane and flushes buffered input; 400 with a routable message when no target is configured. - The amber host-unreachable banner + its 'Wake' / 'Configure WoL' action, and a small config dialog that saves via PUT /api/remote-hosts/:id. - RemoteWakeDeps.resolveRemote: host config is re-resolved for LIVE sessions (throttled + cached), so saving the dialog takes effect without a restart. |
||
|
|
3f0bfde54a |
docs(remote): document the wake-on-LAN invariants; drop wake state on bulk delete
Self-review pass: the input-ladder's two 'buffer' branches were the same three lines, and bulk delete left a session's (bounded, per-random-uuid) wake state behind. Documents the design where the code refers to it - remote-sessions.md section, the architecture invariant, and the CLAUDE.md key pattern. |
||
|
|
88e3faa456 | chore: version packages | ||
|
|
70fc6b32d5 |
docs: record the dup/last input ACK, Shift+drag and right-click copy, and multi-case adopted containers
Three behaviours landed from #375 without their doc entries: the duplicate input ACK now carries `dup:true` and the server's watermark (`docs/reliable-input-delivery.md` still described a bare ACK), Shift+drag and right-click copy in the terminal (the shortcut list did not know them), and one adopted container backing several cases at different in-container directories (the Docker cases paragraph still implied one case per container for adopted containers too). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
942bf37e48 |
fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a custom OpenAI-compatible endpoint by injecting env vars or a config file and restarting the CLI in place. Review of the apply path found four things, two of them destructive. This lands all four plus the smaller items from the same review. 1. Clearing a selection did not clear it. The injected vars reach the CLI via `tmux setenv`, which persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so deleting the keys from the session's envOverrides relaunched the CLI still pointed at the old endpoint, and for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just been deleted. `Session.setCustomModel()` now reports the removed keys, queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys` carries them into `applyEnvOverrides()`, which `setenv -u`s them before re-applying the live overrides, on the same path that already unsets the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket that `setenv -u HOME` hands the next respawn the global HOME back. 2. Applying a model to a local claude session killed the pane. The relaunch was `claude --session-id <id>` and Claude refuses an id that already has a transcript, and unlike the dead-pane respawn this one kills a working pane first. `restartCli()` now pins the live conversation id as the resume id for that respawn when the CLI's launch declares a `fallback` chain, which renders the same `--resume <id> || --session-id <id>` shape the docker and remote pane commands use. Gated on the registry shape, not the CLI id: an entry whose resume id is minted by the CLI itself never declares that chain. 3. pi, omp and grok wrote their config file and then launched without the `--model` that selects it, so the file was ignored. The registry entry now declares `customModelInjection.launchModel` (`custom/{modelId}` for pi and omp, grok's `[model.codeman-custom]` block name), the builder renders it, and `_withCustomModelLaunchModel()` applies it onto the respawn options through `legacyConfigField`, leaving the stored <Mode>Config untouched so a clear falls back to the user's own model. A model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. 4. Remote (SSH) and Docker sessions reported `restarted: true` and changed nothing: their `restartCli()` reattaches the durable tmux rather than relaunching the agent, and the env lands on the local pane. Both are refused with a 400 until those paths are plumbed. Smaller items from the same review: - The selection survives a Codeman restart as the disk-only `__customModel` bookkeeping (endpoint, model, injected key NAMES, config dir, launch model; never the values, which carry the API key). Recovery re-derives the values from the endpoint store through the same apply path the route uses and keeps the bookkeeping even when the endpoint is gone, so a later clear still has keys to unset. - Discovery goes through `webviewFetch()`, so the RESOLVED address is judged by the same egress guard the web-tab proxy uses, and `baseUrl` reuses `webviewUrlSchema` (http(s) only, no embedded credentials, link-local and cloud-metadata addresses refused). undici's `fetch failed` wrapper is unwrapped so the user sees the ECONNREFUSED underneath. - `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session config dir 0700/0600 (pi and omp embed the key literally), and that dir is removed with the session. - `PR.md` is gone from the repo root and the design doc moved to `docs/custom-model-endpoints-plan.md` with the LAN address and the personal name scrubbed; every reference follows. The guide's `authStyle` text matches the shipped schema (`bearer | api-key`, default `bearer`) and says that `customModelEndpointsEnabled` is read by nothing until the picker lands. - `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts` (four real type errors fixed). It is not yet wired into `npm run typecheck` because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json` there is the one-line follow-up. Tests: `test/session-custom-model-restart.test.ts` drives a real Session and fails on the unfixed code for items 1 to 3; the route suite covers item 4 and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets run before the overrides and that a shell-metachar key never reaches tmux. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1e42cb4e2d |
Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses) |