mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
2d3fc65758c7fca40fc0f31b58b406f1e4deb67a
2261
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2d3fc65758 |
feat(custom-model): show real-time llama.cpp backend status in the loading banner
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.
- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
(custom-model-routes.ts): one persistent GET /api/events (SSE)
connection held open per endpoint, parsing logData frames and
keeping the latest source:"upstream" (backend llama-server) line —
filtering out llama-swap's own source:"proxy" request-access lines.
Idle-closed after 30s of no polling, same 20s sweep as the existing
swap-displacement check.
- running-status route now returns logLine alongside the existing
isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
("llama.cpp: <line>", bootlog timestamp/level/component prefix
stripped for display) that stays on the last real thing llama.cpp
said rather than clearing to blank between polls.
⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).
12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
5ddc028a2f |
feat(custom-model): detect and notify when a session's model gets swapped out later
The llama-swap conflict check on the apply/create routes only ever runs at THAT session's own launch/apply moment, and cannot see a swap caused by a DIFFERENT session's later, ordinary use. Confirmed live: a second Codex session picking a different model launched with no warning at all — nothing conflicted at that exact instant — yet it silently evicted the first session's model regardless (llama.cpp runs one model at a time). Reproduced and root-caused via direct API calls against a live test-picker instance rather than guessing. - detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups live sessions with a customModel by endpointId, checks each group's endpoint via GET /running once, and flags a session whose own modelId is no longer in the running list. Read-only, best-effort per endpoint like refreshAllCustomModelHosts's sibling sweep. - Notifies once per displacement via a caller-owned de-dupe Set: a session id is added when displaced, removed once its own model is loaded/ready again, so a later genuinely-new displacement can notify again. - New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS, 20s — much shorter than the 5-minute model-list refresh, since this is time-sensitive) broadcasts a new custom-model:swapped-out SSE event per displacement. De-dupe Set cleared per-session on session cleanup to avoid an unbounded leak. - Frontend: global toast (not tied to the displaced session's tab, since the point is warning before the user types into it) naming the session, its previous model, and what's currently loaded. Chose the "detect after the fact" scope (vs. checking before every message send, which would add a round-trip to every turn on every custom-model session) per explicit user decision after being presented the trade-off. 9 new tests for the detection logic (flag/clear/re-flag cycle, unreachable/deleted endpoints, non-llama-swap servers, multiple sessions on one endpoint). SSE registry bumped 158->159, parity test passing. Typecheck/lint/frontend-syntax clean; full suite shows no new regressions (9 more passing than baseline, matching the new tests; same pre-existing Windows-environment failures). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
470f75b08c |
docs(custom-model): record live findings on codex's model-metadata warning
Investigated the user's report of "Model metadata for <id> not found. Defaulting to fallback metadata..." on every custom-endpoint codex launch, live against the test-picker's llama-swap deployment (codex 0.152.1): - The warning is cosmetic. `codex exec 'reply with just OK'` against the isolated CODEX_HOME still printed the warning and still returned a real reply. - The isolated CODEX_HOME never gets a models_cache.json written into it at all, even after extended real use (inspected a live, actively- used directory) — codex can't reach OpenAI's own hosted model catalog for this session and silently falls back every time, with no local file to create or clean up. There is also no config.toml override for a model's metadata. - Fabricating a fake catalog entry to suppress it would mean copying the SHAPE of OpenAI's own proprietary models_cache.json schema, including real per-model system-prompt content visible in a genuine entry — not something to build for a warning confirmed to have no effect. - More importantly: a real tool-call attempt against the same setup came back as agent_message TEXT (the tool-call JSON printed as the answer) rather than an executable function_call item, confirmed via `codex exec --json`'s raw event stream. Tool execution is what makes codex a coding agent, so it remains not usable for real work regardless of the metadata warning — a more precise, re-verified update to the existing "Responses API protocol gap" finding (which reported a harder Reconnecting/high-demand failure on a different llama-swap deployment; this one answers /v1/responses for plain chat but still can't execute tools). No code changes — recipe/comment/confidence-table documentation only. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
211b872335 |
feat(custom-model): skip Claude Code's first-run wizard on custom-model launches
A fresh, isolated CLAUDE_CONFIG_DIR (used to keep an injected API key away from a stored claude.ai OAuth login) looks like a brand-new Claude Code profile to the CLI, so it replays its ENTIRE first-run sequence on every single launch: the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running with --dangerously-skip-permissions) a one-time bypass-permissions warning — confirmed live, none of which a real, already-onboarded profile shows again. - New registry-declared env-kind field `skipFirstRunPrompts` (alongside apiKeyTrustFile, which it reuses) — claude's entry only, carried through buildCustomModelInjection (pure) into applyCustomModelInjection (IO). - seedFirstRunOnboardingState(): merges hasCompletedOnboarding: true and this session's own projects[workingDir].hasTrustDialogAccepted: true into the same <configDir>/.claude.json the API-key trust file already writes to — other projects and other fields on this session's own entry are left untouched. - seedSkipBypassPermissionsPrompt(): merges skipDangerousModePermissionPrompt: true into <configDir>/settings.json, a separate file, same corrupt-tolerant merge behavior. - applyCustomModelInjection() gains an optional workingDir parameter, threaded from session.workingDir (dedicated apply route) / resolvedCasePath (quick-start route) — boot recovery omits it (a dialog already answered once needs no re-seed on the same, persisted isolated directory). Tests added at the pure-builder, IO-wrapper (including merge-preserves- other-fields and corrupt-file-tolerance cases), and existing directory- listing assertions updated for the new settings.json file. Typecheck/ lint/format clean; full suite shows no new regressions (baseline pre-existing Windows-environment failures unchanged, 8 more passing tests than before — the ones added here). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
2c89359d42 |
fix(custom-model): Cancel/Launch-anyway buttons stacked instead of side by side
Neither dialog's footer had a row layout of its own to override, and .btn-toolbar is display:flex (a block-level flex container with no explicit inline-flex), so with no flex row context each button took its own full-width line and the two stacked vertically. The swap-confirm modal already had a .modal-footer rule (flex-end); the context-warning modal had none at all. Both now share one row-layout rule, centred rather than flex-end per feedback. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
962029bb3d |
fix(custom-model): context-warning/swap-confirm modals hidden behind status banner
Both dialogs can appear while the centred llama-swap status banner is still on screen (right after "Claude started — switching to llama-swap…") — the banner's z-index is 10001, .modal's base z-index is only 1000, so the dialog rendered fully behind it. Reported live against the context-window-too-small modal; the swap-confirm modal has the same structural bug for the same reason, so both get the fix. Also: both messages ARE the modal's whole explanatory content, not a one-line caption under a form field, so .form-hint's 0.65rem caption size read as illegibly small — worst on the multi-sentence context-window explanation. Bumped to 0.85rem/1.5 line-height/--text. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
b45a96358e |
feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.
- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
registry entry declares contextLengthVar (currently only claude) and
the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
(40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
quick-start customModel path) check this before the swap-conflict
check and before launching/restarting anything, returning
{requiresContextWarning, modelId, contextLength, minSafeContextTokens}
— skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
_resolveContextWarningConfirm (session-ui.js), wired into both
_quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
(the path Claude actually uses) ahead of the swap-confirmation check.
Explains the fix in-modal: give the model an explicit larger -c/
--ctx-size in llama-swap instead of relying on --fit-ctx, which
optimizes for the biggest model that fits rather than the biggest
context.
Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
993710263d |
fix(custom-model): stop trusting /props's n_ctx, parse the real context size from /running's cmd
Root cause of the context-overflow regression reported live: "API Error: 400 request (36437 tokens) exceeds the available context size (16384 tokens)". Discovery had stored modelContextLengths.qwen3.8-27b-ud-q4_k_xl = 154112, so CLAUDE_CODE_MAX_CONTEXT_TOKENS told Claude Code it had a huge window and it never compacted - but the real llama-swap server was launched with --fit-ctx 16384 (confirmed against /running's own cmd field) and refused the request right at that real limit. /props?model=<id>'s n_ctx (the field discovery read) is confirmed live to be unreliable for a --fit-ctx-launched backend: it reported 154112 for the same model /running says was launched with --fit-ctx 16384 - appears to report the model's theoretical/trained maximum context, not the runtime- configured one. discoverModels() now parses the REAL configured size straight out of llama-swap's own launch command instead (parseCtxFromCmd(), reading /running's cmd field - --fit-ctx first, then the plain llama.cpp -c/ --ctx-size a hand-written command might use), and only falls back to the old /props probe when cmd states no recognizable flag at all. One /running call now covers every loaded model's context length in a single request, same as it already did for the swap-conflict check and the load trigger. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
7bbe408e44 |
feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout (updated every poll, so every second by default) instead of a static "this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1-3 min) on llama-swap - 47s remaining". If the countdown reaches zero and the model still isn't ready, this is now treated as a real failure rather than a "keep waiting" shrug: - The banner turns into a sticky error (_showCenterStatus gains a `type` option - 'error' drops the spinner and adds a close button, since nothing is "in progress" anymore and a sticky message needs a way to dismiss it), naming the llama-swap server's own logs as where to look for detail. - The session that load was for is closed automatically (closeSession) - requested explicitly: a console left open and pointed at a model that never finished loading is worse than no console at all. Both apply paths now thread the new session's id through to _watchLlamaSwapLoading for this (new required 3rd parameter, after endpointId/modelId). _watchLlamaSwapGeneration's existing stale-call guard extends naturally to this: a superseded call's own eventual timeout recognises it no longer owns the banner and neither shows the error nor closes a session that may by then belong to a different, newer launch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
55dae31530 |
feat(custom-model): estimate model load time from its discovered size
Discovery now also parses a GB figure out of an auto-discovered model's own description (llama-swap writes "Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike context length this needs no /props probe (the figure is right there in /v1/models) so it is populated for every model regardless of loaded state. A hand-configured profile's own description has no such figure and correctly gets no entry. The loading banner (_watchLlamaSwapLoading) now looks this up and, when known, shows it plus a rough estimate from a small size->time matrix (_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) - "Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on llama-swap... this can take a while" - and uses that same estimate's own bracket to scale the banner's default give-up timeout for a very large model, instead of a flat 5 minutes for everything. Explicitly labelled as an UNMEASURED, typical-hardware estimate in every relevant comment - this is not benchmarked against any real endpoint's actual storage/GPU, just a reasonable expectation-setter. A model with no discoverable size (a hand-configured profile) gets no size/estimate shown at all, matching the "never a guess" convention modelContextLengths already established. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
0af233c96c |
fix(custom-model): poll llama-swap readiness every 1s, check immediately, extend the cap
Reported: the "Loading..." banner stayed up past 2 minutes even though llama-swap itself had already finished loading the model. Three fixes: 1. pollIntervalMs default 3000ms -> 1000ms (as asked). 2. The loop now checks readiness IMMEDIATELY on entry rather than sleeping a full interval first - a model that's already ready (a fast load, or a re-apply onto one already loaded) shouldn't sit on "Loading..." at all. 3. maxWaitMs default 120000ms (2 min) -> 300000ms (5 min): a large (20GB+) model reading from disk can genuinely take longer than 2 minutes, which would have looked identical to the reported symptom - "still stuck past the point it should have resolved" - except it would have actually flipped to a "still waiting" warning toast at the 2-minute mark rather than staying on "Loading" indefinitely, so this alone doesn't explain what was reported, but is a real, separate improvement worth making. Also fixes a real, separate bug this surfaced while reasoning through the report: _showCenterStatus's banner is ONE shared, reused DOM node. A second call to _watchLlamaSwapLoading (e.g. switching models again before the first switch's loop had finished) would take over that shared banner, but the FIRST loop was still running and would eventually dismiss or overwrite it once ITS OWN deadline or readiness check resolved - clobbering whatever the second, current loop had put there. A generation counter (_watchLlamaSwapGeneration) now lets each call recognise when it no longer owns the banner and stop touching it silently, rather than only the last call to actually start ever safely reading or writing it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
0929694012 |
fix(custom-model): actually trigger the llama-swap load, not just watch for it
Root cause of "it doesn't look like llama-swap is actually switching the model" (confirmed live: no load_model line in llama-swap's own logs after applying a selection). llama-swap has no "switch model" admin endpoint - the ONLY thing that starts a swap is a real inference request naming the model. Every previous fix (the conflict check, the loading banner) assumed a swap would start on its own; nothing ever actually asked llama-swap to load anything until the launched CLI's first real prompt did, which could be much later than "applying the selection" implied. Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest real request that will start a load - POST <baseUrl>/v1/chat/completions, max_tokens: 1, one throwaway message - fire-and-forget (never awaited by the caller; the frontend's own running-status polling is what actually confirms readiness). Wired into both apply paths (the dedicated restart route and the one-shot quick-start route), fired whenever the target model isn't already the one loaded and ready - a broader condition than the existing swapNeeded (which only gates the "this will evict another session's model" confirmation ask and deliberately stays narrow to that). modelSwapInProgress in both routes' responses now reflects this same broader condition too, so the frontend's loading banner actually correlates with a real in-flight load rather than only firing when something else happened to be loaded already. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
01b32ee6cd |
fix(custom-model): move the switching/loading status to a centred banner
The "Claude started - switching to <endpoint>..." and "Loading <model> on <endpoint>... this can take a while" messages lived in the top-right toast corner along with everything else, easy to miss given they can each sit on screen for well over a minute (a real llama-swap model load). Adds _showCenterStatus() (panels-ui.js): a single, reused, screen-centred banner with a spinner, non-blocking (no backdrop, pointer-events: none on the wrapper) so it never gets in the way of using the app while it's up. Both call sites (_runCustomModelEntryViaRestart's switching message, _watchLlamaSwapLoading's loading message) now use it instead of showToast. Every OTHER status in these two flows - the llama-swap conflict warning already moved to its own modal, apply failures, cancellation, and _watchLlamaSwapLoading's own final "ready"/"still waiting" outcome - stays exactly where it was, in the corner. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
2936ba6e3d |
fix(custom-model): replace the native confirm() popup with an in-app modal
The llama-swap "this will unload it for session X" warning used a native browser confirm() popup, which looks out of place next to the rest of the app's own modals. Adds #customModelSwapConfirmModal (index.html) with Cancel/Switch-anyway buttons, styled to match the app. _confirmModelSwap(message) shows it and returns a promise that resolves true/false the same way confirm() would; _resolveModelSwapConfirm(proceed) (wired to both buttons and the backdrop click) settles it. Both llama-swap conflict call sites (_quickStartWithCustomModelConfirm for the one-shot launch path, _runCustomModelEntryViaRestart for Claude's restart path) now await this instead of calling confirm() directly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
83033b4299 |
fix(custom-model): show a status toast during Claude's native-boot-then-restart window
Claude stays on the launch-then-restart path (see runCustomModelEntry's own comment for why), but with nothing on screen during that window, a native boot that briefly talks to the cloud model read as "the endpoint didn't apply" rather than "the switch hasn't happened yet". A sticky "Claude started - switching to <endpoint>..." toast now covers the whole window from the native launch through the apply call, updated in place (never stacked) as the outcome resolves: dismissed on cancel or failure (replaced by the existing cancellation/error toast), handed off to _watchLlamaSwapLoading's own sticky toast when a model swap is in progress, or updated to the existing "Pointed at ... - restarting" message and auto-dismissed after 3s on a plain success. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
f865f74a0f |
feat(custom-model): launch directly on the endpoint, no restart, for 7 of 8 CLIs
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.
POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.
Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.
Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
fbee1b2d82 |
docs(changeset): add changeset for the Run-menu custom-model picker PR
Covers #430's full scope so far: the picker itself, the model-selection dialog, periodic re-discovery, and the session-busy/toast/CLAUDE_CONFIG_DIR/ context-length/llama-swap-conflict fixes found through live validation against a real llama-swap server. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
bcebc81fcd |
feat(custom-model): detect llama-swap model conflicts before switching
Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.
1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
server) via its own GET /running, which plain llama.cpp has no concept of
at all. New GET /api/model-endpoints/:id/running-status route exposes this
read-only, for the frontend's polling loop below.
2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
what llama-swap currently has loaded. If it differs from the requested
model AND another live session's own customModel selection is actively
using that loaded model, the apply is refused with a
{requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
instead of silently switching. A "confirmed: true" field on the retry
skips the check. Switching with nothing else affected proceeds
immediately, no confirmation asked, only ever when there is something to
warn about.
3. The frontend (runCustomModelEntry) shows a native confirm() naming the
affected session(s) and the model they'd lose, matching this codebase's
existing convention for this class of decision (delete case, kill
session, etc.) rather than a new modal. On a successful apply the response
also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
poll shows a sticky "Loading <model>..." toast via the new running-status
route until llama-swap reports the target model ready (bounded at 2
minutes), so a prompt sent mid-swap reads as "loading", never as silence
or an answer from whatever was loaded a moment before.
Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
25f22b9839 |
test(custom-model): update session-custom-model route test for CLAUDE_CONFIG_DIR isolation
Fixes the CI failure on the last two commits: this route test asserted an exact envKeys list for a claude-mode apply that predates the CLAUDE_CONFIG_DIR isolation fix, so it failed on the new CLAUDE_CONFIG_DIR entry it correctly started appending. Updates the expected list and adds assertions for the isolated config dir path and the pre-seeded .claude.json trust-approval file, matching the behavior added in the two prior commits rather than just tolerating it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
97464bfa27 |
fix(custom-model): pre-approve the injected API key in the isolated Claude config dir
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.
Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
0e8b1981af |
fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker: 1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same config directory and warns about it (confirmed cosmetic - the API key wins for actual requests, verified via a real session's own API Usage Billing line). A custom-model claude session now gets an isolated CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty, no files written into it) so there is nothing to conflict with. projects is symlinked (junction on Windows) back into the real config dir so the response viewer, subagent windows and Read My Mind keep working for that session, best-effort. 2. Context-window overflow. Claude Code assumes a large default context window for a model id it doesn't recognise and never compacts, so a custom endpoint's real, much smaller context (verified live: a 400 exceeding a 16384-token llama-swap model with a stock ~33.7K-token system prompt) silently overflows. Discovery now also learns each model's real context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx), but ONLY for a model llama-swap's own /v1/models response already marks status.value === 'loaded' - never an unloaded one, since llama-swap treats ?model= as a routing hint and probing an unloaded model risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery. A server with no status field at all gets no enrichment rather than a guess; a model not probed this round keeps its previously-learned value until it disappears from the list entirely. Stored per model (CustomModelHost.modelContextLengths) and applied via a new contextLengthVar registry field, set to CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude. Both new fields live on the existing env-kind customModelInjection capability shape, declared only on claude's registry entry - every other CLI's injection is unaffected (pinned by test). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
5c25a52f95 |
fix(custom-model): wait for a freshly launched session to go idle before applying
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.
Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.
Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
409a6e65f9 |
fix(custom-model,toast): surface the real apply error, and make error toasts sticky with a close button
Two related fixes, both needed to actually diagnose 'Session started on the native backend — could not apply the custom endpoint' reports from live testing: 1. runCustomModelEntry()'s apply call went through _apiJson(), which unwraps a success body but SWALLOWS a failure response entirely and returns null — discarding the one thing (error, errorCode) that would tell 'endpoint unreachable' apart from 'not a discovered model', 'remote/Docker session', or a dozen other real causes the apply route already reports distinctly. Switched to _api() so the actual response body is read on failure too, and the toast now includes the real message. 2. showToast() defaulted every toast, error or not, to a 3s auto-dismiss with no way to read it again — exactly what made the above generic message impossible to act on even before the fix above. Error toasts now default to sticky (duration: 0, no auto-dismiss) unless a caller opts into a duration, and every toast — sticky or not — gets an explicit close (x) button, since a sticky toast with no way to dismiss it would just accumulate across repeated failures. Tests: custom-model-run-menu-ui.test.ts's two apply tests updated for the _api() switch (their mocks previously stubbed _apiJson, which the apply call no longer goes through), plus a new test pinning that the real server error string reaches the toast on a failure. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
9a9e542a7d |
fix(custom-model): bound the model-picker dialog's height and make its list scroll
The dialog had no max-height at all, so an endpoint with many discovered models grew it past the viewport with nothing to scroll — reported live as both "takes up the full page" and "the list is truncated", which turn out to be the same bug. Gives #customModelPickModal .modal-content the same bounded-height + scrollable-body shape cronModal's .modal-lg already uses (max-height + flex column on the content, overflow-y:auto + flex:1 on the body), scoped by id rather than folded into the shared .modal-sm class three other modals already use for short, fixed content. max-height: min(70vh, 520px) scales with the viewport (a phone gets 70% of its height; a 4K display never gets a needlessly tall dialog) rather than committing to one fixed pixel value that would be wrong at either end. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
5a9ff07f57 |
feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real llama.cpp server: 1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to apply the endpoint's defaultModelId (or the first discovered model) silently. Now, via the new selectCustomModelEntry() (session-ui.js): - exactly one discovered model launches straight away, same as before - two or more open a new #customModelPickModal listing every discovered model; defaultModelId (if set) is marked but never auto-chosen, since the point of asking is letting ONE launch deliberately differ from the saved default, not just confirming it The endpoint is re-fetched at click time rather than trusting anything cached from the dropdown's own render, since the model list can have changed (the sweep below, or a settings-panel edit) since it opened. runCustomModelEntry() itself — the actual launch, routed through run() for the in-flight lock, snapshot-guarded against applying to the wrong session — is unchanged; it now just always receives an explicit model id from one of these two paths instead of computing one itself. 2. Periodic re-discovery. Every saved endpoint's models now refresh automatically every 5 minutes in the background (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way as the Codex plan-usage poll it sits beside — this.cleanup.setInterval, off under testMode), so a model the server starts or stops serving shows up without another manual "Discover" click. The manual POST .../discover-models route and the new refreshAllCustomModelHosts() sweep (custom-model-routes.ts) now share one pure merge step (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId that no longer appears) rather than two copies that could drift. The sweep is best-effort per host — one endpoint being unreachable on a cycle never blocks the others — and re-reads the store before each host's write, keyed by id, so a concurrent edit or delete from the settings panel always wins over a sweep that started before it. Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated file for the sweep (kept separate from custom-model-routes.test.ts because that file's data dir is shared across every test in it — one temp HOME per FILE, not per test — which would make a sweep-touches-every-host assertion meaningless there). test/custom-model-run-menu-ui.test.ts gained a new describe block driving the real picker modal through JSDOM: single-model bypass, multi-model dialog with the default marked-not-chosen, picking a row closes the modal and launches with that exact model, the endpoint re-fetch, and the two "vanished by click time" toast paths. Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md, docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated — the last of these also caught up two sentences that had gone stale after the draft-review fixes landed (the picker routes through run() now, not a raw run*() call). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
60e1bd52f7 |
fix(custom-model): act on the draft review — unparseable onclick, unwrapped envelope, wrong-session apply, missing lock, no tests
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.
Blockers:
1. Every generated inline onclick was unparseable. JSON.stringify's own
double quotes terminated the double-quoted HTML attribute at the first
one, leaving btn.onclick null on every picker entry and every Discover/
Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
argument, the same idiom deleteCase's onclick already uses four lines
away in session-ui.js. This also closes the live-HTML-injection route
through modelId (server-controlled, from the endpoint's own /v1/models
reply): with quoting intact, a `>` inside it can no longer terminate the
<button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
like every other /api route (server.ts's preSerialization hook applies
to arrays too), so Array.isArray(hosts) was always false in production
and the picker/settings panel silently saw nothing. Both call sites now
go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
returns normally without ever changing activeSessionId, so the apply
step used to silently re-point and restart whatever session the user was
already looking at. runCustomModelEntry() now snapshots activeSessionId
before the launch and requires it to have actually changed.
Majors:
4. Routes the launch through run() itself via a temporary _runMode swap
(never persisted — setRunMode() would sync it to the server) instead of
a parallel hardcoded dispatch table, so a custom-model launch now holds
the same _runInFlight lock every other Run click gets. This also
resolves the "hardcoded runners map contradicts the PR's own design"
minor: dispatch is run()'s own, so a CLI whose customModelInjection
recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
parses markup this module generated itself) for exactly the DOM-level
facts the review said needed no Playwright and no tmux: a generated
button's onclick genuinely compiles and fires, a dangerous modelId never
produces a live element, the envelope unwrap works, the session-changed
guard holds, run() actually gets called (proving the in-flight lock
engages), and _runMode is restored afterward. Confirmed against the
pre-fix code first (reproduces btn.onclick === null exactly) so this
isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
render-index-html.test.ts for the other fixes below.
Minors:
- Generated entries now filter through isCliAvailable(), matching
_refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
(applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
and to settings-modal open) instead of always rendering; the endpoint GET
no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
redactApiKey() replaces the field with a computed apiKeySet: boolean, and
a PUT with no apiKey now keeps the stored one server-side
(applyStoredApiKey()) instead of the client resending a value it was
never given. New tests cover both directions (kept vs. replaced) by
observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
(_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
event, since the real role can resolve after settings were first opened)
— endpoint writes were already admin-only server-side, but the button
used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
(CLAUDE.md already records that exact literal turning the settings
preview into a grey slab on light skins), .run-mode-custom-models gets
the same gap: 2px .run-mode-menu's own flex gap only applies one level
up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
</script> (CliEntry.label is user-clis.json-settable, unlike
__codemanCliAvailable's booleans-only payload) via a new exported
escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
docs/api-reference.md; left the "no zh-CN for the new Models-section
group" minor unaddressed only insofar as the wider Models section (task
routing, thinking effort, etc.) has never had zh-CN coverage either —
everything this PR itself introduces (labels, hints, button text, the
Run-menu's "Custom Endpoints" header) IS translated in i18n.js.
Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
fed6582d3e |
fix(test): strip the custom-model Run-menu picker's injected script too
CI on PR #430 failed test/server-index-title.test.ts's byte-identity check: renderIndexHtml now injects a second unconditional <script> before </head> (window.__codemanCustomModelClis, added alongside the existing __codemanCliAvailable one), and the test only knew to strip the older one before comparing the rendered HTML against the raw template. Strip both. Unlike __codemanCliAvailable (an object, historically injected only where something resolved), the new one is a plain array injected unconditionally, possibly empty, so it needs stripping on every machine, not just one with CLIs installed. Verified the two replace() calls compose correctly against the exact strings server.ts actually produces (simulated in isolation; this box has no tmux, so the real WebServer-backed test file cannot run here at all -- same environment gap noted throughout this PR's review). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
98d26e14d9 |
docs(wiki): document Custom Model Endpoints and the Run-menu picker
New docs/wiki/Custom-Model-Endpoints.md (auto-synced to the live GitHub wiki on push to master, per docs/wiki/Contributing.md) covers turning the feature on, adding an endpoint, the Run-menu picker's one-off-run behaviour, the per-harness confidence table, and what it deliberately does not do yet (remote/Docker sessions, live hot-swap). Linked from the sidebar, from Agent-CLIs.md's "Read next" list plus a short pointer section, and from Settings-Reference.md's Models section. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
25fae9ad10 |
feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment: "generate those entries from the saved profiles rather than a fixed duplicate per harness, and put it in a follow-up PR so this one stays the backend... The Run-menu picker is yours if you want it." Adds the frontend surface the backend has been waiting on: - Run menu: a "Custom Endpoints" section lists one entry per (harness that supports customModelInjection, saved endpoint) pair, e.g. "Claude Code (llama.cpp)". The harness list comes from window.__codemanCustomModelClis, injected at page render straight off the CLI registry's own capabilities (never a hardcoded id list in the frontend), so a CLI whose injection recipe lands later appears with no frontend change. Picking an entry runs that harness's own existing run*() function unmodified (case creation, env overrides, everything, forced to a single instance) and then applies the endpoint's default model to the session it creates via the existing POST /api/sessions/:id/custom-model route. Entries are hidden for a remote/docker active case, since that route already refuses both. - Settings: App Settings -> Models gets a "Custom model endpoints" group wiring up the customModelEndpointsEnabled toggle (declared since #393, read by nothing until now) plus CRUD against the existing /api/model-endpoints routes: list, add/edit (inline form), delete, discover models. - Backend: CustomModelHost gains an optional defaultModelId, the model the picker applies with no further choice per endpoint (one generated menu entry per CLI+endpoint pair, not per CLI+endpoint+model). The route refuses a value that isn't one of the endpoint's own discovered models, and a fresh discovery drops a default that no longer appears rather than carrying an invalid one forward. Docs: docs/custom-model-endpoints.md describes the new picker and settings panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the "backend-only" status note and documents the picker's generation mechanism. Tests: four new route tests cover defaultModelId validation, acceptance, and the drop/keep behaviour across a re-discovery; a new render-index-html test pins the __codemanCustomModelClis injection (present, agent CLIs supporting the capability, antigravity and shell excluded) and its solo-window skip. No browser test was added for the Run-menu picker itself or the settings CRUD panel (this box has no tmux, so the live server used by test:browser/test:mobile could not be exercised here) -- worth a Playwright pass before merge, same as any other frontend PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
bd286bf502 |
docs(wiki): catch the manual up to 1.29.0 and add the three run modes it never had
The wiki was written for seven run modes and never received Grok Build, DeepSeek Harness or OMP. They now appear everywhere the others do: the modes table and per-CLI notes, install commands, environment prefixes, the Quick Start table, the requirements rows, the vocabulary, and every "seven modes" count. The 1.27 to 1.29.0 changes land on the pages that own them: attaching a case to an existing container, multi-case adoption and the copy-a-case picker (Docker Cases); file reads over ssh in remote cases and what stays unavailable (Remote SSH Sessions, Working With Files, Security); single-page app routing, frame recovery, localhost links as tabs and the egress guard (Web Tabs); DeepSeek as the one non-Claude mode with real stop/blocked signals and Approvals items, Codex's own work detection, last-response, the model-endpoint routes and refreshed counts (HTTP API, Driving From An Agent, Hooks, Notifications, Keeping Agents Running, Core Concepts); Shift+drag, right-click copy, Auto Copy, the Ctrl+Z guard, font weight, the vertical rail and its activity sort (Keyboard Shortcuts, Input And Voice, The Dashboard, Settings Reference); the 600px phone cutoff, Codex shift arrows and iPhone Duo (Mobile Guide); the Docker Compose route and its update rule (Installation, Running As A Service); four new symptom entries and a "which CLIs" question (Troubleshooting, FAQ). Custom model endpoints are deliberately left to #430, which adds that page and edits Agent CLIs, Settings Reference and the sidebar; these edits stay out of the regions #430, #428 and #376 touch, and all three still merge cleanly on top. Both READMEs: the web-tab menu entry is labelled "Add URL" in the UI, not "Add dashboard". Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
3248f35081 |
chore: version packages (#437)
* chore: version packages * chore: sync the CLAUDE.md version line to 1.29.1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Codeman maintainer <noreply@anthropic.com>codeman@1.29.1 |
||
|
|
5b920cb43d |
feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream and named tabs wrong five ways (every prompt, every write path, a bare Esc eating the next prompt's first character, pasted newlines as Enter, any CSI clearing the draft) and replaced the whole name, which dropped the case from the tab and reset the w<n> counter. This lands the feature with each of those closed: - First prompt means the first: applyAutoName() flips a placeholder to `auto` whether or not the string changed. nameSource is now the tri-state placeholder | auto | manual; the name setter is the only manual path. - Only user-originated input counts: write()/writeViaMux() take SessionWriteOptions.fromUser, set by the browser WS path and POST /input only, so Ralph, respawn, cron, approvals and the trust-dialog keys can never name a tab. A startMode 'shell' CLI never feeds the tracker (a capability, not an id check); the send-key route feeds trackUserInput() because its line feed bypasses the session. - Prefix form `w3-case: title`: parseSessionPrefix() already renders it as the title with the prefix in the tooltip and the next-session counter still matches it. Composed within MAX_SESSION_NAME_LENGTH. - Tracker rules per key: bare Esc resolves at chunk end; mouse/focus reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R taint the draft so Enter submits nothing rather than a fragment; bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space; the draft keeps its head past 8192 code points; an escape past 64 bytes is abandoned. - Title: slash commands by shape (a path is a prompt), `!` escapes refused, first sentence only past 8 code points ("e.g." is not a title), 72 code points on a word boundary. - Synced `autoNameSessions` setting, default OFF (the prompt reaches mux-sessions.json, session:updated and /api/search), App Settings -> Appearance -> Tabs, read fresh per prompt after the eligibility check. Tests: test/session-auto-name.test.ts (tracker, title, composition, ownership, emit gating), the wiring test (once, prefix, setting off, manual protected), test/routes/session-name-routes.test.ts (PUT /name flips to manual and persists). Verified live on an isolated instance: API and browser-typed prompts name the tab, a second prompt does not, shells and renamed tabs are untouched, nameSource survives a restart. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
c4322513d9 | Merge pull request #376 from shenlvkang-collab/feat/auto-session-names-upstream | ||
|
|
018f0c4160 |
docs(readme): catch both READMEs up to 1.29.0 and repair three merge-damaged lines
DeepSeek Harness joins every CLI list it was missing from (tagline, intro, run-mode table, Multi-CLI bullet with its env prefixes, security allowlist, architecture diagram), and the 1.27 to 1.29.0 features get their bullets: custom model endpoints (HTTP API only, with the verified and gapped CLIs named), web tabs, attaching a case to an existing container, remote SSH file access, the plan-usage chip, the sidebar and activity-sorted rail, font weight, skins and entrance animations, Approvals Inbox, Read My Mind, Claude-login voice dictation, Shift+drag select and right-click copy. The agent guide's rule 7 now counts deepseek among the hook-signalling modes and the recipes read answers through last-response first; the API section carries the new routes and current counts; the download cap reads 2 GB instead of the retired 50 MB; the zerolag package test count is the measured 238. The English file had three spots where the OMP merge of 2026-08-18 left two copies of a line joined without a newline (the Docker credentials bullet, rule 7 of the agent guide, the CLI node of the mermaid diagram). All three are single lines again. The Chinese file was further behind: besides the above it had never received the daemon and service block, the Tailscale install option, the Compose paragraph, the Tab Alerts section, the codeman tui section, the agent-skill walkthrough, the Community section or the closing star paragraph. Those are translated in, so both files now share one section structure. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
88e3faa456 | chore: version packages codeman@1.29.0 | ||
|
|
70fc6b32d5 |
docs: record the dup/last input ACK, Shift+drag and right-click copy, and multi-case adopted containers
Three behaviours landed from #375 without their doc entries: the duplicate input ACK now carries `dup:true` and the server's watermark (`docs/reliable-input-delivery.md` still described a bare ACK), Shift+drag and right-click copy in the terminal (the shortcut list did not know them), and one adopted container backing several cases at different in-container directories (the Docker cases paragraph still implied one case per container for adopted containers too). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
9591b973cf |
fix(docker): carry the owned flag on the wire the way master already does
The cherry-picked "copy an existing case" commit declared a second `CaseInfo.docker.owned` and emitted `owned: true|false` on every docker case, while master had meanwhile shipped the same field from the adopted-container work with a narrower wire shape: `owned` is present only when false, absent means owned. Two declarations failed typecheck, and two emit styles on one response would have made the picker's answer depend on which read path filled it. Keep master's shape at both response sites (the case list and the single-case lookup, which lacked the field entirely), fold the picker's reason for the field into the existing doc comment, and repoint the test that pinned "set on exactly two sites" at the surviving form, adding a negative pin so the duplicate style cannot come back. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
025f061383 |
fix(docker): pre-fill the copied case instead of blanking two fields
The previous version cleared the case name and the in-container directory on the
grounds that they must differ. That left a form with three fields mysteriously
filled and two empty, and turned the most common operation — changing
/srv/app/api to /srv/app/web — into retyping a long path.
Both are now pre-filled, with focus on the in-container directory and the caret
at the end, since the tail is what changes. What stops an unmodified submit is no
longer an empty field but a guard: the values applied are recorded, compared at
submit time, and if nothing changed the reason is stated next to the field and
focus moves to it, without sending a request that is certain to be refused.
The server refuses these anyway (a duplicate case name, a twin case on the same
container and directory) and its errors are clear; but making a round trip to be
told "you forgot to edit the field you are looking at" is worse than saying so on
the spot. The guard only applies when a source case was actually selected, so
filling the adopt form from scratch is unaffected.
⚠️ The status text is written into dockerLinkStatus. My first version referenced
an id that does not exist (dockerAdoptStatus), which made the explanation vanish
silently and left only a toast. The test now extracts that id from the code and
looks it up in index.html, pinning that it must really exist.
(cherry picked from commit
|
||
|
|
7a5543da09 |
feat(docker): add "copy an existing case" to the adopt panel
The backend already lets one adopted container back several cases pointing at
different in-container directories, but using it meant retyping the container
name, host and workspace one by one — exactly the friction that leaves a
capability unused. Picking an existing case from a dropdown now carries those
three over, leaving only the two fields that must differ: the case name and the
in-container directory.
Clearing those two is the point of the feature, not a convenience: keeping the
old name is refused by the server as "case already exists", and keeping the old
directory is refused as "a twin case on the same container and directory". Both
errors are clear, but a form pre-filled with values that are guaranteed to be
rejected is a trap. Focus lands on the in-container directory — the thing the
user came here to change.
⚠️ Only adopted containers are listed (docker.owned === false). A Codeman-built
container's lifecycle belongs to its one case — a second case would be torn out
by that case's recreate or delete — so the server refuses it anyway, and listing
it here would only manufacture a baffling error. `owned` may be absent and absent
means owned, so the test is `!== false`, not truthiness.
CaseInfo.docker gains containerWorkdir and owned for this: the former is the
"which directory does this case use" half of the picker, without which the user
cannot tell what to change it to; the latter backs the filter above. ⚠️ Both
places that build a docker CaseInfo (the list endpoint and the single-case query)
must set them — filling in only one makes the picker work or not depending on
which read path was taken, and a test pins "exactly two".
(cherry picked from commit
|
||
|
|
cbb7f635ff |
feat(docker): let one adopted container back several cases in different dirs
Once a container is adopted, it could not be adopted a second time. But a
container usually holds more than one project directory, and opening a case for
another one had no path forward except starting a second container — precisely
what adoption exists to avoid.
The original reason was in a comment: two cases sharing an adopted container
would make one case's teardown race the other's launch on the same tmux server.
That reason does not hold. The in-container tmux session name is
dockerTmuxSessionName(sessionId), i.e. codeman-dkr-<id8>, keyed by SESSION and
not by case, and buildDockerKillCommand tears down exactly that name, so killing
A never touches B — hosting multiple sessions is what a tmux server is for.
The other three routes into an adopted container's lifecycle do not pass through
here either, confirmed one by one: the stop and remove builders throw outright;
recreate refuses `owned === false` before it even resolves the container name;
and orphan reaping filters on `label=codeman.managed=1`, which a user-built
container does not carry — a structural exclusion.
That leaves exactly three cases worth refusing, none of them tmux-related, split
into the pure, unit-tested classifyAdoptContainerConflict:
- owned-case the container belongs to a Codeman-created case, whose lifecycle
Codeman manages: one recreate or delete there would pull the
container out from under the adopting case.
⚠️ `owned` may be absent and absent means owned (cases predate
the field), so the test is `!== false`, not truthiness.
- other-owner already adopted by a different user. Adoption hands out a shell
inside someone else's container.
- duplicate same container, same directory. The second case would behave
identically to the first, so name the existing one rather than
silently minting a twin. A different in-container directory is
the case this change exists to support and passes.
(cherry picked from commit
|
||
|
|
e5684d0bba |
fix(ui): don't create a compositing layer for a hidden full-screen overlay
`backdrop-filter` promotes an element to its own compositing layer. A
position:fixed full-screen layer that is created and then hidden was measured to
leave a stale hit-test region behind in Chrome: the page renders perfectly, but
pointer events across the viewport go nowhere.
The report came from a long-lived tab connected to a remote server, where a
connection blip shows and then hides #offlineOverlay. The symptoms were a
terminal that would not scroll and, at the same time, an unrelated
click-to-expand that also stopped responding, while a freshly opened tab was
fine; a read-only console command (getComputedStyle + elementFromPoint, both of
which force a hit-test recomputation) then cured it. Two unrelated features
dying together and one read-only command fixing both points at hit-testing
itself rather than at either feature.
So the `backdrop-filter` moves onto the actually-visible selector and the layer
is never created while hidden. Only the two persistent overlays change:
offline-overlay (toggled with [hidden]) and file-preview-overlay (toggled with
.visible). path-picker and path-preview are created and removed by JS, leave
nothing behind, and are untouched.
⚠️ This is an evidence-based inference, not a fix verified by reproduction:
reproducing it needs a long-lived page that has been through a connection blip,
which I could not manufacture in a controlled environment. The guard test pins
both halves — no such property while hidden, and a real blur while shown — so a
later cleanup cannot quietly delete the effect.
(cherry picked from commit
|
||
|
|
c7cc8e28d5 |
fix(sse): stop reloading the whole terminal when a reconnect lands on the same session
handleInit() did not distinguish a first load from an SSE reconnect: it always
cleared the terminal caches in _resetAllAppState() and re-ran selectSession() for
the session that was already on screen. Every reconnect therefore refetched up to
1 MiB of buffer and reset+rewrote xterm. On a link that drops a connection about
once a minute (measured at ~57s intervals against a healthy server) that reads as
the page refreshing itself and throwing away your reading position.
A reconnect that lands back on the still-open session now keeps the terminal
caches and activeSessionId and resyncs through _onSessionNeedsRefresh(). That
path still reloads the buffer, so output produced during the outage is not lost,
but it preserves distance-from-bottom — the same rule #259 established for a
refresh the server triggered rather than the user. The WS is reconnected
explicitly when it is not already on that session, since skipping selectSession()
skips its _connectWs() call.
First load (gen === 1) takes exactly the path it took before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rv24Pk4qzrsDYdVyDyJQmT
(cherry picked from commit
|
||
|
|
01da577053 |
fix(input): recover when the seq counter falls behind the server watermark
Browser input is delivered exactly once by (clientId, seq). The server records a
watermark per clientId and discards anything not above it as a duplicate — but
acknowledged it with an ACK indistinguishable from "applied". The client then
dropped the record from its queue, the UI looked perfectly normal, and the
terminal received nothing at all.
The counter is persisted to localStorage through a debounced write. Kill the page
between "sent" and "persisted" and the restored counter is below the server's
watermark, after which every keystroke lands under it, is discarded, and is
ACKed. Reloading does not help: the clientId is restored from localStorage
alongside that stale counter. Measured on a real session — typing into the same
session from a fresh browser (new clientId, no watermark on the server) worked
perfectly, which is what localised the fault to client state.
Three changes:
- on rejection the server replies {"t":"ia",seq,"dup":true,"last":<watermark>}.
It still ACKs, so the client can drop the record from its queue, but it now
says the input was not applied and supplies the number needed to climb out.
- on `dup` the client lifts its counter above the watermark and re-queues.
⚠️ Only records whose FIRST delivery is being retried are re-sent: a retry
judged duplicate means the mechanism is working (the original did arrive), and
re-sending would type the same text twice.
- the counter is now persisted synchronously. The queue payload can stay
debounced, but the counter is the thing that has to survive a crash, and
leaving it on the lossiest path cancels the only guarantee there is.
⚠️ Reading the watermark is defensive: the session arrives through a structured
port, and a port missing that method must not take the whole input path down —
a throw inside the handler means the ACK is never sent and the record is stuck in
the client queue forever, which is worse than the ambiguity being fixed. A mock
port's test timeout is what exposed this.
(cherry picked from commit
|
||
|
|
631386d3f7 |
fix(cjk): forward Ctrl/Alt-modified navigation keys to the CLI
claude advertises "Jump to bottom (ctrl+End)", so that chord has to actually
reach it. But PASSTHROUGH_KEYS carried only the bare forms (End -> \x1b[F) and
CTRL_KEYS held just six letters (c/d/l/z/a/e), which cannot express End. Ctrl+End
therefore failed in both directions:
- with an empty composer it went out as a bare \x1b[F, the modifier silently
dropped, so the CLI received a plain End;
- with text in the composer the forwarding branch requires empty, so nothing was
forwarded and the browser default applied — the caret jumped to the end of the
draft, which is the "the shortcut now edits my input box" the user saw.
Encode them as CSI 1;<mod><final> instead, and forward Ctrl/Alt-modified
navigation keys whether or not the composer is empty: they are commands for the
CLI, and the composer has no editing semantics for them worth preserving (bare
Home/End still use the old table and edit locally).
⚠️ Bare Shift is deliberately excluded: Shift+arrow selects text in the composer,
a real editing gesture that must stay local. Shift held together with Ctrl/Alt is
still encoded into the modifier mask.
(cherry picked from commit
|
||
|
|
b3a6ba2eb6 |
feat(terminal): make Shift+drag select, and right-click copy the selection
In a native terminal running a TUI with mouse tracking on (claude, codex), Shift
is the "let me select text" modifier: it bypasses the application's mouse
reporting so the emulator selects locally. Users bring that habit here, where it
did nothing — measured, `hasSelection` was already false during a Shift+drag and
no clearSelection call ran at all, because there was never a selection to clear.
The mismatch is that the two Shifts mean different things. xterm reads Shift as
"force selection", but that path is only taken when the application really has
mouse tracking on. The server strips the mouse DECSETs for claude/codex/gemini
(isAltScreenStripMode), so xterm's mouseTrackingMode is permanently `none`, that
branch is unreachable, and Shift instead lands in _onIncrementalClick — which
EXTENDS an existing selection. Extension is a no-op while selectionStart is
empty, so the drag had no anchor.
So plant the anchor xterm is missing. The listener sits on the capture phase of
the `.xterm` root, an ancestor of the `.xterm-screen` that SelectionService binds
to, and therefore runs before xterm's own mousedown; xterm then extends from our
anchor and the drag behaves like any other. Length is 0 so a Shift+click without
a drag does not select a stray character. An existing selection is left alone —
that is a genuine extend gesture, and xterm handles it correctly.
Right-click copies the selection (the mintty/PuTTY convention), completing the
gesture: until now there was nowhere for a finished selection to go. With no
selection the native menu is not hijacked — taking it away while offering
nothing in return is a pure loss.
(cherry picked from commit
|
||
|
|
897a63183f |
chore(typecheck): include the local-LLM harness smoke script
scripts/test-local-llm-harnesses.ts (#393) sits outside tsconfig.json's include, so nothing type-checked it. config/tsconfig.scripts.json pulls it in; npm run typecheck now runs both projects, the way the pr-bot config used to be chained before the bot moved out of the repo. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
942bf37e48 |
fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a custom OpenAI-compatible endpoint by injecting env vars or a config file and restarting the CLI in place. Review of the apply path found four things, two of them destructive. This lands all four plus the smaller items from the same review. 1. Clearing a selection did not clear it. The injected vars reach the CLI via `tmux setenv`, which persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so deleting the keys from the session's envOverrides relaunched the CLI still pointed at the old endpoint, and for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just been deleted. `Session.setCustomModel()` now reports the removed keys, queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys` carries them into `applyEnvOverrides()`, which `setenv -u`s them before re-applying the live overrides, on the same path that already unsets the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket that `setenv -u HOME` hands the next respawn the global HOME back. 2. Applying a model to a local claude session killed the pane. The relaunch was `claude --session-id <id>` and Claude refuses an id that already has a transcript, and unlike the dead-pane respawn this one kills a working pane first. `restartCli()` now pins the live conversation id as the resume id for that respawn when the CLI's launch declares a `fallback` chain, which renders the same `--resume <id> || --session-id <id>` shape the docker and remote pane commands use. Gated on the registry shape, not the CLI id: an entry whose resume id is minted by the CLI itself never declares that chain. 3. pi, omp and grok wrote their config file and then launched without the `--model` that selects it, so the file was ignored. The registry entry now declares `customModelInjection.launchModel` (`custom/{modelId}` for pi and omp, grok's `[model.codeman-custom]` block name), the builder renders it, and `_withCustomModelLaunchModel()` applies it onto the respawn options through `legacyConfigField`, leaving the stored <Mode>Config untouched so a clear falls back to the user's own model. A model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. 4. Remote (SSH) and Docker sessions reported `restarted: true` and changed nothing: their `restartCli()` reattaches the durable tmux rather than relaunching the agent, and the env lands on the local pane. Both are refused with a 400 until those paths are plumbed. Smaller items from the same review: - The selection survives a Codeman restart as the disk-only `__customModel` bookkeeping (endpoint, model, injected key NAMES, config dir, launch model; never the values, which carry the API key). Recovery re-derives the values from the endpoint store through the same apply path the route uses and keeps the bookkeeping even when the endpoint is gone, so a later clear still has keys to unset. - Discovery goes through `webviewFetch()`, so the RESOLVED address is judged by the same egress guard the web-tab proxy uses, and `baseUrl` reuses `webviewUrlSchema` (http(s) only, no embedded credentials, link-local and cloud-metadata addresses refused). undici's `fetch failed` wrapper is unwrapped so the user sees the ECONNREFUSED underneath. - `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session config dir 0700/0600 (pi and omp embed the key literally), and that dir is removed with the session. - `PR.md` is gone from the repo root and the design doc moved to `docs/custom-model-endpoints-plan.md` with the LAN address and the personal name scrubbed; every reference follows. The guide's `authStyle` text matches the shipped schema (`bearer | api-key`, default `bearer`) and says that `customModelEndpointsEnabled` is read by nothing until the picker lands. - `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts` (four real type errors fixed). It is not yet wired into `npm run typecheck` because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json` there is the one-line follow-up. Tests: `test/session-custom-model-restart.test.ts` drives a real Session and fails on the unfixed code for items 1 to 3; the route suite covers item 4 and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets run before the overrides and that a shell-metachar key never reaches tmux. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1e42cb4e2d |
Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses) |
||
|
|
e49c48145b |
fix(files): fail closed on remote symlinks, guard PUT for remote cases, bound ssh fan-out
Follow-up to #421 (remote-case file reads over ssh), addressing the review. Symlink escape on a host without `readlink -f` (blocker). The probe's portable fallback canonicalized only the directory chain and returned the final component unresolved, so on macOS < 12.3 `ws/notes.txt -> ~/.ssh/id_rsa` came back as `.../ws/notes.txt` (with the target's size), passed every containment and blocklist check that runs on `realPath`, and `cat` followed the link. The fallback now walks the directory chain with `cd -P`/`pwd -P` and follows the LAST component with plain `readlink` for a bounded number of hops, and anything it cannot fully resolve (a loop, a readlink failure, the hop cap) is reported with an `x` marker that parses as null, i.e. 404. It never returns the unresolved string. Measured on a real /bin/sh with `readlink -f` shadowed: the pre-fix script reports `/ws/notes.txt`, the fixed one `/secret/id_rsa`; both branches (native and fallback) now agree. `PUT /api/sessions/:id/file-content` never had the remote guard the PR described. It sits ahead of `validateSessionFilePath`, which resolves against the LOCAL filesystem, because with a same-named directory on the Codeman host (an sshfs mount of the remote tree, the documented stop-gap) the write landed on the local twin while the viewer believed it edited the remote file. ssh fan-out is bounded. `src/remote-ssh-limiter.ts` is a document-conversion-limiter-shaped semaphore (default 4, env `CODEMAN_MAX_REMOTE_FILE_SSH`) around every probe and buffered read; the attachment-history list resolves its whole history in ONE batched probe (`probeRemoteAttachmentHistory`, threaded into `registerExternalAttachment({remoteProbes})` so the guards run unchanged) instead of one handshake per entry; and probes chunk at 40 paths because the whole script is one argv string. Terminal output in a remote session is written on the remote host, so a prompt-injected agent printing hundreds of `codeman://attach` links forked one ssh per link, each holding a 20 s timeout, and a 100-entry history re-listed on every attachment:detected tripped OpenSSH's default MaxStartups. Streams are deliberately not counted (one per browser request, held for a whole playback, and gated behind a counted probe anyway). Smaller items from the same review: probe records are NUL-terminated and index-keyed after a leading NUL (a newline in a filename can no longer shift the alignment, and the banner is fenced off without last-N-lines guessing); size comes from `stat -c %s || stat -f %z`; the three IO functions refuse under VITEST instead of opening a connection; an unreachable host now reads as unknown (missing: false) for detected AND external history entries, where external used to fold its 502 into missing; a client that aborted during the guard probe has its body's ssh child reaped (`reply.raw.destroyed` is checked before the close listener is attached); `describeExecError` never returns Node's `Command failed: <ssh line>` message, which carried the identity path and the probe script into a 502 body; and the docs note that `isSensitivePath`'s three home-anchored entries resolve against the Codeman host's home, not the remote one. Tests: the probe script runs on a real /bin/sh with a `readlink` shim that rejects `-f` (the escape, a relative chain through a symlinked directory, a loop, a newline filename, banner chatter that itself looks like a record), the limiter's cap and FIFO order, and route tests for the PUT guard (local twin untouched, no connection), the single batched history probe, the unreachable-host alignment and the aborted-client reap. All four route tests fail against the pre-fix file-routes.ts. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
792a251e35 |
Merge pull request #421 from Randalix/fix/remote-file-access
fix(files): read remote-case previews, downloads and attachments over ssh |