mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-05 15:09:42 +02:00
db9729e1fc548154538aa7cdcff38c78535ef037
126
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
db9729e1fc |
feat(custom-model): remove loading-banner countdown, add manual Cancel
Replaces the size-scaled expected-time estimate + matching auto-timeout with a generic hardware/model-size disclaimer and a user-driven Cancel button, per explicit request. Real load time depends on hardware this feature has no way to know (VRAM, storage speed, GPU contention), so the old estimate/timeout was a guess dressed up as a fact — worse, one that could kill a genuinely slow load partway through on slower hardware. - _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline entirely — polls indefinitely until ready or cancelled, no automatic give-up. Message is now "Loading <model> (<size>) on <endpoint> — this can take a while depending on your hardware and the model size.", with the real llama.cpp log line still on its own second line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/ _formatRemaining (dead code once the countdown is gone) — _lookupModelSizeGB is kept, the GB figure still shows. - _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real "Cancel" button (distinct from the error-type "×" close button, since Cancel has a real consequence) that calls it on click. Caller owns what cancelling actually means, same split as the swap-confirm modal's promise-resolving buttons. - Cancelling dismisses the banner, shows an info toast (not an error — this was deliberate), and closes the session, mirroring what the old timeout used to do automatically but now on the user's own call. - New .center-status-cancel CSS (bordered pill button, distinct from the plain "×" close glyph). Test changes: removed the now-invalid timeout-auto-close/estimate tests, added cancel-flow tests (dismiss/toast-type/session-close, never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests for the new Cancel button (bootAppWithRealCenterStatus, evaluating panels-ui.js instead of stubbing _showCenterStatus, since this button is worth verifying for real rather than just through the stub every other test in the file uses). Typecheck/lint/frontend-syntax clean; full suite shows no new regressions. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
2d3fc65758 |
feat(custom-model): show real-time llama.cpp backend status in the loading banner
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.
- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
(custom-model-routes.ts): one persistent GET /api/events (SSE)
connection held open per endpoint, parsing logData frames and
keeping the latest source:"upstream" (backend llama-server) line —
filtering out llama-swap's own source:"proxy" request-access lines.
Idle-closed after 30s of no polling, same 20s sweep as the existing
swap-displacement check.
- running-status route now returns logLine alongside the existing
isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
("llama.cpp: <line>", bootlog timestamp/level/component prefix
stripped for display) that stays on the last real thing llama.cpp
said rather than clearing to blank between polls.
⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).
12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
b45a96358e |
feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.
- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
registry entry declares contextLengthVar (currently only claude) and
the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
(40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
quick-start customModel path) check this before the swap-conflict
check and before launching/restarting anything, returning
{requiresContextWarning, modelId, contextLength, minSafeContextTokens}
— skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
_resolveContextWarningConfirm (session-ui.js), wired into both
_quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
(the path Claude actually uses) ahead of the swap-confirmation check.
Explains the fix in-modal: give the model an explicit larger -c/
--ctx-size in llama-swap instead of relying on --fit-ctx, which
optimizes for the biggest model that fits rather than the biggest
context.
Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
7bbe408e44 |
feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout (updated every poll, so every second by default) instead of a static "this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1-3 min) on llama-swap - 47s remaining". If the countdown reaches zero and the model still isn't ready, this is now treated as a real failure rather than a "keep waiting" shrug: - The banner turns into a sticky error (_showCenterStatus gains a `type` option - 'error' drops the spinner and adds a close button, since nothing is "in progress" anymore and a sticky message needs a way to dismiss it), naming the llama-swap server's own logs as where to look for detail. - The session that load was for is closed automatically (closeSession) - requested explicitly: a console left open and pointed at a model that never finished loading is worse than no console at all. Both apply paths now thread the new session's id through to _watchLlamaSwapLoading for this (new required 3rd parameter, after endpointId/modelId). _watchLlamaSwapGeneration's existing stale-call guard extends naturally to this: a superseded call's own eventual timeout recognises it no longer owns the banner and neither shows the error nor closes a session that may by then belong to a different, newer launch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
55dae31530 |
feat(custom-model): estimate model load time from its discovered size
Discovery now also parses a GB figure out of an auto-discovered model's own description (llama-swap writes "Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike context length this needs no /props probe (the figure is right there in /v1/models) so it is populated for every model regardless of loaded state. A hand-configured profile's own description has no such figure and correctly gets no entry. The loading banner (_watchLlamaSwapLoading) now looks this up and, when known, shows it plus a rough estimate from a small size->time matrix (_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) - "Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on llama-swap... this can take a while" - and uses that same estimate's own bracket to scale the banner's default give-up timeout for a very large model, instead of a flat 5 minutes for everything. Explicitly labelled as an UNMEASURED, typical-hardware estimate in every relevant comment - this is not benchmarked against any real endpoint's actual storage/GPU, just a reasonable expectation-setter. A model with no discoverable size (a hand-configured profile) gets no size/estimate shown at all, matching the "never a guess" convention modelContextLengths already established. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
0af233c96c |
fix(custom-model): poll llama-swap readiness every 1s, check immediately, extend the cap
Reported: the "Loading..." banner stayed up past 2 minutes even though llama-swap itself had already finished loading the model. Three fixes: 1. pollIntervalMs default 3000ms -> 1000ms (as asked). 2. The loop now checks readiness IMMEDIATELY on entry rather than sleeping a full interval first - a model that's already ready (a fast load, or a re-apply onto one already loaded) shouldn't sit on "Loading..." at all. 3. maxWaitMs default 120000ms (2 min) -> 300000ms (5 min): a large (20GB+) model reading from disk can genuinely take longer than 2 minutes, which would have looked identical to the reported symptom - "still stuck past the point it should have resolved" - except it would have actually flipped to a "still waiting" warning toast at the 2-minute mark rather than staying on "Loading" indefinitely, so this alone doesn't explain what was reported, but is a real, separate improvement worth making. Also fixes a real, separate bug this surfaced while reasoning through the report: _showCenterStatus's banner is ONE shared, reused DOM node. A second call to _watchLlamaSwapLoading (e.g. switching models again before the first switch's loop had finished) would take over that shared banner, but the FIRST loop was still running and would eventually dismiss or overwrite it once ITS OWN deadline or readiness check resolved - clobbering whatever the second, current loop had put there. A generation counter (_watchLlamaSwapGeneration) now lets each call recognise when it no longer owns the banner and stop touching it silently, rather than only the last call to actually start ever safely reading or writing it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
01b32ee6cd |
fix(custom-model): move the switching/loading status to a centred banner
The "Claude started - switching to <endpoint>..." and "Loading <model> on <endpoint>... this can take a while" messages lived in the top-right toast corner along with everything else, easy to miss given they can each sit on screen for well over a minute (a real llama-swap model load). Adds _showCenterStatus() (panels-ui.js): a single, reused, screen-centred banner with a spinner, non-blocking (no backdrop, pointer-events: none on the wrapper) so it never gets in the way of using the app while it's up. Both call sites (_runCustomModelEntryViaRestart's switching message, _watchLlamaSwapLoading's loading message) now use it instead of showToast. Every OTHER status in these two flows - the llama-swap conflict warning already moved to its own modal, apply failures, cancellation, and _watchLlamaSwapLoading's own final "ready"/"still waiting" outcome - stays exactly where it was, in the corner. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
2936ba6e3d |
fix(custom-model): replace the native confirm() popup with an in-app modal
The llama-swap "this will unload it for session X" warning used a native browser confirm() popup, which looks out of place next to the rest of the app's own modals. Adds #customModelSwapConfirmModal (index.html) with Cancel/Switch-anyway buttons, styled to match the app. _confirmModelSwap(message) shows it and returns a promise that resolves true/false the same way confirm() would; _resolveModelSwapConfirm(proceed) (wired to both buttons and the backdrop click) settles it. Both llama-swap conflict call sites (_quickStartWithCustomModelConfirm for the one-shot launch path, _runCustomModelEntryViaRestart for Claude's restart path) now await this instead of calling confirm() directly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
83033b4299 |
fix(custom-model): show a status toast during Claude's native-boot-then-restart window
Claude stays on the launch-then-restart path (see runCustomModelEntry's own comment for why), but with nothing on screen during that window, a native boot that briefly talks to the cloud model read as "the endpoint didn't apply" rather than "the switch hasn't happened yet". A sticky "Claude started - switching to <endpoint>..." toast now covers the whole window from the native launch through the apply call, updated in place (never stacked) as the outcome resolves: dismissed on cancel or failure (replaced by the existing cancellation/error toast), handed off to _watchLlamaSwapLoading's own sticky toast when a model swap is in progress, or updated to the existing "Pointed at ... - restarting" message and auto-dismissed after 3s on a plain success. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
f865f74a0f |
feat(custom-model): launch directly on the endpoint, no restart, for 7 of 8 CLIs
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.
POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.
Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.
Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
bcebc81fcd |
feat(custom-model): detect llama-swap model conflicts before switching
Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.
1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
server) via its own GET /running, which plain llama.cpp has no concept of
at all. New GET /api/model-endpoints/:id/running-status route exposes this
read-only, for the frontend's polling loop below.
2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
what llama-swap currently has loaded. If it differs from the requested
model AND another live session's own customModel selection is actively
using that loaded model, the apply is refused with a
{requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
instead of silently switching. A "confirmed: true" field on the retry
skips the check. Switching with nothing else affected proceeds
immediately, no confirmation asked, only ever when there is something to
warn about.
3. The frontend (runCustomModelEntry) shows a native confirm() naming the
affected session(s) and the model they'd lose, matching this codebase's
existing convention for this class of decision (delete case, kill
session, etc.) rather than a new modal. On a successful apply the response
also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
poll shows a sticky "Loading <model>..." toast via the new running-status
route until llama-swap reports the target model ready (bounded at 2
minutes), so a prompt sent mid-swap reads as "loading", never as silence
or an answer from whatever was loaded a moment before.
Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
5c25a52f95 |
fix(custom-model): wait for a freshly launched session to go idle before applying
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.
Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.
Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
409a6e65f9 |
fix(custom-model,toast): surface the real apply error, and make error toasts sticky with a close button
Two related fixes, both needed to actually diagnose 'Session started on the native backend — could not apply the custom endpoint' reports from live testing: 1. runCustomModelEntry()'s apply call went through _apiJson(), which unwraps a success body but SWALLOWS a failure response entirely and returns null — discarding the one thing (error, errorCode) that would tell 'endpoint unreachable' apart from 'not a discovered model', 'remote/Docker session', or a dozen other real causes the apply route already reports distinctly. Switched to _api() so the actual response body is read on failure too, and the toast now includes the real message. 2. showToast() defaulted every toast, error or not, to a 3s auto-dismiss with no way to read it again — exactly what made the above generic message impossible to act on even before the fix above. Error toasts now default to sticky (duration: 0, no auto-dismiss) unless a caller opts into a duration, and every toast — sticky or not — gets an explicit close (x) button, since a sticky toast with no way to dismiss it would just accumulate across repeated failures. Tests: custom-model-run-menu-ui.test.ts's two apply tests updated for the _api() switch (their mocks previously stubbed _apiJson, which the apply call no longer goes through), plus a new test pinning that the real server error string reaches the toast on a failure. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
5a9ff07f57 |
feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real llama.cpp server: 1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to apply the endpoint's defaultModelId (or the first discovered model) silently. Now, via the new selectCustomModelEntry() (session-ui.js): - exactly one discovered model launches straight away, same as before - two or more open a new #customModelPickModal listing every discovered model; defaultModelId (if set) is marked but never auto-chosen, since the point of asking is letting ONE launch deliberately differ from the saved default, not just confirming it The endpoint is re-fetched at click time rather than trusting anything cached from the dropdown's own render, since the model list can have changed (the sweep below, or a settings-panel edit) since it opened. runCustomModelEntry() itself — the actual launch, routed through run() for the in-flight lock, snapshot-guarded against applying to the wrong session — is unchanged; it now just always receives an explicit model id from one of these two paths instead of computing one itself. 2. Periodic re-discovery. Every saved endpoint's models now refresh automatically every 5 minutes in the background (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way as the Codex plan-usage poll it sits beside — this.cleanup.setInterval, off under testMode), so a model the server starts or stops serving shows up without another manual "Discover" click. The manual POST .../discover-models route and the new refreshAllCustomModelHosts() sweep (custom-model-routes.ts) now share one pure merge step (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId that no longer appears) rather than two copies that could drift. The sweep is best-effort per host — one endpoint being unreachable on a cycle never blocks the others — and re-reads the store before each host's write, keyed by id, so a concurrent edit or delete from the settings panel always wins over a sweep that started before it. Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated file for the sweep (kept separate from custom-model-routes.test.ts because that file's data dir is shared across every test in it — one temp HOME per FILE, not per test — which would make a sweep-touches-every-host assertion meaningless there). test/custom-model-run-menu-ui.test.ts gained a new describe block driving the real picker modal through JSDOM: single-model bypass, multi-model dialog with the default marked-not-chosen, picking a row closes the modal and launches with that exact model, the endpoint re-fetch, and the two "vanished by click time" toast paths. Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md, docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated — the last of these also caught up two sentences that had gone stale after the draft-review fixes landed (the picker routes through run() now, not a raw run*() call). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
60e1bd52f7 |
fix(custom-model): act on the draft review — unparseable onclick, unwrapped envelope, wrong-session apply, missing lock, no tests
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.
Blockers:
1. Every generated inline onclick was unparseable. JSON.stringify's own
double quotes terminated the double-quoted HTML attribute at the first
one, leaving btn.onclick null on every picker entry and every Discover/
Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
argument, the same idiom deleteCase's onclick already uses four lines
away in session-ui.js. This also closes the live-HTML-injection route
through modelId (server-controlled, from the endpoint's own /v1/models
reply): with quoting intact, a `>` inside it can no longer terminate the
<button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
like every other /api route (server.ts's preSerialization hook applies
to arrays too), so Array.isArray(hosts) was always false in production
and the picker/settings panel silently saw nothing. Both call sites now
go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
returns normally without ever changing activeSessionId, so the apply
step used to silently re-point and restart whatever session the user was
already looking at. runCustomModelEntry() now snapshots activeSessionId
before the launch and requires it to have actually changed.
Majors:
4. Routes the launch through run() itself via a temporary _runMode swap
(never persisted — setRunMode() would sync it to the server) instead of
a parallel hardcoded dispatch table, so a custom-model launch now holds
the same _runInFlight lock every other Run click gets. This also
resolves the "hardcoded runners map contradicts the PR's own design"
minor: dispatch is run()'s own, so a CLI whose customModelInjection
recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
parses markup this module generated itself) for exactly the DOM-level
facts the review said needed no Playwright and no tmux: a generated
button's onclick genuinely compiles and fires, a dangerous modelId never
produces a live element, the envelope unwrap works, the session-changed
guard holds, run() actually gets called (proving the in-flight lock
engages), and _runMode is restored afterward. Confirmed against the
pre-fix code first (reproduces btn.onclick === null exactly) so this
isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
render-index-html.test.ts for the other fixes below.
Minors:
- Generated entries now filter through isCliAvailable(), matching
_refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
(applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
and to settings-modal open) instead of always rendering; the endpoint GET
no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
redactApiKey() replaces the field with a computed apiKeySet: boolean, and
a PUT with no apiKey now keeps the stored one server-side
(applyStoredApiKey()) instead of the client resending a value it was
never given. New tests cover both directions (kept vs. replaced) by
observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
(_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
event, since the real role can resolve after settings were first opened)
— endpoint writes were already admin-only server-side, but the button
used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
(CLAUDE.md already records that exact literal turning the settings
preview into a grey slab on light skins), .run-mode-custom-models gets
the same gap: 2px .run-mode-menu's own flex gap only applies one level
up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
</script> (CliEntry.label is user-clis.json-settable, unlike
__codemanCliAvailable's booleans-only payload) via a new exported
escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
docs/api-reference.md; left the "no zh-CN for the new Models-section
group" minor unaddressed only insofar as the wider Models section (task
routing, thinking effort, etc.) has never had zh-CN coverage either —
everything this PR itself introduces (labels, hints, button text, the
Run-menu's "Custom Endpoints" header) IS translated in i18n.js.
Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
25fae9ad10 |
feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment: "generate those entries from the saved profiles rather than a fixed duplicate per harness, and put it in a follow-up PR so this one stays the backend... The Run-menu picker is yours if you want it." Adds the frontend surface the backend has been waiting on: - Run menu: a "Custom Endpoints" section lists one entry per (harness that supports customModelInjection, saved endpoint) pair, e.g. "Claude Code (llama.cpp)". The harness list comes from window.__codemanCustomModelClis, injected at page render straight off the CLI registry's own capabilities (never a hardcoded id list in the frontend), so a CLI whose injection recipe lands later appears with no frontend change. Picking an entry runs that harness's own existing run*() function unmodified (case creation, env overrides, everything, forced to a single instance) and then applies the endpoint's default model to the session it creates via the existing POST /api/sessions/:id/custom-model route. Entries are hidden for a remote/docker active case, since that route already refuses both. - Settings: App Settings -> Models gets a "Custom model endpoints" group wiring up the customModelEndpointsEnabled toggle (declared since #393, read by nothing until now) plus CRUD against the existing /api/model-endpoints routes: list, add/edit (inline form), delete, discover models. - Backend: CustomModelHost gains an optional defaultModelId, the model the picker applies with no further choice per endpoint (one generated menu entry per CLI+endpoint pair, not per CLI+endpoint+model). The route refuses a value that isn't one of the endpoint's own discovered models, and a fresh discovery drops a default that no longer appears rather than carrying an invalid one forward. Docs: docs/custom-model-endpoints.md describes the new picker and settings panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the "backend-only" status note and documents the picker's generation mechanism. Tests: four new route tests cover defaultModelId validation, acceptance, and the drop/keep behaviour across a re-discovery; a new render-index-html test pins the __codemanCustomModelClis injection (present, agent CLIs supporting the capability, antigravity and shell excluded) and its solo-window skip. No browser test was added for the Run-menu picker itself or the settings CRUD panel (this box has no tmux, so the live server used by test:browser/test:mobile could not be exercised here) -- worth a Playwright pass before merge, same as any other frontend PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
025f061383 |
fix(docker): pre-fill the copied case instead of blanking two fields
The previous version cleared the case name and the in-container directory on the
grounds that they must differ. That left a form with three fields mysteriously
filled and two empty, and turned the most common operation — changing
/srv/app/api to /srv/app/web — into retyping a long path.
Both are now pre-filled, with focus on the in-container directory and the caret
at the end, since the tail is what changes. What stops an unmodified submit is no
longer an empty field but a guard: the values applied are recorded, compared at
submit time, and if nothing changed the reason is stated next to the field and
focus moves to it, without sending a request that is certain to be refused.
The server refuses these anyway (a duplicate case name, a twin case on the same
container and directory) and its errors are clear; but making a round trip to be
told "you forgot to edit the field you are looking at" is worse than saying so on
the spot. The guard only applies when a source case was actually selected, so
filling the adopt form from scratch is unaffected.
⚠️ The status text is written into dockerLinkStatus. My first version referenced
an id that does not exist (dockerAdoptStatus), which made the explanation vanish
silently and left only a toast. The test now extracts that id from the code and
looks it up in index.html, pinning that it must really exist.
(cherry picked from commit
|
||
|
|
7a5543da09 |
feat(docker): add "copy an existing case" to the adopt panel
The backend already lets one adopted container back several cases pointing at
different in-container directories, but using it meant retyping the container
name, host and workspace one by one — exactly the friction that leaves a
capability unused. Picking an existing case from a dropdown now carries those
three over, leaving only the two fields that must differ: the case name and the
in-container directory.
Clearing those two is the point of the feature, not a convenience: keeping the
old name is refused by the server as "case already exists", and keeping the old
directory is refused as "a twin case on the same container and directory". Both
errors are clear, but a form pre-filled with values that are guaranteed to be
rejected is a trap. Focus lands on the in-container directory — the thing the
user came here to change.
⚠️ Only adopted containers are listed (docker.owned === false). A Codeman-built
container's lifecycle belongs to its one case — a second case would be torn out
by that case's recreate or delete — so the server refuses it anyway, and listing
it here would only manufacture a baffling error. `owned` may be absent and absent
means owned, so the test is `!== false`, not truthiness.
CaseInfo.docker gains containerWorkdir and owned for this: the former is the
"which directory does this case use" half of the picker, without which the user
cannot tell what to change it to; the latter backs the filter above. ⚠️ Both
places that build a docker CaseInfo (the list endpoint and the single-case query)
must set them — filling in only one makes the picker work or not depending on
which read path was taken, and a test pins "exactly two".
(cherry picked from commit
|
||
|
|
b2b2c767ea |
Merge pull request #361 from timkjr/fix/statusline-injection-opt-out
fix(statusline): inject plan-usage telemetry via ephemeral CLI flag, never disk |
||
|
|
c087d0ae4d |
fix(mobile): raise the phone breakpoint from 430px to 600px
The phone tier stopped at innerWidth < 430 and @media (max-width: 430px), so every current large phone landed in the tablet layout: the 430pt iPhone 14 Pro Max, 15 Plus, 15 Pro Max and 16 Plus, the 440pt iPhone 16 Pro Max and 17 Pro Max, Pixel 6 Pro, 7 Pro and OnePlus 12 Pro, the 448pt Pixel 8 Pro and 9 Pro XL, and the Galaxy Z Fold 5 cover screen at 460. On those devices the header icon row replaced the session pill, the toolbar kept the desktop Run Shell button instead of Enter and the mic, the keyboard accessory bar could never become visible because its .visible rule lives inside the phone block, and the toolbar jumped to the top of the page when the keyboard opened. The new cutoff is 600, the line test/mobile/devices.ts already draws between large phones (430-599) and small tablets (600-767). No physical device sits between 480 and 600, but a phone zoomed out one or two steps in Safari does: a 440pt iPhone at 85% or 75% page zoom reports 518px or 587px and still needs the phone controls, which a 480 cutoff would have taken away. The phone block is max-width: 599px and the tablet block starts at min-width: 600px, so a 600px device is a tablet in CSS and in getDeviceType() alike instead of straddling the boundary the way 430pt phones did. The number changes everywhere it is encoded: JS, CSS, comments, CLAUDE.md, the CI tests that pin the phone block, and the test:mobile helpers. Measurement history that names 430px stays as written. |
||
|
|
d5b75af628 |
fix(statusline): sticky telemetry collection, footer print-through, EOF fix
Responds to Ark0N's review round on the ephemeral-CLI-flag statusline injection rework: - Rebase-detail fixes: registry-gated telemetry eligibility via getCli(mode)?.capabilities.statusLineTelemetry instead of a hardcoded mode === 'claude' check, using the capability flag master's CLI-registry refactor already declares for exactly this purpose. - Design question settled: sticky (a). Rather than persisting the toggle as a new field and threading it through every session-creation path (cron, Ralph Loop API, quick-start), eliminated the per-session field entirely. readPlanUsageTelemetryEnabled() (hooks-config.ts) reads the existing showPlanUsageLimits setting fresh from settings.json at every claude create/respawn (TmuxManager.createSession/respawnPane) - no per-session state to survive a restart, and it applies uniformly to every creation path for free, since they all flow through the same TmuxManager methods. This required fixing a real bug found along the way: showPlanUsageLimits was not actually round-tripping through settings.json on save - settings-ui.js explicitly excluded it from the PUT body as a pure per-device display key. It now flows through normally (both true and false); the load-side per-device merge behavior is unchanged. Removed entirely as a result: the statusLineTelemetry field from CreateSessionSchema/SettingsUpdateSchema, CreateSessionOptions/ RespawnPaneOptions, Session._statusLineTelemetry (this is what makes the restart-persistence bug moot rather than patched), and the frontend send sites. - Footer print-through restored: the no-user-statusline branch of the exporter script now runs the telemetry POST in the foreground so its own stdout becomes the in-terminal footer, falling back to a plain "codeman" marker only on curl failure. - Background-subshell EOF fix: the wrap-a-real-statusline branch closes stdin too, not just stdout/stderr (`>/dev/null 2>&1 </dev/null &`) - the un-redirected subshell process itself, not curl, was what held a reader-to-EOF's pipe open for however long curl took to finish. Added curl --max-time 5 so a hung (not just refused) Codeman cannot wedge the render. Tests: real-shell-execution tests for the footer/EOF fixes (fake curl stand-in on PATH, real sh subprocess spawns, real elapsed-time measurements - verified non-vacuous against a hand-reconstructed old-style script), unit tests for readPlanUsageTelemetryEnabled. Adapted two existing tests whose payloads referenced the removed field. Fixed during independent code review: a stray indentation break and a test exercising the wrong (legacy) exporter code path. Docs synced: CLAUDE.md, docs/usage-limits-display-plan.md (old disk-based section marked superseded, kept for history), docs/architecture-invariants.md. Full suite green: 352 files, 6780 passed, 12 skipped, 0 failed. tsc/lint/format:check/frontend-syntax all clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
344e93c824 |
Merge pull request #386 from irisitymichaelgrundberg/feat/codex-resume
Merging with the phone-overview resumeId fix and the two unified-list doc passages applied on master. |
||
|
|
8ee7926e27 |
feat(agent-cases): tag agent-spawned case dirs and sweep their leftovers
A long orchestration creates one case directory per worker and deleting the sessions never removed them, so ~/codeman-cases accumulated scratch folders that were indistinguishable from real projects. They are now labelled and have a cleanup path. - src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven spawn gets a .codeman-agent-case.json marker (when, by whom, parent session, mode). Only the create branch writes it, so a linked case, a cloned repo or any pre-existing path is never labelled; reading is total, so a malformed marker means "not agent-created" rather than a half-trusted entry. - The signal is the new X-Codeman-Agent-Origin header the skill preamble sets on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body field, falling back to a resolved parentSessionId so a worker spawned by a stale skill copy is still labelled. - GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage badges each case and offers a review-then-delete sweep that names every directory in its confirm and skips any case a live session is working in. Removal stays on the existing DELETE /api/cases/:name. - Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was written per claude session and never removed (236 leftovers measured on a working machine). Now deleted with the session and swept at boot, guarded by a live-session keep set plus a 7-day age floor. Verified end to end on an isolated instance: marker written for header, body and lineage-only spawns, absent with no agent signal and for a pre-existing directory; inUse flipping on session end; badge, sticky bar, confirm and sweep driven in a browser; preamble seeded on create, removed on delete, boot sweep taking only the aged orphans. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2f9663e389 |
Merge branch 'master' into feat/codex-resume
master and this branch both rewrote the two `_claudeSessionId` resets inside
`start()`, so `src/session.ts` conflicted at both of them.
master's commit
|
||
|
|
cfc8fe7e41 |
Merge pull request #381 from mtiller/feat/reverse-proxy-base-url
feat(web): support a reverse-proxy base URL |
||
|
|
7991f481b6 |
Merge pull request #368 from shenlvkang-collab/pr/mobile-add-case-submit
fix(mobile): give the Add Case modal a reachable submit button |
||
|
|
8285fff91c |
feat(codex): list codex conversations and resume them
Codex conversations never appeared in the session list, and the resume path skipped codex, so picking one back up meant finding its thread id by hand and POSTing codexConfig.resumeSessionId to /api/sessions. Two gaps caused it: - The unified list is built from ~/.claude/projects plus omp's own store. Codex writes to neither: its rollouts live in ~/.codex/sessions/<y>/<m>/<d>. - terminal-ui.js sends a continuation only for the CLIs with a "continue most recent" flag. Codex has no such flag — it names a thread by an exact id — and nothing supplied one. Add codex-transcript.ts, the codex analog of omp-transcript.ts, and wire it into gatherUnifiedInputs() beside the omp scan. A rollout row carries `resumeId`, the thread id `codex resume` takes, and the resume path sends it as codexConfig.resumeSessionId. `resumeId` is what keeps the two kinds of row apart: only a transcript scanner sets it, so a LIVE codex row — whose sessionId is Codeman's own uuid — can never ask codex for a thread that does not exist. Three things measured against a real store of 519 rollouts rather than assumed: - Rollouts are far too large to read whole (median 407 KiB, p90 1.3 MiB, max 25 MiB, 381 MiB total), so this reads a 128 KiB head for the identity and the opening prompt and a bounded tail for the most recent one. session_meta is written once and never rewritten, so per-path identity is cached; a warm rescan of that store costs ~75ms against ~470ms cold. - codex 0.152.1 emits no event_msg/user_message rows at all. It writes event_msg/item_completed carrying an item.type of UserMessage. Both shapes are read, plus response_item as a last resort. - That last resort sees injected context, and the first such row is the repo's AGENTS.md every time, so injections are dropped rather than used as titles. Sub-agent threads (thread_source: 'subagent') are left out; codex spawns them for itself and on a real store they outnumber the resumable threads. |
||
|
|
8ad2215118 |
fix(docker): close the three adoption gaps the negative guarantee missed
Review follow-ups to #357. Each is a path that still touched, or still hid, a container Codeman does not own. **Export still mutated it.** The four fail-closed layers cover create/start/ stop/remove, but `POST /api/docker-cases/:name/export` reaches the container twice through neither: a full export `docker commit`s it, and even a workspace-only export `docker pause`s it first for snapshot consistency. Pause freezes the owner's processes for as long as the tar takes, on a container we promised not to touch. Full export is refused for an adopted case (it packages someone else's container, with their logins, into a bundle Codeman hands out); workspace-only keeps working and no longer pauses, accepting a live filesystem the way `tar` does on any running host directory. **A freshly linked OWNED case became unusable.** The run menu now probes the container for its CLIs, and a failed probe hides every agent mode behind the reason. For an adopted case that is right. For an owned one the container does not exist until the first session launches it, so every newly linked Docker case answered `container "codeman-case-x" not found (adoption never creates a container — start it yourself first)` and offered nothing but Shell, for a container the launch chain was about to create itself. A failed probe is recorded only when the case is adopted; `CaseInfo.docker.owned` is on the wire so the frontend can tell them apart. Verified in a browser: owned-with-no- container offers all ten modes and no notice, adopted-but-stopped offers Shell and says why. **Multi-user gating.** Adoption is admin-only, unlike `docker-link` beside it. Linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has already confined to the caller's space; an adopted container's mounts are whatever its owner gave it, so one mounting `/` hands the adopter a shell over the whole host — exactly the workspace scoping multi-user mode exists to enforce. Listing the engine's containers and browsing directories inside an arbitrary one are machine-level reads and follow the docker-HOST policy for the same reason. The preflight is deliberately not admin-only: the run menu fires it for every docker case, so it admits a non-admin for a container already linked to a case they can access, and nothing else. Verified end to end against a real pre-existing root container (alpine + tmux, no bind mounts): adopt, claude session inside it, workspace export, session close and case unlink all left `StartedAt`, `RestartCount`, `Pid` and `Paused` untouched; the pane ran the CONTAINER's claude, without `--dangerously-skip-permissions`; a stopped container was refused at both preflight and launch and was never started. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1 |
||
|
|
3d8ffcb9a2 |
Merge pull request #357 from dignfei/feat/docker-adopt-existing-container
feat(docker): attach a case to an already-running container Conflicts came from work that landed after the PR was opened, and each is resolved onto the newer abstraction rather than by keeping the older code: - `defaultDockerCommandForMode` is registry-driven since #347, so the PR's `runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the refusal is visible only inside the container, so an adopted root container otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and `test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch. - The probe's mode list and its mode -> binary table both duplicated the registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is also what fixes the merge's silent regression: the hand-written list predates `omp`, and the run menu gates every docker case on this probe, so owned containers would have lost that mode. `shell` needs no arm — it declares no binary, so it is dropped from the lookup and reported available regardless. - The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one `missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto it. Its test now pins the single gate instead of counting seven arms. - The create arm keeps #349's swap-limit warning filter, which the adopted arm never reaches; the run-mode list gains `omp` from #353. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1 |
||
|
|
7e4914d991 |
feat(web): support a reverse-proxy base URL (--base-url / CODEMAN_BASE_URL)
Codeman can now be mounted under a sub-path behind a reverse proxy that forwards the prefix unchanged (e.g. https://host/codeman/). Default is `/` (root), which is byte-identical to the historical behavior. Design — few choke points, mirrored ingress/egress: - src/config/base-path.ts: pure single-source normalize/validate/join/strip. - Server ingress: stripBasePath() inside Fastify rewriteUrl, so routes stay declared prefix-agnostic; un-prefixed requests (hooks, health, docker bridge hitting the raw port) pass through unchanged. - Server egress: one onSend hook prepends the base to root-absolute Location headers (covers all redirects). - HTML: renderIndexHtml points <base href> at the mount and injects window.__CODEMAN_BASE__ — ONLY when a base is set (inert at root). - Frontend runtime URLs: CodemanBase.url() route builder in constants.js, applied transparently by a fetch wrapper and explicitly at the EventSource/WebSocket/window.open/<img|iframe|a>-src sites. - sw.js derives its base from self.location; manifest uses relative start_url/scope. - Web-tab proxy: proxyPrefixFor(cap, basePath) is the single base-aware root that cascades to the injected <base>, HTML/attr rewrites, runtimeUrlShim, Set-Cookie Path and Location; capabilityFromReferer strips the base off the browser Referer, while the ingress parsers stay base-agnostic (rewriteUrl already stripped it). --base-url rides the daemon relaunch (buildWebArgs) and the service unit (resolveServicePlan). constants.js is guarded against a missing `window` for isolated unit-test contexts. Tests: test/base-path.test.ts (pure helpers), base-path coverage in webview-proxy/render-index-html/daemon-control; CodemanBase stubbed in the vm-isolated panels-ui test contexts. Docs: Remote-Access.md (sub-path section + nginx example), security-architecture.md env table, CLAUDE.md pattern. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XUkPBxbumnct6qSrx4JDju |
||
|
|
5969a1df96 |
fix(mobile): give the Add Case modal a reachable submit button
mobile.css hides #createCaseModal's .set-foot below 860px, and that modal's header — unlike Settings' — carries no set-head-save. So on a phone the Create/Link button existed nowhere and the modal could not be submitted at all. Adds the header button and drives both together through switchCaseModalTab() and submitCaseModal(), so whichever one is pressed the other shows the same pending state and is equally unclickable. Following the Settings pattern also means Add Case picks up the existing .set-head-actions:has(.set-head-save) tray and .set-head-save sizing with no new CSS; the mobile.css comment that still listed Add Case as a lone-× sheet is corrected to match. |
||
|
|
8e5e207386 |
fix(docker): send the probe body as an object, and explain an unreachable container
The run menu still offered every mode for an attached container. The browser's
actual request showed why:
POST /api/docker-cases/adopt-preflight -> 400
{"error":"Invalid input: expected object, received string"}
_api serializes `body` and sets Content-Type itself, and three call sites each
passed an already-stringified body, so it was encoded twice and the server saw a
JSON string where it expects an object. curl was fine throughout, so nothing in
the server logs pointed at it.
Also fixes the design defect underneath: a failed probe fell through to "do not
gate", which silently offered every mode. When the container has been recreated,
is stopped, or the engine is unreachable, the user sees claude, clicks it, and
it can only fail — with the reason visible nowhere. A failed probe now hides
every agent mode (Shell needs no CLI and stays) and shows the server's own
reason at the top of the menu.
Two static guards switched from a character window to brace matching. They
sliced between two call sites, and _loadRunModeHistory's call appears above its
definition, so the slice came out empty and the assertion verified nothing —
the same trap twice in one file.
|
||
|
|
5452ad5c5a |
feat(docker): add a folder picker to both path fields
Both paths in the adoption form had to be typed. Each gets a Browse button using the same path-input-group markup Link Existing uses, so the two look and behave alike. What they can browse differs, and that is the point. The host workspace path reuses the existing host picker. The container workdir cannot: an adopted container has nothing mounted at a matching host path, so a host listing would be a different filesystem — and getting this field wrong is the source of the opaque OCI chdir error at launch, which makes it the field that most needs to be clickable. Adds a read-only POST /api/docker-cases/browse: one `ls` through docker exec, no writes, no lifecycle, path shell-escaped like every other value. `ls -Ap` marks directories with a trailing slash and keeps names with spaces intact. PathPicker takes an optional fetchListing source rather than being forked: the container variant only swaps where the rows come from, and reuses the rendering, navigation, Up and Choose/Select unchanged. |
||
|
|
3685ad85bc |
fix(docker): stop requiring the CLI on the host for a container session
Attaching a container, picking claude and hitting Run gave one line — `execvp(3) failed.: No such file or directory` — and the run-mode menu offered every mode. Three separate defects, found on a real deployment. TmuxManager.createSession resolved the CLI directory without distinguishing a docker session, so a host with no claude threw, the catch fell back to a direct PTY, and that PTY exec'd the CLI on the HOST. The failure surfaced as a bare execvp error naming nothing. A docker session runs its CLI inside the container; the host does not need it. All eight modes now sit behind a cliRunsInContainer guard, and whether the container has the CLI is settled by the adoption preflight or the image gate before launch. The running check used a bare double quote and command substitution. The whole chain is embedded in an outer `bash -c "…"`, so the unescaped quote closed that string early and the remainder was re-tokenized. It is now a `grep -qx` pipeline using only the single-quote form every other line in the builder already uses. Claude Code refuses --dangerously-skip-permissions as root. Our base image runs a non-root user, so an owned container never hit this; an adopted container's user belongs to its owner and is frequently root, and keeping the flag killed the pane with a message visible only inside the container. The preflight now reports runsAsRoot and the launch chain drops the flag for it. The menu also showed every mode because the container CLI probe only started when the menu opened. It is warmed when the case is selected instead. |
||
|
|
06e7cbe286 |
fix(docker): probe the container's CLIs live instead of trusting attach time
Storing the container's CLIs on the case at attach time left two gaps: a case linked before that field existed has none at all, and a container's CLIs can be installed or removed long after it was linked. A real deployment hit the first one — the host had only codex, the container only claude, and with no stored list the menu still gated on the host and hid the mode that actually worked. The probe now runs when a container case is selected, reusing the existing adopt-preflight endpoint, so there is no new backend surface. Results are cached per case for the page's lifetime, since the menu opens often and the probe is a `docker exec` round trip; a concurrent probe for the same case is deduplicated with an in-flight marker. A failed probe leaves the cache empty, which the caller reads as "unknown" and therefore does not gate. Hiding every mode because one probe failed is worse than offering one that turns out to be missing, which the launch path already refuses with a specific message. The repaint only happens while the menu is still open, so a late answer cannot make the list jump under a user who already closed it. |
||
|
|
2f83a37c6d |
feat(docker): take run-mode availability from the container
The run-mode dropdown hides CLIs that are not installed on the HOST (#201). That is right for local sessions and wrong for a container case, whose agents run inside the container: a host with no claude installed hides the mode while the container ships one, which is exactly what happened on a real deployment. The adoption preflight already probes what the container has, so that result is persisted on the case and surfaced through CaseInfo. Docker cases gate on it; every other case keeps the host probe unchanged. An absent list reads as "do not gate" rather than "nothing available": an owned container runs our base image, which ships every CLI, and treating unknown as empty would leave the menu with Shell alone. |
||
|
|
34c12ca18b |
feat(docker): make the container field a picker you can also type into
Typing a container name from memory is error-prone. The field becomes a native datalist: pick from the engine's containers, type to filter, or type a name that is not listed (the engine may be remote, or the container may not exist yet). A datalist gives all three natively, so no dropdown state machine is introduced. Adds listDockerContainers and GET /api/docker-hosts/:hostId/containers, following the listRemoteCodemanSessions discovery precedent: read-only and never throwing, so an unreachable daemon returns an empty list and the field degrades to plain text instead of erroring. Stopped containers stay in the list, sorted after running ones and labelled. Attaching does require a running container, but hiding stopped ones turns "my container is not in the list" into a dead end, while showing `Exited (137) 8 days ago` says exactly what to fix. |
||
|
|
bb45909169 |
feat(docker): link to container attach from the Create New tab
Attaching lived only on the Docker tab, but the place users look for anything container-shaped is the "Run in an isolated Docker container" checkbox on Create New. A feature nobody can find is a feature nobody has. Adds a one-click link there that switches to the Docker tab, turns the toggle on and focuses the container field. Reuses switchCaseModalTab and the existing sync helper; no new CSS. |
||
|
|
c98a59d709 |
fix(docker): verify the container workdir and end the probe with exit 0
Two defects that only a real container exposes. The probe chained `command -v X && echo X` with semicolons, and a script's exit status is its last command's. A container without the last probed CLI made the whole `sh -lc` exit 1, so a perfectly healthy container with tmux and claude was reported as "could not exec into the container". A missing CLI is data here, not failure, so the script now ends with `exit 0`. containerWorkdir defaulted to hostWorkspacePath. That default holds for an owned container only because the create-time bind mount puts the host directory at that exact path; attaching mounts nothing, so the two are independent facts. A host path absent inside the container makes `docker exec --workdir` fail with an OCI chdir error that surfaces in the pane as a bare "execvp failed". The preflight now proves the directory exists inside the container and refuses at link time. |
||
|
|
bc55b6b0da |
feat(docker): add the attach-an-existing-container panel
The Docker tab gains an "Attach to an existing container" toggle. Ticking it swaps the create-time fields (image, network, advanced) — which describe a `docker create` attaching never runs — for the container name, and routes the submit to the adopt endpoint. Reuses the existing linkDockerCase flow end to end: only the final call differs. The docker-host upsert still applies, since it is what resolves the engine/context/daemon for `docker exec`; its create-time fields are simply never read for an attached case. |
||
|
|
3e1a0e679f |
fix(omp): resume by mode, not silently as claude, and support --continue
resumeHistorySession() never sent mode when recreating a session from a history/session-manager row, so the server default silently opened a plain Claude session for every non-claude row -- reproduced live: OMP rows spawned Claude sessions on click. Thread the row's mode through every call site (welcome list, session manager, mobile overview) and only send the Claude-specific resumeSessionId for claude rows. Codeman has no live PTY-reattach outside server boot, and it's moot for OMP anyway (exiting it kills the pane's only process), so route the non-claude relaunch through each CLI's own continue-most-recent flag instead of a context-free fresh start. OMP never got one: buildOmpCommand only implemented --model/--resume despite omp --help documenting -c/--continue. Added continueSession to OmpConfig end-to-end (type, schema, builder) mirroring the existing opencode/pi/grok/deepseek fields, and wired resumeHistorySession to use it. Verified live: told a real omp session a secret, exited it, closed the tab without killing tmux, relaunched with --continue in the same directory, and had it recall the secret. |
||
|
|
c0423bf560 | fix(omp): complete omp wiring in UI files, skill docs, and tests after rebase | ||
|
|
4f5678fac4 | feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) | ||
|
|
ca5fe1ab3e |
Merge pull request #338 from Ark0N/feat/vertical-rail-detailed-rows
Vertical tab rail: detailed rows (created / working / status), plus a rename-cancel fix |
||
|
|
a628737d1f |
fix(deepseek): review-driven hardening across the harness integration
Fifteen review findings on the dsh mode, the serious ones first: - Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys. _configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into every dsh pane and applyEnvOverrides() lands after it, so a non-granted owner who could redirect the base URL would have the operator's key sent as a bearer credential to a host of their choosing. - Wait registry: until=stop/blocked is refused on docker and remote-SSH dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions). The HERDR triple is set via LOCAL tmux setenv, which crosses neither docker exec nor ssh, so such a session can never post a hook event and the wait burned its whole timeout on every turn. - Approvals: a dsh item is an ALERT, not an answerable card. The answer route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the option parser cannot read a third-party TUI's frames, so an answer was a blind keystroke into a foreign composer), and the push notification carries no Approve/Deny actions for dsh sessions. - Status shim (v3): --seq is forwarded and the server drops stale retried reports inside a 60s window (the TUI retries with backoff, so a retried 'working' could land after 'blocked' and resolve an approval whose dialog was still on screen); 4xx responses exit 0 instead of retrying, so one misconfigured session cannot feed the auth rate-limit bucket until the hook endpoint 429s for the whole instance. - Web-UI server: concurrent starts are serialized through a lock (two racing POSTs used to pick the same port and orphan the winner), and the readiness poll / timeout paths only clear or stop the singleton while it is still theirs. First click actually opens the tab now (refreshWebviews, not the nonexistent loadWebviews). DELETE /api/deepseek/web requires the privileged grant in multi-user mode. - Cron: deepseek jobs run the same two-part launch gate as the HTTP create paths (impl moved into the resolver so all three share it) and no longer stamp a Claude default model on the session. - Parity sweeps: quick-start's docker branch rejects deepSeekConfig like the remote branch; the Ralph auto-enable list gained deepseek; HookEventType gained agent_working; the phone overview run menu filters managed webview records like the desktop menu. - install.sh: the dsh identity probe closes stdin (under curl|bash a child that reads stdin eats the rest of the script), bounds the exec with timeout where available, and is memoized to one scan per install. - Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand identity (it rendered as an unstyled UA-grey button); stale markup comment about the web shortcut rewritten; clamp docs updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c30dfaf0e7 |
fix(deepseek): run the web UI server in the background, not in a shell tab
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a shell tab running the server next to it. The shell was deliberate - the server lived in an ordinary session so it was visible, scrollable, killable and died with its tab, and nothing new had to supervise a long-lived HTTP server. That reasoning was sound and the result was still wrong in use: opening a dashboard should open one tab, and after the first launch the terminal is pure noise. The server moves to a background child process owned by a new `src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session gave away for free is now explicit, which is most of the module: - Exactly one server. A second click reuses the running one instead of racing it for a port; the session flow could not do this at all, because two clicks were simply two sessions. - Restarted when the requested authority changes. `--trusted-host` fences dsh's own /api against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two. Reusing a server fenced for the other origin renders a page whose every call 403s, which reads as a broken dashboard rather than a misconfigured one, so a mismatch restarts instead. - Killed on shutdown. The child is detached so its whole plugin tree can be signalled at once, which also means it would outlive Codeman and hold its port against the next start - the exact EADDRINUSE this feature already got wrong once. - Boot output captured and returned. With no shell tab there is nowhere else for a stack trace to land, so a failed spawn reports its own tail. The endpoint is fenced at the same bar as the profile installer and for the same reason: booting a dsh profile executes the plugin code in it, so this is a privileged action even though it reads as "open a page". `authority` comes from the client (`location.host`) because only the browser knows which origin is in play, and it is regex-confined at the schema boundary - defence in depth behind the argv-array spawn, admitting host:port in the shapes a browser authority can take and nothing readable as a second argument. `GET /api/deepseek/web-port` is gone; port selection moved into the supervisor, which is the thing that knows whether a server is already running. The two client-side probe helpers went with it, since the server now owns the wait. Verified over the tailnet authority end to end: no session is created (session count unchanged, one tab), the server runs on 3081 beside the user's own dsh web on 3080, status reports the tailnet authority, and the proxied dashboard renders with zero 4xx. Full gate green (6148 passed, +6). |
||
|
|
15ae5f5d81 |
fix(deepseek): make the web-UI shortcut pick a free port, verify it, and trust its frame
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.
1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
precisely the port a DeepSeek user is most likely to be serving on already,
so the launch died with EADDRINUSE against the user's own server. The port
now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
free loopback port by BINDING it (a connect probe cannot tell "free" from
"listening but not answering yet").
2. It opened the tab unconditionally. The crashed server left a saved dashboard
pointing at nothing, with the failure only visible in a shell tab nobody had
a reason to look at. The launch now polls the existing webview probe until
the URL answers, and on timeout reports the error naming the shell tab
instead of persisting a dead dashboard.
3. The saved tab was untrusted, so the frame was sandboxed without
`allow-same-origin` and the dashboard was broken twice over: the dsh
client-runtime reads `localStorage` while loading its plugins and died there
("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
every `/api` call no matter which authority `--trusted-host` named. Passing
`location.host` only means anything once the frame actually carries that
origin, so `--trusted-host` had never once done its job. The managed tab is
now created `trusted: true`.
That trade is real and deliberate: a trusted proxied frame is same-origin
with Codeman and can reach Codeman's API. It is defensible only because this
dashboard is an agent harness Codeman just started itself, on loopback, which
can already run code as the user. It is not a precedent for trusting
third-party dashboards, which is why it is set at this one call site rather
than defaulted.
Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.
`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.
Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
|
||
|
|
b330f1d9e8 |
feat(tabs): give the vertical rail the home screen's per-session detail
The vertical tab rail (tabOrientation 'vertical') listed names and nothing
else, while the rich sidebar and both home screens already answered the
question a docked column exists to answer: which of these sessions wants me
next, and how long has it been like that. The rail is a docked column too, so
it now draws the same row.
- New per-device setting tabRailDetail ('rich' | 'simple', default rich),
App Settings -> Appearance -> Tabs, in SettingsUpdateSchema + displayKeys and
stamped as data-tab-rail-detail by the pre-paint script, so a detailed rail
does not flash through simple rows on every load.
- ONE gate for both vertical surfaces: isRichTabRows() =
isSessionSidebarRich() || isTabRailRich(). The row model, the markup and the
20s in-place clock are the existing rich-sidebar ones, classified by
_mobileOverviewState/_mobileOverviewSince, so the rail, the sidebar, the
desktop home rail and the phone overview cannot disagree about what
"working" means or which stamp measures it.
- Detail rides on its OWN attribute, exactly as the sidebar's does, so every
existing [data-tab-orientation='vertical'] rule keeps matching both variants
untouched. A flip of detail ALONE still forces a full render (the stamps line
is emitted by the row template, not toggled by CSS) and re-runs
applyTabWrapSettings(), which owns the folder line and is now rail-aware.
- CSS: every rich paint rule gains a rail twin as a COMMA-GROUPED selector,
never :is() - an :is() list takes its most specific argument, which would
lift the sidebar arm from (0,3,1) to the rail's (0,5,1) and let these rules
outrank things they never used to.
- Width is why there are thresholds. At 256px the stamps line ellipsizes
mid-word, the same reason the rich sidebar is 300px, so a rail that has never
been sized defaults to 320 (RICH_DEFAULT_WIDTH, the existing Wide preset,
which also keeps the settings select on a named choice). A width the user has
chosen is never overridden: below 288px the created stamp is dropped rather
than truncated (tab-rail-tight, CSS only) and below 240px the rows go back to
simple (tab-rail-compact, which re-renders).
- The rich clock is armed and disarmed by applyTabOrientation() as well as
applySessionListLayout(); a leaked interval would rewrite stamps in a list
that no longer has any.
Also fixes a data-loss bug in the inline tab rename that predates the rail and
reproduces in every layout, header strip included: Escape set the input to ''
and blurred it, and the blur handler commits - so cancelling a rename PUT an
empty name, and the tab fell back to its folder label (measured against a live
server: ["rail-alpha","","rail-gamma"]). Escape now calls cancelRename(), which
invalidates the edit so the blur that follows the input's removal is a no-op.
Tests: rail-detail gate, the three ways it turns back off (simple, compact,
horizontal), sidebar-wins, render-on-detail-flip and the plumbing/CSS guards in
test/session-list-layout.test.ts; the rename cancel in test/inline-rename.test.ts
(browser suite), pinned by running it against the old code first. Verified live
against a real server on an isolated instance: detailed/simple/compact/header/
sidebar variants, click-select, the ... menu, inline rename, Alt+N, the in-place
stamp tick and a full settings-picker round-trip including reload.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
4cda150493 |
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/ antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI as a Codeman web tab. DeepSeek is wired unlike its siblings in three ways, each of which is the reason for a design decision rather than an accident: 1. The agent is a PROFILE, not the binary. `dsh` is a launcher over $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and `base` -- the interactive terminal front door is always a third-party plugin. So availability is two questions: `isDeepSeekAvailable()` (binary) and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run button gates on the latter, because reporting only the binary would spawn a pane that dies on arrival. When the binary is present but no profile is, the run menu offers to install one (POST /api/deepseek/install-profile). 2. The permission switch is an env var, not a flag. The harness has no command-line permission option; its sandbox/approval rows read DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access). Exported via `tmux setenv`, never on the spawn line. Absent = the harness's own workspace-write, which still asks, so the multi-user clamp is the only-if-sent branch and clamps to workspace-write, never read-only. 3. It is the only non-claude mode that passes hooksAvailableForMode(), and it earned that. The terminal front door reports idle/working/blocked to a supervising process over a generic env-gated contract; a generated shim (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each report to /api/hook-event as stop / agent_working / permission_prompt. So a dsh session gets definitive respawn triggers, real wait-endpoint signals and real Approvals Inbox items instead of output-stabilization guesswork. `agent_working` is new (157th SSE constant) and joins APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its alert at once. The resolver needs the strictest identity probe of the family: `dsh` is not merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell), so `dsh --help` must print the harness's own banner before a candidate is handed a spawn line. Model is deliberately not a session field -- it is a composition entry in the profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider keys named by a settings-file `apiKeyEnv` stay out, which is pi's 34-provider-key problem in a new shape. Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the status endpoint's two-part answer, the no-profile refusal, the profile bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the permission mode injected via setenv, and the full status bridge -- a send-and-wait returned signal "stop" from a real turn, and blocked/working created and cleared an Approvals Inbox item. Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md (decisions + honest gaps). Tests: test/deepseek-mode.test.ts, test/deepseek-cli-resolver.test.ts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c173ae0264 |
Merge remote-tracking branch 'origin/master' into worktree-grok-mode
# Conflicts: # src/web/public/app.js |