Commit Graph
344 Commits
Author SHA1 Message Date
DevvynandClaude Sonnet 5 8520925e76 docs(custom-model): bring CLAUDE.md and api-reference.md up to date
Full documentation review pass across the branch's 30 commits.
CLAUDE.md's Custom Model Endpoint Profiles entry hadn't been touched
since the initial backend+picker cut (3 early commits) despite 27
follow-up commits adding real behavior — it described restart-in-place
as universal (now claude-only; 7 other CLIs launch one-shot) and
claimed codex's Responses-API gap as a flat protocol break (now
re-verified as a more precise tool-calling gap). Corrected both and
added a new paragraph covering everything landed since: the llama-swap
conflict check, the after-the-fact swap-displacement sweep, the
/running-cmd-based context-length fix, the context-window floor
warning, skipFirstRunPrompts, the real-time /api/events-based log
status, and the countdown-to-Cancel-button change.

docs/api-reference.md's custom-model-endpoints section was missing the
running-status route, the requiresConfirmation/requiresContextWarning
response shapes, and POST /api/quick-start's customModel field
entirely (the primary launch path for 7 of 8 supported CLIs) — added
all three. Also fixed a real markdown bug in custom-model-endpoints.md:
an inline code span (`POST <baseUrl>/v1/chat/completions`) split across
a line break, which CommonMark renders with the line ending collapsed
to a space, so it displayed as ".../v1/chat/ completions" with a
spurious space inside the path.

Verified: origin/master and upstream/master are both already an
ancestor of this branch (identical at bd286bf5, no new commits since
this branch was cut) — nothing to merge, no conflicts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:18:04 +08:00
DevvynandClaude Sonnet 5 db9729e1fc feat(custom-model): remove loading-banner countdown, add manual Cancel
Replaces the size-scaled expected-time estimate + matching auto-timeout
with a generic hardware/model-size disclaimer and a user-driven Cancel
button, per explicit request. Real load time depends on hardware this
feature has no way to know (VRAM, storage speed, GPU contention), so
the old estimate/timeout was a guess dressed up as a fact — worse, one
that could kill a genuinely slow load partway through on slower
hardware.

- _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline
  entirely — polls indefinitely until ready or cancelled, no automatic
  give-up. Message is now "Loading <model> (<size>) on <endpoint> —
  this can take a while depending on your hardware and the model
  size.", with the real llama.cpp log line still on its own second
  line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/
  _formatRemaining (dead code once the countdown is gone) —
  _lookupModelSizeGB is kept, the GB figure still shows.
- _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real
  "Cancel" button (distinct from the error-type "×" close button,
  since Cancel has a real consequence) that calls it on click. Caller
  owns what cancelling actually means, same split as the swap-confirm
  modal's promise-resolving buttons.
- Cancelling dismisses the banner, shows an info toast (not an error —
  this was deliberate), and closes the session, mirroring what the old
  timeout used to do automatically but now on the user's own call.
- New .center-status-cancel CSS (bordered pill button, distinct from
  the plain "×" close glyph).

Test changes: removed the now-invalid timeout-auto-close/estimate
tests, added cancel-flow tests (dismiss/toast-type/session-close,
never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests
for the new Cancel button (bootAppWithRealCenterStatus, evaluating
panels-ui.js instead of stubbing _showCenterStatus, since this button
is worth verifying for real rather than just through the stub every
other test in the file uses). Typecheck/lint/frontend-syntax clean;
full suite shows no new regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:00:50 +08:00
DevvynandClaude Sonnet 5 2d3fc65758 feat(custom-model): show real-time llama.cpp backend status in the loading banner
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.

- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
  (custom-model-routes.ts): one persistent GET /api/events (SSE)
  connection held open per endpoint, parsing logData frames and
  keeping the latest source:"upstream" (backend llama-server) line —
  filtering out llama-swap's own source:"proxy" request-access lines.
  Idle-closed after 30s of no polling, same 20s sweep as the existing
  swap-displacement check.
- running-status route now returns logLine alongside the existing
  isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
  ("llama.cpp: <line>", bootlog timestamp/level/component prefix
  stripped for display) that stays on the last real thing llama.cpp
  said rather than clearing to blank between polls.

⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).

12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 12:22:45 +08:00
DevvynandClaude Sonnet 5 5ddc028a2f feat(custom-model): detect and notify when a session's model gets swapped out later
The llama-swap conflict check on the apply/create routes only ever runs
at THAT session's own launch/apply moment, and cannot see a swap caused
by a DIFFERENT session's later, ordinary use. Confirmed live: a second
Codex session picking a different model launched with no warning at
all — nothing conflicted at that exact instant — yet it silently
evicted the first session's model regardless (llama.cpp runs one model
at a time). Reproduced and root-caused via direct API calls against a
live test-picker instance rather than guessing.

- detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups
  live sessions with a customModel by endpointId, checks each group's
  endpoint via GET /running once, and flags a session whose own modelId
  is no longer in the running list. Read-only, best-effort per endpoint
  like refreshAllCustomModelHosts's sibling sweep.
- Notifies once per displacement via a caller-owned de-dupe Set: a
  session id is added when displaced, removed once its own model is
  loaded/ready again, so a later genuinely-new displacement can notify
  again.
- New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
  20s — much shorter than the 5-minute model-list refresh, since this
  is time-sensitive) broadcasts a new custom-model:swapped-out SSE
  event per displacement. De-dupe Set cleared per-session on session
  cleanup to avoid an unbounded leak.
- Frontend: global toast (not tied to the displaced session's tab,
  since the point is warning before the user types into it) naming the
  session, its previous model, and what's currently loaded.

Chose the "detect after the fact" scope (vs. checking before every
message send, which would add a round-trip to every turn on every
custom-model session) per explicit user decision after being presented
the trade-off.

9 new tests for the detection logic (flag/clear/re-flag cycle,
unreachable/deleted endpoints, non-llama-swap servers, multiple
sessions on one endpoint). SSE registry bumped 158->159, parity test
passing. Typecheck/lint/frontend-syntax clean; full suite shows no new
regressions (9 more passing than baseline, matching the new tests;
same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 11:09:24 +08:00
DevvynandClaude Sonnet 5 470f75b08c docs(custom-model): record live findings on codex's model-metadata warning
Investigated the user's report of "Model metadata for <id> not found.
Defaulting to fallback metadata..." on every custom-endpoint codex
launch, live against the test-picker's llama-swap deployment (codex
0.152.1):

- The warning is cosmetic. `codex exec 'reply with just OK'` against the
  isolated CODEX_HOME still printed the warning and still returned a
  real reply.
- The isolated CODEX_HOME never gets a models_cache.json written into
  it at all, even after extended real use (inspected a live, actively-
  used directory) — codex can't reach OpenAI's own hosted model catalog
  for this session and silently falls back every time, with no local
  file to create or clean up. There is also no config.toml override for
  a model's metadata.
- Fabricating a fake catalog entry to suppress it would mean copying the
  SHAPE of OpenAI's own proprietary models_cache.json schema, including
  real per-model system-prompt content visible in a genuine entry — not
  something to build for a warning confirmed to have no effect.
- More importantly: a real tool-call attempt against the same setup came
  back as agent_message TEXT (the tool-call JSON printed as the answer)
  rather than an executable function_call item, confirmed via
  `codex exec --json`'s raw event stream. Tool execution is what makes
  codex a coding agent, so it remains not usable for real work regardless
  of the metadata warning — a more precise, re-verified update to the
  existing "Responses API protocol gap" finding (which reported a harder
  Reconnecting/high-demand failure on a different llama-swap deployment;
  this one answers /v1/responses for plain chat but still can't execute
  tools).

No code changes — recipe/comment/confidence-table documentation only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 10:33:48 +08:00
DevvynandClaude Sonnet 5 211b872335 feat(custom-model): skip Claude Code's first-run wizard on custom-model launches
A fresh, isolated CLAUDE_CONFIG_DIR (used to keep an injected API key
away from a stored claude.ai OAuth login) looks like a brand-new Claude
Code profile to the CLI, so it replays its ENTIRE first-run sequence on
every single launch: the theme picker, the security-notes screen, the
per-project "trust this folder?" dialog, and (running with
--dangerously-skip-permissions) a one-time bypass-permissions warning —
confirmed live, none of which a real, already-onboarded profile shows
again.

- New registry-declared env-kind field `skipFirstRunPrompts` (alongside
  apiKeyTrustFile, which it reuses) — claude's entry only, carried
  through buildCustomModelInjection (pure) into
  applyCustomModelInjection (IO).
- seedFirstRunOnboardingState(): merges hasCompletedOnboarding: true and
  this session's own projects[workingDir].hasTrustDialogAccepted: true
  into the same <configDir>/.claude.json the API-key trust file already
  writes to — other projects and other fields on this session's own
  entry are left untouched.
- seedSkipBypassPermissionsPrompt(): merges
  skipDangerousModePermissionPrompt: true into <configDir>/settings.json,
  a separate file, same corrupt-tolerant merge behavior.
- applyCustomModelInjection() gains an optional workingDir parameter,
  threaded from session.workingDir (dedicated apply route) /
  resolvedCasePath (quick-start route) — boot recovery omits it
  (a dialog already answered once needs no re-seed on the same,
  persisted isolated directory).

Tests added at the pure-builder, IO-wrapper (including merge-preserves-
other-fields and corrupt-file-tolerance cases), and existing directory-
listing assertions updated for the new settings.json file. Typecheck/
lint/format clean; full suite shows no new regressions (baseline
pre-existing Windows-environment failures unchanged, 8 more passing
tests than before — the ones added here).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 09:21:36 +08:00
DevvynandClaude Sonnet 5 b45a96358e feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.

- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
  registry entry declares contextLengthVar (currently only claude) and
  the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
  (40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
  quick-start customModel path) check this before the swap-conflict
  check and before launching/restarting anything, returning
  {requiresContextWarning, modelId, contextLength, minSafeContextTokens}
  — skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
  _resolveContextWarningConfirm (session-ui.js), wired into both
  _quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
  (the path Claude actually uses) ahead of the swap-confirmation check.
  Explains the fix in-modal: give the model an explicit larger -c/
  --ctx-size in llama-swap instead of relying on --fit-ctx, which
  optimizes for the biggest model that fits rather than the biggest
  context.

Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 07:36:57 +08:00
DevvynandClaude Sonnet 5 993710263d fix(custom-model): stop trusting /props's n_ctx, parse the real context size from /running's cmd
Root cause of the context-overflow regression reported live: "API Error: 400
request (36437 tokens) exceeds the available context size (16384 tokens)".
Discovery had stored modelContextLengths.qwen3.8-27b-ud-q4_k_xl = 154112,
so CLAUDE_CODE_MAX_CONTEXT_TOKENS told Claude Code it had a huge window and
it never compacted - but the real llama-swap server was launched with
--fit-ctx 16384 (confirmed against /running's own cmd field) and refused
the request right at that real limit.

/props?model=<id>'s n_ctx (the field discovery read) is confirmed live to
be unreliable for a --fit-ctx-launched backend: it reported 154112 for the
same model /running says was launched with --fit-ctx 16384 - appears to
report the model's theoretical/trained maximum context, not the runtime-
configured one.

discoverModels() now parses the REAL configured size straight out of
llama-swap's own launch command instead (parseCtxFromCmd(), reading
/running's cmd field - --fit-ctx first, then the plain llama.cpp -c/
--ctx-size a hand-written command might use), and only falls back to the
old /props probe when cmd states no recognizable flag at all. One /running
call now covers every loaded model's context length in a single request,
same as it already did for the swap-conflict check and the load trigger.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 20:47:12 +08:00
DevvynandClaude Sonnet 5 7bbe408e44 feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout
(updated every poll, so every second by default) instead of a static
"this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically
~1-3 min) on llama-swap - 47s remaining".

If the countdown reaches zero and the model still isn't ready, this is now
treated as a real failure rather than a "keep waiting" shrug:
- The banner turns into a sticky error (_showCenterStatus gains a `type`
  option - 'error' drops the spinner and adds a close button, since nothing
  is "in progress" anymore and a sticky message needs a way to dismiss it),
  naming the llama-swap server's own logs as where to look for detail.
- The session that load was for is closed automatically (closeSession) -
  requested explicitly: a console left open and pointed at a model that
  never finished loading is worse than no console at all. Both apply paths
  now thread the new session's id through to _watchLlamaSwapLoading for
  this (new required 3rd parameter, after endpointId/modelId).

_watchLlamaSwapGeneration's existing stale-call guard extends naturally to
this: a superseded call's own eventual timeout recognises it no longer owns
the banner and neither shows the error nor closes a session that may by
then belong to a different, newer launch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 20:22:02 +08:00
DevvynandClaude Sonnet 5 55dae31530 feat(custom-model): estimate model load time from its discovered size
Discovery now also parses a GB figure out of an auto-discovered model's own
description (llama-swap writes "Auto-discovered 16.35 GB - parameters
auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike
context length this needs no /props probe (the figure is right there in
/v1/models) so it is populated for every model regardless of loaded state.
A hand-configured profile's own description has no such figure and
correctly gets no entry.

The loading banner (_watchLlamaSwapLoading) now looks this up and, when
known, shows it plus a rough estimate from a small size->time matrix
(_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) -
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on
llama-swap... this can take a while" - and uses that same estimate's own
bracket to scale the banner's default give-up timeout for a very large
model, instead of a flat 5 minutes for everything. Explicitly labelled as
an UNMEASURED, typical-hardware estimate in every relevant comment - this
is not benchmarked against any real endpoint's actual storage/GPU, just a
reasonable expectation-setter. A model with no discoverable size (a
hand-configured profile) gets no size/estimate shown at all, matching the
"never a guess" convention modelContextLengths already established.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 19:58:46 +08:00
DevvynandClaude Sonnet 5 0929694012 fix(custom-model): actually trigger the llama-swap load, not just watch for it
Root cause of "it doesn't look like llama-swap is actually switching the
model" (confirmed live: no load_model line in llama-swap's own logs after
applying a selection). llama-swap has no "switch model" admin endpoint - the
ONLY thing that starts a swap is a real inference request naming the model.
Every previous fix (the conflict check, the loading banner) assumed a swap
would start on its own; nothing ever actually asked llama-swap to load
anything until the launched CLI's first real prompt did, which could be
much later than "applying the selection" implied.

Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest
real request that will start a load - POST <baseUrl>/v1/chat/completions,
max_tokens: 1, one throwaway message - fire-and-forget (never awaited by
the caller; the frontend's own running-status polling is what actually
confirms readiness). Wired into both apply paths (the dedicated restart
route and the one-shot quick-start route), fired whenever the target model
isn't already the one loaded and ready - a broader condition than the
existing swapNeeded (which only gates the "this will evict another
session's model" confirmation ask and deliberately stays narrow to that).
modelSwapInProgress in both routes' responses now reflects this same
broader condition too, so the frontend's loading banner actually correlates
with a real in-flight load rather than only firing when something else
happened to be loaded already.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:53:34 +08:00
DevvynandClaude Sonnet 5 f865f74a0f feat(custom-model): launch directly on the endpoint, no restart, for 7 of 8 CLIs
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.

POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.

Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.

Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:04:10 +08:00
DevvynandClaude Sonnet 5 97464bfa27 fix(custom-model): pre-approve the injected API key in the isolated Claude config dir
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.

Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 13:08:07 +08:00
DevvynandClaude Sonnet 5 0e8b1981af fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker:

1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still
   coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same
   config directory and warns about it (confirmed cosmetic - the API key
   wins for actual requests, verified via a real session's own API Usage
   Billing line). A custom-model claude session now gets an isolated
   CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty,
   no files written into it) so there is nothing to conflict with. projects
   is symlinked (junction on Windows) back into the real config dir so the
   response viewer, subagent windows and Read My Mind keep working for that
   session, best-effort.

2. Context-window overflow. Claude Code assumes a large default context
   window for a model id it doesn't recognise and never compacts, so a
   custom endpoint's real, much smaller context (verified live: a 400
   exceeding a 16384-token llama-swap model with a stock ~33.7K-token system
   prompt) silently overflows. Discovery now also learns each model's real
   context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx),
   but ONLY for a model llama-swap's own /v1/models response already marks
   status.value === 'loaded' - never an unloaded one, since llama-swap
   treats ?model= as a routing hint and probing an unloaded model risks
   triggering an actual, slow, GPU-swapping load as a side effect of
   read-only discovery. A server with no status field at all gets no
   enrichment rather than a guess; a model not probed this round keeps its
   previously-learned value until it disappears from the list entirely.
   Stored per model (CustomModelHost.modelContextLengths) and applied via a
   new contextLengthVar registry field, set to
   CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude.

Both new fields live on the existing env-kind customModelInjection
capability shape, declared only on claude's registry entry - every other
CLI's injection is unaffected (pinned by test).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 12:08:26 +08:00
DevvynandClaude Sonnet 5 5c25a52f95 fix(custom-model): wait for a freshly launched session to go idle before applying
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.

Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.

Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 10:38:37 +08:00
DevvynandClaude Sonnet 5 5a9ff07f57 feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:

1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
   apply the endpoint's defaultModelId (or the first discovered model)
   silently. Now, via the new selectCustomModelEntry() (session-ui.js):
   - exactly one discovered model launches straight away, same as before
   - two or more open a new #customModelPickModal listing every discovered
     model; defaultModelId (if set) is marked but never auto-chosen, since
     the point of asking is letting ONE launch deliberately differ from
     the saved default, not just confirming it
   The endpoint is re-fetched at click time rather than trusting anything
   cached from the dropdown's own render, since the model list can have
   changed (the sweep below, or a settings-panel edit) since it opened.
   runCustomModelEntry() itself — the actual launch, routed through run()
   for the in-flight lock, snapshot-guarded against applying to the wrong
   session — is unchanged; it now just always receives an explicit model
   id from one of these two paths instead of computing one itself.

2. Periodic re-discovery. Every saved endpoint's models now refresh
   automatically every 5 minutes in the background
   (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
   as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
   off under testMode), so a model the server starts or stops serving shows
   up without another manual "Discover" click. The manual POST
   .../discover-models route and the new refreshAllCustomModelHosts()
   sweep (custom-model-routes.ts) now share one pure merge step
   (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
   that no longer appears) rather than two copies that could drift. The
   sweep is best-effort per host — one endpoint being unreachable on a
   cycle never blocks the others — and re-reads the store before each
   host's write, keyed by id, so a concurrent edit or delete from the
   settings panel always wins over a sweep that started before it.

Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.

Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 08:50:18 +08:00
DevvynandClaude Sonnet 5 60e1bd52f7 fix(custom-model): act on the draft review — unparseable onclick, unwrapped envelope, wrong-session apply, missing lock, no tests
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.

Blockers:

1. Every generated inline onclick was unparseable. JSON.stringify's own
   double quotes terminated the double-quoted HTML attribute at the first
   one, leaving btn.onclick null on every picker entry and every Discover/
   Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
   argument, the same idiom deleteCase's onclick already uses four lines
   away in session-ui.js. This also closes the live-HTML-injection route
   through modelId (server-controlled, from the endpoint's own /v1/models
   reply): with quoting intact, a `>` inside it can no longer terminate the
   <button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
   like every other /api route (server.ts's preSerialization hook applies
   to arrays too), so Array.isArray(hosts) was always false in production
   and the picker/settings panel silently saw nothing. Both call sites now
   go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
   returns normally without ever changing activeSessionId, so the apply
   step used to silently re-point and restart whatever session the user was
   already looking at. runCustomModelEntry() now snapshots activeSessionId
   before the launch and requires it to have actually changed.

Majors:

4. Routes the launch through run() itself via a temporary _runMode swap
   (never persisted — setRunMode() would sync it to the server) instead of
   a parallel hardcoded dispatch table, so a custom-model launch now holds
   the same _runInFlight lock every other Run click gets. This also
   resolves the "hardcoded runners map contradicts the PR's own design"
   minor: dispatch is run()'s own, so a CLI whose customModelInjection
   recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
   against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
   parses markup this module generated itself) for exactly the DOM-level
   facts the review said needed no Playwright and no tmux: a generated
   button's onclick genuinely compiles and fires, a dangerous modelId never
   produces a live element, the envelope unwrap works, the session-changed
   guard holds, run() actually gets called (proving the in-flight lock
   engages), and _runMode is restored afterward. Confirmed against the
   pre-fix code first (reproduces btn.onclick === null exactly) so this
   isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
   render-index-html.test.ts for the other fixes below.

Minors:

- Generated entries now filter through isCliAvailable(), matching
  _refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
  (applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
  and to settings-modal open) instead of always rendering; the endpoint GET
  no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
  redactApiKey() replaces the field with a computed apiKeySet: boolean, and
  a PUT with no apiKey now keeps the stored one server-side
  (applyStoredApiKey()) instead of the client resending a value it was
  never given. New tests cover both directions (kept vs. replaced) by
  observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
  (_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
  event, since the real role can resolve after settings were first opened)
  — endpoint writes were already admin-only server-side, but the button
  used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
  toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
  route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
  that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
  (CLAUDE.md already records that exact literal turning the settings
  preview into a grey slab on light skins), .run-mode-custom-models gets
  the same gap: 2px .run-mode-menu's own flex gap only applies one level
  up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
  </script> (CliEntry.label is user-clis.json-settable, unlike
  __codemanCliAvailable's booleans-only payload) via a new exported
  escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
  docs/api-reference.md; left the "no zh-CN for the new Models-section
  group" minor unaddressed only insofar as the wider Models section (task
  routing, thinking effort, etc.) has never had zh-CN coverage either —
  everything this PR itself introduces (labels, hints, button text, the
  Run-menu's "Custom Endpoints" header) IS translated in i18n.js.

Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
DevvynandClaude Sonnet 5 98d26e14d9 docs(wiki): document Custom Model Endpoints and the Run-menu picker
New docs/wiki/Custom-Model-Endpoints.md (auto-synced to the live GitHub
wiki on push to master, per docs/wiki/Contributing.md) covers turning the
feature on, adding an endpoint, the Run-menu picker's one-off-run
behaviour, the per-harness confidence table, and what it deliberately does
not do yet (remote/Docker sessions, live hot-swap). Linked from the
sidebar, from Agent-CLIs.md's "Read next" list plus a short pointer
section, and from Settings-Reference.md's Models section.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
DevvynandClaude Sonnet 5 25fae9ad10 feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment:
"generate those entries from the saved profiles rather than a fixed
duplicate per harness, and put it in a follow-up PR so this one stays the
backend... The Run-menu picker is yours if you want it."

Adds the frontend surface the backend has been waiting on:

- Run menu: a "Custom Endpoints" section lists one entry per (harness that
  supports customModelInjection, saved endpoint) pair, e.g.
  "Claude Code (llama.cpp)". The harness list comes from
  window.__codemanCustomModelClis, injected at page render straight off the
  CLI registry's own capabilities (never a hardcoded id list in the
  frontend), so a CLI whose injection recipe lands later appears with no
  frontend change. Picking an entry runs that harness's own existing run*()
  function unmodified (case creation, env overrides, everything, forced to
  a single instance) and then applies the endpoint's default model to the
  session it creates via the existing POST /api/sessions/:id/custom-model
  route. Entries are hidden for a remote/docker active case, since that
  route already refuses both.
- Settings: App Settings -> Models gets a "Custom model endpoints" group
  wiring up the customModelEndpointsEnabled toggle (declared since #393,
  read by nothing until now) plus CRUD against the existing
  /api/model-endpoints routes: list, add/edit (inline form), delete,
  discover models.
- Backend: CustomModelHost gains an optional defaultModelId, the model the
  picker applies with no further choice per endpoint (one generated menu
  entry per CLI+endpoint pair, not per CLI+endpoint+model). The route
  refuses a value that isn't one of the endpoint's own discovered models,
  and a fresh discovery drops a default that no longer appears rather than
  carrying an invalid one forward.

Docs: docs/custom-model-endpoints.md describes the new picker and settings
panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the
"backend-only" status note and documents the picker's generation mechanism.

Tests: four new route tests cover defaultModelId validation, acceptance,
and the drop/keep behaviour across a re-discovery; a new render-index-html
test pins the __codemanCustomModelClis injection (present, agent CLIs
supporting the capability, antigravity and shell excluded) and its
solo-window skip. No browser test was added for the Run-menu picker itself
or the settings CRUD panel (this box has no tmux, so the live server used
by test:browser/test:mobile could not be exercised here) -- worth a
Playwright pass before merge, same as any other frontend PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
Codeman maintainer bd286bf502 docs(wiki): catch the manual up to 1.29.0 and add the three run modes it never had
The wiki was written for seven run modes and never received Grok Build, DeepSeek
Harness or OMP. They now appear everywhere the others do: the modes table and
per-CLI notes, install commands, environment prefixes, the Quick Start table, the
requirements rows, the vocabulary, and every "seven modes" count.

The 1.27 to 1.29.0 changes land on the pages that own them: attaching a case to an
existing container, multi-case adoption and the copy-a-case picker (Docker Cases);
file reads over ssh in remote cases and what stays unavailable (Remote SSH Sessions,
Working With Files, Security); single-page app routing, frame recovery, localhost
links as tabs and the egress guard (Web Tabs); DeepSeek as the one non-Claude mode
with real stop/blocked signals and Approvals items, Codex's own work detection,
last-response, the model-endpoint routes and refreshed counts (HTTP API, Driving
From An Agent, Hooks, Notifications, Keeping Agents Running, Core Concepts);
Shift+drag, right-click copy, Auto Copy, the Ctrl+Z guard, font weight, the vertical
rail and its activity sort (Keyboard Shortcuts, Input And Voice, The Dashboard,
Settings Reference); the 600px phone cutoff, Codex shift arrows and iPhone Duo
(Mobile Guide); the Docker Compose route and its update rule (Installation, Running
As A Service); four new symptom entries and a "which CLIs" question (Troubleshooting,
FAQ).

Custom model endpoints are deliberately left to #430, which adds that page and edits
Agent CLIs, Settings Reference and the sidebar; these edits stay out of the regions
#430, #428 and #376 touch, and all three still merge cleanly on top.

Both READMEs: the web-tab menu entry is labelled "Add URL" in the UI, not
"Add dashboard".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 19:05:59 +02:00
Codeman maintainer 5b920cb43d feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream
and named tabs wrong five ways (every prompt, every write path, a bare Esc
eating the next prompt's first character, pasted newlines as Enter, any CSI
clearing the draft) and replaced the whole name, which dropped the case from
the tab and reset the w<n> counter. This lands the feature with each of those
closed:

- First prompt means the first: applyAutoName() flips a placeholder to
  `auto` whether or not the string changed. nameSource is now the tri-state
  placeholder | auto | manual; the name setter is the only manual path.
- Only user-originated input counts: write()/writeViaMux() take
  SessionWriteOptions.fromUser, set by the browser WS path and POST /input
  only, so Ralph, respawn, cron, approvals and the trust-dialog keys can
  never name a tab. A startMode 'shell' CLI never feeds the tracker (a
  capability, not an id check); the send-key route feeds trackUserInput()
  because its line feed bypasses the session.
- Prefix form `w3-case: title`: parseSessionPrefix() already renders it as
  the title with the prefix in the tooltip and the next-session counter
  still matches it. Composed within MAX_SESSION_NAME_LENGTH.
- Tracker rules per key: bare Esc resolves at chunk end; mouse/focus
  reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R
  taint the draft so Enter submits nothing rather than a fragment;
  bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space;
  the draft keeps its head past 8192 code points; an escape past 64 bytes
  is abandoned.
- Title: slash commands by shape (a path is a prompt), `!` escapes
  refused, first sentence only past 8 code points ("e.g." is not a title),
  72 code points on a word boundary.
- Synced `autoNameSessions` setting, default OFF (the prompt reaches
  mux-sessions.json, session:updated and /api/search), App Settings ->
  Appearance -> Tabs, read fresh per prompt after the eligibility check.

Tests: test/session-auto-name.test.ts (tracker, title, composition,
ownership, emit gating), the wiring test (once, prefix, setting off,
manual protected), test/routes/session-name-routes.test.ts (PUT /name
flips to manual and persists). Verified live on an isolated instance: API
and browser-typed prompts name the tab, a second prompt does not, shells
and renamed tabs are untouched, nameSource survives a restart.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 17:59:16 +02:00
Codeman maintainer c4322513d9 Merge pull request #376 from shenlvkang-collab/feat/auto-session-names-upstream 2026-09-15 17:19:57 +02:00
Codeman maintainer 70fc6b32d5 docs: record the dup/last input ACK, Shift+drag and right-click copy, and multi-case adopted containers
Three behaviours landed from #375 without their doc entries: the
duplicate input ACK now carries `dup:true` and the server's watermark
(`docs/reliable-input-delivery.md` still described a bare ACK), Shift+drag
and right-click copy in the terminal (the shortcut list did not know
them), and one adopted container backing several cases at different
in-container directories (the Docker cases paragraph still implied one
case per container for adopted containers too).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:56:19 +02:00
Codeman maintainer 942bf37e48 fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a
custom OpenAI-compatible endpoint by injecting env vars or a config file
and restarting the CLI in place. Review of the apply path found four
things, two of them destructive. This lands all four plus the smaller
items from the same review.

1. Clearing a selection did not clear it. The injected vars reach the CLI
   via `tmux setenv`, which persists at the tmux-session level and is
   inherited by `respawn-pane` (measured: `setenv FOO bar` survived two
   successive `respawn-pane -k`), so deleting the keys from the session's
   envOverrides relaunched the CLI still pointed at the old endpoint, and
   for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just
   been deleted. `Session.setCustomModel()` now reports the removed keys,
   queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys`
   carries them into `applyEnvOverrides()`, which `setenv -u`s them before
   re-applying the live overrides, on the same path that already unsets
   the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket
   that `setenv -u HOME` hands the next respawn the global HOME back.

2. Applying a model to a local claude session killed the pane. The
   relaunch was `claude --session-id <id>` and Claude refuses an id that
   already has a transcript, and unlike the dead-pane respawn this one
   kills a working pane first. `restartCli()` now pins the live
   conversation id as the resume id for that respawn when the CLI's launch
   declares a `fallback` chain, which renders the same
   `--resume <id> || --session-id <id>` shape the docker and remote pane
   commands use. Gated on the registry shape, not the CLI id: an entry
   whose resume id is minted by the CLI itself never declares that chain.

3. pi, omp and grok wrote their config file and then launched without the
   `--model` that selects it, so the file was ignored. The registry entry
   now declares `customModelInjection.launchModel` (`custom/{modelId}` for
   pi and omp, grok's `[model.codeman-custom]` block name), the builder
   renders it, and `_withCustomModelLaunchModel()` applies it onto the
   respawn options through `legacyConfigField`, leaving the stored
   <Mode>Config untouched so a clear falls back to the user's own model.
   A model id the CLI's `model` token pattern cannot carry is refused
   with a 400 rather than silently dropped by the argv engine.

4. Remote (SSH) and Docker sessions reported `restarted: true` and changed
   nothing: their `restartCli()` reattaches the durable tmux rather than
   relaunching the agent, and the env lands on the local pane. Both are
   refused with a 400 until those paths are plumbed.

Smaller items from the same review:

- The selection survives a Codeman restart as the disk-only `__customModel`
  bookkeeping (endpoint, model, injected key NAMES, config dir, launch
  model; never the values, which carry the API key). Recovery re-derives
  the values from the endpoint store through the same apply path the route
  uses and keeps the bookkeeping even when the endpoint is gone, so a
  later clear still has keys to unset.
- Discovery goes through `webviewFetch()`, so the RESOLVED address is
  judged by the same egress guard the web-tab proxy uses, and `baseUrl`
  reuses `webviewUrlSchema` (http(s) only, no embedded credentials,
  link-local and cloud-metadata addresses refused). undici's `fetch failed`
  wrapper is unwrapped so the user sees the ECONNREFUSED underneath.
- `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session
  config dir 0700/0600 (pi and omp embed the key literally), and that dir
  is removed with the session.
- `PR.md` is gone from the repo root and the design doc moved to
  `docs/custom-model-endpoints-plan.md` with the LAN address and the
  personal name scrubbed; every reference follows. The guide's `authStyle`
  text matches the shipped schema (`bearer | api-key`, default `bearer`)
  and says that `customModelEndpointsEnabled` is read by nothing until
  the picker lands.
- `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts`
  (four real type errors fixed). It is not yet wired into `npm run typecheck`
  because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json`
  there is the one-line follow-up.

Tests: `test/session-custom-model-restart.test.ts` drives a real Session and
fails on the unfixed code for items 1 to 3; the route suite covers item 4
and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets
run before the overrides and that a shell-metachar key never reaches tmux.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:46:28 +02:00
Codeman maintainer 1e42cb4e2d Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses)
2026-09-14 23:46:27 +02:00
Codeman maintainer e49c48145b fix(files): fail closed on remote symlinks, guard PUT for remote cases, bound ssh fan-out
Follow-up to #421 (remote-case file reads over ssh), addressing the review.

Symlink escape on a host without `readlink -f` (blocker). The probe's
portable fallback canonicalized only the directory chain and returned the
final component unresolved, so on macOS < 12.3 `ws/notes.txt -> ~/.ssh/id_rsa`
came back as `.../ws/notes.txt` (with the target's size), passed every
containment and blocklist check that runs on `realPath`, and `cat` followed
the link. The fallback now walks the directory chain with `cd -P`/`pwd -P`
and follows the LAST component with plain `readlink` for a bounded number of
hops, and anything it cannot fully resolve (a loop, a readlink failure, the
hop cap) is reported with an `x` marker that parses as null, i.e. 404. It
never returns the unresolved string. Measured on a real /bin/sh with
`readlink -f` shadowed: the pre-fix script reports `/ws/notes.txt`, the fixed
one `/secret/id_rsa`; both branches (native and fallback) now agree.

`PUT /api/sessions/:id/file-content` never had the remote guard the PR
described. It sits ahead of `validateSessionFilePath`, which resolves against
the LOCAL filesystem, because with a same-named directory on the Codeman host
(an sshfs mount of the remote tree, the documented stop-gap) the write landed
on the local twin while the viewer believed it edited the remote file.

ssh fan-out is bounded. `src/remote-ssh-limiter.ts` is a
document-conversion-limiter-shaped semaphore (default 4, env
`CODEMAN_MAX_REMOTE_FILE_SSH`) around every probe and buffered read; the
attachment-history list resolves its whole history in ONE batched probe
(`probeRemoteAttachmentHistory`, threaded into
`registerExternalAttachment({remoteProbes})` so the guards run unchanged)
instead of one handshake per entry; and probes chunk at 40 paths because the
whole script is one argv string. Terminal output in a remote session is
written on the remote host, so a prompt-injected agent printing hundreds of
`codeman://attach` links forked one ssh per link, each holding a 20 s
timeout, and a 100-entry history re-listed on every attachment:detected
tripped OpenSSH's default MaxStartups. Streams are deliberately not counted
(one per browser request, held for a whole playback, and gated behind a
counted probe anyway).

Smaller items from the same review: probe records are NUL-terminated and
index-keyed after a leading NUL (a newline in a filename can no longer shift
the alignment, and the banner is fenced off without last-N-lines guessing);
size comes from `stat -c %s || stat -f %z`; the three IO functions refuse
under VITEST instead of opening a connection; an unreachable host now reads
as unknown (missing: false) for detected AND external history entries, where
external used to fold its 502 into missing; a client that aborted during the
guard probe has its body's ssh child reaped (`reply.raw.destroyed` is checked
before the close listener is attached); `describeExecError` never returns
Node's `Command failed: <ssh line>` message, which carried the identity path
and the probe script into a 502 body; and the docs note that
`isSensitivePath`'s three home-anchored entries resolve against the Codeman
host's home, not the remote one.

Tests: the probe script runs on a real /bin/sh with a `readlink` shim that
rejects `-f` (the escape, a relative chain through a symlinked directory, a
loop, a newline filename, banner chatter that itself looks like a record),
the limiter's cap and FIFO order, and route tests for the PUT guard (local
twin untouched, no connection), the single batched history probe, the
unreachable-host alignment and the aborted-client reap. All four route tests
fail against the pre-fix file-routes.ts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:42:06 +02:00
Codeman maintainer 792a251e35 Merge pull request #421 from Randalix/fix/remote-file-access
fix(files): read remote-case previews, downloads and attachments over ssh
2026-09-14 23:42:06 +02:00
Codeman maintainer 6dc27ae727 docs(webview): record the lost-frame page as the third unauthenticated 200, and the inline-style limit
The lost-frame recovery page is answered ahead of the credential checks in both
auth hooks, which makes it the third unauthenticated 200 beside the two hook
routes, and the only one decided by request headers alone. CLAUDE.md's security
table listed exactly two, and docs/web-tabs.md is not where anyone auditing that
looks, so it now has a row in the table and a fourth property in
docs/security-architecture.md section 10b, including the `/` carve-out and its
credential-free condition. Both state the property that comes with it: a
non-browser client can set those headers, so an unauthenticated caller can tell a
registered route (401) from a non-route (200) and enumerate the route table,
accepted because the routes are public in docs/api-reference.md.

docs/web-tabs.md gains the landing-page case in layer 6 and a Known limits entry:
masking trades away the Referer safety net, only HTML is rewritten server-side,
and a root-absolute url() inside an inline <style> block has the masked document
as its Referer, so it 404s where the Referer fallback used to rescue it. External
stylesheets are unaffected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:38:54 +02:00
Codeman maintainer b0dddc9c57 Merge pull request #402 from shenlvkang-collab/pr/webview-route-masking
fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads
2026-09-14 23:38:54 +02:00
Codeman maintainer 2bda191471 docs(docker): describe the root start and drop, and keep the override file out of the image
docker/README.md and docs/docker-compose.md now say that the container starts
as root, corrects a daemon-created bind source and drops to PUID:PGID with
setpriv, which capabilities that needs, and that a compose file written
elsewhere must carry them. The README's PowerShell example runs Compose from
inside docker/ so the override file is discovered, instead of the `-f
docker/docker-compose.yaml` form its own Local customisation section warns
silently drops it, and the reverse-proxy section no longer asks for an override
file now that docker-compose.yaml forwards CODEMAN_ALLOWED_HOSTS itself.

.dockerignore excludes docker-compose.override.* everywhere: it is the
documented home for host-specific settings and rode `COPY . .` into the image,
the same shape as the docker/.env exclusion above it (verified with a scratch
build context: the override files and docker/.env are absent, .env.example and
the compose file present).

CLAUDE.md's Compose paragraph carries the corrected cap list, the writability
probe, and the two traps behind it (KILL is for tini, the CLI prefix is
appended to PATH).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:37:04 +02:00
Codeman maintainer a29e1f61ef Merge pull request #377 from opticon454/bugfix-docker-user-perms
fix(docker): bind-mount ownership, Compose override discovery, and the default runtime account
2026-09-14 23:37:04 +02:00
Randalix 63aafdf274 fix(files): serve remote-case attachments, the path a click takes outside the case
A clicked path that points OUTSIDE the case directory goes through the attachment
routes (the frontend's `_isExternalPreviewPath` sends every absolute path not under
`workingDir` to `POST /attachments`), and those had the same local-`fs` assumption
as file-raw: `realpathSync`/`fs.stat` on a path that only exists on the remote host,
so the file never opened — the case the #415 report was actually about.

- `registerExternalAttachment()` accepts `remote` and resolves through
  `remoteProbePaths` (canonical path, size/mtime, kind, plus the workspace root for
  the confinement check). Everything around it — blocklist, extension allowlist,
  workspace confinement, registry/dedupe — is now shared by both branches, so the
  remote path cannot drift from the local one.
- The by-id routes (`raw`, `preview`, `thumbnail`), the metadata poll and the
  attachment history list resolve over ssh too. `raw` streams with the same
  Range contract as file-raw; `preview` (office) and `thumbnail` answer 400 for a
  remote record; an unreachable host answers 502, a vanished file 404.
- Which host a record is read from follows the SESSION, never the path string: the
  same absolute path is a different file on each host, and a remote session never
  falls back to a local file with that name.
- Codex generated artifacts keep force-workspace confinement for a remote case: the
  well-known artifact directories are anchored at THIS host's home, so only a file
  inside the remote workspace is trusted.

Still local-only by design: writes, office conversion, thumbnails, the file
tree/picker and tail-file.
2026-09-14 17:06:42 +02:00
Ark0N 0c00a40530 Merge pull request #407 from Ark0N/feat/iphone-duo
iPhone Duo support: fold-aware dialogs, and a fold is no longer mistaken for the keyboard
2026-09-14 16:10:38 +02:00
Codeman maintainer 21dcec5d24 test(mobile): follow the 600px phone cut on the Duo branch
Rebased over #390, which moved the phone tier's cutoff from 430px to
600px. The palette's compound fold rule now lives in the 600-768px band
mobile.css pads, the cascade samples the palette inside that band, and
the closed iPhone Duo (466pt) is a phone rather than a small tablet while
the open one (626pt) stays a tablet. Comments in both stylesheets, the
device registry, CLAUDE.md and architecture-invariants say 600.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 16:02:02 +02:00
Codeman maintainer 707ea345eb fix(statusline): GET /api/settings never writes, and a save sends the collection switch only on a flip
Two follow-ups to #361's sticky telemetry switch.

GET /api/settings reconciled an absent showPlanUsageLimits by persisting
true, but readJsonConfig() answers {} for ANY read failure (a parse
error, EACCES, EMFILE, a read landing inside PUT's non-atomic write), not
only ENOENT, and every page load calls this route, so one unlucky read
replaced the whole settings file with a one-key file. The route is a
plain read again and the default moved into the reader:
readPlanUsageTelemetryEnabled() treats an absent key as ON, the same way
readWorkspaceHooksEnabled() does, which is what the desktop chip already
shows for an install that never touched the setting.

saveAppSettings() sent showPlanUsageLimits on every save. The chip
defaults OFF on handhelds, so a phone saving its font size persisted
false and switched collection off for every desktop, whose chip then
went stale with no error anywhere. The key is now stripped like the
other per-device display keys and re-added only when the save FLIPS the
chip relative to what the device had (planUsageCollectionFlip), so an
explicit toggle on any device still writes it in either direction.

Tests pin both: the GET route with a mocked filesystem (absent, missing,
EACCES, garbage, explicit), the reader default, and the flip helper plus
its wiring in saveAppSettings.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:59:00 +02:00
Ark0N b2b2c767ea Merge pull request #361 from timkjr/fix/statusline-injection-opt-out
fix(statusline): inject plan-usage telemetry via ephemeral CLI flag, never disk
2026-09-14 15:58:48 +02:00
Codeman maintainer 2f9fc72252 docs(mobile): record the fold cascade traps and the keyboard-free baseline
CLAUDE.md's folding-devices rule gains the two new invariants (a shape change
with the keyboard up baselines to window.innerHeight; a base gutter overridden
by a later @media block needs its own zero-base fold restatement, and a
compound rule written against a mobile.css shorthand is scoped to that band)
plus the architecture-invariants pointer it lacked; the new Folding devices
section there carries the mechanisms and the measurements. The device count
is 138 since the two Duo profiles landed (68 Playwright + 70 custom).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:57:08 +02:00
Codeman maintainer a5cf1f6005 docs(cli-registry): name the real tests and fields the catalogue docs point at
Three instructions a future contributor would follow literally were stale after
the last review round: the "Adding a CLI" checklist sent the agent-image reason to
AGENT_IMAGE_SPECIAL_CASES, a constant that no longer exists (it is
discovery.install.agentImageLayer on the entry in stock.ts), the trust-boundary
paragraph credited the embedded-commands pin to the invariants test when it is
test/cli-catalog-sync.test.ts, and install.sh claimed "the parity test" pinned the
DeepSeek Harness banner when no test did. That pin now exists: the invariants test
asserts the script's grep literal and the registry's discovery.identity.regex agree
on "DeepSeek Harness", and the comment names it.

docs/docker-cases.md separated the two reasons a CLI stays out of the shared npm
layer (no npmPackage at all versus an agentImageLayer entry), which it had folded
into one, and architecture-invariants no longer lists the agent image's CLI set by
hand.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:55:20 +02:00
Ark0N e6e5a62d9b Merge pull request #380 from opticon454/feature/cli-catalog-consumers
feat(cli-registry): drive install.sh and the Docker agent image from the CLI catalogue
2026-09-14 15:55:08 +02:00
Codeman maintainer c9c8ffddde test(mobile): read PHONE_MAX as an exclusive bound everywhere, drop the stale 430px baselines
Follow-up to #390. PHONE_MAX had become 599, an inclusive bound, while
three of its four consumers still read it as exclusive (width < PHONE_MAX
for phone); the one site that switched to <= disagreed with
getDeviceType(). It is 600 again with < at every site. The breakpoint
table in docs/mobile-testing-report.md says 600, and the three 430px
visual baselines are removed: they depict the tablet tier now, and the
visual suite recreates a missing baseline on its next run on the machine
that owns them. device-matrix.test.ts is also run through Prettier, which
the commit hook demanded and the format gate (src/ only) never did.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:53:22 +02:00
Codeman maintainer c2d019d956 chore: move the maintainer PR bot out of this repository
`scripts/pr-bot/` was maintainer tooling, not part of the server, the CLI or the
npm package: a Telegram bot that reviews open pull requests in Codeman sessions
and reports to the maintainer. It now lives in its own private repository and
keeps running unchanged, as a client of Codeman's HTTP API like any other.

It moved because it grew a second watcher, for GitHub Discussions, and shipping
that here would mean publishing the briefs it hands its review sessions, the
judgement calls in them and its safety model. None of that helps anyone
installing Codeman, and all of it is easier to change when it is not a public
interface. The move cost nothing structurally: the whole tree depended on one
external package plus Node builtins.

What this removes from the repo, and nothing else: the sources, their three test
files, `config/tsconfig.pr-bot.json`, `docs/pr-bot.md`, the `pr-bot` npm script,
the bot's globs in the typecheck/lint/format scripts, and its knip entry. CLAUDE.md
keeps a short pointer in place of the section, because the bot still constrains
work in here: it takes the `prbot-<n>` and `dscbot-<n>` session names on the local
Codeman, holds clones under `~/.codeman/pr-bot/`, and fetches pull-request heads
into `refs/pr-bot/*` of this checkout, which it must never check out or reset.

The CHANGELOG entries from 1.25.0 and earlier still describe it. That is history
rather than drift, and is left alone.

Verified after the removal: typecheck, lint and format:check clean, and the suite
passes 6843 tests across 357 files, which is the previous run minus exactly the
70 tests that moved out with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 15:24:58 +02:00
Randalix 013a5d9cc8 fix(files): read remote-case file previews and downloads over ssh
A remote case's workingDir is an absolute path on the remote host, but the
file read routes resolved it with local `fs`: `validateSessionFilePath`'s
realpathSync fails for a path that does not exist on the Codeman host, so
every preview of an agent-written file answered "File not found" (#415).

Add src/remote-files.ts as the single remote-read layer, built on the same
buildSshConnectionArgs() the launch uses:

- remoteProbePaths(): ONE round trip returning realpath + stat for the
  requested path AND the workspace root, so containment is checked against a
  remotely canonicalized root (a symlinked remotePath is ordinary).
- remoteCreateReadStream(): streams the body (cat, or tail -c +N | head -c L
  for a Range) with nothing buffered in memory, and reaps the ssh child when
  the response ends so an aborted download cannot orphan it.
- remoteReadFile(): bounded read for file-content.

file-raw, file-content, file-preview and file-thumbnail now share one local/
remote target resolution. Guards keep their local strength: lexical pre-check,
remote realpath, workspace containment, sensitive-path blocklist, and the size
cap applied to the remote size before any bytes are read. An unreachable host
answers 502 with the remote reason instead of a misleading 404. Nothing is ever
copied to the Codeman host and there is NO local fallback (an sshfs mount of
the same tree must not shadow the remote bytes).

Deliberately unchanged: writes (edit=1 / PUT now answer 400 explicitly while
the viewer hides its Edit affordance), office previews, thumbnails, file tree,
picker, external attachment registration and tail-file stay local-only.
2026-09-14 14:54:41 +02:00
Codeman maintainer fc098aaab2 docs(plugin): say to pick one install route, since plugin and user-level skill list twice
Measured with both installed: a fresh Claude Code lists `codeman` (the
user-level or per-case copy) and `codeman:codeman` (the plugin). Neither
shadows the other and both work, so this is noise rather than breakage, but
the README, the wiki and the plugin README now say to choose one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:31:26 +02:00
Codeman maintainer f6c08118dc feat(skill): ship the codeman agent skill as a Claude Code plugin from the repo's own marketplace
`.claude-plugin/marketplace.json` at the repo root makes
`/plugin marketplace add Ark0N/Codeman` work, and the one plugin it lists is
the repo itself (`source: "./"`), whose one component is `skills/codeman/`.
So `/plugin install codeman@codeman` is a third install route next to
`npx skills add` and `codeman skill install`, and the skill shows up in the
plugin directories that index Claude Code marketplaces.

Both manifests carry package.json's version: `scripts/sync-plugin-version.mjs`
rewrites them inside `version-packages`, right after `changeset version`, and
`test/plugin-manifest.test.ts` pins the equality, the skill's frontmatter name
(without it the installed skill would be named after a versioned cache dir),
and that no other plugin component (`commands/`, `agents/`, `hooks/`,
`.mcp.json`, `settings.json`) appears at the repo root, since an install would
silently ship it.

Verified with `claude plugin validate` (one expected warning: CLAUDE.md at a
plugin root is not plugin context) and a local marketplace add, install,
details, uninstall cycle against a clean checkout of this commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:24:35 +02:00
Codeman maintainer 48f30f3055 style: drop em-dashes from the text added in c2114615
House style, and these land in the changelog. Only the sentences added in the
previous commit are touched; the em-dashes in contributor text and in the
pre-existing COD-54/COD-115 comments are left alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:46:24 +02:00
Codeman maintainer c211461500 fix: merge-time follow-ups for #409, #404 and #399
#409 (Claude truecolor). The changeset becomes the changelog, and its premise
does not hold on tmux 3.2 or newer. Measured here on tmux 3.4: `default-terminal`
sits at its compiled default of `tmux-256color`, a live claude pane reports
`TERM=tmux-256color`, and supports-color reads that as 256 colors, where
rgb(55,55,55) lands on ESC[48;5;237m — visible, just not the color the theme
named. The invisible block the PR describes needs TERM to resolve to a 16-color
entry: tmux older than 3.2, or a ~/.tmux.conf setting `default-terminal screen`,
which Codeman's own tmux server does read (it passes no -f). Both the changeset
and the invariants paragraph now say that, so the next report here gets paired
with the reporter's tmux -V instead of being read as universal. The change itself
stands on the simpler argument: claude was one of two entries not asking for
truecolor while twelve do.

Also reorders buildClaudeEnv(). It applied the registry's unset/exports AFTER the
whole env was built, so a clis.json entry naming CODEMAN_HOOK_SECRET_FILE or PATH
would strip it on the direct-PTY path while the tmux pane kept it — buildEnvExports()
emits `...cliEnv` ahead of `export CODEMAN_MUX=1` and cannot. The block now runs
first and Codeman's own keys are assigned on top, matching the pane.

#404 (Ctrl+Z trap). Adds the missing changeset, and records what the trap does
not cover: an agent CLI already holds its tty with ISIG off (verified on three
live panes: `susp = ^Z -isig -icanon`), so this is defence for the startup window
rather than a fix for the steady state, and two input paths still reach the PTY
unfiltered — the mobile accessory bar's one-shot Ctrl and the CJK textarea.

#399 (path picker sort). The server sorts by name and cuts at 500, so the client
sorting those 500 by date gives "the newest of the first 500 by name", which is
wrong in exactly the >500-entry folder the date sort exists for. The status line
now says "(first 500 by name)" so the cut is legible, with the reasoning parked
on _sortEntries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:45:41 +02:00
Ark0N e0ebbbdc91 Merge pull request #409 from irisitymichaelgrundberg/fix/claude-truecolor-in-panes
fix(terminal): let Claude use truecolor so its themed backgrounds render
2026-09-14 12:35:28 +02:00
Ark0N 8c237223b0 Merge pull request #399 from shenlvkang-collab/pr/path-picker-sort-jump
feat(files): let the path picker jump to a typed path and sort by name or date
2026-09-14 12:35:23 +02:00
Michael GrundbergandClaude Opus 5 dae2ac580f fix(terminal): read the colour env from the registry on every local spawn path
buildClaudeEnv(), the direct-PTY fallback taken when mux creation fails, now
reads getCli('claude').env and applies its unset and exports lists. It used to
delete COLORTERM and CLAUDECODE from a hand-maintained list of its own, which
left it contradicting the registry entry that the tmux pane and the attach
client both read. An engine value needing a mux name has nothing to resolve
against on this path, so it is skipped rather than guessed.

Claude no longer unsets NO_COLOR. The invisible-background bug does not need
it, and unsetting it overrides a preference the user set deliberately, so a
user who exports NO_COLOR globally keeps monochrome panes. The other seven
truecolor CLIs still unset it; that inconsistency is intentional and the
comment on the entry says so.

The invariants doc gains a Terminal colour env paragraph under Session launch
modes, where a reader looking up Claude will find it — the previous sentence
sat under a heading that lists only the non-Claude CLIs. It now says the lists
are the stock catalog and a clis.json override replaces them wholesale, and
that the declarations reach the tmux pane, its attach client and the direct
PTY but not a remote pane, whose command carries no env exports at all. Docker
hands COLORTERM=truecolor to every mode, including the two the registry says
must unset it.

The changeset named six peer CLIs and there are seven: deepseek also exports
truecolor. A test beside the existing OpenCode assertion pins the new
behaviour, so a future registry edit cannot make the backgrounds vanish again
in silence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 14:42:30 +02:00
Michael GrundbergandClaude Opus 5 7767b16d4f fix(terminal): let Claude use truecolor so its themed backgrounds render
Claude draws the user's own messages as a block of background color, and
inside a Codeman pane that block was invisible. tmux hands each pane
TERM=screen, which supports-color reads as 16 colors, and Claude's registry
entry deleted COLORTERM on top of that. Claude therefore quantized every RGB
color its theme asked for down to the basic palette, where rgb(55, 55, 55)
and every other dark background becomes ESC[40m, the terminal's own black.
Changing the color in a custom Claude theme moved nothing on screen.

Claude now exports COLORTERM=truecolor and unsets NO_COLOR, matching codex,
gemini, antigravity, pi, grok and omp. CLAUDECODE stays unset, because Claude
reads it as a signal that it is running nested inside itself. Both the tmux
session and the attach client read this one registry entry, so they cannot
disagree.

PR #3 introduced the unset in February, citing xterm.js#484 for the claim
that xterm.js mishandles truecolor. xterm.js closed that issue in April 2019,
Codeman now depends on @xterm/xterm 6, and TmuxManager already sets
terminal-overrides ",*:Tc" on its own tmux server, so 24-bit color reaches
the browser today for every CLI that asks for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-13 12:15:28 +02:00