Commit Graph
4 Commits
Author SHA1 Message Date
DevvynandClaude Sonnet 5 bcebc81fcd feat(custom-model): detect llama-swap model conflicts before switching
Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.

1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
   server) via its own GET /running, which plain llama.cpp has no concept of
   at all. New GET /api/model-endpoints/:id/running-status route exposes this
   read-only, for the frontend's polling loop below.

2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
   what llama-swap currently has loaded. If it differs from the requested
   model AND another live session's own customModel selection is actively
   using that loaded model, the apply is refused with a
   {requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
   instead of silently switching. A "confirmed: true" field on the retry
   skips the check. Switching with nothing else affected proceeds
   immediately, no confirmation asked, only ever when there is something to
   warn about.

3. The frontend (runCustomModelEntry) shows a native confirm() naming the
   affected session(s) and the model they'd lose, matching this codebase's
   existing convention for this class of decision (delete case, kill
   session, etc.) rather than a new modal. On a successful apply the response
   also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
   poll shows a sticky "Loading <model>..." toast via the new running-status
   route until llama-swap reports the target model ready (bounded at 2
   minutes), so a prompt sent mid-swap reads as "loading", never as silence
   or an answer from whatever was loaded a moment before.

Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 14:35:07 +08:00
DevvynandClaude Sonnet 5 25f22b9839 test(custom-model): update session-custom-model route test for CLAUDE_CONFIG_DIR isolation
Fixes the CI failure on the last two commits: this route test asserted an
exact envKeys list for a claude-mode apply that predates the
CLAUDE_CONFIG_DIR isolation fix, so it failed on the new CLAUDE_CONFIG_DIR
entry it correctly started appending. Updates the expected list and adds
assertions for the isolated config dir path and the pre-seeded
.claude.json trust-approval file, matching the behavior added in the two
prior commits rather than just tolerating it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 13:23:54 +08:00
Codeman maintainer 942bf37e48 fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a
custom OpenAI-compatible endpoint by injecting env vars or a config file
and restarting the CLI in place. Review of the apply path found four
things, two of them destructive. This lands all four plus the smaller
items from the same review.

1. Clearing a selection did not clear it. The injected vars reach the CLI
   via `tmux setenv`, which persists at the tmux-session level and is
   inherited by `respawn-pane` (measured: `setenv FOO bar` survived two
   successive `respawn-pane -k`), so deleting the keys from the session's
   envOverrides relaunched the CLI still pointed at the old endpoint, and
   for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just
   been deleted. `Session.setCustomModel()` now reports the removed keys,
   queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys`
   carries them into `applyEnvOverrides()`, which `setenv -u`s them before
   re-applying the live overrides, on the same path that already unsets
   the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket
   that `setenv -u HOME` hands the next respawn the global HOME back.

2. Applying a model to a local claude session killed the pane. The
   relaunch was `claude --session-id <id>` and Claude refuses an id that
   already has a transcript, and unlike the dead-pane respawn this one
   kills a working pane first. `restartCli()` now pins the live
   conversation id as the resume id for that respawn when the CLI's launch
   declares a `fallback` chain, which renders the same
   `--resume <id> || --session-id <id>` shape the docker and remote pane
   commands use. Gated on the registry shape, not the CLI id: an entry
   whose resume id is minted by the CLI itself never declares that chain.

3. pi, omp and grok wrote their config file and then launched without the
   `--model` that selects it, so the file was ignored. The registry entry
   now declares `customModelInjection.launchModel` (`custom/{modelId}` for
   pi and omp, grok's `[model.codeman-custom]` block name), the builder
   renders it, and `_withCustomModelLaunchModel()` applies it onto the
   respawn options through `legacyConfigField`, leaving the stored
   <Mode>Config untouched so a clear falls back to the user's own model.
   A model id the CLI's `model` token pattern cannot carry is refused
   with a 400 rather than silently dropped by the argv engine.

4. Remote (SSH) and Docker sessions reported `restarted: true` and changed
   nothing: their `restartCli()` reattaches the durable tmux rather than
   relaunching the agent, and the env lands on the local pane. Both are
   refused with a 400 until those paths are plumbed.

Smaller items from the same review:

- The selection survives a Codeman restart as the disk-only `__customModel`
  bookkeeping (endpoint, model, injected key NAMES, config dir, launch
  model; never the values, which carry the API key). Recovery re-derives
  the values from the endpoint store through the same apply path the route
  uses and keeps the bookkeeping even when the endpoint is gone, so a
  later clear still has keys to unset.
- Discovery goes through `webviewFetch()`, so the RESOLVED address is
  judged by the same egress guard the web-tab proxy uses, and `baseUrl`
  reuses `webviewUrlSchema` (http(s) only, no embedded credentials,
  link-local and cloud-metadata addresses refused). undici's `fetch failed`
  wrapper is unwrapped so the user sees the ECONNREFUSED underneath.
- `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session
  config dir 0700/0600 (pi and omp embed the key literally), and that dir
  is removed with the session.
- `PR.md` is gone from the repo root and the design doc moved to
  `docs/custom-model-endpoints-plan.md` with the LAN address and the
  personal name scrubbed; every reference follows. The guide's `authStyle`
  text matches the shipped schema (`bearer | api-key`, default `bearer`)
  and says that `customModelEndpointsEnabled` is read by nothing until
  the picker lands.
- `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts`
  (four real type errors fixed). It is not yet wired into `npm run typecheck`
  because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json`
  there is the one-line follow-up.

Tests: `test/session-custom-model-restart.test.ts` drives a real Session and
fails on the unfixed code for items 1 to 3; the route suite covers item 4
and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets
run before the overrides and that a shell-metachar key never reaches tmux.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:46:28 +02:00
DevvynandClaude Sonnet 5 41416566aa feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).

- Registry: capabilities.customModelInjection per CLI entry (env /
  configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
  model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
  custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
  (POST /api/sessions/:id/custom-model), reusing the existing
  respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
  CLI's privilegedEnvKeys, closing a pre-existing gap where several were
  already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
  against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
  every CLI's injected values through a real HTTP shape

Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
  a top-level model string + [model_providers.custom]); fixing it then
  surfaced a genuine, documented protocol incompatibility (Codex only
  speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
  don't implement)
- Claude Code's async session-title-generation call validates
  ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
  hangs the whole -p invocation on an unrecognized name; documented for
  chunk 6, worked around in the standalone script only (--bare is NOT
  safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
  and api-key headers) reliably hung a real server; removed the option
  entirely rather than just changing the default

Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00