mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a custom OpenAI-compatible endpoint by injecting env vars or a config file and restarting the CLI in place. Review of the apply path found four things, two of them destructive. This lands all four plus the smaller items from the same review. 1. Clearing a selection did not clear it. The injected vars reach the CLI via `tmux setenv`, which persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so deleting the keys from the session's envOverrides relaunched the CLI still pointed at the old endpoint, and for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just been deleted. `Session.setCustomModel()` now reports the removed keys, queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys` carries them into `applyEnvOverrides()`, which `setenv -u`s them before re-applying the live overrides, on the same path that already unsets the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket that `setenv -u HOME` hands the next respawn the global HOME back. 2. Applying a model to a local claude session killed the pane. The relaunch was `claude --session-id <id>` and Claude refuses an id that already has a transcript, and unlike the dead-pane respawn this one kills a working pane first. `restartCli()` now pins the live conversation id as the resume id for that respawn when the CLI's launch declares a `fallback` chain, which renders the same `--resume <id> || --session-id <id>` shape the docker and remote pane commands use. Gated on the registry shape, not the CLI id: an entry whose resume id is minted by the CLI itself never declares that chain. 3. pi, omp and grok wrote their config file and then launched without the `--model` that selects it, so the file was ignored. The registry entry now declares `customModelInjection.launchModel` (`custom/{modelId}` for pi and omp, grok's `[model.codeman-custom]` block name), the builder renders it, and `_withCustomModelLaunchModel()` applies it onto the respawn options through `legacyConfigField`, leaving the stored <Mode>Config untouched so a clear falls back to the user's own model. A model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. 4. Remote (SSH) and Docker sessions reported `restarted: true` and changed nothing: their `restartCli()` reattaches the durable tmux rather than relaunching the agent, and the env lands on the local pane. Both are refused with a 400 until those paths are plumbed. Smaller items from the same review: - The selection survives a Codeman restart as the disk-only `__customModel` bookkeeping (endpoint, model, injected key NAMES, config dir, launch model; never the values, which carry the API key). Recovery re-derives the values from the endpoint store through the same apply path the route uses and keeps the bookkeeping even when the endpoint is gone, so a later clear still has keys to unset. - Discovery goes through `webviewFetch()`, so the RESOLVED address is judged by the same egress guard the web-tab proxy uses, and `baseUrl` reuses `webviewUrlSchema` (http(s) only, no embedded credentials, link-local and cloud-metadata addresses refused). undici's `fetch failed` wrapper is unwrapped so the user sees the ECONNREFUSED underneath. - `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session config dir 0700/0600 (pi and omp embed the key literally), and that dir is removed with the session. - `PR.md` is gone from the repo root and the design doc moved to `docs/custom-model-endpoints-plan.md` with the LAN address and the personal name scrubbed; every reference follows. The guide's `authStyle` text matches the shipped schema (`bearer | api-key`, default `bearer`) and says that `customModelEndpointsEnabled` is read by nothing until the picker lands. - `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts` (four real type errors fixed). It is not yet wired into `npm run typecheck` because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json` there is the one-line follow-up. Tests: `test/session-custom-model-restart.test.ts` drives a real Session and fails on the unfixed code for items 1 to 3; the route suite covers item 4 and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets run before the overrides and that a shell-metachar key never reaches tmux. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
@@ -1,296 +0,0 @@
|
||||
# feat: Custom Model Endpoint Profiles (local or cloud, all harnesses)
|
||||
|
||||
> **⭐ Shout-out up front:** this feature was partly inspired by — and is a
|
||||
> great fit for — **[Ark0N/Qwen5090](https://github.com/Ark0N/Qwen5090)**,
|
||||
> the maintainer's other project: a one-click Windows / one-command Linux
|
||||
> installer that stands up Qwen3.8-27B locally on an RTX 5090 behind an
|
||||
> OpenAI-compatible API (vLLM / NInfer / llama.cpp). Once this feature lands,
|
||||
> pointing Codeman at a Qwen5090 box is just adding one endpoint entry — no
|
||||
> extra code, no special-casing. Qwen5090 already wires up DeepSeek Harness
|
||||
> and Claude Code as local coding agents itself, which is basically this
|
||||
> feature's idea in miniature. 🙂
|
||||
|
||||
**Status: open for review.** The backend/CLI-injection side (chunks 1-5,
|
||||
7-8) is built, tested, and verified end-to-end against a real server;
|
||||
chunk 6 (frontend toolbar UI) is not yet built, and gemini/deepseek remain
|
||||
documented, unresolved gaps rather than working paths — see
|
||||
[Status](#status) below for exactly what's done and what's still open.
|
||||
|
||||
---
|
||||
|
||||
## What
|
||||
|
||||
Adds a settings-gated (default **OFF**) way to point any Codeman-supported
|
||||
harness — Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP, or
|
||||
Antigravity — at a **custom OpenAI-compatible endpoint** instead of its
|
||||
native cloud backend, for a given session. "Custom endpoint" covers both:
|
||||
|
||||
- **Local hardware**: llama.cpp, Ollama, vLLM, a home GPU rig, or
|
||||
purpose-built on-prem boxes like NVIDIA DGX Spark, AMD Strix Halo
|
||||
(Ryzen AI Max) mini-PCs, or the [Qwen5090](https://github.com/Ark0N/Qwen5090)
|
||||
setup above.
|
||||
- **Cloud**: Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
|
||||
company's self-hosted gateway.
|
||||
|
||||
The user adds an endpoint by base URL (+ optional API key), Codeman
|
||||
discovers its available models via `GET /v1/models`, and a new toolbar
|
||||
picker lets them apply one of those models to a session — which then
|
||||
restarts that session's CLI process pointed at the endpoint.
|
||||
|
||||
**Why discovery instead of asking the user to type a model name:** it turns
|
||||
"go read your inference server's docs to find the exact model identifier it
|
||||
expects" into "pick from a list Codeman already fetched" — one less place
|
||||
for a user to get a name/casing wrong and have a harness fail with an
|
||||
opaque "model not found." It also means this feature works unmodified
|
||||
against **multi-model hosting setups**, not just a single-model server: a
|
||||
gateway like **[llama-swap](https://github.com/mostlygeek/llama-swap)**
|
||||
(or vLLM/LiteLLM/Ollama serving several loaded/loadable models behind one
|
||||
`/v1/models` list) already advertises every model it can hot-swap to, so
|
||||
the toolbar picker becomes a live menu of everything that endpoint can
|
||||
serve — no per-model endpoint entries, no separate configuration step,
|
||||
just "add the gateway once, everything behind it shows up."
|
||||
|
||||
## Why
|
||||
|
||||
The maintainer pays for a Claude Code subscription but also runs a capable
|
||||
local model. Every harness Codeman drives already _has_ its own mechanism
|
||||
for pointing at a custom endpoint (env vars for Claude, a JSON config blob
|
||||
for opencode, a TOML file for Codex, etc.) — Codeman just never exposed a
|
||||
UI for it. Full motivation, the per-CLI recipe table, and the on-prem
|
||||
hardware use cases are written up in **[`deployment_plan.md`](deployment_plan.md)**.
|
||||
|
||||
## How
|
||||
|
||||
- **`src/config/cli-registry/{types,schema,stock}.ts`** — new
|
||||
`capabilities.customModelInjection` field per CLI entry, one of four
|
||||
kinds: `env` (Claude, Gemini, Grok, DeepSeek), `configContentEnv`
|
||||
(opencode, reusing its existing `OPENCODE_CONFIG_CONTENT` mechanism),
|
||||
`configDir` (Codex/Pi/OMP — writes an isolated config file, never touches
|
||||
the user's real one), or `unsupported` (Antigravity — no known mechanism,
|
||||
toolbar entry stays disabled). Declared data-driven per the repo's
|
||||
existing "never branch on CLI id" rule.
|
||||
- **`src/custom-model-injection.ts`** — pure function turning
|
||||
`(CliEntry, endpoint, modelId)` into the real env vars / config content.
|
||||
No IO; a caller writes `configDir` files to disk.
|
||||
- **`src/custom-model-hosts.ts`** + **`src/web/routes/custom-model-routes.ts`** —
|
||||
read/write-array endpoint store (`~/.codeman/custom-model-hosts.json`,
|
||||
same shape as `remote-hosts.ts`) and `GET/POST/PUT/DELETE
|
||||
/api/model-endpoints` + `POST /:id/discover-models`, admin-gated in
|
||||
multi-user mode, SSRF-guarded via the same `isBlockedWebviewUrl()` check
|
||||
web tabs use.
|
||||
- **`src/web/schemas.ts`** — `customModelEndpointsEnabled` (synced, default
|
||||
OFF) + the endpoint payload schema.
|
||||
- **`scripts/test-local-llm-harnesses.ts`** — standalone smoke-test script
|
||||
that spawns each real CLI binary one-shot against a real endpoint and
|
||||
checks it can answer "hello world," independent of the web UI. Reads
|
||||
defaults from a gitignored `scripts/local-llm-test.config.json` (see the
|
||||
committed `.example.json`) so real IPs/keys never land in git.
|
||||
|
||||
### A finding along the way: multi-user privilege hardening
|
||||
|
||||
Building this surfaced that several env vars (`GOOGLE_GEMINI_BASE_URL`,
|
||||
`GROK_BASE_URL`, `CODEX_HOME`, `PI_CONFIG_DIR`, `OPENCODE_CONFIG_CONTENT`,
|
||||
etc.) were **already** reachable via the generic `envOverrides` API today,
|
||||
pre-existing this PR, because Codeman's env allowlist is prefix-based and
|
||||
global. A non-granted multi-user owner could already redirect a session's
|
||||
endpoint/credentials via a plain `envOverrides` field. This PR adds all of
|
||||
them to their CLI's `privilegedEnvKeys` (the existing clamp mechanism
|
||||
`DEEPSEEK_BASE_URL` already used), closing that gap rather than widening it.
|
||||
`CODEX_HOME` and `PI_CONFIG_DIR` are flagged as extra-sensitive: a
|
||||
redirected config dir can restate approval/sandbox policy or, for Pi,
|
||||
redirect to a dir Pi will execute `.pi/extensions` TypeScript from.
|
||||
|
||||
Claude is the deliberate exception: `ANTHROPIC_*` is **not** added to
|
||||
Claude's allowed env prefixes at all, so it stays reachable only through
|
||||
the dedicated, admin-configured, SSRF-guarded custom-model route — never
|
||||
through a plain client-supplied `envOverrides`.
|
||||
|
||||
## Status
|
||||
|
||||
Built in reviewable chunks; ✅ = done and verified (typecheck + lint +
|
||||
format + tests green), ⬜ = not started.
|
||||
|
||||
- ✅ **1. Registry types** — `customModelInjection` capability shape
|
||||
- ✅ **2. Pure injection builder** — `custom-model-injection.ts` + 15 unit tests
|
||||
- ✅ **3. Endpoint store + CRUD routes** — `custom-model-hosts.ts`,
|
||||
`custom-model-routes.ts`, discovery + SSRF guard, 7 route tests
|
||||
- ✅ **4. Settings + security hardening** — `customModelEndpointsEnabled`
|
||||
flag, `privilegedEnvKeys` additions across 7 CLI entries
|
||||
- ✅ **5. Session integration** — `session.customModel` state field,
|
||||
`session.setCustomModel()`/`session.restartCli()` (a generalized,
|
||||
de-restricted `reattachRemote()` reusing the existing `respawn-pane -k`
|
||||
primitive), `POST /api/sessions/:id/custom-model` restart route. 5 new
|
||||
route tests; the existing `session.test.ts`/`session-cleanup.test.ts`
|
||||
suites can't run at all on this Windows dev box (no local `tmux` —
|
||||
confirmed identical on unmodified `master`, not a regression), which is
|
||||
exactly why the container test below matters.
|
||||
- ⬜ **6. Frontend** — settings group, toolbar picker, tab badge
|
||||
- ✅ **7. Mock-server contract tests** — `test/fixtures/mock-openai-server.ts`
|
||||
and `test/custom-model-injection-contract.test.ts`, 10 tests replaying
|
||||
every CLI's real injected values through an HTTP call shaped the way that
|
||||
CLI sends it, against an in-process fake server
|
||||
- ✅ **8. Docs** — `docs/custom-model-endpoints.md` (user guide, HTTP-API-only
|
||||
until chunk 6 lands) + a CLAUDE.md pointer bullet
|
||||
|
||||
Also done outside the chunk list: the standalone
|
||||
`scripts/test-local-llm-harnesses.ts` smoke-test script (now dynamic —
|
||||
reads the live CLI registry rather than a hand-maintained harness list) +
|
||||
its gitignored config file, the on-prem-hardware use-case writeup in
|
||||
`deployment_plan.md` (DGX Spark, Strix Halo, Qwen5090), a
|
||||
`codeman/agent:llm-test` Docker image (all 9 CLI binaries, built from
|
||||
`docker/agent.Dockerfile`), and a **completed real end-to-end run of all 9
|
||||
harnesses** against a live llama-swap server — see Testing below.
|
||||
|
||||
## Testing performed so far
|
||||
|
||||
- `npm run typecheck` — clean after every chunk
|
||||
- `npm run lint` / `npx prettier --check` — clean
|
||||
- `npm test -- test/cli-registry test/custom-model-injection.test.ts
|
||||
test/custom-model-injection-contract.test.ts test/routes/custom-model-routes.test.ts
|
||||
test/routes/session-custom-model.test.ts test/routes/external-cli-bypass-clamp.test.ts` —
|
||||
245+ tests passing, including the existing multi-user clamp suite (no
|
||||
regressions from the `privilegedEnvKeys` additions)
|
||||
- Refactored `scripts/test-local-llm-harnesses.ts` (now `npx tsx`-run, was
|
||||
plain `.mjs`) to import `enabledClis()` and `buildCustomModelInjection()`
|
||||
directly from source instead of keeping a second hand-maintained copy of
|
||||
every CLI's env/config shape — a registry change now needs zero edits to
|
||||
the test script. Extracted the config-dir-write logic shared with the
|
||||
production route into `custom-model-injection-apply.ts` so both places
|
||||
call exactly one implementation.
|
||||
- **Real end-to-end run against the maintainer's live llama-swap server**
|
||||
(`http://10.10.11.241:8080`), inside `codeman/agent:llm-test` (all 9 CLI
|
||||
binaries, built via `docker/agent.Dockerfile`), against the smallest
|
||||
available model (`qwen3.5-0.8b-ud-q8_k_xl`, 1.1GB — picked by parsing the
|
||||
server's own reported model sizes). **Full 9-harness result: claude,
|
||||
opencode, pi, grok, omp all PASS with a genuine "hello world" reply
|
||||
round-tripped through the real endpoint; codex FAILs for a confirmed
|
||||
protocol reason (not a bug — see below); gemini and deepseek reach the
|
||||
server but fail for reasons not yet root-caused; antigravity SKIPs (no
|
||||
known mechanism); all correctly classified by the now-dynamic
|
||||
`scripts/test-local-llm-harnesses.ts`, which reads the live CLI registry
|
||||
rather than a hand-maintained harness list.** Real findings, not
|
||||
simulated:
|
||||
- **opencode: PASS.** Genuinely round-tripped a "hello world" reply
|
||||
through the real endpoint.
|
||||
- **pi: PASS, after two real bugs found and fixed.** `PI_CONFIG_DIR` does
|
||||
nothing for pi at all (grepped pi's entire bundled JS source — the
|
||||
string appears nowhere); the real redirect is the child process's own
|
||||
`HOME`, since pi hardcodes `~/.pi/agent/models.json` with no dedicated
|
||||
override. Separately, pi's `models` field must be an **array** of
|
||||
`{id}` objects, not an object keyed by id (confirmed against pi's own
|
||||
bundled `docs/models.md`) — the object shape silently loaded zero
|
||||
models. Also needs an explicit `--model custom/<id>` on invocation.
|
||||
- **grok: PASS, after the original recipe turned out to be flat-out
|
||||
wrong**, not just unverified — the env-var recipe in this table's first
|
||||
draft (`GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) produced "Not signed
|
||||
in" against a real binary. Researched xAI's actual docs and corrected
|
||||
to a `config.toml` with a `[model.<name>]` block redirected via
|
||||
`GROK_HOME`, with the key riding as an `env_key`-named env var — then
|
||||
confirmed working end-to-end.
|
||||
- **omp: PASS**, after the same two fixes as pi (array-shaped `models`,
|
||||
`HOME`-redirect instead of `PI_CONFIG_DIR`) plus `--model custom/<id>`.
|
||||
Unverified against omp's own official docs (none are bundled in the
|
||||
install), but empirically confirmed working live.
|
||||
- **gemini: confirmed broken, unresolved after real investigation.**
|
||||
Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an
|
||||
undocumented `AuthType.GATEWAY` path with validation requirements a
|
||||
live run never satisfies (`Invalid auth method selected`, regardless of
|
||||
key format). Tried and ruled out: a Google-format dummy key,
|
||||
`GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE`
|
||||
override, and a hand-written `settings.json`. `--skip-trust` is a real,
|
||||
separate fix for a different symptom (an untrusted-folder check
|
||||
silently overriding `--approval-mode yolo`) and is kept, but does not
|
||||
touch this auth failure. Left as an open, documented gap rather than
|
||||
claimed as working.
|
||||
- **deepseek: confirmed reaching the server, still failing, unresolved.**
|
||||
A real run returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)`
|
||||
consistently — the env vars are read (the request reaches the network
|
||||
rather than failing locally), but the root cause was not identified in
|
||||
the time available. By analogy with codex's Responses-API gap, `dsh`
|
||||
may expect DeepSeek's own API response shape rather than a generic
|
||||
OpenAI-compatible one, but this was not confirmed by reading dsh's own
|
||||
bundled source the way the pi/grok questions were resolved. Documented
|
||||
as best-effort/unknown, matching its pre-existing lowest confidence tag.
|
||||
- **codex: real bug found and fixed.** The recipe's TOML shape
|
||||
(`[model].default`) was rejected by a real codex binary ("invalid
|
||||
type: map, expected a string") — codex wants a top-level `model`
|
||||
string plus `[model_providers.custom]`, and the API key rides as an
|
||||
`env_key`-named env var, never a literal TOML field. Fixed in
|
||||
`custom-model-injection.ts`, the standalone script, and both test
|
||||
suites. **Then a second, deeper finding**: codex only speaks the
|
||||
Responses API now (`wire_api = "responses"`, the only value it accepts
|
||||
since dropping `"chat"` support in Feb 2026) — a real run against the
|
||||
now-correctly-shaped config still failed (`Reconnecting...` × 5, then
|
||||
"high demand" errors) because llama-swap doesn't implement
|
||||
`/v1/responses`. This is a genuine, currently-unresolved protocol
|
||||
incompatibility, not a bug in this PR's code — documented prominently
|
||||
in `deployment_plan.md`'s confidence table.
|
||||
- **claude: PASS, after two real bugs found and fixed.** (1) Claude
|
||||
Code's async session-title-generation call also uses
|
||||
`ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude's own
|
||||
internal recognized-model list, printing `[claude-code:unrecognized_model]`
|
||||
and, in `-p` mode, hanging the whole invocation rather than just
|
||||
warning. `--settings '{"autoTitle":false}'` does NOT
|
||||
stop it (confirmed); `--bare` does — the warning still prints, but the
|
||||
real prompt now runs and returns the real answer. ⚠️ `--bare` is only
|
||||
safe for this standalone one-shot test script — it also disables hooks,
|
||||
LSP, plugin sync, and CLAUDE.md auto-discovery, so it must NEVER be
|
||||
applied to a real interactive Codeman session (which depends on hooks
|
||||
for idle detection, trust-dialog auto-accept, etc.). Whether an
|
||||
INTERACTIVE session with a custom model hits the same hang (vs. just a
|
||||
background warning, which would be harmless) is untested — flagged as
|
||||
an open item for chunk 5/6, not assumed either way. (2) A separate,
|
||||
genuinely nasty bug in the test script itself: a `POST` issued right
|
||||
after a `GET` in the same Node process reliably HUNG indefinitely
|
||||
against this real server (reproduced repeatedly: GET alone ~30ms, POST
|
||||
alone ~1-2s, GET-then-immediate-POST times out completely; a 2s pause
|
||||
between them fixed it every time) — looks like Node's fetch/undici
|
||||
reusing a pooled keep-alive connection the server doesn't handle
|
||||
cleanly for a second request right behind a first. Fixed with a 2s
|
||||
pause between the script's discovery GET and its baseline POST. This
|
||||
is a tooling-correctness fix (affects the script's own baseline check),
|
||||
not a claim about how any CLI's own HTTP client behaves.
|
||||
- **Also found and fixed**: an earlier design sent BOTH `Authorization:
|
||||
Bearer` and `api-key` auth header conventions on every discovery/
|
||||
baseline request, on the theory that an unused header is harmless.
|
||||
Live-tested against the real server, sending both reliably HUNG the
|
||||
request (reproduced 3×: either header alone ~500-600ms, both together
|
||||
no response inside 15s). Removed the `'both'` option entirely from
|
||||
`CustomModelAuthStyle` (was `'bearer' | 'api-key' | 'both'`, now just
|
||||
the first two, default `'bearer'`) — in the schema, the store type, the
|
||||
discovery route, and the standalone script (`--auth-style` flag added).
|
||||
This was a real, currently-shipped-in-this-PR bug fixed before it ever
|
||||
reached anyone, not a pre-existing one.
|
||||
|
||||
## Not yet done / open questions for review
|
||||
|
||||
- **Chunk 6 (frontend)** — settings group, toolbar picker, tab badge — is
|
||||
still entirely unbuilt; the feature is currently HTTP-API-only (see
|
||||
`docs/custom-model-endpoints.md`).
|
||||
- **Gemini is confirmed broken end-to-end** (`Invalid auth method
|
||||
selected`, traced to an undocumented `GATEWAY` AuthType gemini-cli
|
||||
selects once `GOOGLE_GEMINI_BASE_URL` is set) — needs upstream
|
||||
investigation before it can be called supported. Documented in full in
|
||||
`deployment_plan.md`'s confidence table rather than silently shipped as
|
||||
working.
|
||||
- **DeepSeek is confirmed reaching the server but failing** with a
|
||||
consistent `HTTP_404`, root cause not identified — documented as
|
||||
best-effort/unknown, same as its pre-existing lowest confidence tag.
|
||||
- **Codex cannot work against a plain OpenAI-Chat-Completions server**
|
||||
(llama.cpp/llama-swap/Ollama/vLLM's default) — it only speaks the
|
||||
Responses API since Feb 2026. This is an external protocol
|
||||
incompatibility, not something this PR can fix; codex support is real
|
||||
only against a Responses-API-compatible endpoint.
|
||||
- Antigravity has no known mechanism at all and stays unsupported.
|
||||
- Chunk 5's session-restart design needs a careful look before merge:
|
||||
switching a session's endpoint restarts its CLI process in place
|
||||
(confirmed acceptable with the maintainer — these harnesses read
|
||||
endpoint config at process start, not per-turn). Whether an INTERACTIVE
|
||||
claude session with a custom model hits the same async-title-generation
|
||||
hang the standalone script worked around with `--bare` (vs. just a
|
||||
harmless background warning) is untested and should be checked before
|
||||
calling claude's chunk 5 support done — `--bare` itself must never be
|
||||
applied to a real interactive session, since it disables hooks Codeman
|
||||
depends on.
|
||||
|
||||
🤖 Generated with [Claude Code](https://claude.com/claude-code)
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"extends": "../tsconfig.json",
|
||||
"compilerOptions": {
|
||||
"rootDir": "..",
|
||||
"noEmit": true,
|
||||
"declaration": false,
|
||||
"declarationMap": false,
|
||||
"sourceMap": false
|
||||
},
|
||||
"include": ["../scripts/test-local-llm-harnesses.ts"]
|
||||
}
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## Context
|
||||
|
||||
Devvyn pays for Claude Code but also runs a capable local model behind an
|
||||
The author pays for Claude Code but also runs a capable local model behind an
|
||||
OpenAI-compatible server (llama.cpp) — and wants the same mechanism to work
|
||||
against a **cloud** OpenAI-compatible endpoint too (e.g. Azure AI Foundry's
|
||||
OpenAI-compatible inference endpoint, OpenRouter, a self-hosted gateway).
|
||||
@@ -60,15 +60,15 @@ also notable for already wiring up DeepSeek Harness and Claude Code as
|
||||
coding agents against that local server itself, which is effectively the
|
||||
same "point a Codeman-supported harness at a local endpoint" idea this
|
||||
feature is generalizing — worth using as a real-world reference/test target
|
||||
once chunk 5 (session integration) exists, alongside Devvyn's own llama.cpp
|
||||
once chunk 5 (session integration) exists, alongside the author's own llama.cpp
|
||||
box.
|
||||
|
||||
Each harness has its own (different-shaped) mechanism for pointing at a
|
||||
custom OpenAI-compatible base URL + model — env vars for Claude, a JSON
|
||||
config blob for opencode, a TOML file for Codex, etc. Devvyn gave the
|
||||
config blob for opencode, a TOML file for Codex, etc. The author gave the
|
||||
starting recipes for those three; the rest (Gemini, Pi, Grok, DeepSeek, OMP,
|
||||
Antigravity) were researched for this plan and are flagged by confidence
|
||||
below. A real end-to-end pass against Devvyn's own llama-swap server
|
||||
below. A real end-to-end pass against the author's own llama-swap server
|
||||
(`scripts/test-local-llm-harnesses.ts`, inside a `codeman/agent:llm-test`
|
||||
Docker image with all 9 CLIs installed) then confirmed **claude and
|
||||
opencode work end-to-end**, corrected a real Codex config.toml schema bug
|
||||
@@ -89,7 +89,7 @@ The feature must be:
|
||||
or a model discovered from one of the configured custom endpoints.
|
||||
- Picking a custom-endpoint model for an **already-running session restarts
|
||||
that session's CLI process** with the injected env/config pointed at that
|
||||
endpoint (confirmed with Devvyn — these harnesses read endpoint config at
|
||||
endpoint (confirmed with the maintainer — these harnesses read endpoint config at
|
||||
process start, not per-turn, so a live hot-swap isn't possible).
|
||||
- **New sessions always default back to the harness's native cloud backend.**
|
||||
A custom-endpoint selection is a per-session override, not a sticky global
|
||||
@@ -269,7 +269,7 @@ mode, same as remote/docker hosts.
|
||||
|
||||
## Mock-server validation strategy (CI-runnable, no real CLI binaries needed)
|
||||
|
||||
Spawning nine real CLI binaries in CI isn't realistic, and neither Devvyn's
|
||||
Spawning nine real CLI binaries in CI isn't realistic, and neither the author's
|
||||
llama.cpp box nor a real cloud subscription can be a CI dependency. So the
|
||||
injection _logic_ gets a tier of automated coverage that sits between the
|
||||
pure unit tests and the live manual checks in Verification:
|
||||
@@ -342,7 +342,7 @@ pure unit tests and the live manual checks in Verification:
|
||||
Codeman's UI entirely, and is DYNAMIC (reads `enabledClis()` + calls the
|
||||
real `buildCustomModelInjection()`, so a future registry change is picked
|
||||
up automatically with zero edits to the script). Already run to
|
||||
completion against Devvyn's llama-swap server (`http://10.10.11.241:8080`,
|
||||
completion against the author's llama-swap server (a LAN address,
|
||||
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
|
||||
claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected**
|
||||
(Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED**
|
||||
@@ -9,7 +9,7 @@ services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
|
||||
company gateway) — anything answering `GET /v1/models` and
|
||||
`POST /v1/chat/completions` in the standard shape. Design doc, per-CLI
|
||||
recipe confidence table, and security reasoning:
|
||||
[`deployment_plan.md`](../deployment_plan.md).
|
||||
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
|
||||
|
||||
> **Status**: backend is implemented and tested (registry capability, the
|
||||
> injection engine, the endpoint store + discovery route, the session
|
||||
@@ -21,7 +21,10 @@ recipe confidence table, and security reasoning:
|
||||
## Turning it on
|
||||
|
||||
App Settings → Agents & CLIs → **Custom Model Endpoints** (synced setting
|
||||
`customModelEndpointsEnabled`, default **OFF**). The API equivalent:
|
||||
`customModelEndpointsEnabled`, default **OFF**). Until the toolbar picker
|
||||
lands, nothing reads this setting: the HTTP routes below work whether it is
|
||||
on or off, and it exists now only so the picker has a switch to hang off
|
||||
when it ships. The API equivalent:
|
||||
|
||||
```bash
|
||||
curl -sk -X PUT https://localhost:3000/api/settings \
|
||||
@@ -38,9 +41,14 @@ curl -sk -X POST https://localhost:3000/api/model-endpoints \
|
||||
```
|
||||
|
||||
`apiKey` is optional (most local servers don't check it). `authStyle`
|
||||
(`bearer` | `api-key` | `both`, default `both`) controls which auth header
|
||||
convention discovery uses — `both` works whether the endpoint is llama.cpp
|
||||
(ignores the header) or a cloud gateway like Azure (wants `api-key`).
|
||||
(`bearer` | `api-key`, default `bearer`) controls which auth header
|
||||
convention discovery uses: `bearer` is `Authorization: Bearer <key>`
|
||||
(llama.cpp, OpenAI-compatible servers, most gateways), `api-key` is the
|
||||
`api-key: <key>` header Azure AI Foundry wants. There is deliberately no
|
||||
"send both" option: measured against a real llama-swap server, a request
|
||||
carrying both headers hung indefinitely. `baseUrl` must be `http(s)`, carry
|
||||
no embedded credentials, and may not point at a link-local or cloud-metadata
|
||||
address; discovery re-checks the address the name actually resolves to.
|
||||
|
||||
Discover its available models:
|
||||
|
||||
@@ -63,17 +71,36 @@ curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
||||
```
|
||||
|
||||
This computes the CLI-specific env vars / config for that session's mode
|
||||
(see the recipe table in `deployment_plan.md`) and **restarts the session's
|
||||
(see the recipe table in `custom-model-endpoints-plan.md`) and **restarts the session's
|
||||
CLI process in place** — same pane, same tmux session, fresh env. That
|
||||
restart is necessary, not incidental: every supported harness reads its
|
||||
endpoint config at process start, not per-turn, so there is no live
|
||||
hot-swap. Clear back to the harness's native cloud default with:
|
||||
hot-swap. A Claude session is relaunched with `--resume <conversation> ||
|
||||
--session-id <id>`, so it continues the conversation it was on; pi, omp and
|
||||
grok are relaunched with the `--model` value that selects the injected
|
||||
provider (`custom/<modelId>` for pi and omp, `codeman-custom` for grok),
|
||||
since for those three the config file alone does not switch the model.
|
||||
**Remote (SSH) and Docker sessions are refused** (400) for now: their restart
|
||||
reattaches the durable remote/in-container tmux rather than relaunching the
|
||||
agent, so the selection would report success and change nothing.
|
||||
|
||||
Clear back to the harness's native cloud default with:
|
||||
|
||||
```bash
|
||||
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
||||
-H 'Content-Type: application/json' -d '{"clear": true}'
|
||||
```
|
||||
|
||||
Clearing also removes the env vars the selection injected from the tmux
|
||||
session (they persist there and would otherwise be inherited by the
|
||||
relaunched CLI) and deletes the per-session config directory
|
||||
(`~/.codeman/custom-model-configs/<sessionId>`, written 0600 because pi and
|
||||
omp embed the API key in it). That directory is also removed when the
|
||||
session is deleted. The selection survives a Codeman restart: the endpoint
|
||||
id, model and injected key NAMES are persisted, the values are re-derived
|
||||
from the endpoint store on recovery, and the pane keeps running against the
|
||||
endpoint in between because tmux retains its environment.
|
||||
|
||||
**New sessions always default back to the harness's native backend.** A
|
||||
custom-endpoint selection is a per-session choice, never a sticky global
|
||||
default — starting a fresh session doesn't inherit whatever the last one was
|
||||
@@ -101,7 +128,7 @@ automatically). Results:
|
||||
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
|
||||
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
|
||||
|
||||
See the confidence table in `deployment_plan.md` for the full detail behind
|
||||
See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind
|
||||
each result. `scripts/test-local-llm-harnesses.ts` is the standalone script
|
||||
used to check a harness against a real endpoint outside the web UI
|
||||
entirely; see its own `--help` for usage.
|
||||
@@ -115,6 +142,6 @@ non-granted multi-user owner cannot set one directly via the generic
|
||||
`envOverrides` API field — only through this feature's own route, which
|
||||
computes the value from an admin-configured, SSRF-guarded endpoint rather
|
||||
than trusting arbitrary client input. See the "Multi-user security
|
||||
hardening" section of `deployment_plan.md` for the full reasoning; several
|
||||
hardening" section of `custom-model-endpoints-plan.md` for the full reasoning; several
|
||||
of these were reachable via the generic `envOverrides` field even before
|
||||
this feature existed, and building this surfaced and closed that gap.
|
||||
|
||||
@@ -401,8 +401,8 @@ async function baselineCheck(
|
||||
signal: AbortSignal.timeout(timeoutMs),
|
||||
});
|
||||
if (!res.ok) throw new Error(`HTTP ${res.status}`);
|
||||
const body = await res.json();
|
||||
const ids: string[] = (body.data ?? []).map((m: { id: string }) => m.id);
|
||||
const body = (await res.json()) as { data?: Array<{ id: string }> };
|
||||
const ids: string[] = (body.data ?? []).map((m) => m.id);
|
||||
console.log(`GET /v1/models -> ${ids.length ? ids.join(', ') : '(empty list)'}`);
|
||||
if (!discoveredModel && ids.length) discoveredModel = ids[0];
|
||||
} catch (err) {
|
||||
@@ -436,7 +436,7 @@ async function baselineCheck(
|
||||
signal: AbortSignal.timeout(timeoutMs),
|
||||
});
|
||||
if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
|
||||
const body = await res.json();
|
||||
const body = (await res.json()) as { choices?: Array<{ message?: { content?: string } }> };
|
||||
const reply: string = body.choices?.[0]?.message?.content ?? '';
|
||||
if (!reply.trim()) throw new Error('empty reply');
|
||||
console.log(`POST /v1/chat/completions -> "${reply.trim().slice(0, 200)}"`);
|
||||
@@ -619,8 +619,8 @@ async function main(): Promise<void> {
|
||||
// an enabled agent CLI) — it's the `unsupported` capability check in runHarness that
|
||||
// skips it, not an exclusion here.
|
||||
const allEntries = enabledClis().filter((e) => e.kind === 'agent');
|
||||
const byId = new Map(allEntries.map((e) => [e.id, e]));
|
||||
const ids = opts.only ?? [...byId.keys()];
|
||||
const byId = new Map<string, CliEntry>(allEntries.map((e) => [e.id as string, e]));
|
||||
const ids: string[] = opts.only ?? [...byId.keys()];
|
||||
const unknownIds = ids.filter((id) => !byId.has(id));
|
||||
if (unknownIds.length) {
|
||||
console.error(`${TAG} unknown harness id(s): ${unknownIds.join(', ')}`);
|
||||
|
||||
@@ -261,6 +261,19 @@ const echoSchema = z
|
||||
})
|
||||
.strict();
|
||||
|
||||
/**
|
||||
* `capabilities.customModelInjection.launchModel`: the `model` launch-param value that
|
||||
* selects the injected provider, with `{modelId}` standing for the chosen id. Bounded to
|
||||
* the characters the `model`/`model-pi` token patterns accept plus the placeholder braces,
|
||||
* so a template can never smuggle a token the argv engine would have to quote.
|
||||
*/
|
||||
const launchModelTemplate = z
|
||||
.string()
|
||||
.min(1)
|
||||
.max(120)
|
||||
.regex(/^[a-zA-Z0-9._\-/:{}]+$/)
|
||||
.optional();
|
||||
|
||||
const capabilitiesSchema = z
|
||||
.object({
|
||||
external: z.boolean(),
|
||||
@@ -326,6 +339,7 @@ const capabilitiesSchema = z
|
||||
// Empty is valid: deepseek's model routing is a profile-composition concern, not
|
||||
// an env var, so it declares baseUrl/apiKey injection with no model var at all.
|
||||
modelVars: z.array(envName).max(8),
|
||||
launchModel: launchModelTemplate,
|
||||
})
|
||||
.strict(),
|
||||
z
|
||||
@@ -333,6 +347,7 @@ const capabilitiesSchema = z
|
||||
kind: z.literal('configContentEnv'),
|
||||
envVar: envName,
|
||||
template: z.literal('opencode-json'),
|
||||
launchModel: launchModelTemplate,
|
||||
})
|
||||
.strict(),
|
||||
z
|
||||
@@ -341,6 +356,7 @@ const capabilitiesSchema = z
|
||||
dirEnvVar: envName,
|
||||
fileName: z.string().min(1).max(80),
|
||||
template: z.enum(['codex-toml', 'pi-models-json', 'omp-models-yml', 'grok-toml']),
|
||||
launchModel: launchModelTemplate,
|
||||
})
|
||||
.strict(),
|
||||
z.object({ kind: z.literal('unsupported') }).strict(),
|
||||
|
||||
@@ -228,7 +228,7 @@ const CLAUDE: CliEntry = {
|
||||
privilegedParams: [],
|
||||
// ANTHROPIC_* is NOT in allowedPrefixes/allowedKeys above (deliberately — see the
|
||||
// allowedPrefixes comment nearby), so these are unreachable via plain envOverrides
|
||||
// today; listed here only so the dedicated custom-model route (deployment_plan.md
|
||||
// today; listed here only so the dedicated custom-model route (docs/custom-model-endpoints-plan.md
|
||||
// chunk 5) clamps them for a non-granted multi-user owner the same way every other
|
||||
// CLI's injection vars are clamped, the day that route widens who can set them.
|
||||
privilegedEnvKeys: [
|
||||
@@ -239,7 +239,7 @@ const CLAUDE: CliEntry = {
|
||||
'ANTHROPIC_DEFAULT_OPUS_MODEL',
|
||||
],
|
||||
gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } },
|
||||
// Custom Model Endpoint Profiles (deployment_plan.md) — verified by hand against a real
|
||||
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — verified by hand against a real
|
||||
// llama.cpp server. Claude reads these at process start only, so switching requires a
|
||||
// respawn, never a live hot-swap.
|
||||
customModelInjection: {
|
||||
@@ -781,6 +781,10 @@ const PI: CliEntry = {
|
||||
dirEnvVar: 'HOME',
|
||||
fileName: '.pi/agent/models.json',
|
||||
template: 'pi-models-json',
|
||||
// Writing models.json is not enough: without `--model custom/<id>` pi stays on its
|
||||
// own default provider and fails with "No API key found for the selected model"
|
||||
// (confirmed live). `custom` is the provider name pi-models-json declares.
|
||||
launchModel: 'custom/{modelId}',
|
||||
},
|
||||
// HOME is not `PI_`-prefixed, so unlike the old (wrong) PI_CONFIG_DIR guess this was
|
||||
// never reachable via the generic envOverrides allowlist at all — listed here anyway,
|
||||
@@ -895,6 +899,10 @@ const GROK: CliEntry = {
|
||||
dirEnvVar: 'GROK_HOME',
|
||||
fileName: 'config.toml',
|
||||
template: 'grok-toml',
|
||||
// The `[model.<name>]` block the grok-toml template writes; `--model <name>` is what
|
||||
// selects it (GROK_CUSTOM_MODEL_NAME in custom-model-injection.ts, pinned equal by
|
||||
// test/custom-model-injection.test.ts so the two cannot drift).
|
||||
launchModel: 'codeman-custom',
|
||||
},
|
||||
// GROK_HOME already matches the GROK_ allowedPrefix above, so it was ALREADY
|
||||
// reachable via plain envOverrides before this feature existed — same reasoning
|
||||
@@ -1189,6 +1197,9 @@ const OMP: CliEntry = {
|
||||
dirEnvVar: 'HOME',
|
||||
fileName: '.omp/agent/models.yml',
|
||||
template: 'omp-models-yml',
|
||||
// Same as pi: omp's own default model has no credential, so without an explicit
|
||||
// `--model custom/<id>` it never reaches the injected provider at all.
|
||||
launchModel: 'custom/{modelId}',
|
||||
},
|
||||
},
|
||||
overlays: {
|
||||
|
||||
@@ -460,7 +460,7 @@ export interface CliCapabilities {
|
||||
/**
|
||||
* How this CLI is pointed at a user-supplied custom OpenAI-compatible
|
||||
* endpoint (local, e.g. llama.cpp, or cloud, e.g. Azure AI Foundry) — the
|
||||
* Custom Model Endpoint Profiles feature (`deployment_plan.md`). Declared
|
||||
* Custom Model Endpoint Profiles feature (`docs/custom-model-endpoints-plan.md`). Declared
|
||||
* per entry, never branched on id, same as every other capability here.
|
||||
*
|
||||
* `env`: plain env vars (claude's `ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY`/
|
||||
@@ -478,7 +478,7 @@ export interface CliCapabilities {
|
||||
* because those env vars are not grok's real custom-endpoint mechanism at
|
||||
* all. The real one is a `[model.<name>]` block in a `config.toml` under
|
||||
* `GROK_HOME` (verified against xAI's own docs), same shape as codex/pi/
|
||||
* omp — this is why the confidence table in deployment_plan.md exists:
|
||||
* omp — this is why the confidence table in docs/custom-model-endpoints-plan.md exists:
|
||||
* "researched" web docs can still be plausible-sounding and wrong.
|
||||
*
|
||||
* Every env var name this introduces that can redirect a session's
|
||||
@@ -486,15 +486,26 @@ export interface CliCapabilities {
|
||||
* `DEEPSEEK_BASE_URL` — a non-granted multi-user owner redirecting a
|
||||
* session to their own endpoint is a credential-exfiltration path, not
|
||||
* just a mischief redirect.
|
||||
*
|
||||
* `launchModel` is the value the entry's own `model` launch param must carry
|
||||
* for the CLI to SELECT the injected provider, as a template where
|
||||
* `{modelId}` is the chosen model id. Writing the config file is not enough
|
||||
* for pi and omp (`--model custom/<id>`, or the CLI stays on its own default
|
||||
* provider and reports "No API key found for the selected model") or for
|
||||
* grok (`--model codeman-custom`, the `[model.<name>]` block the config
|
||||
* declares). Absent = the config alone selects the model (claude's env vars,
|
||||
* opencode's blob, codex's top-level `model` key). Applied by the session's
|
||||
* respawn options through the entry's `legacyConfigField`, never by id.
|
||||
*/
|
||||
customModelInjection:
|
||||
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[] }
|
||||
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json' }
|
||||
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[]; launchModel?: string }
|
||||
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string }
|
||||
| {
|
||||
kind: 'configDir';
|
||||
dirEnvVar: string;
|
||||
fileName: string;
|
||||
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' | 'grok-toml';
|
||||
launchModel?: string;
|
||||
}
|
||||
| { kind: 'unsupported' };
|
||||
}
|
||||
|
||||
@@ -1,8 +1,11 @@
|
||||
/**
|
||||
* @fileoverview Read/write-array store for user-configured custom OpenAI-compatible
|
||||
* model endpoints (local or cloud — deployment_plan.md). Same shape as
|
||||
* `remote-hosts.ts` / `webview-store.ts`: `~/.codeman/custom-model-hosts.json`
|
||||
* holding a plain array, read/written whole.
|
||||
* model endpoints (local or cloud — docs/custom-model-endpoints-plan.md). Same
|
||||
* shape as `remote-hosts.ts` / `webview-store.ts`: `~/.codeman/custom-model-hosts.json`
|
||||
* holding a plain array, read/written whole. The file can hold API keys, so it is
|
||||
* written 0600 via tmp+rename like `intents.json` (`mode` on `writeFile` applies only
|
||||
* to a file being created; the rename is what keeps an existing file's bytes and
|
||||
* mode from ever being observable half-written or world-readable).
|
||||
*/
|
||||
|
||||
import { existsSync, mkdirSync } from 'node:fs';
|
||||
@@ -55,5 +58,8 @@ export async function readCustomModelHosts(configDir: string): Promise<CustomMod
|
||||
|
||||
export async function writeCustomModelHosts(configDir: string, hosts: CustomModelHost[]): Promise<void> {
|
||||
if (!existsSync(configDir)) mkdirSync(configDir, { recursive: true });
|
||||
await fs.writeFile(customModelHostsPath(configDir), JSON.stringify(hosts, null, 2));
|
||||
const target = customModelHostsPath(configDir);
|
||||
const tmp = `${target}.${process.pid}.tmp`;
|
||||
await fs.writeFile(tmp, JSON.stringify(hosts, null, 2), { mode: 0o600 });
|
||||
await fs.rename(tmp, target);
|
||||
}
|
||||
|
||||
@@ -6,13 +6,24 @@
|
||||
* Before this existed, the route and the standalone script each carried
|
||||
* their own copy of this logic, which is exactly the kind of drift the CLI
|
||||
* registry's "declare once, consume everywhere" design exists to prevent —
|
||||
* see deployment_plan.md and the "dynamic to support cli-registry changes"
|
||||
* requirement it was written against.
|
||||
* see docs/custom-model-endpoints-plan.md and the "dynamic to support
|
||||
* cli-registry changes" requirement it was written against.
|
||||
*/
|
||||
|
||||
import { mkdirSync, writeFileSync, rmSync } from 'node:fs';
|
||||
import { chmodSync, mkdirSync, writeFileSync, rmSync } from 'node:fs';
|
||||
import { join, dirname } from 'node:path';
|
||||
import type { ConfigDirInjection } from './custom-model-injection.js';
|
||||
import { dataPath } from './config/instance.js';
|
||||
import type { CliEntry } from './config/cli-registry/types.js';
|
||||
import {
|
||||
buildCustomModelInjection,
|
||||
type ConfigDirInjection,
|
||||
type CustomModelEndpoint,
|
||||
} from './custom-model-injection.js';
|
||||
|
||||
/** Where a session's isolated `configDir`-kind files live: never the user's real CLI config path. */
|
||||
export function customModelConfigDir(sessionId: string): string {
|
||||
return join(dataPath('custom-model-configs'), sessionId);
|
||||
}
|
||||
|
||||
/**
|
||||
* Writes a `ConfigDirInjection`'s files under `baseDir` and returns the full
|
||||
@@ -21,12 +32,18 @@ import type { ConfigDirInjection } from './custom-model-injection.js';
|
||||
* name). Never touches anything outside `baseDir` — the caller is
|
||||
* responsible for choosing an isolated directory (never the user's real
|
||||
* `~/.codex`, `~/.pi`, etc.).
|
||||
*
|
||||
* pi and omp embed the API key literally in the file, so the tree is written
|
||||
* 0700/0600 like every other secret-bearing file under `~/.codeman`; the chmod
|
||||
* covers a re-apply onto a file that already exists (`mode` only applies at
|
||||
* creation).
|
||||
*/
|
||||
export function applyConfigDirInjection(baseDir: string, injection: ConfigDirInjection): Record<string, string> {
|
||||
for (const file of injection.files) {
|
||||
const filePath = join(baseDir, file.relPath);
|
||||
mkdirSync(dirname(filePath), { recursive: true });
|
||||
writeFileSync(filePath, file.content, 'utf8');
|
||||
mkdirSync(dirname(filePath), { recursive: true, mode: 0o700 });
|
||||
writeFileSync(filePath, file.content, { encoding: 'utf8', mode: 0o600 });
|
||||
chmodSync(filePath, 0o600);
|
||||
}
|
||||
return { [injection.dirEnvVar]: baseDir, ...injection.extraEnv };
|
||||
}
|
||||
@@ -40,3 +57,40 @@ export function removeConfigDir(dir: string | undefined): void {
|
||||
// best-effort cleanup only
|
||||
}
|
||||
}
|
||||
|
||||
/** What applying an endpoint to a session yields, ready for `Session.setCustomModel()`. */
|
||||
export interface AppliedCustomModel {
|
||||
envOverrides: Record<string, string>;
|
||||
envKeys: string[];
|
||||
configDir?: string;
|
||||
launchModel?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Compute (and for the `configDir` kind, write) everything a session needs to run
|
||||
* against `endpoint`/`modelId`. Returns undefined for a CLI with no mechanism.
|
||||
*
|
||||
* Idempotent on purpose: the boot-recovery path calls it again for a session that
|
||||
* was already pointed at an endpoint, so the config files are rewritten in place
|
||||
* (same content) and the env values, which are never persisted because they carry
|
||||
* the API key, are re-derived from the endpoint store instead.
|
||||
*/
|
||||
export function applyCustomModelInjection(
|
||||
entry: Pick<CliEntry, 'capabilities'>,
|
||||
endpoint: CustomModelEndpoint,
|
||||
modelId: string,
|
||||
sessionId: string
|
||||
): AppliedCustomModel | undefined {
|
||||
const injection = buildCustomModelInjection(entry, endpoint, modelId);
|
||||
if (injection.kind === 'unsupported') return undefined;
|
||||
if (injection.kind === 'env') {
|
||||
return {
|
||||
envOverrides: injection.envOverrides,
|
||||
envKeys: Object.keys(injection.envOverrides),
|
||||
launchModel: injection.launchModel,
|
||||
};
|
||||
}
|
||||
const configDir = customModelConfigDir(sessionId);
|
||||
const envOverrides = applyConfigDirInjection(configDir, injection);
|
||||
return { envOverrides, envKeys: Object.keys(envOverrides), configDir, launchModel: injection.launchModel };
|
||||
}
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
/**
|
||||
* @fileoverview Pure builder for the Custom Model Endpoint Profiles feature
|
||||
* (deployment_plan.md): turns a CLI registry entry's
|
||||
* (docs/custom-model-endpoints-plan.md): turns a CLI registry entry's
|
||||
* `capabilities.customModelInjection` declaration, a configured endpoint,
|
||||
* and a chosen model id into the concrete env vars / config-file content
|
||||
* that would redirect that CLI's session at the endpoint.
|
||||
@@ -20,7 +20,7 @@
|
||||
* Chat-Completions server (llama.cpp, llama-swap, most local setups) does
|
||||
* NOT implement the Responses API — so codex may still fail at the
|
||||
* PROTOCOL level even with a correctly-shaped config file. That gap is
|
||||
* real and current, not a stale warning; see deployment_plan.md. The rest
|
||||
* real and current, not a stale warning; see docs/custom-model-endpoints-plan.md. The rest
|
||||
* (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION flags
|
||||
* confirmed against real installed binaries' own `--help` output, but
|
||||
* their custom-endpoint env/config conventions remain web-researched,
|
||||
@@ -42,6 +42,8 @@ export interface EnvInjection {
|
||||
kind: 'env';
|
||||
/** Ready to merge into a session's envOverrides. */
|
||||
envOverrides: Record<string, string>;
|
||||
/** See {@link ConfigDirInjection.launchModel}. */
|
||||
launchModel?: string;
|
||||
}
|
||||
|
||||
export interface ConfigDirInjection {
|
||||
@@ -57,6 +59,13 @@ export interface ConfigDirInjection {
|
||||
* or the config points at a credential that was never actually set.
|
||||
*/
|
||||
extraEnv?: Record<string, string>;
|
||||
/**
|
||||
* The value the CLI's `model` launch param must carry for it to SELECT the injected
|
||||
* provider (pi/omp: `custom/<modelId>`; grok: the `[model.<name>]` block name). Absent
|
||||
* when the config alone selects the model. Rendered from the registry entry's
|
||||
* `customModelInjection.launchModel` template, never hand-built per CLI.
|
||||
*/
|
||||
launchModel?: string;
|
||||
}
|
||||
|
||||
export interface UnsupportedInjection {
|
||||
@@ -93,17 +102,21 @@ export function buildCustomModelInjection(
|
||||
[cap.apiKeyVar]: apiKey,
|
||||
};
|
||||
for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId;
|
||||
return { kind: 'env', envOverrides };
|
||||
return withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId);
|
||||
}
|
||||
|
||||
case 'configContentEnv': {
|
||||
const content = renderConfigContent(cap.template, endpoint, modelId, apiKey);
|
||||
return { kind: 'env', envOverrides: { [cap.envVar]: content } };
|
||||
return withLaunchModel({ kind: 'env', envOverrides: { [cap.envVar]: content } }, cap.launchModel, modelId);
|
||||
}
|
||||
|
||||
case 'configDir': {
|
||||
const { content, extraEnv } = renderConfigFile(cap.template, endpoint, modelId, apiKey);
|
||||
return { kind: 'configDir', dirEnvVar: cap.dirEnvVar, files: [{ relPath: cap.fileName, content }], extraEnv };
|
||||
return withLaunchModel(
|
||||
{ kind: 'configDir', dirEnvVar: cap.dirEnvVar, files: [{ relPath: cap.fileName, content }], extraEnv },
|
||||
cap.launchModel,
|
||||
modelId
|
||||
);
|
||||
}
|
||||
|
||||
case 'unsupported':
|
||||
@@ -111,6 +124,16 @@ export function buildCustomModelInjection(
|
||||
}
|
||||
}
|
||||
|
||||
/** Render a `launchModel` template (`{modelId}` = the chosen id) onto an injection result. */
|
||||
function withLaunchModel<T extends EnvInjection | ConfigDirInjection>(
|
||||
result: T,
|
||||
template: string | undefined,
|
||||
modelId: string
|
||||
): T {
|
||||
if (!template) return result;
|
||||
return { ...result, launchModel: template.split('{modelId}').join(modelId) };
|
||||
}
|
||||
|
||||
function renderConfigContent(
|
||||
template: 'opencode-json',
|
||||
endpoint: CustomModelEndpoint,
|
||||
|
||||
@@ -123,6 +123,13 @@ export interface RespawnPaneOptions {
|
||||
resumeSessionId?: string;
|
||||
/** Extra env vars exported before launching the CLI (preserved across respawns). */
|
||||
envOverrides?: Record<string, string>;
|
||||
/**
|
||||
* Env vars to REMOVE from the tmux session (`setenv -u`) before `envOverrides` is
|
||||
* applied. `setenv` persists at the tmux-session level and is inherited by
|
||||
* `respawn-pane`, so a key that merely disappears from `envOverrides` stays set
|
||||
* for the relaunched CLI; clearing a custom-model selection has to name it.
|
||||
*/
|
||||
unsetEnvKeys?: string[];
|
||||
/** Claude CLI effort level (preserved across respawns, injected via `--settings`) */
|
||||
effort?: EffortLevel;
|
||||
/** Original tmux history-limit retained for config parity; respawn cannot resize the existing pane. */
|
||||
|
||||
@@ -254,7 +254,14 @@ export class SessionManager extends EventEmitter {
|
||||
// future reader of state.json.
|
||||
const state = session.toState();
|
||||
const envOverrides = session.getEnvOverridesForPersist();
|
||||
const toStore = envOverrides ? { ...state, __envOverrides: envOverrides } : state;
|
||||
// __customModel: same convention, the disk-only bookkeeping of a custom-model
|
||||
// selection (env KEYS, config dir, launch model; never the injected values).
|
||||
const customModel = session.getCustomModelForPersist();
|
||||
const toStore = {
|
||||
...state,
|
||||
...(envOverrides ? { __envOverrides: envOverrides } : {}),
|
||||
...(customModel ? { __customModel: customModel } : {}),
|
||||
};
|
||||
this.store.setSession(session.id, toStore as SessionState);
|
||||
}
|
||||
|
||||
|
||||
+92
-18
@@ -48,6 +48,8 @@ import {
|
||||
type OpenCodeConfig,
|
||||
type CodexConfig,
|
||||
type EffortLevel,
|
||||
type CustomModelBookkeeping,
|
||||
type CustomModelSelection,
|
||||
type GeminiConfig,
|
||||
type AntigravityConfig,
|
||||
type PiConfig,
|
||||
@@ -577,13 +579,19 @@ export class Session extends EventEmitter {
|
||||
// the CLAUDE_CODE_EFFORT_LEVEL env var, which would hard-lock the session.
|
||||
private _effort: EffortLevel | undefined;
|
||||
|
||||
// Custom Model Endpoint Profiles (deployment_plan.md). `envKeys` and `configDir` are
|
||||
// internal bookkeeping ONLY (never surfaced via toState()/customModel getter): they are
|
||||
// what setCustomModel() needs to undo a previous injection (remove exactly the env keys
|
||||
// it added, delete a previous isolated config dir) without guessing what it once wrote.
|
||||
private _customModel:
|
||||
| { endpointId: string; modelId: string; label?: string; envKeys: string[]; configDir?: string }
|
||||
| undefined;
|
||||
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md). `envKeys`,
|
||||
// `configDir` and `launchModel` are internal bookkeeping ONLY (never surfaced via
|
||||
// toState()/the customModel getter): they are what setCustomModel() needs to undo a
|
||||
// previous injection (remove exactly the env keys it added, delete a previous isolated
|
||||
// config dir) without guessing what it once wrote. Persisted disk-only (`__customModel`).
|
||||
private _customModel: CustomModelBookkeeping | undefined;
|
||||
|
||||
// Env keys a retired custom-model selection injected that the NEXT respawn must
|
||||
// `tmux setenv -u`. Deleting a key from `_envOverrides` alone does nothing to the
|
||||
// tmux session, which keeps every `setenv` and hands it to `respawn-pane`, so the
|
||||
// relaunched CLI would come back still pointed at the old endpoint (measured, see
|
||||
// TmuxManager.applyEnvOverrides). Drained after a successful respawn.
|
||||
private _pendingEnvUnsets = new Set<string>();
|
||||
|
||||
// tmux history-limit (scrollback lines) allocated when this session's pane is created.
|
||||
private readonly _tmuxHistoryLimit: number;
|
||||
@@ -1246,38 +1254,63 @@ export class Session extends EventEmitter {
|
||||
}
|
||||
}
|
||||
|
||||
// Custom Model Endpoint Profiles (deployment_plan.md) — public-safe subset only
|
||||
// (never envKeys/configDir, which are internal bookkeeping for setCustomModel below).
|
||||
get customModel(): { endpointId: string; modelId: string; label?: string } | undefined {
|
||||
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — public-safe
|
||||
// subset only (never envKeys/configDir/launchModel, the bookkeeping for setCustomModel).
|
||||
get customModel(): CustomModelSelection | undefined {
|
||||
if (!this._customModel) return undefined;
|
||||
const { endpointId, modelId, label } = this._customModel;
|
||||
return { endpointId, modelId, label };
|
||||
}
|
||||
|
||||
/**
|
||||
* The full selection incl. bookkeeping, for state.json ONLY (`__customModel`, the
|
||||
* same disk-only convention as `getEnvOverridesForPersist()`). Carries no env values,
|
||||
* so nothing secret lands on disk; recovery re-derives them from the endpoint store.
|
||||
* Without this a Codeman restart left the pane on the custom endpoint (tmux keeps
|
||||
* its `setenv`s) while `customModel` came back undefined, so the state was wrong and
|
||||
* clearing had nothing to unset. Must NOT be included in any API-bound serializer.
|
||||
*/
|
||||
getCustomModelForPersist(): CustomModelBookkeeping | undefined {
|
||||
return this._customModel ? { ...this._customModel, envKeys: [...this._customModel.envKeys] } : undefined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Update this session's custom-model selection and merge the endpoint's injected env
|
||||
* vars into `_envOverrides` — first UNDOING whatever the previous selection injected
|
||||
* (removing exactly those env keys), so switching endpoints, or clearing back to the
|
||||
* harness's native cloud default, never leaves a stale key behind. Synchronous and
|
||||
* side-effect-free beyond mutating state, matching `setNice`/`setColor` above — this
|
||||
* class does no file IO, so it returns the PREVIOUS `configDir` (if any) for the
|
||||
* class does no file IO, so it reports the PREVIOUS `configDir` (if any) for the
|
||||
* caller to clean up on disk (custom-model-injection.ts's configDir kind).
|
||||
*
|
||||
* Keys the previous selection injected that the new one does not re-set are queued
|
||||
* for `tmux setenv -u` on the next respawn (`_pendingEnvUnsets`, threaded through
|
||||
* `_buildRespawnPaneOptions().unsetEnvKeys`): the tmux session inherits every
|
||||
* `setenv` into `respawn-pane`, so dropping them from the map alone would relaunch
|
||||
* the CLI still pointed at the old endpoint — and for the `configDir` kinds, at a
|
||||
* `HOME`/`CODEX_HOME`/`GROK_HOME` the caller has just deleted.
|
||||
*/
|
||||
setCustomModel(
|
||||
next: { endpointId: string; modelId: string; label?: string; envKeys: string[]; configDir?: string } | undefined,
|
||||
next: CustomModelBookkeeping | undefined,
|
||||
envOverrides?: Record<string, string>
|
||||
): string | undefined {
|
||||
): { removedEnvKeys: string[]; previousConfigDir: string | undefined } {
|
||||
const previousConfigDir = this._customModel?.configDir;
|
||||
const removedEnvKeys: string[] = [];
|
||||
if (this._customModel) {
|
||||
for (const key of this._customModel.envKeys) {
|
||||
if (this._envOverrides) delete this._envOverrides[key];
|
||||
removedEnvKeys.push(key);
|
||||
this._pendingEnvUnsets.add(key);
|
||||
}
|
||||
}
|
||||
this._customModel = next;
|
||||
this._customModel = next ? { ...next, envKeys: [...next.envKeys] } : undefined;
|
||||
if (envOverrides && Object.keys(envOverrides).length > 0) {
|
||||
this._envOverrides = { ...(this._envOverrides ?? {}), ...envOverrides };
|
||||
// A key the new selection sets again does not need an unset (applyEnvOverrides
|
||||
// would set it right back anyway); keep the list to what actually goes away.
|
||||
for (const key of Object.keys(envOverrides)) this._pendingEnvUnsets.delete(key);
|
||||
}
|
||||
return previousConfigDir;
|
||||
return { removedEnvKeys, previousConfigDir };
|
||||
}
|
||||
|
||||
// Token tracking getters and setters
|
||||
@@ -1660,6 +1693,7 @@ export class Session extends EventEmitter {
|
||||
console.error('[Session] Failed to respawn pane, will create new session');
|
||||
needsNewSession = true;
|
||||
} else {
|
||||
this._pendingEnvUnsets.clear();
|
||||
// Wait a moment for the respawned process to fully start
|
||||
await new Promise((resolve) => setTimeout(resolve, MUX_STARTUP_DELAY_MS));
|
||||
}
|
||||
@@ -1756,7 +1790,7 @@ export class Session extends EventEmitter {
|
||||
/**
|
||||
* Kill and relaunch this session's CLI process IN PLACE — same pane, same tmux
|
||||
* session, fresh env/args from current state. Custom Model Endpoint Profiles
|
||||
* (deployment_plan.md) is the first caller: after `setCustomModel()` merges new
|
||||
* (docs/custom-model-endpoints-plan.md) is the first caller: after `setCustomModel()` merges new
|
||||
* env vars into `_envOverrides`, the running CLI process still has the OLD env
|
||||
* (inherited at its own process start, not live-reloaded), so switching a
|
||||
* session's model/endpoint requires this restart to actually take effect.
|
||||
@@ -1783,11 +1817,26 @@ export class Session extends EventEmitter {
|
||||
}
|
||||
|
||||
this._pinOmpRespawnId();
|
||||
const newPid = await mux.respawnPane(this._buildRespawnPaneOptions());
|
||||
const options = this._buildRespawnPaneOptions();
|
||||
// Unlike the dead-pane respawn, this one kills a WORKING pane whose conversation
|
||||
// already has a transcript, and a CLI that launches with `--session-id <id>` refuses
|
||||
// an id that is already in use (claude: `Error: Session ID ... is already in use.`),
|
||||
// which turned an endpoint switch into a dead pane and a lost session. A launch that
|
||||
// declares a `fallback` chain renders `resume || new` once a resume id is set, the
|
||||
// same `--resume <id> || --session-id <id>` shape the docker and remote pane commands
|
||||
// already use, so pin the live conversation id for THIS respawn only. The registry
|
||||
// shape is the gate, not the CLI's name: an entry whose resume id is minted by the
|
||||
// CLI itself (codex/pi/omp/grok) never declares that chain, and its resume field is
|
||||
// read from its own `<Mode>Config` rather than this top-level one anyway.
|
||||
if (!options.resumeSessionId && getCli(this.mode)?.launch.chain === 'fallback') {
|
||||
options.resumeSessionId = this._claudeSessionId ?? this.id;
|
||||
}
|
||||
const newPid = await mux.respawnPane(options);
|
||||
if (!newPid) {
|
||||
console.error('[Session] restartCli: respawnPane failed for', this._muxSession.muxName);
|
||||
return false;
|
||||
}
|
||||
this._pendingEnvUnsets.clear();
|
||||
console.log('[Session] restartCli: restarted CLI for', this._muxSession.muxName, 'pid', newPid);
|
||||
return true;
|
||||
}
|
||||
@@ -1799,7 +1848,7 @@ export class Session extends EventEmitter {
|
||||
* the spawn path.
|
||||
*/
|
||||
private _buildRespawnPaneOptions(): import('./mux-interface.js').RespawnPaneOptions {
|
||||
return {
|
||||
const options: import('./mux-interface.js').RespawnPaneOptions = {
|
||||
sessionId: this.id,
|
||||
workingDir: this.workingDir,
|
||||
mode: this.mode,
|
||||
@@ -1827,12 +1876,37 @@ export class Session extends EventEmitter {
|
||||
ompConfig: this._ompConfig,
|
||||
resumeSessionId: this._resumeSessionId,
|
||||
envOverrides: this._envOverrides,
|
||||
unsetEnvKeys: this._pendingEnvUnsets.size > 0 ? [...this._pendingEnvUnsets] : undefined,
|
||||
effort: this._effort,
|
||||
historyLimit: this._tmuxHistoryLimit,
|
||||
remote: this._remote,
|
||||
docker: this._docker,
|
||||
owner: this._owner,
|
||||
};
|
||||
return this._withCustomModelLaunchModel(options);
|
||||
}
|
||||
|
||||
/**
|
||||
* Force the custom-model selection's `launchModel` (pi/omp `custom/<id>`, grok's
|
||||
* `[model.<name>]` block name) onto the CLI's `model` launch param. Where that param
|
||||
* lives is registry DATA — the entry's `legacyConfigField` (`piConfig`, `grokConfig`,
|
||||
* ...) or the top-level `model` for an entry that declares none — so this stays a
|
||||
* generic reader rather than a branch per CLI. Applied on the OPTIONS only: the stored
|
||||
* `<Mode>Config` keeps whatever model the user chose at create, which is exactly what a
|
||||
* later clear must fall back to.
|
||||
*/
|
||||
private _withCustomModelLaunchModel(
|
||||
options: import('./mux-interface.js').RespawnPaneOptions
|
||||
): import('./mux-interface.js').RespawnPaneOptions {
|
||||
const launchModel = this._customModel?.launchModel;
|
||||
if (!launchModel) return options;
|
||||
const entry = getCli(this.mode);
|
||||
if (!entry) return options;
|
||||
const field = entry.launch.legacyConfigField;
|
||||
if (!field) return { ...options, model: launchModel };
|
||||
const bag = options as unknown as Record<string, unknown>;
|
||||
const existing = (bag[field] ?? {}) as Record<string, unknown>;
|
||||
return { ...options, [field]: { ...existing, model: launchModel } };
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
+26
-11
@@ -1738,21 +1738,34 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
|
||||
* Key validation is strict (`/^[A-Z_][A-Z0-9_]*$/`) as defense-in-depth against
|
||||
* shell-metachar injection even if upstream schema check is bypassed.
|
||||
*/
|
||||
private applyEnvOverrides(muxName: string, envOverrides?: Record<string, string>): void {
|
||||
private applyEnvOverrides(muxName: string, envOverrides?: Record<string, string>, unsetKeys?: string[]): void {
|
||||
const VALID_KEY = /^[A-Z_][A-Z0-9_]*$/;
|
||||
// Legacy cleanup: pre-0.7.2 set CLAUDE_CODE_EFFORT_LEVEL via setenv, which persists
|
||||
// on the tmux session and hard-locks /effort switching in every respawned pane.
|
||||
// Effort now flows as a `--settings` soft default (see buildEffortSettingsFlag),
|
||||
// so unconditionally unset the stale var before applying current overrides.
|
||||
try {
|
||||
execSync(`${this.tmux()} setenv -t ${shellescape(muxName)} -u CLAUDE_CODE_EFFORT_LEVEL`, {
|
||||
timeout: EXEC_TIMEOUT_MS,
|
||||
stdio: ['pipe', 'pipe', 'pipe'],
|
||||
});
|
||||
} catch {
|
||||
/* Non-critical — var may not exist */
|
||||
//
|
||||
// The caller's own unsets ride the same path, and run BEFORE the overrides are
|
||||
// (re)applied: a key that is both unset and present in `envOverrides` ends up set,
|
||||
// so a stale unset can never clobber a live value. Removing a key from the map is
|
||||
// not enough on its own — `setenv` persists at the tmux-session level and is
|
||||
// inherited by `respawn-pane`, measured: `setenv FOO bar` survived two successive
|
||||
// `respawn-pane -k`. Clearing a custom-model selection is what needs this.
|
||||
for (const key of ['CLAUDE_CODE_EFFORT_LEVEL', ...(unsetKeys ?? [])]) {
|
||||
if (!VALID_KEY.test(key)) {
|
||||
console.warn(`[TmuxManager] Skipping invalid env unset key: ${JSON.stringify(key)}`);
|
||||
continue;
|
||||
}
|
||||
try {
|
||||
execSync(`${this.tmux()} setenv -t ${shellescape(muxName)} -u ${key}`, {
|
||||
timeout: EXEC_TIMEOUT_MS,
|
||||
stdio: ['pipe', 'pipe', 'pipe'],
|
||||
});
|
||||
} catch {
|
||||
/* Non-critical — var may not exist */
|
||||
}
|
||||
}
|
||||
if (!envOverrides) return;
|
||||
const VALID_KEY = /^[A-Z_][A-Z0-9_]*$/;
|
||||
for (const [key, value] of Object.entries(envOverrides)) {
|
||||
if (!value) continue; // Skip empty — nothing to set
|
||||
if (!VALID_KEY.test(key)) {
|
||||
@@ -2208,6 +2221,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
|
||||
ompConfig,
|
||||
resumeSessionId,
|
||||
envOverrides,
|
||||
unsetEnvKeys,
|
||||
effort,
|
||||
remote,
|
||||
docker,
|
||||
@@ -2269,8 +2283,9 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
|
||||
);
|
||||
this._configureStatusLineUserCommand(muxName, userStatusLineCommand);
|
||||
|
||||
// Re-apply user env overrides before respawn so the new shell inherits them.
|
||||
this.applyEnvOverrides(muxName, envOverrides);
|
||||
// Re-apply user env overrides before respawn so the new shell inherits them,
|
||||
// dropping the ones the caller retired first (see applyEnvOverrides).
|
||||
this.applyEnvOverrides(muxName, envOverrides, unsetEnvKeys);
|
||||
|
||||
// -c /tmp + cd bounce — see createSession() for rationale (stale FUSE state).
|
||||
const launchCmd = remote || docker ? fullCmd : `cd ${JSON.stringify(workingDir)} && ${fullCmd}`;
|
||||
|
||||
+30
-5
@@ -558,6 +558,29 @@ export interface SessionAttachmentHistoryItem {
|
||||
/**
|
||||
* Current state of a session
|
||||
*/
|
||||
/** The public half of a session's custom-model selection (on the wire, in `SessionState`). */
|
||||
export interface CustomModelSelection {
|
||||
endpointId: string;
|
||||
modelId: string;
|
||||
label?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* The full custom-model selection a session keeps: the public selection plus the
|
||||
* bookkeeping `Session.setCustomModel()` needs to UNDO it later without guessing what
|
||||
* it once wrote. Persisted to state.json only as the disk-only `__customModel` field
|
||||
* (never broadcast); the injected env VALUES are not in here at all, since they carry
|
||||
* the endpoint's API key, and are re-derived from the endpoint store on recovery.
|
||||
*/
|
||||
export interface CustomModelBookkeeping extends CustomModelSelection {
|
||||
/** Env keys the selection injected into the session's envOverrides / tmux session. */
|
||||
envKeys: string[];
|
||||
/** Isolated per-session config directory written for a `configDir`-kind CLI. */
|
||||
configDir?: string;
|
||||
/** Value forced onto the CLI's `model` launch param (pi/omp `custom/<id>`, grok's block name). */
|
||||
launchModel?: string;
|
||||
}
|
||||
|
||||
export interface SessionState {
|
||||
/** Unique session identifier */
|
||||
id: string;
|
||||
@@ -678,12 +701,14 @@ export interface SessionState {
|
||||
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
|
||||
effort?: EffortLevel;
|
||||
/**
|
||||
* Custom Model Endpoint Profiles (deployment_plan.md): the custom OpenAI-compatible
|
||||
* endpoint (local or cloud) this session's CLI is currently pointed at, if any.
|
||||
* Undefined = the harness's native cloud default. No secrets here — the endpoint's
|
||||
* base URL/api key live only in Session._envOverrides, never in this public state.
|
||||
* Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): the custom
|
||||
* OpenAI-compatible endpoint (local or cloud) this session's CLI is currently pointed
|
||||
* at, if any. Undefined = the harness's native cloud default. No secrets here — the
|
||||
* endpoint's base URL/api key live only in Session._envOverrides, never in this public
|
||||
* state. The internal half (which env keys were injected, which config dir was
|
||||
* written) is {@link CustomModelBookkeeping}, persisted disk-only like `__envOverrides`.
|
||||
*/
|
||||
customModel?: { endpointId: string; modelId: string; label?: string };
|
||||
customModel?: CustomModelSelection;
|
||||
/** Sanitized per-session attachment history. */
|
||||
attachmentHistory?: SessionAttachmentHistoryItem[];
|
||||
/**
|
||||
|
||||
@@ -1,16 +1,17 @@
|
||||
/**
|
||||
* @fileoverview Custom Model Endpoint Profiles CRUD + discovery
|
||||
* (deployment_plan.md). Endpoints are machine-level infra, like remote/docker
|
||||
* hosts, so writes are admin-only in multi-user mode
|
||||
* (docs/custom-model-endpoints-plan.md). Endpoints are machine-level infra,
|
||||
* like remote/docker hosts, so writes are admin-only in multi-user mode
|
||||
* (`case-routes.ts`'s `/api/remote-hosts` is the pattern this mirrors).
|
||||
*
|
||||
* Discovery (`POST /:id/discover-models`) fetches `${baseUrl}/v1/models`.
|
||||
* `isBlockedWebviewUrl()` is the same synchronous hostname/link-local/cloud-
|
||||
* metadata check `webview-egress-policy.ts` uses for saved dashboard URLs —
|
||||
* reused here as a save-time and discover-time guard. It does NOT re-check
|
||||
* the DNS-RESOLVED address the way `webviewFetch()`'s undici lookup hook
|
||||
* does; wiring that dispatcher-level guard here is a followup, not done in
|
||||
* this pass, since this route is already admin-only in multi-user mode.
|
||||
* Discovery (`POST /:id/discover-models`) fetches `${baseUrl}/v1/models`
|
||||
* through `webviewFetch()` (`webview-egress.ts`), the same guarded dispatcher
|
||||
* the web-tab proxy uses: `baseUrl` is refused at save time by the schema's
|
||||
* hostname check (link-local / cloud-metadata literals and names), and the
|
||||
* undici lookup hook refuses a name that RESOLVES into one of those ranges at
|
||||
* connect time, redirects included — a save-time hostname check alone would
|
||||
* let `models.example` resolve to 169.254.169.254 later. The endpoint is
|
||||
* admin-configured, so this is defence in depth rather than the only gate.
|
||||
*/
|
||||
|
||||
import type { FastifyInstance, FastifyRequest } from 'fastify';
|
||||
@@ -19,6 +20,7 @@ import { isAdmin, parseBody } from '../route-helpers.js';
|
||||
import { isMultiUserMode } from '../../config/multiuser.js';
|
||||
import { getDataDir } from '../../config/instance.js';
|
||||
import { isBlockedWebviewUrl } from '../webview-egress-policy.js';
|
||||
import { egressBlockedReason, webviewFetch } from '../webview-egress.js';
|
||||
import { CustomModelHostSchema } from '../schemas.js';
|
||||
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../../custom-model-hosts.js';
|
||||
|
||||
@@ -40,7 +42,7 @@ async function discoverModels(host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' |
|
||||
if (apiKey && style === 'bearer') headers.Authorization = `Bearer ${apiKey}`;
|
||||
if (apiKey && style === 'api-key') headers['api-key'] = apiKey;
|
||||
|
||||
const res = await fetch(`${host.baseUrl.replace(/\/+$/, '')}/v1/models`, {
|
||||
const res = await webviewFetch(new URL(`${host.baseUrl.replace(/\/+$/, '')}/v1/models`), {
|
||||
headers,
|
||||
signal: AbortSignal.timeout(DISCOVER_TIMEOUT_MS),
|
||||
});
|
||||
@@ -49,6 +51,21 @@ async function discoverModels(host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' |
|
||||
return (body.data ?? []).map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0);
|
||||
}
|
||||
|
||||
/**
|
||||
* undici reports every network failure as `TypeError('fetch failed', { cause })`, with the
|
||||
* useful part (`connect ECONNREFUSED 127.0.0.1:8080`) one level down; surface the deepest
|
||||
* message so the user sees the refused connection, not the wrapper.
|
||||
*/
|
||||
function describeFetchError(err: unknown): string {
|
||||
let message = err instanceof Error ? err.message : String(err);
|
||||
let current: unknown = err;
|
||||
for (let depth = 0; depth < 5 && current instanceof Error && current.cause !== undefined; depth++) {
|
||||
current = current.cause;
|
||||
if (current instanceof Error && current.message) message = current.message;
|
||||
}
|
||||
return message;
|
||||
}
|
||||
|
||||
export function registerCustomModelRoutes(app: FastifyInstance): void {
|
||||
app.get('/api/model-endpoints', async (req) =>
|
||||
isMultiUserMode() && !isAdmin(req) ? [] : readCustomModelHosts(CODEMAN_CONFIG_DIR)
|
||||
@@ -118,9 +135,10 @@ export function registerCustomModelRoutes(app: FastifyInstance): void {
|
||||
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next);
|
||||
return { success: true, data: { models } };
|
||||
} catch (err) {
|
||||
const blocked = egressBlockedReason(err);
|
||||
return createErrorResponse(
|
||||
ApiErrorCode.OPERATION_FAILED,
|
||||
`Could not reach endpoint: ${err instanceof Error ? err.message : String(err)}`
|
||||
blocked ? `Endpoint refused: ${blocked}` : `Could not reach endpoint: ${describeFetchError(err)}`
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -54,8 +54,8 @@ import {
|
||||
CustomModelSelectionSchema,
|
||||
} from '../schemas.js';
|
||||
import { readCustomModelHosts } from '../../custom-model-hosts.js';
|
||||
import { buildCustomModelInjection } from '../../custom-model-injection.js';
|
||||
import { applyConfigDirInjection, removeConfigDir } from '../../custom-model-injection-apply.js';
|
||||
import { applyCustomModelInjection, removeConfigDir } from '../../custom-model-injection-apply.js';
|
||||
import { matchesPattern } from '../../config/cli-registry/patterns.js';
|
||||
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
|
||||
import { TabLayoutValidationError } from '../../tab-layout.js';
|
||||
import {
|
||||
@@ -1156,7 +1156,7 @@ export function registerSessionRoutes(
|
||||
return { color: session.color };
|
||||
});
|
||||
|
||||
// ========== Custom Model Endpoint Profiles (deployment_plan.md) ==========
|
||||
// ========== Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) ==========
|
||||
//
|
||||
// Applies (or clears) a session's custom OpenAI-compatible endpoint selection and
|
||||
// RESTARTS the pane's CLI process — these harnesses read endpoint config at process
|
||||
@@ -1165,17 +1165,30 @@ export function registerSessionRoutes(
|
||||
// (chunk 3's CRUD routes), never raw client-supplied env — that's what keeps this
|
||||
// route safe to let any session owner call for their own session, unlike the
|
||||
// generic envOverrides field the privilegedEnvKeys clamp exists to guard.
|
||||
//
|
||||
// ⚠️ Local sessions only for now. A remote session's `restartCli()` renders
|
||||
// `ssh ... tmux new-session -A`, which reattaches the durable remote tmux rather than
|
||||
// restarting the agent, and the env lands on the LOCAL pane running ssh, which
|
||||
// forwards nothing; docker is the same attach-or-create shape. Both used to answer
|
||||
// `restarted: true` and change nothing, so they are refused until those paths are
|
||||
// plumbed (the env would have to ride the remote/in-container launch command).
|
||||
app.post('/api/sessions/:id/custom-model', async (req) => {
|
||||
const { id } = req.params as { id: string };
|
||||
const body = parseBody(CustomModelSelectionSchema, req.body, 'Invalid request body');
|
||||
const session = findSessionOrFail(ctx, id, req);
|
||||
|
||||
if (session.remote || session.docker) {
|
||||
return createErrorResponse(
|
||||
ApiErrorCode.INVALID_INPUT,
|
||||
'Custom model endpoints are not supported for remote (SSH) or Docker sessions yet'
|
||||
);
|
||||
}
|
||||
if (session.isBusy()) {
|
||||
return createErrorResponse(ApiErrorCode.SESSION_BUSY, 'Session is busy');
|
||||
}
|
||||
|
||||
if ('clear' in body) {
|
||||
const previousConfigDir = session.setCustomModel(undefined);
|
||||
const { previousConfigDir } = session.setCustomModel(undefined);
|
||||
removeConfigDir(previousConfigDir);
|
||||
const restarted = await session.restartCli();
|
||||
persistAndBroadcastSession(ctx, session);
|
||||
@@ -1196,32 +1209,41 @@ export function registerSessionRoutes(
|
||||
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
|
||||
}
|
||||
|
||||
const injection = buildCustomModelInjection(entry, endpoint, body.modelId);
|
||||
|
||||
let envOverrides: Record<string, string>;
|
||||
let envKeys: string[];
|
||||
let configDir: string | undefined;
|
||||
|
||||
if (injection.kind === 'env') {
|
||||
envOverrides = injection.envOverrides;
|
||||
envKeys = Object.keys(injection.envOverrides);
|
||||
} else if (injection.kind === 'configDir') {
|
||||
// Isolated per-session dir — never the user's real CLI config path.
|
||||
configDir = join(dataPath('custom-model-configs'), session.id);
|
||||
envOverrides = applyConfigDirInjection(configDir, injection);
|
||||
envKeys = Object.keys(envOverrides);
|
||||
} else {
|
||||
// 'unsupported' is already handled above; this keeps the switch exhaustive.
|
||||
// A CLI whose config alone cannot select the model also gets its `model` launch param
|
||||
// forced (pi/omp `custom/<id>`, grok's block name). The argv engine DROPS a token that
|
||||
// fails its pattern rather than quoting it, which would silently launch the CLI on its
|
||||
// own default provider again, so refuse an id the pattern cannot carry up front.
|
||||
const modelSpec = entry.launch.params.model;
|
||||
const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id);
|
||||
if (!applied) {
|
||||
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
|
||||
}
|
||||
if (
|
||||
applied.launchModel !== undefined &&
|
||||
modelSpec?.type === 'token' &&
|
||||
!matchesPattern(modelSpec.pattern, applied.launchModel)
|
||||
) {
|
||||
removeConfigDir(applied.configDir);
|
||||
return createErrorResponse(
|
||||
ApiErrorCode.INVALID_INPUT,
|
||||
`Model id ${JSON.stringify(body.modelId)} cannot be passed to ${session.mode} on its command line`
|
||||
);
|
||||
}
|
||||
|
||||
const previousConfigDir = session.setCustomModel(
|
||||
{ endpointId: endpoint.id, modelId: body.modelId, label: endpoint.label, envKeys, configDir },
|
||||
envOverrides
|
||||
const { previousConfigDir } = session.setCustomModel(
|
||||
{
|
||||
endpointId: endpoint.id,
|
||||
modelId: body.modelId,
|
||||
label: endpoint.label,
|
||||
envKeys: applied.envKeys,
|
||||
configDir: applied.configDir,
|
||||
launchModel: applied.launchModel,
|
||||
},
|
||||
applied.envOverrides
|
||||
);
|
||||
// Clean up the OLD config dir on disk, unless the new one happens to reuse the same
|
||||
// path (same session, configDir kind again) — never delete the dir we just wrote.
|
||||
if (previousConfigDir && previousConfigDir !== configDir) {
|
||||
if (previousConfigDir && previousConfigDir !== applied.configDir) {
|
||||
removeConfigDir(previousConfigDir);
|
||||
}
|
||||
|
||||
|
||||
+27
-24
@@ -739,29 +739,6 @@ export const RemoteHostSchema = z.object({
|
||||
commands: RemoteCommandOverridesSchema,
|
||||
});
|
||||
|
||||
// Custom Model Endpoint Profiles (deployment_plan.md) — a user-configured custom
|
||||
// OpenAI-compatible endpoint, local (llama.cpp) or cloud (Azure AI Foundry, etc.).
|
||||
export const CustomModelHostSchema = z.object({
|
||||
id: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
|
||||
label: z.string().min(1).max(100),
|
||||
baseUrl: z.string().url().max(2048),
|
||||
apiKey: z.string().max(4096).optional(),
|
||||
// No 'both': live-tested against a real server, sending both auth header
|
||||
// conventions on one request reliably HANGS it — see custom-model-hosts.ts.
|
||||
authStyle: z.enum(['bearer', 'api-key']).optional(),
|
||||
models: z.array(z.string().max(200)).max(200).optional(),
|
||||
lastDiscoveredAt: z.string().max(64).optional(),
|
||||
});
|
||||
|
||||
/** POST /api/sessions/:id/custom-model — apply or clear a session's custom-model selection. */
|
||||
export const CustomModelSelectionSchema = z.union([
|
||||
z.object({
|
||||
endpointId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
|
||||
modelId: z.string().min(1).max(200),
|
||||
}),
|
||||
z.object({ clear: z.literal(true) }),
|
||||
]);
|
||||
|
||||
export const RemoteCaseLinkSchema = z.object({
|
||||
name: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid case name format'),
|
||||
hostId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid remote host id'),
|
||||
@@ -1261,7 +1238,7 @@ export const SettingsUpdateSchema = z
|
||||
*/
|
||||
readMyMindEnabled: z.boolean().optional(),
|
||||
/**
|
||||
* Custom Model Endpoint Profiles (deployment_plan.md): the toolbar picker that lets a
|
||||
* Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): the toolbar picker that lets a
|
||||
* session point at a user-configured custom OpenAI-compatible endpoint (local or
|
||||
* cloud) instead of its native cloud backend. SYNCED, default OFF — endpoint entry,
|
||||
* discovery, and the extra toolbar surface are all opt-in.
|
||||
@@ -1918,3 +1895,29 @@ export const WebviewUpdateSchema = WebviewBaseSchema.partial();
|
||||
|
||||
/** POST /api/webviews/probe: reachability + framing check for the editor's Test button. */
|
||||
export const WebviewProbeSchema = z.object({ url: webviewUrlSchema });
|
||||
|
||||
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — a
|
||||
// user-configured custom OpenAI-compatible endpoint, local (llama.cpp) or cloud
|
||||
// (Azure AI Foundry, etc.). Lives below `webviewUrlSchema` because `baseUrl` IS that
|
||||
// schema: http(s) only, a real hostname, no embedded credentials, and the link-local /
|
||||
// cloud-metadata refusal, the same bar a saved dashboard URL has to clear.
|
||||
export const CustomModelHostSchema = z.object({
|
||||
id: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
|
||||
label: z.string().min(1).max(100),
|
||||
baseUrl: webviewUrlSchema,
|
||||
apiKey: z.string().max(4096).optional(),
|
||||
// No 'both': live-tested against a real server, sending both auth header
|
||||
// conventions on one request reliably HANGS it — see custom-model-hosts.ts.
|
||||
authStyle: z.enum(['bearer', 'api-key']).optional(),
|
||||
models: z.array(z.string().max(200)).max(200).optional(),
|
||||
lastDiscoveredAt: z.string().max(64).optional(),
|
||||
});
|
||||
|
||||
/** POST /api/sessions/:id/custom-model — apply or clear a session's custom-model selection. */
|
||||
export const CustomModelSelectionSchema = z.union([
|
||||
z.object({
|
||||
endpointId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
|
||||
modelId: z.string().min(1).max(200),
|
||||
}),
|
||||
z.object({ clear: z.literal(true) }),
|
||||
]);
|
||||
|
||||
@@ -67,6 +67,10 @@ import {
|
||||
import { imageWatcher } from '../image-watcher.js';
|
||||
import { workflowRunWatcher, summarizeRun } from '../workflow-run-watcher.js';
|
||||
import { attachmentRegistry, buildFileThumbnailRoute, registerExternalAttachment } from '../attachment-registry.js';
|
||||
import { getCli } from '../config/cli-registry/registry.js';
|
||||
import { readCustomModelHosts } from '../custom-model-hosts.js';
|
||||
import { applyCustomModelInjection, customModelConfigDir, removeConfigDir } from '../custom-model-injection-apply.js';
|
||||
import type { CustomModelBookkeeping } from '../types/session.js';
|
||||
import { registerGeneratedArtifactAttachment } from '../generated-artifact-attachments.js';
|
||||
import {
|
||||
buildDetectedAttachmentHistoryItem,
|
||||
@@ -1181,6 +1185,34 @@ export class WebServer extends EventEmitter {
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Recovery half of Custom Model Endpoint Profiles: the env values a selection injects
|
||||
* are never persisted (they carry the API key), so they are computed again from the
|
||||
* endpoint store, through the SAME apply path the route uses. Undefined when the
|
||||
* endpoint is gone or the CLI is unregistered: the bookkeeping is still restored so
|
||||
* the selection can be cleared, and the pane keeps running on tmux's retained env.
|
||||
*/
|
||||
private async _rebuildCustomModelEnv(
|
||||
session: Session,
|
||||
saved: CustomModelBookkeeping
|
||||
): Promise<Record<string, string> | undefined> {
|
||||
const entry = getCli(session.mode);
|
||||
if (!entry) return undefined;
|
||||
const endpoint = (await readCustomModelHosts(getDataDir())).find((h) => h.id === saved.endpointId);
|
||||
if (!endpoint) {
|
||||
console.warn(
|
||||
`[WebServer] custom-model endpoint ${saved.endpointId} no longer exists; selection kept for clearing`
|
||||
);
|
||||
return undefined;
|
||||
}
|
||||
try {
|
||||
return applyCustomModelInjection(entry, endpoint, saved.modelId, session.id)?.envOverrides;
|
||||
} catch (err) {
|
||||
console.warn('[WebServer] Failed to rebuild custom-model env on recovery:', err);
|
||||
return undefined;
|
||||
}
|
||||
}
|
||||
|
||||
/** Persists full session state including respawn config to state.json */
|
||||
private _persistSessionStateNow(session: Session): void {
|
||||
// See session-manager.updateSessionState: __envOverrides is an internal disk-only
|
||||
@@ -1190,10 +1222,14 @@ export class WebServer extends EventEmitter {
|
||||
// __attachmentHistory keeps the private (externalPath-bearing) history on disk,
|
||||
// separate from the sanitized public attachmentHistory in toState().
|
||||
const attachmentHistory = session.getAttachmentHistoryForPersist();
|
||||
// __customModel keeps the selection's bookkeeping (injected env KEYS, config dir,
|
||||
// launch model; never the values) so recovery can restore and later clear it.
|
||||
const customModel = session.getCustomModelForPersist();
|
||||
const state = {
|
||||
...base,
|
||||
...(envOverrides ? { __envOverrides: envOverrides } : {}),
|
||||
...(attachmentHistory ? { __attachmentHistory: attachmentHistory } : {}),
|
||||
...(customModel ? { __customModel: customModel } : {}),
|
||||
} as SessionState;
|
||||
const controller = this.respawnControllers.get(session.id);
|
||||
if (controller) {
|
||||
@@ -1396,6 +1432,9 @@ export class WebServer extends EventEmitter {
|
||||
// come back to a loader whose file we deleted.
|
||||
if (killMux) {
|
||||
void removeAgentSessionPreamble(sessionId);
|
||||
// The per-session custom-model config dir carries the endpoint's API key (pi and
|
||||
// omp embed it literally); it must not outlive the session it was written for.
|
||||
removeConfigDir(customModelConfigDir(sessionId));
|
||||
}
|
||||
await session.stop(killMux);
|
||||
this.sessions.delete(sessionId);
|
||||
@@ -2853,6 +2892,7 @@ export class WebServer extends EventEmitter {
|
||||
// Note: a legacy CLAUDE_CODE_EFFORT_LEVEL entry is auto-migrated to `effort`
|
||||
// by the Session constructor (env var would hard-lock /effort switching).
|
||||
const savedEnvOverrides = (savedState as { __envOverrides?: Record<string, string> })?.__envOverrides;
|
||||
const savedCustomModel = (savedState as { __customModel?: CustomModelBookkeeping })?.__customModel;
|
||||
// Prefer the private (externalPath-bearing) history; fall back to the
|
||||
// sanitized public copy for sessions persisted before that split.
|
||||
const savedAttachmentHistory =
|
||||
@@ -2915,6 +2955,16 @@ export class WebServer extends EventEmitter {
|
||||
parentSessionId: savedState?.parentSessionId,
|
||||
});
|
||||
|
||||
// Custom-model selection survives the restart. The tmux session still carries
|
||||
// the injected `setenv`s (that is what kept the pane on the endpoint across the
|
||||
// restart), but `_envOverrides` is rebuilt from a persist that deliberately
|
||||
// excludes them, so re-derive the values from the endpoint store and re-write
|
||||
// the isolated config dir; an endpoint that has since been deleted still gets
|
||||
// the bookkeeping restored, which is what a later clear needs to unset.
|
||||
if (savedCustomModel) {
|
||||
session.setCustomModel(savedCustomModel, await this._rebuildCustomModelEnv(session, savedCustomModel));
|
||||
}
|
||||
|
||||
// Update session name if it was a "Restored:" placeholder or doesn't match saved name
|
||||
if (savedState?.name && muxSession.name !== savedState.name) {
|
||||
this.mux.updateSessionName(muxSession.sessionId, savedState.name);
|
||||
|
||||
@@ -35,6 +35,26 @@ function expectRejected(mutate: (entry: Record<string, unknown>) => void, becaus
|
||||
expect(result.success, `expected rejection: ${because}`).toBe(false);
|
||||
}
|
||||
|
||||
describe('customModelInjection.launchModel', () => {
|
||||
it('rejects a template with characters the argv engine would have to quote', () => {
|
||||
expectRejected((e) => {
|
||||
const caps = e.capabilities as Record<string, unknown>;
|
||||
caps.customModelInjection = { ...(caps.customModelInjection as object), launchModel: 'custom/{modelId} --yolo' };
|
||||
}, 'a space in the launch-model template');
|
||||
expectRejected((e) => {
|
||||
const caps = e.capabilities as Record<string, unknown>;
|
||||
caps.customModelInjection = { ...(caps.customModelInjection as object), launchModel: '' };
|
||||
}, 'an empty launch-model template');
|
||||
});
|
||||
|
||||
it('accepts the placeholder form the stock entries use', () => {
|
||||
const entry = baseEntry();
|
||||
const caps = entry.capabilities as Record<string, unknown>;
|
||||
caps.customModelInjection = { ...(caps.customModelInjection as object), launchModel: 'custom/{modelId}' };
|
||||
expect(CliEntrySchema.safeParse(entry).success).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('the shipped catalog', () => {
|
||||
it('validates every stock entry exactly as shipped', () => {
|
||||
// If this fails, the catalog cannot load at all — every other test here is downstream.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
/**
|
||||
* @fileoverview Contract tests for Custom Model Endpoint Profiles
|
||||
* (deployment_plan.md chunk 7): for every CLI with a `customModelInjection`
|
||||
* (docs/custom-model-endpoints-plan.md chunk 7): for every CLI with a `customModelInjection`
|
||||
* capability, build the real injection via `buildCustomModelInjection()`,
|
||||
* then replay those exact values through an HTTP request shaped the way that
|
||||
* CLI is documented to send it, against the in-process mock server
|
||||
@@ -8,7 +8,7 @@
|
||||
* request at the injected base URL, with the injected API key in the
|
||||
* expected header, and the injected model id in the body.
|
||||
*
|
||||
* LIMITATION (stated here and in deployment_plan.md, not left implicit): this
|
||||
* LIMITATION (stated here and in docs/custom-model-endpoints-plan.md, not left implicit): this
|
||||
* proves "if the CLI honors its documented env/config contract, it will hit
|
||||
* the right endpoint with the right model." It does NOT prove the real CLI
|
||||
* binary actually reads that env var / config file the way its docs say —
|
||||
@@ -148,7 +148,7 @@ describe('custom-model-injection contract (mock server)', () => {
|
||||
// SDK appends the path itself. Whether each of these TWO CLIs' own OpenAI-compatible
|
||||
// client expects the var to already include /v1 (the common OpenAI-SDK convention) or
|
||||
// appends it itself is genuinely CLI-specific and UNVERIFIED (see the confidence table
|
||||
// in deployment_plan.md) — these tests model the common OpenAI-SDK convention (base_url
|
||||
// in docs/custom-model-endpoints-plan.md) — these tests model the common OpenAI-SDK convention (base_url
|
||||
// ends in /v1) since that's the more likely behavior for an OpenAI-compatible client,
|
||||
// but that assumption should be corrected here the moment it's checked against a real
|
||||
// binary. (grok WAS in this group too, until live-testing showed the whole `env` recipe
|
||||
|
||||
@@ -9,7 +9,12 @@
|
||||
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import { getCli } from '../src/config/cli-registry/index.js';
|
||||
import { buildCustomModelInjection, withV1Suffix, type CustomModelEndpoint } from '../src/custom-model-injection.js';
|
||||
import {
|
||||
buildCustomModelInjection,
|
||||
withV1Suffix,
|
||||
GROK_CUSTOM_MODEL_NAME,
|
||||
type CustomModelEndpoint,
|
||||
} from '../src/custom-model-injection.js';
|
||||
|
||||
const endpoint: CustomModelEndpoint = {
|
||||
id: 'ep1',
|
||||
@@ -156,3 +161,28 @@ describe('buildCustomModelInjection', () => {
|
||||
expect(result).toEqual({ kind: 'unsupported' });
|
||||
});
|
||||
});
|
||||
|
||||
describe('launchModel (the model launch param that selects the injected provider)', () => {
|
||||
it('pi and omp get --model custom/<modelId>: the config file alone leaves them on their default provider', () => {
|
||||
for (const id of ['pi', 'omp']) {
|
||||
const result = buildCustomModelInjection(entryOrThrow(id), endpoint, 'qwen3.5-0.8b');
|
||||
if (result.kind !== 'configDir') throw new Error('unreachable');
|
||||
expect(result.launchModel, id).toBe('custom/qwen3.5-0.8b');
|
||||
}
|
||||
});
|
||||
|
||||
it('grok gets the [model.<name>] block name, pinned to the constant the config template writes', () => {
|
||||
const result = buildCustomModelInjection(entryOrThrow('grok'), endpoint, 'qwen3');
|
||||
if (result.kind !== 'configDir') throw new Error('unreachable');
|
||||
expect(result.launchModel).toBe(GROK_CUSTOM_MODEL_NAME);
|
||||
expect(result.files[0].content).toContain(`[model.${GROK_CUSTOM_MODEL_NAME}]`);
|
||||
});
|
||||
|
||||
it('CLIs whose config selects the model on its own declare none', () => {
|
||||
for (const id of ['claude', 'opencode', 'codex', 'gemini', 'deepseek']) {
|
||||
const result = buildCustomModelInjection(entryOrThrow(id), endpoint, 'qwen3');
|
||||
if (result.kind === 'unsupported') throw new Error('unreachable');
|
||||
expect(result.launchModel, id).toBeUndefined();
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
Vendored
+1
-1
@@ -1,6 +1,6 @@
|
||||
/**
|
||||
* @fileoverview In-process fake OpenAI/Anthropic-compatible HTTP server for the
|
||||
* Custom Model Endpoint Profiles contract tests (deployment_plan.md chunk 7).
|
||||
* Custom Model Endpoint Profiles contract tests (docs/custom-model-endpoints-plan.md chunk 7).
|
||||
*
|
||||
* No external deps — plain `node:http`. Captures every request it receives
|
||||
* (method, path, headers, parsed JSON body) so a test can assert the injected
|
||||
|
||||
@@ -330,21 +330,42 @@ export class MockSession extends EventEmitter {
|
||||
this.color = c;
|
||||
});
|
||||
|
||||
/** Custom Model Endpoint Profiles (deployment_plan.md) */
|
||||
/** Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) */
|
||||
customModel: { endpointId: string; modelId: string; label?: string } | undefined = undefined;
|
||||
private _mockCustomModelConfigDir: string | undefined;
|
||||
remote: unknown = undefined;
|
||||
docker: unknown = undefined;
|
||||
private _mockCustomModel:
|
||||
| {
|
||||
endpointId: string;
|
||||
modelId: string;
|
||||
label?: string;
|
||||
envKeys: string[];
|
||||
configDir?: string;
|
||||
launchModel?: string;
|
||||
}
|
||||
| undefined;
|
||||
setCustomModel = vi.fn(
|
||||
(
|
||||
next: { endpointId: string; modelId: string; label?: string; envKeys: string[]; configDir?: string } | undefined,
|
||||
next:
|
||||
| {
|
||||
endpointId: string;
|
||||
modelId: string;
|
||||
label?: string;
|
||||
envKeys: string[];
|
||||
configDir?: string;
|
||||
launchModel?: string;
|
||||
}
|
||||
| undefined,
|
||||
_envOverrides?: Record<string, string>
|
||||
): string | undefined => {
|
||||
const previous = this._mockCustomModelConfigDir;
|
||||
this._mockCustomModelConfigDir = next?.configDir;
|
||||
): { removedEnvKeys: string[]; previousConfigDir: string | undefined } => {
|
||||
const previous = this._mockCustomModel;
|
||||
this._mockCustomModel = next;
|
||||
this.customModel = next ? { endpointId: next.endpointId, modelId: next.modelId, label: next.label } : undefined;
|
||||
return previous;
|
||||
return { removedEnvKeys: previous?.envKeys ?? [], previousConfigDir: previous?.configDir };
|
||||
}
|
||||
);
|
||||
restartCli = vi.fn(async () => true);
|
||||
getCustomModelForPersist = vi.fn(() => this._mockCustomModel);
|
||||
|
||||
/** Stub for sendInput */
|
||||
sendInput = vi.fn();
|
||||
|
||||
@@ -1,18 +1,34 @@
|
||||
/**
|
||||
* @fileoverview Route tests for Custom Model Endpoint Profiles CRUD + discovery.
|
||||
*
|
||||
* Discovery is mocked at `webviewFetch()` (webview-egress.ts), NOT at the global
|
||||
* `fetch`: the route deliberately goes through the guarded undici dispatcher whose
|
||||
* lookup hook refuses a name that resolves into a link-local / cloud-metadata range,
|
||||
* so a global-fetch stub that still satisfied these tests would mean the guard had
|
||||
* been bypassed.
|
||||
* Port: N/A (app.inject, no real port needed)
|
||||
*/
|
||||
import { describe, it, expect, vi, afterEach } from 'vitest';
|
||||
import { registerCustomModelRoutes } from '../../src/web/routes/custom-model-routes.js';
|
||||
import { webviewFetch } from '../../src/web/webview-egress.js';
|
||||
import { createRouteTestHarness } from './_route-test-utils.js';
|
||||
|
||||
vi.mock('../../src/web/webview-egress.js', async () => {
|
||||
const actual = await vi.importActual<typeof import('../../src/web/webview-egress.js')>(
|
||||
'../../src/web/webview-egress.js'
|
||||
);
|
||||
return { ...actual, webviewFetch: vi.fn() };
|
||||
});
|
||||
|
||||
const fetchMock = vi.mocked(webviewFetch);
|
||||
|
||||
async function setup() {
|
||||
return createRouteTestHarness(registerCustomModelRoutes);
|
||||
}
|
||||
|
||||
describe('custom model endpoint CRUD', () => {
|
||||
afterEach(() => {
|
||||
vi.unstubAllGlobals();
|
||||
fetchMock.mockReset();
|
||||
});
|
||||
|
||||
it('starts empty', async () => {
|
||||
@@ -88,18 +104,18 @@ describe('custom model endpoint CRUD', () => {
|
||||
payload: { id: 'ep1', label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'k' },
|
||||
});
|
||||
|
||||
const fetchMock = vi.fn(async (url: string, init?: RequestInit) => {
|
||||
expect(url).toBe('http://localhost:8080/v1/models');
|
||||
fetchMock.mockImplementation(async (url: URL, init?: RequestInit) => {
|
||||
expect(url.href).toBe('http://localhost:8080/v1/models');
|
||||
const headers = init?.headers as Record<string, string>;
|
||||
// Exactly ONE auth header — never both (a real server hung when sent both).
|
||||
expect(headers.Authorization).toBe('Bearer k');
|
||||
expect(headers['api-key']).toBeUndefined();
|
||||
return new Response(JSON.stringify({ data: [{ id: 'qwen3' }, { id: 'llama3' }] }), { status: 200 });
|
||||
});
|
||||
vi.stubGlobal('fetch', fetchMock);
|
||||
|
||||
const res = await app.inject({ method: 'POST', url: '/api/model-endpoints/ep1/discover-models' });
|
||||
expect(res.json().data.models).toEqual(['qwen3', 'llama3']);
|
||||
expect(fetchMock).toHaveBeenCalledTimes(1);
|
||||
|
||||
// Data dir is shared across this WHOLE test file (one temp HOME per file, not per
|
||||
// test — test/setup.ts), so find by id rather than assuming index 0.
|
||||
@@ -117,13 +133,12 @@ describe('custom model endpoint CRUD', () => {
|
||||
payload: { id: 'ep-azure', label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'k', authStyle: 'api-key' },
|
||||
});
|
||||
|
||||
const fetchMock = vi.fn(async (_url: string, init?: RequestInit) => {
|
||||
fetchMock.mockImplementation(async (_url: URL, init?: RequestInit) => {
|
||||
const headers = init?.headers as Record<string, string>;
|
||||
expect(headers['api-key']).toBe('k');
|
||||
expect(headers.Authorization).toBeUndefined();
|
||||
return new Response(JSON.stringify({ data: [] }), { status: 200 });
|
||||
});
|
||||
vi.stubGlobal('fetch', fetchMock);
|
||||
|
||||
await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-azure/discover-models' });
|
||||
expect(fetchMock).toHaveBeenCalledTimes(1);
|
||||
@@ -136,11 +151,9 @@ describe('custom model endpoint CRUD', () => {
|
||||
url: '/api/model-endpoints',
|
||||
payload: { id: 'ep-err', label: 'A', baseUrl: 'http://localhost:8080' },
|
||||
});
|
||||
vi.stubGlobal(
|
||||
'fetch',
|
||||
vi.fn(async () => {
|
||||
throw new Error('connect ECONNREFUSED');
|
||||
})
|
||||
// undici's shape: a bare `fetch failed` with the real reason one level down.
|
||||
fetchMock.mockRejectedValue(
|
||||
new TypeError('fetch failed', { cause: new Error('connect ECONNREFUSED 127.0.0.1:8080') })
|
||||
);
|
||||
|
||||
const res = await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-err/discover-models' });
|
||||
@@ -148,4 +161,36 @@ describe('custom model endpoint CRUD', () => {
|
||||
expect(res.json().errorCode).toBe('OPERATION_FAILED');
|
||||
expect(res.json().error).toContain('ECONNREFUSED');
|
||||
});
|
||||
|
||||
it('names the egress refusal when the endpoint resolves into a blocked range', async () => {
|
||||
const { app } = await setup();
|
||||
await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/model-endpoints',
|
||||
payload: { id: 'ep-meta', label: 'A', baseUrl: 'http://models.example:8080' },
|
||||
});
|
||||
const { WebviewEgressBlockedError } = await vi.importActual<typeof import('../../src/web/webview-egress.js')>(
|
||||
'../../src/web/webview-egress.js'
|
||||
);
|
||||
fetchMock.mockRejectedValue(
|
||||
new TypeError('fetch failed', { cause: new WebviewEgressBlockedError('resolves to 169.254.169.254') })
|
||||
);
|
||||
|
||||
const res = await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-meta/discover-models' });
|
||||
expect(res.json().success).toBe(false);
|
||||
expect(res.json().error).toMatch(/refused.*169\.254\.169\.254/);
|
||||
});
|
||||
|
||||
it('refuses a baseUrl with embedded credentials or a non-http scheme at save time', async () => {
|
||||
const { app } = await setup();
|
||||
for (const baseUrl of ['http://user:pw@host:8080', 'ftp://host/models', 'http://169.254.169.254']) {
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/model-endpoints',
|
||||
payload: { id: 'bad', label: 'A', baseUrl },
|
||||
});
|
||||
expect(res.json().success, baseUrl).toBe(false);
|
||||
expect(res.json().errorCode, baseUrl).toBe('INVALID_INPUT');
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
/**
|
||||
* @fileoverview Tests for POST /api/sessions/:id/custom-model (deployment_plan.md
|
||||
* @fileoverview Tests for POST /api/sessions/:id/custom-model (docs/custom-model-endpoints-plan.md
|
||||
* chunk 5 — applying/clearing a session's custom model endpoint + CLI restart).
|
||||
* Port: N/A (app.inject, no real port needed)
|
||||
*/
|
||||
@@ -8,6 +8,8 @@ import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
|
||||
import { createRouteTestHarness } from './_route-test-utils.js';
|
||||
import { getDataDir } from '../../src/config/instance.js';
|
||||
import { writeCustomModelHosts, type CustomModelHost } from '../../src/custom-model-hosts.js';
|
||||
import { existsSync, statSync } from 'node:fs';
|
||||
import { join } from 'node:path';
|
||||
|
||||
const CLAUDE_ENDPOINT: CustomModelHost = {
|
||||
id: 'ep1',
|
||||
@@ -102,6 +104,103 @@ describe('POST /api/sessions/:id/custom-model', () => {
|
||||
expect(res.json().errorCode).toBe('OPERATION_FAILED');
|
||||
});
|
||||
|
||||
it('refuses a remote (SSH) session before touching it: restartCli would only reattach the remote tmux', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
const session = ctx.sessions.get('test-session-1')!;
|
||||
session.mode = 'claude';
|
||||
session.remote = { hostId: 'h1', remotePath: '/srv/case' };
|
||||
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/sessions/test-session-1/custom-model',
|
||||
payload: { endpointId: 'ep1', modelId: 'qwen3' },
|
||||
});
|
||||
|
||||
expect(res.json().success).toBe(false);
|
||||
expect(res.json().errorCode).toBe('INVALID_INPUT');
|
||||
expect(res.json().error).toMatch(/remote/i);
|
||||
expect(session.setCustomModel).not.toHaveBeenCalled();
|
||||
expect(session.restartCli).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('refuses a Docker session the same way, for clear as well as apply', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
const session = ctx.sessions.get('test-session-1')!;
|
||||
session.mode = 'claude';
|
||||
session.docker = { containerName: 'codeman-case' };
|
||||
|
||||
for (const payload of [{ endpointId: 'ep1', modelId: 'qwen3' }, { clear: true }]) {
|
||||
const res = await app.inject({ method: 'POST', url: '/api/sessions/test-session-1/custom-model', payload });
|
||||
expect(res.json().success).toBe(false);
|
||||
expect(res.json().errorCode).toBe('INVALID_INPUT');
|
||||
}
|
||||
expect(session.setCustomModel).not.toHaveBeenCalled();
|
||||
expect(session.restartCli).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('pi: writes the config dir AND forces --model custom/<id>, since the file alone does not select the model', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
const session = ctx.sessions.get('test-session-1')!;
|
||||
session.mode = 'pi';
|
||||
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/sessions/test-session-1/custom-model',
|
||||
payload: { endpointId: 'ep1', modelId: 'qwen3.5-0.8b' },
|
||||
});
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
const [next, envOverrides] = session.setCustomModel.mock.calls[0];
|
||||
expect(next.launchModel).toBe('custom/qwen3.5-0.8b');
|
||||
expect(next.configDir).toBe(join(getDataDir(), 'custom-model-configs', 'test-session-1'));
|
||||
expect(envOverrides.HOME).toBe(next.configDir);
|
||||
const written = join(next.configDir, '.pi', 'agent', 'models.json');
|
||||
expect(existsSync(written)).toBe(true);
|
||||
// pi embeds the key literally, so the file is private to the server account.
|
||||
expect(statSync(written).mode & 0o777).toBe(0o600);
|
||||
expect(session.restartCli).toHaveBeenCalledTimes(1);
|
||||
});
|
||||
|
||||
it('refuses a model id the CLI cannot carry on its command line instead of launching without it', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
const session = ctx.sessions.get('test-session-1')!;
|
||||
session.mode = 'pi';
|
||||
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/sessions/test-session-1/custom-model',
|
||||
payload: { endpointId: 'ep1', modelId: 'qwen 3 with spaces' },
|
||||
});
|
||||
|
||||
expect(res.json().success).toBe(false);
|
||||
expect(res.json().errorCode).toBe('INVALID_INPUT');
|
||||
expect(session.setCustomModel).not.toHaveBeenCalled();
|
||||
expect(session.restartCli).not.toHaveBeenCalled();
|
||||
// The config dir written before the check is cleaned up again.
|
||||
expect(existsSync(join(getDataDir(), 'custom-model-configs', 'test-session-1'))).toBe(false);
|
||||
});
|
||||
|
||||
it('clear removes the previous config dir the session reports', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
const session = ctx.sessions.get('test-session-1')!;
|
||||
session.mode = 'pi';
|
||||
await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/sessions/test-session-1/custom-model',
|
||||
payload: { endpointId: 'ep1', modelId: 'qwen3' },
|
||||
});
|
||||
const dir = join(getDataDir(), 'custom-model-configs', 'test-session-1');
|
||||
expect(existsSync(dir)).toBe(true);
|
||||
|
||||
const res = await app.inject({
|
||||
method: 'POST',
|
||||
url: '/api/sessions/test-session-1/custom-model',
|
||||
payload: { clear: true },
|
||||
});
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(existsSync(dir)).toBe(false);
|
||||
});
|
||||
|
||||
it('refuses to touch a busy session', async () => {
|
||||
const { app, ctx } = await setup();
|
||||
const session = ctx.sessions.get('test-session-1')!;
|
||||
|
||||
@@ -0,0 +1,183 @@
|
||||
/**
|
||||
* @fileoverview Custom Model Endpoint Profiles, the Session half of applying and
|
||||
* clearing a selection against a live pane (`Session.setCustomModel()` +
|
||||
* `Session.restartCli()`), pinned against the three ways the first cut broke a
|
||||
* working session:
|
||||
*
|
||||
* 1. Clearing did not clear. The injected vars reach the CLI via `tmux setenv`,
|
||||
* which persists at the tmux-session level and is inherited by `respawn-pane`,
|
||||
* so removing them from `_envOverrides` relaunched the CLI still pointed at the
|
||||
* endpoint (and, for the configDir kinds, at a `HOME`/`CODEX_HOME` that had just
|
||||
* been deleted). The retired keys must ride `RespawnPaneOptions.unsetEnvKeys`.
|
||||
* 2. Applying to a local claude session killed the pane: the relaunch was
|
||||
* `claude --session-id <id>` and Claude refuses an id that already has a
|
||||
* transcript, so it needs the `--resume <id> || --session-id <id>` shape the
|
||||
* docker and remote pane commands use, i.e. a pinned resume id.
|
||||
* 3. pi/omp/grok wrote their config file and then launched without the `--model`
|
||||
* that selects it, so the file was ignored.
|
||||
*
|
||||
* Drives a real `Session` against the in-memory tmux layer vitest substitutes,
|
||||
* spying on `respawnPane` to read the options the relaunch would get.
|
||||
* Port: N/A.
|
||||
*/
|
||||
import { mkdirSync, rmSync } from 'node:fs';
|
||||
import { homedir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import { afterEach, describe, expect, it, vi } from 'vitest';
|
||||
|
||||
import { Session } from '../src/session.js';
|
||||
import { TmuxManager } from '../src/tmux-manager.js';
|
||||
import type { MuxSession, SessionMode } from '../src/types.js';
|
||||
|
||||
const workingDir = join(homedir(), 'codeman-cases', 'custom-model-restart');
|
||||
const sessions: Session[] = [];
|
||||
|
||||
afterEach(() => {
|
||||
for (const s of sessions.splice(0)) s.stop();
|
||||
rmSync(workingDir, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
function liveSession(mode: SessionMode, extra: Record<string, unknown> = {}) {
|
||||
mkdirSync(workingDir, { recursive: true });
|
||||
const mux = new TmuxManager();
|
||||
const muxSession: MuxSession = {
|
||||
sessionId: 'placeholder',
|
||||
muxName: 'codeman-cafe0001',
|
||||
pid: 1,
|
||||
createdAt: Date.now(),
|
||||
workingDir,
|
||||
mode,
|
||||
attached: false,
|
||||
};
|
||||
const session = new Session({ workingDir, mode, mux, useMux: true, muxSession, ...extra });
|
||||
sessions.push(session);
|
||||
vi.spyOn(mux, 'muxSessionExists').mockReturnValue(true);
|
||||
const respawn = vi.spyOn(mux, 'respawnPane').mockResolvedValue(4242);
|
||||
return { session, respawn };
|
||||
}
|
||||
|
||||
const CLAUDE_KEYS = ['ANTHROPIC_BASE_URL', 'ANTHROPIC_API_KEY', 'ANTHROPIC_DEFAULT_SONNET_MODEL'];
|
||||
const claudeEnv = {
|
||||
ANTHROPIC_BASE_URL: 'http://box:8080',
|
||||
ANTHROPIC_API_KEY: 'k',
|
||||
ANTHROPIC_DEFAULT_SONNET_MODEL: 'q',
|
||||
};
|
||||
|
||||
describe('clearing a selection unsets what it injected', () => {
|
||||
it('queues the retired keys for setenv -u on the next respawn and drops them from the overrides', async () => {
|
||||
const { session, respawn } = liveSession('claude', { envOverrides: { CLAUDE_CODE_KEEP: '1' } });
|
||||
session.setCustomModel({ endpointId: 'ep', modelId: 'q', envKeys: CLAUDE_KEYS }, claudeEnv);
|
||||
|
||||
const result = session.setCustomModel(undefined);
|
||||
expect(result.removedEnvKeys).toEqual(CLAUDE_KEYS);
|
||||
expect(session.customModel).toBeUndefined();
|
||||
|
||||
expect(await session.restartCli()).toBe(true);
|
||||
const options = respawn.mock.calls[0][0];
|
||||
expect(options.unsetEnvKeys).toEqual(CLAUDE_KEYS);
|
||||
expect(options.envOverrides).toEqual({ CLAUDE_CODE_KEEP: '1' });
|
||||
|
||||
// Drained once the respawn succeeded: the next relaunch has nothing to unset.
|
||||
respawn.mockClear();
|
||||
await session.restartCli();
|
||||
expect(respawn.mock.calls[0][0].unsetEnvKeys).toBeUndefined();
|
||||
});
|
||||
|
||||
it('switching endpoints re-sets the shared keys instead of unsetting them', async () => {
|
||||
const { session, respawn } = liveSession('claude');
|
||||
session.setCustomModel({ endpointId: 'a', modelId: 'q', envKeys: CLAUDE_KEYS }, claudeEnv);
|
||||
const next = { ...claudeEnv, ANTHROPIC_BASE_URL: 'http://other:8080' };
|
||||
session.setCustomModel({ endpointId: 'b', modelId: 'q', envKeys: CLAUDE_KEYS }, next);
|
||||
|
||||
await session.restartCli();
|
||||
const options = respawn.mock.calls[0][0];
|
||||
expect(options.unsetEnvKeys).toBeUndefined();
|
||||
expect(options.envOverrides?.ANTHROPIC_BASE_URL).toBe('http://other:8080');
|
||||
});
|
||||
|
||||
it('reports the previous config dir so the caller can delete it, and never leaks bookkeeping on the wire', () => {
|
||||
const { session } = liveSession('pi');
|
||||
session.setCustomModel(
|
||||
{ endpointId: 'ep', modelId: 'q', envKeys: ['HOME'], configDir: '/tmp/cfg-1', launchModel: 'custom/q' },
|
||||
{ HOME: '/tmp/cfg-1' }
|
||||
);
|
||||
expect(session.customModel).toEqual({ endpointId: 'ep', modelId: 'q', label: undefined });
|
||||
expect(session.toState().customModel).toEqual({ endpointId: 'ep', modelId: 'q', label: undefined });
|
||||
expect(session.getCustomModelForPersist()).toEqual({
|
||||
endpointId: 'ep',
|
||||
modelId: 'q',
|
||||
label: undefined,
|
||||
envKeys: ['HOME'],
|
||||
configDir: '/tmp/cfg-1',
|
||||
launchModel: 'custom/q',
|
||||
});
|
||||
expect(session.setCustomModel(undefined).previousConfigDir).toBe('/tmp/cfg-1');
|
||||
});
|
||||
});
|
||||
|
||||
describe('restartCli() must not kill a working pane', () => {
|
||||
it('claude: pins the live conversation id so the relaunch renders --resume <id> || --session-id <id>', async () => {
|
||||
const { session, respawn } = liveSession('claude');
|
||||
await session.restartCli();
|
||||
const options = respawn.mock.calls[0][0];
|
||||
expect(options.resumeSessionId).toBe(session.claudeSessionId);
|
||||
expect(options.resumeSessionId).toBe(session.id);
|
||||
// The pin is per-respawn: nothing about the session's own resume id changed.
|
||||
expect(session.toState().resumeSessionId).toBeUndefined();
|
||||
});
|
||||
|
||||
it('claude: an explicit resume id from a resume-from-history launch wins over the pin', async () => {
|
||||
const RESUMED = '01a060f0-0361-7f91-abde-b283020db0d7';
|
||||
const { session, respawn } = liveSession('claude', { resumeSessionId: RESUMED });
|
||||
await session.restartCli();
|
||||
expect(respawn.mock.calls[0][0].resumeSessionId).toBe(RESUMED);
|
||||
});
|
||||
|
||||
it('pi: no top-level resume pin, its resume id is minted by the CLI and lives in piConfig', async () => {
|
||||
const { session, respawn } = liveSession('pi');
|
||||
await session.restartCli();
|
||||
expect(respawn.mock.calls[0][0].resumeSessionId).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
describe('launchModel reaches the CLI through its own launch param', () => {
|
||||
it('pi: the selection forces piConfig.model on the respawn options, leaving the stored config alone', async () => {
|
||||
const { session, respawn } = liveSession('pi', { piConfig: { model: 'anthropic/claude-x', thinking: 'low' } });
|
||||
session.setCustomModel(
|
||||
{ endpointId: 'ep', modelId: 'qwen3', envKeys: ['HOME'], configDir: '/tmp/cfg', launchModel: 'custom/qwen3' },
|
||||
{ HOME: '/tmp/cfg' }
|
||||
);
|
||||
await session.restartCli();
|
||||
expect(respawn.mock.calls[0][0].piConfig).toEqual({ model: 'custom/qwen3', thinking: 'low' });
|
||||
// The user's own choice survives underneath, which is what a clear falls back to.
|
||||
expect(session.toState().piConfig).toEqual({ model: 'anthropic/claude-x', thinking: 'low' });
|
||||
|
||||
session.setCustomModel(undefined);
|
||||
respawn.mockClear();
|
||||
await session.restartCli();
|
||||
expect(respawn.mock.calls[0][0].piConfig).toEqual({ model: 'anthropic/claude-x', thinking: 'low' });
|
||||
});
|
||||
|
||||
it('grok: the block name lands in grokConfig.model', async () => {
|
||||
const { session, respawn } = liveSession('grok');
|
||||
session.setCustomModel(
|
||||
{
|
||||
endpointId: 'ep',
|
||||
modelId: 'qwen3',
|
||||
envKeys: ['GROK_HOME'],
|
||||
configDir: '/tmp/cfg',
|
||||
launchModel: 'codeman-custom',
|
||||
},
|
||||
{ GROK_HOME: '/tmp/cfg' }
|
||||
);
|
||||
await session.restartCli();
|
||||
expect(respawn.mock.calls[0][0].grokConfig).toEqual({ model: 'codeman-custom' });
|
||||
});
|
||||
|
||||
it('claude: a selection without a launchModel leaves the model param untouched', async () => {
|
||||
const { session, respawn } = liveSession('claude');
|
||||
session.setCustomModel({ endpointId: 'ep', modelId: 'q', envKeys: CLAUDE_KEYS }, claudeEnv);
|
||||
await session.restartCli();
|
||||
expect(respawn.mock.calls[0][0].model).toBeUndefined();
|
||||
});
|
||||
});
|
||||
@@ -77,6 +77,9 @@ vi.mock('../src/session.js', () => {
|
||||
};
|
||||
}
|
||||
|
||||
getCustomModelForPersist() {
|
||||
return undefined;
|
||||
}
|
||||
getEnvOverridesForPersist() {
|
||||
return undefined;
|
||||
}
|
||||
|
||||
@@ -853,6 +853,52 @@ describe('TmuxManager (unit)', () => {
|
||||
nonTestManager.destroy();
|
||||
}
|
||||
});
|
||||
|
||||
it('unsets retired env keys on the tmux session BEFORE re-applying the live overrides', async () => {
|
||||
// `tmux setenv` persists at the session level and is inherited by `respawn-pane`, so
|
||||
// a key that merely disappears from envOverrides comes back in the relaunched CLI
|
||||
// (measured: `setenv FOO bar` survived two `respawn-pane -k`). Clearing a custom-model
|
||||
// selection names the keys to drop; a key both dropped and re-set must end up SET.
|
||||
const NonTestTmuxManager = await importWithTmuxCommandsEnabled();
|
||||
const nonTestManager = new NonTestTmuxManager();
|
||||
nonTestManager.registerSession({
|
||||
sessionId: 'respawn5678',
|
||||
muxName: 'codeman-abcd5678',
|
||||
pid: 1000,
|
||||
createdAt: Date.now(),
|
||||
workingDir: '/tmp',
|
||||
mode: 'shell',
|
||||
attached: false,
|
||||
});
|
||||
|
||||
try {
|
||||
const pid = await nonTestManager.respawnPane({
|
||||
sessionId: 'respawn5678',
|
||||
workingDir: '/tmp',
|
||||
mode: 'shell',
|
||||
envOverrides: { CLAUDE_CODE_KEEP: '1', ANTHROPIC_API_KEY: 'again' },
|
||||
unsetEnvKeys: ['ANTHROPIC_BASE_URL', 'ANTHROPIC_API_KEY', 'not-a-key; rm -rf /'],
|
||||
});
|
||||
expect(pid).toBe(4242);
|
||||
|
||||
const setenvCalls = mockedExecSync.mock.calls
|
||||
.map(([cmd]) => cmd)
|
||||
.filter((cmd): cmd is string => typeof cmd === 'string' && cmd.includes(" setenv -t 'codeman-abcd5678'"));
|
||||
const unsetBase = setenvCalls.findIndex((cmd) => cmd.endsWith(' -u ANTHROPIC_BASE_URL'));
|
||||
const unsetKey = setenvCalls.findIndex((cmd) => cmd.endsWith(' -u ANTHROPIC_API_KEY'));
|
||||
const setKeep = setenvCalls.findIndex((cmd) => cmd.includes(' CLAUDE_CODE_KEEP '));
|
||||
const setKey = setenvCalls.findIndex((cmd) => cmd.includes(' ANTHROPIC_API_KEY ') && !cmd.includes(' -u '));
|
||||
expect(unsetBase).toBeGreaterThanOrEqual(0);
|
||||
expect(unsetKey).toBeGreaterThanOrEqual(0);
|
||||
expect(setKeep).toBeGreaterThan(unsetBase);
|
||||
// Re-set AFTER its own unset, so the live value wins.
|
||||
expect(setKey).toBeGreaterThan(unsetKey);
|
||||
// The shell-metachar key never reaches tmux at all.
|
||||
expect(setenvCalls.some((cmd) => cmd.includes('rm -rf'))).toBe(false);
|
||||
} finally {
|
||||
nonTestManager.destroy();
|
||||
}
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
Reference in New Issue
Block a user