mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
627b76739cddc49f63855575e13af8eb97b9577d
97
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bf73a84732 | fix(docker): address git identity review | ||
|
|
8d358aaa26 | feat(docker): configure static git identity | ||
|
|
95a3b87062 |
chore: drop changesets already released in 1.33.1
The pnpm and uv/uvx changesets describe work upstream shipped in 1.33.1 (#485, #487), so keeping them would repeat those notes in the next release. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
8cef31086b |
fix(docker): persist CLIs installed from Settings across container updates
The image sets NPM_CONFIG_PREFIX=/opt/codeman-cli, which is image content, so Update-Codeman.sh discarded every npm-installed CLI (dsh, pi). In the Compose container, POST /api/clis/:id/install now installs into ~/.local on the persistent home mount, and ~/.local/bin is appended to the image PATH. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
55790964b7 | Merge remote-tracking branch 'upstream/master' into feature/docker-uv-uvx | ||
|
|
8841bcc93f |
feat(cases): refresh the case picker and add search to Manage (#483)
The Run bar's case picker only loaded /api/cases at page load, so folders deleted or created on disk stayed listed until a reload. It now refetches on open and every 5 seconds while open, repainting only when the list changed and falling back to another case if the selected one was removed. The Manage tab of Add Case gains a search box filtering by name or path. Reorder arrows are disabled while a filter is active so a swap cannot involve a hidden case. Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
77ba41f8da |
feat(docker): install uv/uvx, libsecret-1-0 and pnpm (#487)
* fix(docker): install pnpm in the Compose server image `dsh plugin` spawns a literal `pnpm` with no npm fallback, so the Run menu's "DeepSeek - add a terminal profile" button failed with `dsh: pnpm not found on PATH` (exit 127) on the server image. The agent image already installs pnpm for the same reason (#352). Pin pnpm@12.6.0 in the runtime-writable CLI prefix and note it in the DeepSeek doc. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n * feat(docker): install uv and uvx in server and agent images Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n * feat(docker): install libsecret-1-0 for the Azure DevOps MCP Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n * feat(docker): add sudo to the agent image Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n * feat(docker): install sudo with passwordless access for the agent user Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n * Revert "feat(docker): install sudo with passwordless access for the agent user" This reverts commit |
||
|
|
d6c3386102 |
Revert "feat(docker): add sudo to the agent image"
This reverts commit
|
||
|
|
10c263a5b8 |
Revert "feat(docker): install sudo with passwordless access for the agent user"
This reverts commit
|
||
|
|
b070c9ee65 |
feat(docker): install sudo with passwordless access for the agent user
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
e98127a804 |
feat(docker): add sudo to the agent image
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
3e3a4612e6 |
feat(docker): install libsecret-1-0 for the Azure DevOps MCP
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
a5283c565d |
feat(docker): install uv and uvx in server and agent images
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
b46588f247 |
fix(docker): install pnpm in the Compose server image
`dsh plugin` spawns a literal `pnpm` with no npm fallback, so the Run menu's "DeepSeek - add a terminal profile" button failed with `dsh: pnpm not found on PATH` (exit 127) on the server image. The agent image already installs pnpm for the same reason (#352). Pin pnpm@12.6.0 in the runtime-writable CLI prefix and note it in the DeepSeek doc. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
0a52a99ca9 |
feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (
|
||
|
|
dd230b0b6e |
fix(test): strip every inherited CODEMAN_* var and move quick-start off 3099 (#479)
Split out of #476. A Docker Compose deployment exports CODEMAN_CASES_PATH, which bypasses the temp HOME, so route tests wrote into the real case root. Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
02e40f506b |
fix(docker): gate gh/az seeding on its switch; no shared git sign-in for non-admin clones
Addresses the review on #472. - CRED_STORES: `.config/gh` and `.azure` now carry `enabledByEnv` (CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ), and resolveDockerCredentialArtifacts skips a store unless that variable is exactly `1`, read at container create. A host that merely has ~/.config/gh/hosts.yml or a plaintext MSAL cache no longer copies them into every case container. Tests: the default environment seeds neither even with the files present, and each store follows only its own switch. - Multi-user mode: a non-admin's Clone Repo clone and preflight run with `git -c credential.helper=` (GIT_NO_CREDENTIAL_HELPERS, placed before the subcommand), so the server account's helpers are never lent to them. Verified against a real private repo that it also clears the URL-scoped credential.<url>.helper entries, and that public clones still work. Tests: the argv in test/git-clone.test.ts, and the route decision (non-admin cleared; admin and single-user kept) in test/routes/case-clone-credential-helpers.test.ts. - Docs: recreate the case container to pick up seeds (docker/README.md, Docker-Cases wiki, docker-cases.md); the multi-user behaviour in docker/README.md and security-architecture.md; "functionally unchanged" instead of "unchanged" for an image built with both switches off (server.Dockerfile comment, README, changeset). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw |
||
|
|
5cf5a45438 |
feat(docker): opt-in gh + az CLIs with git credential helpers for private repos
Add Case -> Clone Repo could only reach public repositories in the Docker deployment. This lets a deployment opt in to the GitHub CLI and the Azure CLI (+ azure-devops extension) as git credential helpers. Codeman itself still collects no credentials. - server.Dockerfile / agent.Dockerfile: CODEMAN_INSTALL_GH / CODEMAN_INSTALL_AZ build args (0 or 1, default 0; anything else stops the build). Off leaves no apt repository, package, extension, helper script or credential entry, so a default build is unchanged. On installs from the vendors' apt repositories and configures system gitconfig helpers: github.com / gist.github.com -> `gh auth git-credential`, dev.azure.com / *.visualstudio.com -> new docker/git-credential-azure-cli (an Entra ID token from `az account get-access-token`, or AZURE_DEVOPS_EXT_PAT). A helper whose CLI is not signed in prints nothing, so a private clone still fails fast. - The extension lives in AZURE_EXTENSION_DIR outside HOME (/opt/codeman-az-extensions, runtime-owned; /opt/az-extensions, gid-0 group-writable in the agent image). - Hosts turn them on in docker-compose.override.yml: `build: args:` for the server image, `environment:` CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ for the agent image. build-agent-image.mjs and the in-app auto-build share one env -> ARG table (pinned by the parity test) and pass nothing when unset. docker-compose.yaml is untouched; .env.example only gains a comment, so the self-updater's environment gate sees no new keys. - Docker cases seed the gh sign-in (~/.config/gh/hosts.yml, config.yml) and the az sign-in files from ~/.azure per file, read-only, like pi/grok. - The Clone Repo AUTH_REQUIRED message says how to sign the server's git in instead of claiming private repositories cannot be cloned. - Docs: docker/README.md "Private repositories", docker-compose.md, docker-cases.md, the Quick-Start / Core-Concepts / Docker-Cases wiki pages, security-architecture.md, architecture-invariants.md, changeset. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw |
||
|
|
a2dcc91ddf |
fix(docker): add and correct the cross-checkout collision guard for Update-Codeman.sh
Same guard as Start-Codeman.sh's own (docs/docker-self-update.md-adjacent incident, 2026-09-21): docker-compose.yaml hard-codes `name: codeman`, so a second checkout run without COMPOSE_PROJECT_NAME resolves to the SAME Compose project as any other checkout on the host. It has to live here too, not just in Start-Codeman.sh: this script's own --no-cache build and its `down`/`down --volumes` both run BEFORE the handoff at the bottom of the file, so Start-Codeman.sh's copy of the guard would only fire after this script's own destructive calls already ran — and its default `down --volumes` is more destructive than Start-Codeman.sh's own targeted refresh, clearing every named volume the resolved project has. Also fixes a real bug the same guard shipped with: under `set -o pipefail`, `grep -v` legitimately exits 1 when nothing survives the filter (the ordinary, no-collision case), and without `|| true` on the pipeline that non-zero status propagates through the command substitution and `set -e` aborts the WHOLE script at the guard — every time, collision or not. Caught only by actually executing the guard end-to-end against a stub `docker` (the existing smoke-test harness), never by a static text/regex check on the source; the stub's `config --format json` response was also fixed to pretty-print like real Compose does, since a compact one-liner silently resolved project_name to empty and exercised neither script's guard the way production output does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
5b4878df3b |
fix(docker): address Ark0N's PR review on Update-Codeman.sh — fix the handoff, the build/down ordering, and default-clear the build volumes
Three blockers, all fixed and verified by actually running the script (not just string-matching it): 1. `exec "$script_dir/Start-Codeman.sh"` failed EACCES/exit 126 on every checkout, since Start-Codeman.sh is committed non-executable (100644) — the same fact my own second commit on this branch established. Fixed to `exec bash "$script_dir/Start-Codeman.sh"`. 2. `down` ran before `build --no-cache`, so Codeman and every session it was running were offline for the entire rebuild, and a build failure left the stack down with nothing to bring it back — the exact ordering mistake Start-Codeman.sh's own "Build BEFORE taking the stack down" comment exists to prevent. Reordered to build, then down, then hand off. 3. The default path could throw the rebuild away: codeman-node-modules/ codeman-dist only re-seed from the image while EMPTY, Start-Codeman.sh only clears them when it detects the checkout's HEAD or package-lock.json moved, and neither condition is true for the Dockerfile-only change this script exists for — so a plain `bash docker/Update-Codeman.sh` rebuilt an image whose fresh node_modules/dist then sat unused behind the old volumes. Made clearing them the default; `--keep-volumes` opts out (replaces the old `--volumes`/`-v` flag, which is no longer needed since clearing is now the default). Smaller items from the same review, also fixed: - The --no-cache build now derives PUID/PGID from CODEMAN_APPDATA_PATH's owner first, via the identical owner_of() helper Start-Codeman.sh uses (parity-tested) — without it, the build used Compose's default 1000:1000 regardless of the real appdata owner (99:100 on the Unraid layout docker/README.md documents), and Start-Codeman.sh's own correctly-PUID'd build during the handoff would then rebuild those layers anyway, so the --no-cache image never actually shipped. - docker/README.md's "rebuilds ... only when it detects ... moved" wrongly described BOTH the rebuild and the volume-clearing as conditional; Start-Codeman.sh rebuilds on every start, only the volume-clearing is conditional. Corrected, and reworded around the new default. - --help/-h now prints usage and exits 0 instead of falling into the unrecognised-argument branch. - "the ONLY named volumes this stack declares" now says docker-compose.yaml specifically, since a docker-compose.override.yml could add more. New tests: PUID/PGID derivation parity with Start-Codeman.sh's owner_of(), --help handling, and — the one that actually catches blocker #1, which five source-string-matching tests did not — a real end-to-end smoke test: a synthetic deployment, a stub `docker` on PATH logging every invocation, the real script executed via a real subprocess. Confirms the real command sequence (build --no-cache, then down --volumes or plain down, then evidence the handoff genuinely ran Start-Codeman.sh) and that a working handoff fails honestly at Start-Codeman.sh's own later check rather than with EACCES. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
3b714446b4 |
fix(docker): drop the wrong executable-bit assertion for Update-Codeman.sh
Start-Codeman.sh, its sibling and the script it hands off to, is itself committed non-executable (100644) upstream — it's documented and invoked as `bash docker/Start-Codeman.sh`, never `./docker/Start-Codeman.sh`. The "is executable" test I'd added for Update-Codeman.sh asserted the opposite convention, which the file correctly does not follow. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
9ba90a674a |
chore(docker): add Update-Codeman.sh for scripted major-update rebuilds
docker/README.md and docs/docker-self-update.md both already point operators at "stop the stack, rebuild, restart" for anything the in-app updater refuses to apply (a changed server.Dockerfile, a changed docker-compose.yaml, or a new required .env key) — but that was a manual, hand-typed procedure with no script of its own, unlike every other start/update path this deployment has. docker/Update-Codeman.sh scripts it: `docker compose down`, then an unconditional `docker compose build --no-cache` (a major update should be certain of what actually ships, not reuse whatever layers happened to be cached), then hands off to the existing Start-Codeman.sh for the same careful PUID/PGID, override-file and fingerprint handling every other start already goes through — rather than reimplementing any of that by hand and risking it drifting out of step. An optional --volumes/-v flag also removes the codeman-node-modules/ codeman-dist named volumes, the scripted form of the "Resetting the build artefacts" procedure docs/docker-self-update.md already documents by hand. Safe: those two are the only named volumes this stack declares; application data and case workspaces are host bind mounts, never touched by `docker compose down` either way. Docs updated: a "Major updates" section in docker/README.md, and a pointer from docs/docker-self-update.md's existing "Resetting the build artefacts" troubleshooting entry. Tests: extended test/docker-entrypoint.test.ts (the existing home for Start-Codeman.sh's own static checks) with a bash -n parse check, the down-before-build-before-handoff ordering, the --volumes flag's effect, unrecognised-argument handling, and byte-for-byte agreement with Start-Codeman.sh's own override-file resolution logic (so `down` here and `up` there can never target different Compose files). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
d3ee9f23c2 |
fix(cli-registry): correct accent colours, and a real gemini/antigravity/omp rendering bug
Two related fixes, found while re-measuring stock.ts's `accent` field
against the actual rendered UI (docs/cli-registry.md flags this field as
"transcribed, not authoritative — re-measure before wiring one up"):
1. A real, user-visible bug: `.btn-toolbar.btn-run.mode-gemini`,
`.mode-antigravity` and `.mode-omp` had no override rule inside the
`html:not([data-skin="og"])` block, unlike codex/pi/grok/deepseek, which
do. The generic `.btn-toolbar.btn-run` rule in that block resolves at
higher specificity than the base sheet's per-mode pair, so all three
rendered as plain claude-blue on every skin except `og` — including
`daylight-blue`, which is the actual DEFAULT skin for a fresh install
(index.html's pre-paint script), not an edge case. Added the three
missing rules, sourced from each CLI's own already-designed og-skin
colours (no new colours invented), mirroring the exact pattern
pi/grok/deepseek already use. Also corrected the stale comment on the
pi rule, which claimed this was still broken for gemini/antigravity.
2. `stock.ts`'s `accent` field was simply wrong for most CLIs — e.g. claude
was registered as Anthropic's brand orange (#d97757) while its button
renders blue, antigravity was registered purple while it renders cyan,
pi was registered green while it renders pink. Measured each CLI's real
`border-color` from its own `.mode-<id>` rule on the og skin (the
cleanest single representative hex each entry's gradient resolves
around) and corrected all 9 non-shell entries to match. `accent` has no
reader yet (confirmed via the DECLARED_FOR_LATER guard test), so this
changes no rendered output — it's a data-accuracy fix, matching the
registry's own "transcribed, not authoritative" warning taken literally.
Also fixed a false claim in types.ts's doc comment for the field
("CSS derives every per-CLI gradient from it via --cli-accent") — no
such CSS variable exists anywhere in the codebase.
Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format:check/
check:public-assets/check:frontend-syntax all clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
|
||
|
|
d9e6ebb20a |
fix(cli-registry): address round-2 review on #458 — count-based allowlist, RUN_MODE_LAUNCH drift guard
Three of Ark0N's four "will take at merge" items, applied instead since they were straightforward to do properly: 1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on <file>::<expression> (fixed last round) closed the line-shift problem but opened a new one: every stock id was already allowlisted for session-ui.js in the `mode === '<id>'` form, so a BRAND NEW branch reusing that exact expression anywhere in the file passed unnoticed. Reproduced live (`if (this.mode === 'codex')` injected into runOpenCode()) — stayed green under the old version. Each allowlist entry now carries the exact count of approved call sites, and a new test asserts actual-vs-declared count for every key; a mismatch in either direction is real (higher = new unreviewed branch riding in on an existing approval, lower = a reviewed site was removed and the entry is now stale). Reproduced again against the fix: same injection now fails with an exact diagnostic (expected 2, found 3). 2. Added test/run-mode-launch-table-drift.test.ts. RUN_MODE_LAUNCH restates four things stock.ts already owns (label, install command, supportsCustomModel, the external-mode key set), and they agree today with nothing enforcing it. supportsCustomModel is the dangerous one: the Run-menu picker's rows come from the server-injected window.__codemanCustomModelClis (built from capabilities.customModelInjection.kind), so a CLI gaining a real injection recipe later would be OFFERED in the picker while _runCliMode silently drops the customModel field for it — the session launches on the vendor's cloud while the UI claims the local endpoint. Drives the real session-ui.js via JSDOM and compares RUN_MODE_LAUNCH against STOCK_CLIS on all four axes. 3. Inlined the "Open Question 7 in PR-B2.md" references in the allowlist reasons — PR-B2.md is a local planning doc, never part of the committed tree, so the reference was dead on arrival for anyone reading the repo. Points at the PR #458 review thread instead. 4. Added a sentence to docs/cli-registry.md naming the new frontend guard alongside the backend one it mirrors. Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format/ check:frontend-syntax all clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
88e5b7b200 |
fix(custom-model): address Ark0N's PR review — client-side probe timeout, defer "last used" past confirmation, docs, zh-CN
Four things from the maintainer's review on PR #459, all fixed: 1. Bound _getCustomModelCurrentlyLoaded's probe client-side (~800ms via Promise.race, on top of — never instead of — the route's own 5s server-side timeout). Without it, an asleep/firewalled endpoint behind a saved model list left the picker completely invisible for up to 5s after the Run menu had already closed, with no spinner or toast. `timeoutMs` is an optional param (default 800, real callers never pass it) so a test can drive it in milliseconds, same pattern as `_watchLlamaSwapLoading`'s own `pollIntervalMs` — this code runs in a JSDOM window's own realm, whose setTimeout vi.useFakeTimers() cannot patch. 2. "Last used" is now written only once a launch actually applies, never on the mere click. It moved out of runCustomModelEntry (unconditional) and into each path's own success point: _quickStartWithCustomModelConfirm after the final post succeeds, and _runCustomModelEntryViaRestart right after the apply's success check. A context-window-warning decline means this exact model cannot work with this CLI at all, so the old unconditional write would promote, next time the picker opened, the one model guaranteed to fail again. 3. Documented the promotion/tag precedence and the new codeman:customModelLastUsed:<mode>:<endpointId> localStorage key in both CLAUDE.md's Custom Model Endpoint Profiles section and docs/custom-model-endpoints.md's Run-menu picker section. 4. Added zh-CN entries for "Currently loaded" and "Last used" in i18n.js, next to this modal's existing "Choose a model"/"Custom Endpoints" pair. New tests: the client-side timeout (endpoint that never answers, one that answers within the bound, and a rejected-after-timeout probe settling quietly), and "last used" recording on success vs. NOT recording on either confirmation's decline, for both the restart and one-shot paths. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
73607663fd |
fix(custom-model): guard the picker's async open against a slower, superseded probe
Code review (high effort) on the previous commit found a real race: making _openCustomModelPickModal async (it now awaits the currently-loaded-model probe before rendering) meant a second, faster call for a different endpoint could render first, only for the first call's slower probe to resolve afterwards and overwrite the modal with the wrong endpoint's model list — while _pendingCustomModelPick (set synchronously, before either await) still named the second, correct endpoint. Picking a model in that state would launch/apply the wrong model on the wrong endpoint. Fixed with the same mutable-generation-counter guard _watchLlamaSwapLoading already uses for an identical async-superseded-by- newer-call shape: every DOM write, including _pendingCustomModelPick itself, is deferred until after the awaited probe, and a call that finds its generation already superseded bails out untouched instead of clobbering whatever a newer call already rendered. Added a regression test driving two overlapping opens with a controlled promise so the earlier, slower probe resolves after the later, faster one renders, asserting the late response is a no-op. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
458ca578e7 |
feat(custom-model): promote the currently-loaded/last-used model in the Run-menu picker
Custom Model Endpoint Profiles' "which model" picker (session-ui.js's _openCustomModelPickModal) always listed models in their raw discovery order, so on a host with several downloaded GGUFs the user had to remember (or eyeball the "Default" tag) which one llama-swap actually had hot before picking — the whole point of the picker being fast is undone if it makes you think first. The picker now promotes exactly one model to the top of the list: - If llama-swap reports a model from this host's own list `ready` right now (via the existing GET /api/model-endpoints/:id/running-status route), that model is promoted and tagged "Currently loaded" — it's what a launch attaches to with zero wait. - Otherwise, the last model actually launched on this exact (harness, endpoint) pair is promoted and tagged "Last used", read from a new per-device localStorage key (codeman:customModelLastUsed:<mode>:<endpointId>), written by runCustomModelEntry on every launch attempt regardless of outcome. - A plain (non-llama-swap) OpenAI-compatible server, an unreachable endpoint, or a loaded-but-not-yet-ready model never promotes anything — the rest of the list keeps its discovery order. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
2df9355367 |
fix(cli-registry): address PR B2 review — fix two test guards, drop unused catalogue
Two required fixes from Ark0N's review of #458: 1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on <file>::<line>::<expression>. A single inserted line anywhere above an entry shifted every subsequent line number, so all 21 entries went stale simultaneously and the same 21 branches were reported as "new" — on a file six other open PRs also touch. Dropped the line number from the key (<file>::<expression>, matching the backend guard's own design), which collapses 21 line-keyed entries to 11 or-collapse where the same expression recurs at multiple call sites in the same file. 2. test/run-mode-ui.test.ts's terminal-ownership guard scanned method bodies via `^ {2}async (run[A-Za-z]*)\(\) \{$`, which matched the 8 one-line run<Mode>() wrappers PR B2 introduced but not _runCliMode(mode), where the real logic (and the actual risk the guard exists to catch) now lives. Fixed the regex to `^ {2}async (_?run[A-Za-z]*)\(\w*\) \{$` and added _runCliMode to the sanity list. Same-class fix in test/opencode-resize.test.ts, which had the identical blind spot via runOpenCode.toString(). Both reproduced live before fixing (inserted the same comment line; added this.terminal.clear() to _runCliMode) to confirm the bug, then confirmed the fix catches it and the suite stays green otherwise. Also resolves Open Question 2 by dropping window.__codemanCliCatalog entirely: nothing consumed it, and a registry DECLARED_FOR_LATER field costs nothing until read while an unconsumed script tag on every page render is a different trade. Reverts Phase 1 cleanly — server.ts's injection, shortBadge back in types.ts's DECLARED_FOR_LATER list and the pinned guard test, and the three associated render-index-html.test.ts / server-index-title.test.ts assertions. Full gate: 405 files / 7717 tests / 0 failures (net unchanged), typecheck/ lint/format clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
cd64b0a3f7 |
feat(cli-registry): drive the run-menu frontend from the CLI catalogue (PR B2)
PR #380 (PR B) held back the frontend half of the CLI registry refactor, explicitly deferring window.__codemanCliCatalog and making session-ui.js / mobile-overview.js catalogue-driven as "PR B2". - Inject window.__codemanCliCatalog in renderIndexHtml(), following the existing __codemanCustomModelClis pattern (escapeScriptJson-guarded, resolved per-request). Reading CliEntry.shortBadge here is what makes it genuinely read, so it drops out of types.ts's DECLARED_FOR_LATER list. - Consolidate session-ui.js's 8 near-duplicate run<Mode>() launch functions (opencode/codex/gemini/antigravity/pi/omp/grok/deepseek) into one shared _runCliMode() plus a local RUN_MODE_LAUNCH config table. The 8 method names stay as thin wrappers (index.html calls them by name; tests assert on the name). Also collapses a duplicated 8-way isAltMode/isExternalCli OR-chain (same expression, copy-pasted twice in openSessionOptions) into one EXTERNAL_CLI_MODES check. - Add test/frontend-cli-no-id-branching.test.ts, a guard scoped to session-ui.js/mobile-overview.js only (not the rest of src/web/public/, which stays explicitly out of scope per CLAUDE.md), mirroring the backend's own no-id-branching guard. mobile-overview.js and the wiring of accent/echo/wheelForward/ keyboardAccessory were investigated and deliberately left alone: the first is already a single, tested, gated table (not duplicated logic); the second set belongs to terminal-ui.js/keyboard-accessory.js/styles.css, files outside this PR's mandate. Verified on a tmux-capable devbox (this sandbox has no tmux): full CI gate at 405 files / 7717 tests / 0 failures, typecheck clean, 94 targeted tests covering exact per-CLI wire-body shapes unmodified and passing, and a live anti-vacuity check on the new guard (injected a real branch, confirmed it fails, reverted, confirmed green). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
2d573d8a34 | Merge branch 'master' of https://github.com/Ark0N/Codeman into followups | ||
|
|
5fc391a47c |
fix(custom-model): address fourth pre-merge review + merge upstream master (Ark0N)
Merged upstream/master (22 commits: reboot-restore recovery feature,
terminal keycode229 recovery work, install.sh/CLI-catalog generator
changes, CHANGELOG/version bump to 1.30.0) into this branch. No
conflicts; git auto-merged every overlapping file (CLAUDE.md,
docs/api-reference.md, app.js, index.html, styles.css, routes/index.ts,
session-routes.ts, schemas.ts, server.ts).
Two required fixes from the latest review:
1. privilegedEnvKeys widening (stock.ts) changes behaviour outside this
feature. The reviewer decided to keep both CLAUDE_CODE_MAX_CONTEXT_TOKENS
and CLAUDE_CONFIG_DIR listed (types.ts's rule that every traffic-
redirecting var this feature introduces must appear there stays
literally true), and asked for the real consequences documented
instead of hidden:
- Corrected session-env-clamp.ts's fileoverview, which stated the
opposite of what the code now does (reboot-restore's clamp call
used to be able to strip nothing for claude; it now strips a
persisted CLAUDE_CONFIG_DIR for a non-granted owner).
- Corrected the rationale comments in stock.ts: privilegedEnvKeys
has exactly one consumer (ownerClampedEnvKeys, feeding the
generic envOverrides clamp on create/quick-start/reboot-restore),
not the custom-model routes.
- Added a CLAUDE.md line to the CLAUDE_CONFIG_DIR gotcha covering
the admin-only-in-multi-user-mode and reboot-restore-strips-it
consequences.
- Added a "Claude multi-user clamp" test next to the existing
DeepSeek/OMP ones, pinning the new stripping behaviour.
2. GET .../running-status (custom-model-routes.ts) no longer passes
the raw llama-swap `cmd` field (the literal launch line, which can
carry model paths and --api-key) to the browser -- the frontend
only ever reads model/state, cmd exists solely for server-side
parseCtxFromCmd() during discovery. Added a test asserting the
response never contains cmd or a planted secret.
Also regenerated config/clis.stock.json and install.sh's catalogue
block (npm run generate:cli-catalog) to clear drift introduced by the
upstream merge, since it was failing the sync check.
Left to the reviewer, as they said they'd take at merge: the two
"comments pointing at removed code" cleanups, the two stale CLAUDE.md
counts, and the small items list (mode==='claude' frontend branch,
isCliAvailable() unknown-id gap, shared confirmed flag ordering,
one-shot cancel toast severity, pumpLlamaSwapLogTail buffer cap).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
|
||
|
|
55a80eab86 | Merge branch 'master' of https://github.com/Ark0N/Codeman into followups | ||
|
|
56209e7829 | Merge remote-tracking branch 'upstream/master' into feature/run-menu-custom-model-picker | ||
|
|
afb6754453 |
fix(custom-model): address third pre-merge review (Ark0N)
Blocker 1: the loading banner hides itself ~200ms after it reopens. - _showCenterStatus reuses one shared DOM node; dismiss() scheduled el.hidden = true 200ms later with nothing to cancel it. On the Claude path, switchingToast.dismiss() is followed by one same- origin request (5-30ms locally) before _watchLlamaSwapLoading opens the new banner -- well inside that window -- so the stale timer fired against the shared node and hid the fresh banner, leaving the whole model-load wait with no progress text, no log line and no reachable Cancel button. - Fixed by parking the pending timeout on the element and clearing it at the top of _showCenterStatus. Added a regression test that reproduces the exact repro (open, dismiss, reopen 20ms later, advance past 200ms) alongside the existing Cancel-button DOM tests; confirmed it fails without the fix and passes with it. Blocker 2: the swap-conflict warning named other users' sessions. - Both affectedSessions scans (POST .../custom-model and quick-start) walked the whole session map with no ownership filter, so in multi- user mode a non-admin pointing their own session at a shared endpoint learned another user's session name and id -- which with autoNameSessions on is that user's own prompt. - The swap is still blocked pending confirmation regardless of ownership (a foreign session is just as real a disruption); only which ones get NAMED back to the caller is scoped, via the already-imported canAccessOwned. Added a two-owner test to test/routes/session-custom-model.test.ts covering both the foreign-owner (blocked, not named) and same-owner (named) cases. Smaller ride-along fixes: - server.ts boot recovery now passes contextLength into applyCustomModelInjection, so CLAUDE_CODE_MAX_CONTEXT_TOKENS is correctly rebuilt into _envOverrides after a restart instead of surviving only because tmux retains the old setenv. - pumpLlamaSwapLogTail's finally now deletes by IDENTITY, not just by key, so an aborted pump finishing after a newer entry was created for the same endpoint can no longer delete that newer entry and orphan its connection. - docs/custom-model-endpoints.md now notes that clearing a custom model removes injected keys by name, including CLAUDE_CONFIG_DIR -- so a session that also had CLAUDE_CONFIG_DIR set via envOverrides (the per-client-account case) silently falls back to the default account on clear. Left for later, as flagged in the review itself: the quick-start case-scaffolding/cancel ordering (real behavioural reordering across a large handler, too risky to make without a live re-test), and retiring runCustomModelEntry's mode === 'claude' branch behind a launchStrategy registry field (explicitly deferred by the reviewer to "the next one"). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R |
||
|
|
9982a1325f |
fix(custom-model): address second pre-merge review (Ark0N)
Blocker: .center-status-banner never actually disappears.
- Add `.center-status-banner[hidden] { display: none; }`, same trap as
`.home-sessions[hidden]`: the author-level `display: flex` beat the
UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
card stayed laid out at `opacity: 0` with its text/cancel/close
children still `pointer-events: auto` -- an invisible 442x67 click
blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
banner (10001) and the swap-confirm/context-warning modals (10010)
in CLAUDE.md's Z-index layers list.
Stale wording pointed at the reverted sticky-toast default:
- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
`.toast-message` comment in styles.css all still said "toasts
default to sticky" after
|
||
|
|
1f32128ca9 |
fix(custom-model): address PR #430 pre-merge review (Ark0N)
Four blockers from the 2026-09-18 review: - PUT /api/model-endpoints/:id now merges modelContextLengths/ modelSizesGB back in from the stored record instead of trusting the editor's body, so renaming an endpoint or changing its default model no longer silently drops the context-window floor check and CLAUDE_CODE_MAX_CONTEXT_TOKENS injection. - custom-model:swapped-out is now session-scoped (added to SESSION_PREFIXES) instead of broadcasting to every connected client. - The quick-start custom-model path now hands setCustomModel() only the endpoint's own injected env vars, not the full merged set, matching the restart-in-place path — the full set put CLAUDE_CODE_EFFORT_LEVEL back after the Session constructor had already stripped it. - The quick-start launchModel override for pi/grok/omp is now applied generically via the registry's legacyConfigField, mirroring Session._withCustomModelLaunchModel, instead of three hardcoded mode === '<id>' branches a future CLI's injection recipe would miss. Also scopes the sticky-toast default (item 5): reverted the blanket "all error toasts are sticky" default, which had no container cap or eviction, back to a flat 3s; the one message that needs a moment to read (a failed custom-model apply) now passes an explicit duration: 0 at its own call site. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R |
||
|
|
e2034177c5 |
fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.
Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
published dependencies) into a scratch dir purely to read
dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
same endpoint. dsh's own error template ("DeepSeek API error (HTTP
${status})") reproduces the originally-reported
"dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.
- New registry field `appendV1Suffix` (env kind only, deepseek's entry
alone — claude/gemini must NOT get it, since claude was already
confirmed working against the unmodified baseUrl). When set,
buildCustomModelInjection runs endpoint.baseUrl through the same
withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
use, instead of writing it verbatim.
Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.
2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
8520925e76 |
docs(custom-model): bring CLAUDE.md and api-reference.md up to date
Full documentation review pass across the branch's 30 commits.
CLAUDE.md's Custom Model Endpoint Profiles entry hadn't been touched
since the initial backend+picker cut (3 early commits) despite 27
follow-up commits adding real behavior — it described restart-in-place
as universal (now claude-only; 7 other CLIs launch one-shot) and
claimed codex's Responses-API gap as a flat protocol break (now
re-verified as a more precise tool-calling gap). Corrected both and
added a new paragraph covering everything landed since: the llama-swap
conflict check, the after-the-fact swap-displacement sweep, the
/running-cmd-based context-length fix, the context-window floor
warning, skipFirstRunPrompts, the real-time /api/events-based log
status, and the countdown-to-Cancel-button change.
docs/api-reference.md's custom-model-endpoints section was missing the
running-status route, the requiresConfirmation/requiresContextWarning
response shapes, and POST /api/quick-start's customModel field
entirely (the primary launch path for 7 of 8 supported CLIs) — added
all three. Also fixed a real markdown bug in custom-model-endpoints.md:
an inline code span (`POST <baseUrl>/v1/chat/completions`) split across
a line break, which CommonMark renders with the line ending collapsed
to a space, so it displayed as ".../v1/chat/ completions" with a
spurious space inside the path.
Verified: origin/master and upstream/master are both already an
ancestor of this branch (identical at
|
||
|
|
db9729e1fc |
feat(custom-model): remove loading-banner countdown, add manual Cancel
Replaces the size-scaled expected-time estimate + matching auto-timeout with a generic hardware/model-size disclaimer and a user-driven Cancel button, per explicit request. Real load time depends on hardware this feature has no way to know (VRAM, storage speed, GPU contention), so the old estimate/timeout was a guess dressed up as a fact — worse, one that could kill a genuinely slow load partway through on slower hardware. - _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline entirely — polls indefinitely until ready or cancelled, no automatic give-up. Message is now "Loading <model> (<size>) on <endpoint> — this can take a while depending on your hardware and the model size.", with the real llama.cpp log line still on its own second line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/ _formatRemaining (dead code once the countdown is gone) — _lookupModelSizeGB is kept, the GB figure still shows. - _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real "Cancel" button (distinct from the error-type "×" close button, since Cancel has a real consequence) that calls it on click. Caller owns what cancelling actually means, same split as the swap-confirm modal's promise-resolving buttons. - Cancelling dismisses the banner, shows an info toast (not an error — this was deliberate), and closes the session, mirroring what the old timeout used to do automatically but now on the user's own call. - New .center-status-cancel CSS (bordered pill button, distinct from the plain "×" close glyph). Test changes: removed the now-invalid timeout-auto-close/estimate tests, added cancel-flow tests (dismiss/toast-type/session-close, never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests for the new Cancel button (bootAppWithRealCenterStatus, evaluating panels-ui.js instead of stubbing _showCenterStatus, since this button is worth verifying for real rather than just through the stub every other test in the file uses). Typecheck/lint/frontend-syntax clean; full suite shows no new regressions. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
2d3fc65758 |
feat(custom-model): show real-time llama.cpp backend status in the loading banner
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.
- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
(custom-model-routes.ts): one persistent GET /api/events (SSE)
connection held open per endpoint, parsing logData frames and
keeping the latest source:"upstream" (backend llama-server) line —
filtering out llama-swap's own source:"proxy" request-access lines.
Idle-closed after 30s of no polling, same 20s sweep as the existing
swap-displacement check.
- running-status route now returns logLine alongside the existing
isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
("llama.cpp: <line>", bootlog timestamp/level/component prefix
stripped for display) that stays on the last real thing llama.cpp
said rather than clearing to blank between polls.
⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).
12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
5ddc028a2f |
feat(custom-model): detect and notify when a session's model gets swapped out later
The llama-swap conflict check on the apply/create routes only ever runs at THAT session's own launch/apply moment, and cannot see a swap caused by a DIFFERENT session's later, ordinary use. Confirmed live: a second Codex session picking a different model launched with no warning at all — nothing conflicted at that exact instant — yet it silently evicted the first session's model regardless (llama.cpp runs one model at a time). Reproduced and root-caused via direct API calls against a live test-picker instance rather than guessing. - detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups live sessions with a customModel by endpointId, checks each group's endpoint via GET /running once, and flags a session whose own modelId is no longer in the running list. Read-only, best-effort per endpoint like refreshAllCustomModelHosts's sibling sweep. - Notifies once per displacement via a caller-owned de-dupe Set: a session id is added when displaced, removed once its own model is loaded/ready again, so a later genuinely-new displacement can notify again. - New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS, 20s — much shorter than the 5-minute model-list refresh, since this is time-sensitive) broadcasts a new custom-model:swapped-out SSE event per displacement. De-dupe Set cleared per-session on session cleanup to avoid an unbounded leak. - Frontend: global toast (not tied to the displaced session's tab, since the point is warning before the user types into it) naming the session, its previous model, and what's currently loaded. Chose the "detect after the fact" scope (vs. checking before every message send, which would add a round-trip to every turn on every custom-model session) per explicit user decision after being presented the trade-off. 9 new tests for the detection logic (flag/clear/re-flag cycle, unreachable/deleted endpoints, non-llama-swap servers, multiple sessions on one endpoint). SSE registry bumped 158->159, parity test passing. Typecheck/lint/frontend-syntax clean; full suite shows no new regressions (9 more passing than baseline, matching the new tests; same pre-existing Windows-environment failures). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
470f75b08c |
docs(custom-model): record live findings on codex's model-metadata warning
Investigated the user's report of "Model metadata for <id> not found. Defaulting to fallback metadata..." on every custom-endpoint codex launch, live against the test-picker's llama-swap deployment (codex 0.152.1): - The warning is cosmetic. `codex exec 'reply with just OK'` against the isolated CODEX_HOME still printed the warning and still returned a real reply. - The isolated CODEX_HOME never gets a models_cache.json written into it at all, even after extended real use (inspected a live, actively- used directory) — codex can't reach OpenAI's own hosted model catalog for this session and silently falls back every time, with no local file to create or clean up. There is also no config.toml override for a model's metadata. - Fabricating a fake catalog entry to suppress it would mean copying the SHAPE of OpenAI's own proprietary models_cache.json schema, including real per-model system-prompt content visible in a genuine entry — not something to build for a warning confirmed to have no effect. - More importantly: a real tool-call attempt against the same setup came back as agent_message TEXT (the tool-call JSON printed as the answer) rather than an executable function_call item, confirmed via `codex exec --json`'s raw event stream. Tool execution is what makes codex a coding agent, so it remains not usable for real work regardless of the metadata warning — a more precise, re-verified update to the existing "Responses API protocol gap" finding (which reported a harder Reconnecting/high-demand failure on a different llama-swap deployment; this one answers /v1/responses for plain chat but still can't execute tools). No code changes — recipe/comment/confidence-table documentation only. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
211b872335 |
feat(custom-model): skip Claude Code's first-run wizard on custom-model launches
A fresh, isolated CLAUDE_CONFIG_DIR (used to keep an injected API key away from a stored claude.ai OAuth login) looks like a brand-new Claude Code profile to the CLI, so it replays its ENTIRE first-run sequence on every single launch: the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running with --dangerously-skip-permissions) a one-time bypass-permissions warning — confirmed live, none of which a real, already-onboarded profile shows again. - New registry-declared env-kind field `skipFirstRunPrompts` (alongside apiKeyTrustFile, which it reuses) — claude's entry only, carried through buildCustomModelInjection (pure) into applyCustomModelInjection (IO). - seedFirstRunOnboardingState(): merges hasCompletedOnboarding: true and this session's own projects[workingDir].hasTrustDialogAccepted: true into the same <configDir>/.claude.json the API-key trust file already writes to — other projects and other fields on this session's own entry are left untouched. - seedSkipBypassPermissionsPrompt(): merges skipDangerousModePermissionPrompt: true into <configDir>/settings.json, a separate file, same corrupt-tolerant merge behavior. - applyCustomModelInjection() gains an optional workingDir parameter, threaded from session.workingDir (dedicated apply route) / resolvedCasePath (quick-start route) — boot recovery omits it (a dialog already answered once needs no re-seed on the same, persisted isolated directory). Tests added at the pure-builder, IO-wrapper (including merge-preserves- other-fields and corrupt-file-tolerance cases), and existing directory- listing assertions updated for the new settings.json file. Typecheck/ lint/format clean; full suite shows no new regressions (baseline pre-existing Windows-environment failures unchanged, 8 more passing tests than before — the ones added here). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
2c89359d42 |
fix(custom-model): Cancel/Launch-anyway buttons stacked instead of side by side
Neither dialog's footer had a row layout of its own to override, and .btn-toolbar is display:flex (a block-level flex container with no explicit inline-flex), so with no flex row context each button took its own full-width line and the two stacked vertically. The swap-confirm modal already had a .modal-footer rule (flex-end); the context-warning modal had none at all. Both now share one row-layout rule, centred rather than flex-end per feedback. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
962029bb3d |
fix(custom-model): context-warning/swap-confirm modals hidden behind status banner
Both dialogs can appear while the centred llama-swap status banner is still on screen (right after "Claude started — switching to llama-swap…") — the banner's z-index is 10001, .modal's base z-index is only 1000, so the dialog rendered fully behind it. Reported live against the context-window-too-small modal; the swap-confirm modal has the same structural bug for the same reason, so both get the fix. Also: both messages ARE the modal's whole explanatory content, not a one-line caption under a form field, so .form-hint's 0.65rem caption size read as illegibly small — worst on the multi-sentence context-window explanation. Bumped to 0.85rem/1.5 line-height/--text. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
b45a96358e |
feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.
- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
registry entry declares contextLengthVar (currently only claude) and
the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
(40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
quick-start customModel path) check this before the swap-conflict
check and before launching/restarting anything, returning
{requiresContextWarning, modelId, contextLength, minSafeContextTokens}
— skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
_resolveContextWarningConfirm (session-ui.js), wired into both
_quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
(the path Claude actually uses) ahead of the swap-confirmation check.
Explains the fix in-modal: give the model an explicit larger -c/
--ctx-size in llama-swap instead of relying on --fit-ctx, which
optimizes for the biggest model that fits rather than the biggest
context.
Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
993710263d |
fix(custom-model): stop trusting /props's n_ctx, parse the real context size from /running's cmd
Root cause of the context-overflow regression reported live: "API Error: 400 request (36437 tokens) exceeds the available context size (16384 tokens)". Discovery had stored modelContextLengths.qwen3.8-27b-ud-q4_k_xl = 154112, so CLAUDE_CODE_MAX_CONTEXT_TOKENS told Claude Code it had a huge window and it never compacted - but the real llama-swap server was launched with --fit-ctx 16384 (confirmed against /running's own cmd field) and refused the request right at that real limit. /props?model=<id>'s n_ctx (the field discovery read) is confirmed live to be unreliable for a --fit-ctx-launched backend: it reported 154112 for the same model /running says was launched with --fit-ctx 16384 - appears to report the model's theoretical/trained maximum context, not the runtime- configured one. discoverModels() now parses the REAL configured size straight out of llama-swap's own launch command instead (parseCtxFromCmd(), reading /running's cmd field - --fit-ctx first, then the plain llama.cpp -c/ --ctx-size a hand-written command might use), and only falls back to the old /props probe when cmd states no recognizable flag at all. One /running call now covers every loaded model's context length in a single request, same as it already did for the swap-conflict check and the load trigger. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
7bbe408e44 |
feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout (updated every poll, so every second by default) instead of a static "this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1-3 min) on llama-swap - 47s remaining". If the countdown reaches zero and the model still isn't ready, this is now treated as a real failure rather than a "keep waiting" shrug: - The banner turns into a sticky error (_showCenterStatus gains a `type` option - 'error' drops the spinner and adds a close button, since nothing is "in progress" anymore and a sticky message needs a way to dismiss it), naming the llama-swap server's own logs as where to look for detail. - The session that load was for is closed automatically (closeSession) - requested explicitly: a console left open and pointed at a model that never finished loading is worse than no console at all. Both apply paths now thread the new session's id through to _watchLlamaSwapLoading for this (new required 3rd parameter, after endpointId/modelId). _watchLlamaSwapGeneration's existing stale-call guard extends naturally to this: a superseded call's own eventual timeout recognises it no longer owns the banner and neither shows the error nor closes a session that may by then belong to a different, newer launch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
55dae31530 |
feat(custom-model): estimate model load time from its discovered size
Discovery now also parses a GB figure out of an auto-discovered model's own description (llama-swap writes "Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike context length this needs no /props probe (the figure is right there in /v1/models) so it is populated for every model regardless of loaded state. A hand-configured profile's own description has no such figure and correctly gets no entry. The loading banner (_watchLlamaSwapLoading) now looks this up and, when known, shows it plus a rough estimate from a small size->time matrix (_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) - "Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on llama-swap... this can take a while" - and uses that same estimate's own bracket to scale the banner's default give-up timeout for a very large model, instead of a flat 5 minutes for everything. Explicitly labelled as an UNMEASURED, typical-hardware estimate in every relevant comment - this is not benchmarked against any real endpoint's actual storage/GPU, just a reasonable expectation-setter. A model with no discoverable size (a hand-configured profile) gets no size/estimate shown at all, matching the "never a guess" convention modelContextLengths already established. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
0af233c96c |
fix(custom-model): poll llama-swap readiness every 1s, check immediately, extend the cap
Reported: the "Loading..." banner stayed up past 2 minutes even though llama-swap itself had already finished loading the model. Three fixes: 1. pollIntervalMs default 3000ms -> 1000ms (as asked). 2. The loop now checks readiness IMMEDIATELY on entry rather than sleeping a full interval first - a model that's already ready (a fast load, or a re-apply onto one already loaded) shouldn't sit on "Loading..." at all. 3. maxWaitMs default 120000ms (2 min) -> 300000ms (5 min): a large (20GB+) model reading from disk can genuinely take longer than 2 minutes, which would have looked identical to the reported symptom - "still stuck past the point it should have resolved" - except it would have actually flipped to a "still waiting" warning toast at the 2-minute mark rather than staying on "Loading" indefinitely, so this alone doesn't explain what was reported, but is a real, separate improvement worth making. Also fixes a real, separate bug this surfaced while reasoning through the report: _showCenterStatus's banner is ONE shared, reused DOM node. A second call to _watchLlamaSwapLoading (e.g. switching models again before the first switch's loop had finished) would take over that shared banner, but the FIRST loop was still running and would eventually dismiss or overwrite it once ITS OWN deadline or readiness check resolved - clobbering whatever the second, current loop had put there. A generation counter (_watchLlamaSwapGeneration) now lets each call recognise when it no longer owns the banner and stop touching it silently, rather than only the last call to actually start ever safely reading or writing it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
0929694012 |
fix(custom-model): actually trigger the llama-swap load, not just watch for it
Root cause of "it doesn't look like llama-swap is actually switching the model" (confirmed live: no load_model line in llama-swap's own logs after applying a selection). llama-swap has no "switch model" admin endpoint - the ONLY thing that starts a swap is a real inference request naming the model. Every previous fix (the conflict check, the loading banner) assumed a swap would start on its own; nothing ever actually asked llama-swap to load anything until the launched CLI's first real prompt did, which could be much later than "applying the selection" implied. Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest real request that will start a load - POST <baseUrl>/v1/chat/completions, max_tokens: 1, one throwaway message - fire-and-forget (never awaited by the caller; the frontend's own running-status polling is what actually confirms readiness). Wired into both apply paths (the dedicated restart route and the one-shot quick-start route), fired whenever the target model isn't already the one loaded and ready - a broader condition than the existing swapNeeded (which only gates the "this will evict another session's model" confirmation ask and deliberately stays narrow to that). modelSwapInProgress in both routes' responses now reflects this same broader condition too, so the frontend's loading banner actually correlates with a real in-flight load rather than only firing when something else happened to be loaded already. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
01b32ee6cd |
fix(custom-model): move the switching/loading status to a centred banner
The "Claude started - switching to <endpoint>..." and "Loading <model> on <endpoint>... this can take a while" messages lived in the top-right toast corner along with everything else, easy to miss given they can each sit on screen for well over a minute (a real llama-swap model load). Adds _showCenterStatus() (panels-ui.js): a single, reused, screen-centred banner with a spinner, non-blocking (no backdrop, pointer-events: none on the wrapper) so it never gets in the way of using the app while it's up. Both call sites (_runCustomModelEntryViaRestart's switching message, _watchLlamaSwapLoading's loading message) now use it instead of showToast. Every OTHER status in these two flows - the llama-swap conflict warning already moved to its own modal, apply failures, cancellation, and _watchLlamaSwapLoading's own final "ready"/"still waiting" outcome - stays exactly where it was, in the corner. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
2936ba6e3d |
fix(custom-model): replace the native confirm() popup with an in-app modal
The llama-swap "this will unload it for session X" warning used a native browser confirm() popup, which looks out of place next to the rest of the app's own modals. Adds #customModelSwapConfirmModal (index.html) with Cancel/Switch-anyway buttons, styled to match the app. _confirmModelSwap(message) shows it and returns a promise that resolves true/false the same way confirm() would; _resolveModelSwapConfirm(proceed) (wired to both buttons and the backdrop click) settles it. Both llama-swap conflict call sites (_quickStartWithCustomModelConfirm for the one-shot launch path, _runCustomModelEntryViaRestart for Claude's restart path) now await this instead of calling confirm() directly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
b1db5515d7 | Merge branch 'master' of https://github.com/Ark0N/Codeman into followups | ||
|
|
83033b4299 |
fix(custom-model): show a status toast during Claude's native-boot-then-restart window
Claude stays on the launch-then-restart path (see runCustomModelEntry's own comment for why), but with nothing on screen during that window, a native boot that briefly talks to the cloud model read as "the endpoint didn't apply" rather than "the switch hasn't happened yet". A sticky "Claude started - switching to <endpoint>..." toast now covers the whole window from the native launch through the apply call, updated in place (never stacked) as the outcome resolves: dismissed on cancel or failure (replaced by the existing cancellation/error toast), handed off to _watchLlamaSwapLoading's own sticky toast when a model swap is in progress, or updated to the existing "Pointed at ... - restarting" message and auto-dismissed after 3s on a plain success. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
f865f74a0f |
feat(custom-model): launch directly on the endpoint, no restart, for 7 of 8 CLIs
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.
POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.
Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.
Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
fbee1b2d82 |
docs(changeset): add changeset for the Run-menu custom-model picker PR
Covers #430's full scope so far: the picker itself, the model-selection dialog, periodic re-discovery, and the session-busy/toast/CLAUDE_CONFIG_DIR/ context-length/llama-swap-conflict fixes found through live validation against a real llama-swap server. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
bcebc81fcd |
feat(custom-model): detect llama-swap model conflicts before switching
Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.
1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
server) via its own GET /running, which plain llama.cpp has no concept of
at all. New GET /api/model-endpoints/:id/running-status route exposes this
read-only, for the frontend's polling loop below.
2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
what llama-swap currently has loaded. If it differs from the requested
model AND another live session's own customModel selection is actively
using that loaded model, the apply is refused with a
{requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
instead of silently switching. A "confirmed: true" field on the retry
skips the check. Switching with nothing else affected proceeds
immediately, no confirmation asked, only ever when there is something to
warn about.
3. The frontend (runCustomModelEntry) shows a native confirm() naming the
affected session(s) and the model they'd lose, matching this codebase's
existing convention for this class of decision (delete case, kill
session, etc.) rather than a new modal. On a successful apply the response
also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
poll shows a sticky "Loading <model>..." toast via the new running-status
route until llama-swap reports the target model ready (bounded at 2
minutes), so a prompt sent mid-swap reads as "loading", never as silence
or an answer from whatever was loaded a moment before.
Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
25f22b9839 |
test(custom-model): update session-custom-model route test for CLAUDE_CONFIG_DIR isolation
Fixes the CI failure on the last two commits: this route test asserted an exact envKeys list for a claude-mode apply that predates the CLAUDE_CONFIG_DIR isolation fix, so it failed on the new CLAUDE_CONFIG_DIR entry it correctly started appending. Updates the expected list and adds assertions for the isolated config dir path and the pre-seeded .claude.json trust-approval file, matching the behavior added in the two prior commits rather than just tolerating it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
97464bfa27 |
fix(custom-model): pre-approve the injected API key in the isolated Claude config dir
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.
Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
0e8b1981af |
fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker: 1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same config directory and warns about it (confirmed cosmetic - the API key wins for actual requests, verified via a real session's own API Usage Billing line). A custom-model claude session now gets an isolated CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty, no files written into it) so there is nothing to conflict with. projects is symlinked (junction on Windows) back into the real config dir so the response viewer, subagent windows and Read My Mind keep working for that session, best-effort. 2. Context-window overflow. Claude Code assumes a large default context window for a model id it doesn't recognise and never compacts, so a custom endpoint's real, much smaller context (verified live: a 400 exceeding a 16384-token llama-swap model with a stock ~33.7K-token system prompt) silently overflows. Discovery now also learns each model's real context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx), but ONLY for a model llama-swap's own /v1/models response already marks status.value === 'loaded' - never an unloaded one, since llama-swap treats ?model= as a routing hint and probing an unloaded model risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery. A server with no status field at all gets no enrichment rather than a guess; a model not probed this round keeps its previously-learned value until it disappears from the list entirely. Stored per model (CustomModelHost.modelContextLengths) and applied via a new contextLengthVar registry field, set to CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude. Both new fields live on the existing env-kind customModelInjection capability shape, declared only on claude's registry entry - every other CLI's injection is unaffected (pinned by test). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
5c25a52f95 |
fix(custom-model): wait for a freshly launched session to go idle before applying
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.
Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.
Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
409a6e65f9 |
fix(custom-model,toast): surface the real apply error, and make error toasts sticky with a close button
Two related fixes, both needed to actually diagnose 'Session started on the native backend — could not apply the custom endpoint' reports from live testing: 1. runCustomModelEntry()'s apply call went through _apiJson(), which unwraps a success body but SWALLOWS a failure response entirely and returns null — discarding the one thing (error, errorCode) that would tell 'endpoint unreachable' apart from 'not a discovered model', 'remote/Docker session', or a dozen other real causes the apply route already reports distinctly. Switched to _api() so the actual response body is read on failure too, and the toast now includes the real message. 2. showToast() defaulted every toast, error or not, to a 3s auto-dismiss with no way to read it again — exactly what made the above generic message impossible to act on even before the fix above. Error toasts now default to sticky (duration: 0, no auto-dismiss) unless a caller opts into a duration, and every toast — sticky or not — gets an explicit close (x) button, since a sticky toast with no way to dismiss it would just accumulate across repeated failures. Tests: custom-model-run-menu-ui.test.ts's two apply tests updated for the _api() switch (their mocks previously stubbed _apiJson, which the apply call no longer goes through), plus a new test pinning that the real server error string reaches the toast on a failure. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
9a9e542a7d |
fix(custom-model): bound the model-picker dialog's height and make its list scroll
The dialog had no max-height at all, so an endpoint with many discovered models grew it past the viewport with nothing to scroll — reported live as both "takes up the full page" and "the list is truncated", which turn out to be the same bug. Gives #customModelPickModal .modal-content the same bounded-height + scrollable-body shape cronModal's .modal-lg already uses (max-height + flex column on the content, overflow-y:auto + flex:1 on the body), scoped by id rather than folded into the shared .modal-sm class three other modals already use for short, fixed content. max-height: min(70vh, 520px) scales with the viewport (a phone gets 70% of its height; a 4K display never gets a needlessly tall dialog) rather than committing to one fixed pixel value that would be wrong at either end. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
5a9ff07f57 |
feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real llama.cpp server: 1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to apply the endpoint's defaultModelId (or the first discovered model) silently. Now, via the new selectCustomModelEntry() (session-ui.js): - exactly one discovered model launches straight away, same as before - two or more open a new #customModelPickModal listing every discovered model; defaultModelId (if set) is marked but never auto-chosen, since the point of asking is letting ONE launch deliberately differ from the saved default, not just confirming it The endpoint is re-fetched at click time rather than trusting anything cached from the dropdown's own render, since the model list can have changed (the sweep below, or a settings-panel edit) since it opened. runCustomModelEntry() itself — the actual launch, routed through run() for the in-flight lock, snapshot-guarded against applying to the wrong session — is unchanged; it now just always receives an explicit model id from one of these two paths instead of computing one itself. 2. Periodic re-discovery. Every saved endpoint's models now refresh automatically every 5 minutes in the background (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way as the Codex plan-usage poll it sits beside — this.cleanup.setInterval, off under testMode), so a model the server starts or stops serving shows up without another manual "Discover" click. The manual POST .../discover-models route and the new refreshAllCustomModelHosts() sweep (custom-model-routes.ts) now share one pure merge step (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId that no longer appears) rather than two copies that could drift. The sweep is best-effort per host — one endpoint being unreachable on a cycle never blocks the others — and re-reads the store before each host's write, keyed by id, so a concurrent edit or delete from the settings panel always wins over a sweep that started before it. Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated file for the sweep (kept separate from custom-model-routes.test.ts because that file's data dir is shared across every test in it — one temp HOME per FILE, not per test — which would make a sweep-touches-every-host assertion meaningless there). test/custom-model-run-menu-ui.test.ts gained a new describe block driving the real picker modal through JSDOM: single-model bypass, multi-model dialog with the default marked-not-chosen, picking a row closes the modal and launches with that exact model, the endpoint re-fetch, and the two "vanished by click time" toast paths. Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md, docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated — the last of these also caught up two sentences that had gone stale after the draft-review fixes landed (the picker routes through run() now, not a raw run*() call). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
60e1bd52f7 |
fix(custom-model): act on the draft review — unparseable onclick, unwrapped envelope, wrong-session apply, missing lock, no tests
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.
Blockers:
1. Every generated inline onclick was unparseable. JSON.stringify's own
double quotes terminated the double-quoted HTML attribute at the first
one, leaving btn.onclick null on every picker entry and every Discover/
Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
argument, the same idiom deleteCase's onclick already uses four lines
away in session-ui.js. This also closes the live-HTML-injection route
through modelId (server-controlled, from the endpoint's own /v1/models
reply): with quoting intact, a `>` inside it can no longer terminate the
<button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
like every other /api route (server.ts's preSerialization hook applies
to arrays too), so Array.isArray(hosts) was always false in production
and the picker/settings panel silently saw nothing. Both call sites now
go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
returns normally without ever changing activeSessionId, so the apply
step used to silently re-point and restart whatever session the user was
already looking at. runCustomModelEntry() now snapshots activeSessionId
before the launch and requires it to have actually changed.
Majors:
4. Routes the launch through run() itself via a temporary _runMode swap
(never persisted — setRunMode() would sync it to the server) instead of
a parallel hardcoded dispatch table, so a custom-model launch now holds
the same _runInFlight lock every other Run click gets. This also
resolves the "hardcoded runners map contradicts the PR's own design"
minor: dispatch is run()'s own, so a CLI whose customModelInjection
recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
parses markup this module generated itself) for exactly the DOM-level
facts the review said needed no Playwright and no tmux: a generated
button's onclick genuinely compiles and fires, a dangerous modelId never
produces a live element, the envelope unwrap works, the session-changed
guard holds, run() actually gets called (proving the in-flight lock
engages), and _runMode is restored afterward. Confirmed against the
pre-fix code first (reproduces btn.onclick === null exactly) so this
isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
render-index-html.test.ts for the other fixes below.
Minors:
- Generated entries now filter through isCliAvailable(), matching
_refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
(applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
and to settings-modal open) instead of always rendering; the endpoint GET
no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
redactApiKey() replaces the field with a computed apiKeySet: boolean, and
a PUT with no apiKey now keeps the stored one server-side
(applyStoredApiKey()) instead of the client resending a value it was
never given. New tests cover both directions (kept vs. replaced) by
observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
(_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
event, since the real role can resolve after settings were first opened)
— endpoint writes were already admin-only server-side, but the button
used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
(CLAUDE.md already records that exact literal turning the settings
preview into a grey slab on light skins), .run-mode-custom-models gets
the same gap: 2px .run-mode-menu's own flex gap only applies one level
up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
</script> (CliEntry.label is user-clis.json-settable, unlike
__codemanCliAvailable's booleans-only payload) via a new exported
escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
docs/api-reference.md; left the "no zh-CN for the new Models-section
group" minor unaddressed only insofar as the wider Models section (task
routing, thinking effort, etc.) has never had zh-CN coverage either —
everything this PR itself introduces (labels, hints, button text, the
Run-menu's "Custom Endpoints" header) IS translated in i18n.js.
Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
fed6582d3e |
fix(test): strip the custom-model Run-menu picker's injected script too
CI on PR #430 failed test/server-index-title.test.ts's byte-identity check: renderIndexHtml now injects a second unconditional <script> before </head> (window.__codemanCustomModelClis, added alongside the existing __codemanCliAvailable one), and the test only knew to strip the older one before comparing the rendered HTML against the raw template. Strip both. Unlike __codemanCliAvailable (an object, historically injected only where something resolved), the new one is a plain array injected unconditionally, possibly empty, so it needs stripping on every machine, not just one with CLIs installed. Verified the two replace() calls compose correctly against the exact strings server.ts actually produces (simulated in isolation; this box has no tmux, so the real WebServer-backed test file cannot run here at all -- same environment gap noted throughout this PR's review). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
98d26e14d9 |
docs(wiki): document Custom Model Endpoints and the Run-menu picker
New docs/wiki/Custom-Model-Endpoints.md (auto-synced to the live GitHub wiki on push to master, per docs/wiki/Contributing.md) covers turning the feature on, adding an endpoint, the Run-menu picker's one-off-run behaviour, the per-harness confidence table, and what it deliberately does not do yet (remote/Docker sessions, live hot-swap). Linked from the sidebar, from Agent-CLIs.md's "Read next" list plus a short pointer section, and from Settings-Reference.md's Models section. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
25fae9ad10 |
feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment: "generate those entries from the saved profiles rather than a fixed duplicate per harness, and put it in a follow-up PR so this one stays the backend... The Run-menu picker is yours if you want it." Adds the frontend surface the backend has been waiting on: - Run menu: a "Custom Endpoints" section lists one entry per (harness that supports customModelInjection, saved endpoint) pair, e.g. "Claude Code (llama.cpp)". The harness list comes from window.__codemanCustomModelClis, injected at page render straight off the CLI registry's own capabilities (never a hardcoded id list in the frontend), so a CLI whose injection recipe lands later appears with no frontend change. Picking an entry runs that harness's own existing run*() function unmodified (case creation, env overrides, everything, forced to a single instance) and then applies the endpoint's default model to the session it creates via the existing POST /api/sessions/:id/custom-model route. Entries are hidden for a remote/docker active case, since that route already refuses both. - Settings: App Settings -> Models gets a "Custom model endpoints" group wiring up the customModelEndpointsEnabled toggle (declared since #393, read by nothing until now) plus CRUD against the existing /api/model-endpoints routes: list, add/edit (inline form), delete, discover models. - Backend: CustomModelHost gains an optional defaultModelId, the model the picker applies with no further choice per endpoint (one generated menu entry per CLI+endpoint pair, not per CLI+endpoint+model). The route refuses a value that isn't one of the endpoint's own discovered models, and a fresh discovery drops a default that no longer appears rather than carrying an invalid one forward. Docs: docs/custom-model-endpoints.md describes the new picker and settings panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the "backend-only" status note and documents the picker's generation mechanism. Tests: four new route tests cover defaultModelId validation, acceptance, and the drop/keep behaviour across a re-discovery; a new render-index-html test pins the __codemanCustomModelClis injection (present, agent CLIs supporting the capability, antigravity and shell excluded) and its solo-window skip. No browser test was added for the Run-menu picker itself or the settings CRUD panel (this box has no tmux, so the live server used by test:browser/test:mobile could not be exercised here) -- worth a Playwright pass before merge, same as any other frontend PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
3f2928ae73 |
chore(cli-registry): clean up dead code and stale claims left after #380
Addresses the "left as they are"/"worth knowing" items Ark0N named when merging #380 (the CLI-catalogue-driven install.sh + Docker agent image PR), none of which were correctness-blocking but all of which were real: - Removed install.sh's dead _cli_index/check_cli/get_cli_path helpers: the catalogue-driven menu and hints stopped calling them and nothing else ever did. - The generator no longer emits CLI_KIND/CLI_NPM, two bash arrays install.sh never read (the .mjs/docker-hosts.ts producers already read the JSON catalogue's kind/npmPackage fields directly, so only the bash copies were dead). - detect_all_clis now skips a disabled entry's probe entirely instead of running it and filtering the result downstream. No stock entry ships disabled today, so this closes a latent inefficiency before it is a latent bug rather than fixing an observed one. - The install hint for a launcherProfile entry (DeepSeek today) now explains in one line why it's a docs link and not a command: its own docs page documents `npm install -g @deepseek-ai/dsh`, which installs the launcher only and can't drive a pane, the exact trap the menu already avoids by withholding the command. Driven by a new generated CLI_LAUNCHER_ONLY array (from discovery.launcherProfile), not an id check, so any future launcherProfile entry gets the same caveat free. - Corrected the non-interactive-default comment: on a wget-only host, Claude's curl one-liner is filtered out of the offered list first, so the default becomes whichever npm-based entry sorts earliest instead (Codex today), not always Claude. Behaviour is unchanged — it was already printed, never silent — only the comment overclaimed. Tests: extended test/install-sh-invariants.test.ts with a positive guard for the new array and the trimmed array list, a negative guard that CLI_KIND/CLI_NPM/the three dead helpers cannot come back, and two real-bash tests (driven the same way the existing skip-menu tests are) proving a disabled entry is genuinely never probed rather than merely filtered after the fact. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011WzDjJnbK7zug8iQWnCc9z |
||
|
|
a0628a40e8 |
fix(cli-registry): address maintainer review on #380
Rebased onto current master (the one real conflict was the import line
in docker-hosts.ts Ark0N flagged; kept both), then addressed every
point from the review:
**1. Rebase.** Done — this branch now sits on current upstream/master.
**2. Agent-image special cases are data now, not an id-keyed table
outside stock.ts.** `AGENT_IMAGE_SPECIAL_CASE_IDS`/`AGENT_IMAGE_SPECIAL_CASES`
are gone. `CliDiscovery.install.agentImageLayer?: { kind: 'dedicated';
reason: string }` is a field on the registry entry itself (pi,
deepseek), `reason` is required by schema.ts, both producers
(docker-hosts.ts and cli-catalog.mjs) filter on its presence instead
of an id, and the coverage test reads it from the generated catalogue.
Also added the npm-package-name validation to the TS producer, which
only the .mjs one had — same SAFE_PACKAGE regex, duplicated
(necessarily, one side can't import the other) and now pinned
byte-identical by a new parity test.
**3. Changeset said five, it's eight.** (Not nine — see the DeepSeek
point below, which changes the true count.) Reworded to state it
structurally rather than pin a number that will go stale again.
Then the four behavior-changing findings:
- **DeepSeek was offered as a normal install option but can't actually
drive a pane.** `npm install -g @deepseek-ai/dsh` installs the
launcher only; DeepSeek ships no profile that can run standalone.
The generator now emits an empty install command for any
`launcherProfile` entry, so install.sh's menu (which requires a
non-empty command) skips it and falls through to its docs URL hint
instead — matching what the old hand-written code did before this
PR replaced it.
- **wget-only hosts lost every automatic install, including the npm
ones that never needed curl.** The menu-building loop now filters
PER ENTRY (only a command starting with `curl ` is held back) rather
than wiping the whole menu when DOWNLOADER != curl.
- **The DISPLAY/TRUSTED split and the catalogue refresh didn't hold up
under review** (refresh's only real write was the label; it ran
before the Node existence check; its own eval-detection test was
tripped by the word "eval'd" in a comment). Dropped entirely per
your own recommendation — embedded catalogue only, no network
fetch, no second array. install-sh-invariants.test.ts now asserts
the refresh/DISPLAY machinery does not exist rather than testing its
internals.
The three take-or-leave items, applied:
- `dsh_banner_probe`'s bash 3.2 empty-array bug: `${runner[@]}` →
`${runner[@]+"${runner[@]}"}`. Verified live in a real `bash:3.2.57`
container with `timeout` removed from PATH — crashed before, clean
now, full `detect_all_clis` path exercised end to end.
- `docker-agent-image-coverage.test.ts` now anchors on each layer's
`<binary> --version` proof line instead of `Dockerfile.includes(binary)`,
which stayed true if a layer were deleted but its comment survived.
- Doc drift: docs/docker-cases.md (four → five, and now describes the
data field), docker/agent.Dockerfile's "other four CLIs" comment (no
longer a magic number — CLI_NPM_PACKAGES is generated and can grow),
CLAUDE.md's install.sh size (104KB → ~112KB) and its stale mention of
the now-dropped refresh.
Verified: tsc clean, prettier clean, the full targeted suite (142
tests across the 8 affected files) green, and the full `npm test` gate
diffed BY TEST NAME against a clean upstream/master baseline run on
this same machine — identical 201-name failure set both sides (168
tests / 67 files, all pre-existing Windows-environment noise: symlinks,
PTY spawning, POSIX permission bits — none of it touching anything
this PR changes), zero new failures either side of the diff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
|
||
|
|
c5c015d648 |
docs(cli-registry): document the catalogue's consumers and the trust boundary
Adds a "Consumers outside the server" section covering the two generated artifacts, why each exists (neither install.sh nor a .mjs can import TypeScript), what is deliberately NOT exported and why, the three-rule install command trust boundary, and the bash 3.2 constraint with the offset/length window shape it forces. The adding-a-CLI checklist gains the regenerate step, since forgetting it is how the installer would keep detecting the old set while the server offers the new one — the drift this change removes, one level out. docs/docker-cases.md gains how CLI_NPM_PACKAGES is derived, why it reads the stock catalogue and not the merged registry, and a table of the four documented Dockerfile special cases with their reasons. CLAUDE.md gains a command row and names the generated block, the bash 3.2 rule and the trust boundary in its install.sh paragraph. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12 |
||
|
|
7af4dbc0f8 |
feat(docker): derive the agent image's npm CLI list from the catalogue
docker/agent.Dockerfile hardcoded the four npm-published CLIs it installs, one of the several lists that had to be kept in step with the registry by hand. It now takes them as `ARG CLI_NPM_PACKAGES`, supplied by scripts/build-agent-image.mjs from config/clis.stock.json, with the default set to today's list so a bare `docker build` still produces the same image. The arg is expanded unquoted because word splitting is what turns the list into several arguments, which is exactly why every token is validated against ^[@A-Za-z0-9][@A-Za-z0-9/._-]*$ on the producing side; a package name carrying a space or a metacharacter is refused rather than reaching the RUN line. Verified by building the layer: four packages in, four arguments out, and the default still applies with no arg. The list is filtered on each entry's `enabled` flag — the field whose absence was the maintainer's §3 finding, where a CLI shipping disabled still got baked into every image. No stock entry is disabled today, so that assertion would pass vacuously; a unit test feeds the pure helper a fabricated disabled entry so the fix is covered now rather than the first time someone ships one. ⚠️ It reads the STOCK catalogue, never the merged registry. A user's ~/.codeman/clis.json must not change what is inside an image tagged codeman/agent:base, or two machines holding that tag hold different images. Four CLIs keep hand-written layers because the registry cannot describe what makes them special: pi's --ignore-scripts, deepseek's pnpm companion and dsh-tui profile, and the three standalone installers. Rather than extend the schema for a Docker-only benefit, the coverage test requires each to carry a written reason AND still be present, so an exclusion cannot quietly become an omission. There are two producers of this command line and there have to be — the .mjs cannot import TypeScript, and src/docker-hosts.ts builds the same argv for the in-app auto-build — so a parity test pins them together, package list, arg pairs and rendered argv. Their order is pinned too: a different order is a different RUN string and so a needless cache miss between the two build paths. docker/server.Dockerfile is deliberately NOT edited (PRs #373 and #377 both modify it); its narrower list is asserted as a declared omission list instead, so the divergence is reviewable without touching the file. Also fixes the in-app hint at index.html, which the new coverage test caught still omitting omp. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12 |
||
|
|
1ca35095e7 |
refactor(install): drive CLI detection, the install menu and hints from the catalogue
install.sh carried nine search-path arrays, eighteen near-identical
check_<cli>/get_<cli>_path functions, and three separately hand-maintained
enumerations of all nine CLIs. They had to agree and did not: upstream
|
||
|
|
7d6f612ef5 |
feat(cli-registry): generate a CLI catalogue for install.sh and the Docker build
Two consumers of the registry cannot import TypeScript: `install.sh`, which runs via `curl | bash` before any checkout exists, and `scripts/build-agent-image.mjs`. Both currently hand-maintain their own CLI lists, and both have already drifted. `scripts/generate-cli-catalog.mts` (`npm run generate:cli-catalog`, plus a `--check` mode) emits from `STOCK_CLIS`: - `config/clis.stock.json` for the `.mjs` and the tests. It carries `enabled` — the field the earlier attempt omitted, which is how a disabled CLI's npm package still got baked into every agent image. - a marker-delimited block inside `install.sh`, embedded rather than fetched. The embedded copy is the FULL catalogue on purpose: the earlier design fetched it and fell back to a hardcoded two-CLI list, degrading silently on an empty response. There is no degraded mode to fall into now. The block is bash 3.2 safe: parallel indexed arrays, no associative arrays, no namerefs, no mapfile. Variable-length lists use OFFSET/LENGTH windows into one flat array rather than a delimiter, so a $HOME containing a space needs no IFS handling and `shell` (no binaries) gets length 0 and is never iterated. Search paths are emitted dir-major, matching the probe order the hand-written arrays use and `test/install-sh-detection-parity.test.ts` pins. Only fields the two consumers need are exported. `launch`/`env`/`capabilities`/ `overlays` are spawn-time concerns the server alone interprets, and a test asserts they never leak into the artifact. `main()` sits behind an `isMainModule()` guard so the sync test can import the renderers. Without it, importing the module would rewrite the artifacts as a side effect of checking them — passing always, guarding never. This commit adds the block; it does not yet delete the hand-written arrays, so the detection pin keeps measuring both against each other. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12 |
||
|
|
84f71e5704 |
test(install): pin install.sh's CLI detection paths before generating them
PR B replaces nine hand-written `*_SEARCH_PATHS` arrays in install.sh with one block generated from `STOCK_CLIS`. This lands FIRST, against the hand-written arrays, so the replacement has something to be measured against. The arrays are not uniform, which is why "generate them from the registry" is a claim rather than an obvious truth: claude alone has `~/.claude/local`, opencode alone has `~/go/bin`, opencode/codex/gemini/pi/omp carry `~/.bun/bin` while dsh/grok/agy do not, and omp's `~/.omp/bin` sits second rather than first. A generated list that silently narrows leaves a user with that CLI installed being told no AI CLI was found — upstream `b6d0f1fa` is that bug, fixed for omp by hand after it shipped. The test asserts a three-way identity: the pinned literals equal what install.sh contains today, AND equal `searchDirs x binaries` from the registry, dir-major so the probe ORDER is pinned too and not just the set. Both halves were verified to fail independently — dropping one path from install.sh fails the first, changing one `searchDirs` entry fails the second — because a pin that cannot fail is worse than no pin. A fourth case asserts every stock CLI with a binary is covered, which is the omp bug restated so it cannot recur silently. No production code changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12 |
||
|
|
b6f75b87f5 |
fix(custom-model): don't clamp DEEPSEEK_API_KEY as a privileged env key
CI caught a real regression: DEEPSEEK_API_KEY was added to deepseek's privilegedEnvKeys alongside DEEPSEEK_BASE_URL on the theory that "the pair travels together," but that contradicts the documented and tested design (clampEnvOverridesForOwner()'s own docstring in session-routes.ts) — a non-granted owner supplying their OWN DeepSeek key removes privilege rather than granting it, since the exfiltration vector is the BASE URL (which redirects the server's own forwarded key to a foreign host), not the key itself. Removed it from the list; test/deepseek-mode.test.ts's existing two clamp tests now pass again. Also swapped that test's "unrelated override" example off CODEX_HOME, which the earlier commit in this same PR legitimately made privileged (closing a real pre-existing gap, documented in PR.md) — so it stopped being a valid "unrelated" example the moment that fix landed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3 |
||
|
|
e18499aa67 |
docs(pr): drop the draft/WIP framing now that the PR is submitted
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3 |
||
|
|
61779745aa |
test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
|
||
|
|
41416566aa |
feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. Covers local hardware (llama.cpp, Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter). Off by default (customModelEndpointsEnabled, synced, default OFF). - Registry: capabilities.customModelInjection per CLI entry (env / configContentEnv / configDir / unsupported kinds) - Pure injection builder (custom-model-injection.ts) turning an endpoint + model id into the real env vars / config content per CLI - Endpoint store + CRUD routes (custom-model-hosts.ts, custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded - Session integration: Session.setCustomModel()/restartCli() (POST /api/sessions/:id/custom-model), reusing the existing respawn-pane -k primitive to restart the CLI process with new env - Multi-user hardening: every new redirect-capable env var added to its CLI's privilegedEnvKeys, closing a pre-existing gap where several were already reachable via the generic envOverrides field's prefix allowlist - Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries against a real endpoint outside the web UI, independent of tmux/sessions - Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying every CLI's injected values through a real HTTP shape Real end-to-end validation against a live llama-swap server (inside a codeman/agent:llm-test Docker image with all 9 CLI binaries) found and fixed three real bugs before they shipped: - Codex's config.toml schema was wrong ([model].default table instead of a top-level model string + [model_providers.custom]); fixing it then surfaced a genuine, documented protocol incompatibility (Codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement) - Claude Code's async session-title-generation call validates ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and hangs the whole -p invocation on an unrecognized name; documented for chunk 6, worked around in the standalone script only (--bare is NOT safe for a real interactive session, which needs hooks) - The discovery route's authStyle: 'both' option (send both Authorization and api-key headers) reliably hung a real server; removed the option entirely rather than just changing the default Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see PR.md and deployment_plan.md for the full chunk breakdown and confidence table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3 |
||
|
|
c179daf869 |
fix(docker): re-assert /opt/codeman-cli ownership every start, not just at build
/opt/codeman-cli is chowned to PUID:PGID once, at image build time, from the PUID/PGID build args. That bake only happens when the image is actually rebuilt (`docker compose up --build`, which Start-Codeman.sh always does) — a deployment that runs the compose file directly instead (Unraid's Compose Manager, a native systemd unit, any plain `docker compose up`/`restart`) can change PUID/PGID in .env and restart without ever rebuilding. The container then runs as the NEW uid via entrypoint's setpriv (Linux needs no /etc/passwd entry to setuid to an arbitrary number) while the CLI directory is still owned by the OLD one baked into the image layer — silently breaking the self-update-a-CLI- in-place fix that directory exists for. Unlike HOME/CODEMAN_CASES_PATH, this one is pure image content Codeman itself populated, never host data that might legitimately belong to someone else, so there is no ownership to be careful about — it is always correct for it to be owned by whoever the container is about to run as. Re-assert it unconditionally on every start. Verified live: built an image with PUID=99/PGID=100, ran it with PUID=1234/PGID=4321 (no rebuild, simulating a changed .env restarted directly), confirmed /opt/codeman-cli ends up 1234:4321-owned and is genuinely writable by the running process. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
ae32daf135 |
fix(docker): address maintainer review on #377
Two real bugs the review caught, both verified live against a real build on the Unraid host: 1. entrypoint.sh's chown fired on ANY ownership mismatch, not just a directory the daemon itself created root-owned. A host tree legitimately owned by some other account - an existing CODEMAN_CASES_PATH the README already allows pointing at a normal projects directory, or appdata under a different PUID/PGID convention than the one in use - got silently recursively re-owned with one log line to explain it. Now gated on the target actually being root-owned; anything else is a clean refusal naming the directory, its owner, and PUID/PGID. Start-Codeman.sh also now pre-creates CODEMAN_CASES_PATH the same way it already did CODEMAN_APPDATA_PATH, so Compose never has to materialise a missing bind source as root in the first place - the in-container chown becomes a safety net, not the primary mechanism. 2. The CLI-update chown (chown -R .../node_modules /usr/local/bin) handed the runtime account write access to entrypoint.sh itself (root-owned, executed as root on every container start with CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node binary - owning the DIRECTORY is enough to rename it aside and drop a replacement, which would let a compromised session arrange for its own script to run as root at the next restart. The four CLIs now install into a dedicated /opt/codeman-cli prefix (NPM_CONFIG_PREFIX); only that directory is chowned, /usr/local stays root-owned throughout. Smaller fixes from the same review: - Start-Codeman.sh's volume-refresh label filter wasn't project-scoped: a second Compose stack on the same host sharing the `codeman-dist` volume KEY could have had ITS volume deleted. Added a com.docker.compose.project filter, resolved from this stack's own `compose config --format json`. - Override-file precedence was backwards (checked .yaml before .yml; Compose actually prefers .yml) - swapped, plus a warning when both exist. - entrypoint.sh's setpriv now also passes --bounding-set -all, so CapBnd actually clears post-drop rather than just CapPrm/CapEff. - A comment on git_head_commit() noting it returns nothing for a worktree checkout (.git as a file), consistent with the script's existing -d .git convention elsewhere. - Doc drift: CLAUDE.md's Docker Compose section still described the old pre-created-and-chowned-by-hand model and didn't mention the root-then-drop entrypoint; the state-files list was missing docker-build-source.json; docs/docker-compose.md and docker/.env.example still had the pre-rename `Coding/codeman` path in one place each. Verified end to end against a real build on the Unraid host: a root-owned bind source is corrected as before; a directory owned by neither root nor PUID:PGID is refused rather than silently rewritten; a correctly-owned directory is left alone entirely; the four CLIs resolve via PATH from /opt/codeman-cli while /usr/local/bin, /usr/local/lib/node_modules and entrypoint.sh itself stay root-owned; CapBnd is fully cleared post-drop. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
8fe3f34fc5 |
fix(docker): detect and refresh stale build-artefact volumes
codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded from the image only while empty, so a rebuilt image's fresh dist/ node_modules sat unused behind old volume content until something cleared it. The in-app self-updater never hit this (it rebuilds INSIDE the running container, into the very volume already in use), but a `docker compose build` triggered from outside it — Start-Codeman.sh, after a manual `git pull` — did: the container came back up looking unchanged, serving stale compiled routes against current source. Start-Codeman.sh now compares the checkout's HEAD commit and package-lock.json hash against a recorded marker (docker-build-source.json) and clears just the affected volume(s) before its own --build when either moved. The in-place self-update path writes that same marker after a successful build, so the two mechanisms agree on what the volumes currently reflect — without it, the next plain Start-Codeman.sh run would see the HEAD self-update just checked out, not recognise it as already accounted for, and wipe the volumes self-update just correctly rebuilt right back to the older baked image. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
89e2cb5814 |
fix(docker): let the runtime account update its own global CLIs
The four CLIs (claude, gemini, codex, opencode) are npm-installed globally as root during the image build, before the unprivileged runtime account exists. A session running as that account (e.g. a codex-mode terminal) then hits EACCES the moment it tries to update one in place, because npm renames the old package directory aside before installing the new one, which needs write access to the parent (/usr/local/lib/node_modules), not just the target package. Chown that tree plus /usr/local/bin's CLI symlinks to PUID:PGID in the same step that creates/renames the runtime account. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
d38bf33a69 |
docs(docker): document the reverse-proxy host allowlist
CODEMAN_ALLOWED_HOSTS is a real, documented application setting (the Host- header allowlist in network-auth-policy.ts), but docker-compose.yaml does not forward it from .env into the container - Compose only passes through variables explicitly listed under environment:, and this is not one of them. Set without that passthrough, any request through a reverse proxy is rejected with 403 Forbidden: host not allowed before it reaches any handler, and nothing in the Docker deployment docs said why. Document the variable and the override needed to forward it, using the Local customisation mechanism already described above it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
9702126046 |
chore(docker): name the default runtime account codeman
CODEMAN_RUNTIME_USER defaulted to `opencode`, which no longer matches the project and is confusing in a deployment whose every other identifier is codeman. Rename the default in .env.example and in the Dockerfile ARG that mirrors it, and correct the example comment that referred to /home/opencode/codeman-cases. Also drop the `Coding/` component from the example application-data path. CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH now suggest /mnt/user/appdata/codeman and its codeman-cases child, matching the account name and removing a directory level that meant nothing outside the original author's host. README.md is updated to match, including the chown example. The npm package `opencode-ai` and the references to the OpenCode CLI are deliberately left alone: those name a different tool, not this account. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
748bbf5423 |
fix(docker): honour docker-compose.override.yml in Start-Codeman.sh
Naming a Compose file with -f disables Compose's automatic discovery of the override file, so Start-Codeman.sh silently ignored docker-compose.override.yml. Any local customisation placed in the conventional override file was dropped without warning, and the only way to notice was to inspect the running container. Collect the -f arguments into an array, append the override file when one is present, and reuse that array for the final launch so the two cannot drift apart again. Both .yml and .yaml are checked, in Compose's own precedence order, and the chosen file is reported on startup. Document the override file in docker/README.md, including the two things that are easy to get wrong: it is ignored when -f is passed without naming it, and it cannot remove a key such as ports, which Compose concatenates. Add the override file to .gitignore. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
10876aa440 |
fix(docker): correct bind-mount ownership before dropping privileges
Compose binds CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH from the host. When either path does not exist yet - a first run, a cleared application-data directory, a restored backup - the Docker daemon creates it owned by root. The server runs unprivileged as CODEMAN_RUNTIME_USER, so it cannot create its own state directory, and the container restarts forever on: Failed to start web server: EACCES: permission denied, mkdir '/home/<user>/.codeman' Start-Codeman.sh already worked around this by preparing the directory on the host, so the failure only appears when Compose is run directly, which the README documents as a supported path. Add docker/entrypoint.sh, which starts as root, corrects the ownership of both bind mounts, then drops to PUID:PGID with setpriv. The Dockerfile's USER instruction is replaced by that entrypoint and CMD is unchanged. docker-compose.yaml adds back only the four capabilities the chown and the privilege drop require, so cap_drop: ALL continues to remove everything else. Two guards keep existing deployments working: - A container started with an explicit `user:` is left alone. The entrypoint execs straight through, with no elevation and no chown. - A chown that fails is a warning, not an error. Bind mounts backed by NFS, CIFS or a rootless daemon can refuse chown while remaining perfectly writable, and those deployments must keep starting. PUID and PGID are also exported as runtime environment defaults so the image behaves correctly when run without Compose, rather than depending on build args alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
06febfa032 |
fix(file-picker): default the case picker to Codeman Cases, not Home
The "Link Existing" case picker opens with an empty path and no sessionId, so the browse endpoint's fallback root picked whichever root happened to be first in the list — which was always `Home`. On the native default that's harmless (~/codeman-cases nests inside Home anyway), but a Docker deployment binds CODEMAN_APPDATA_PATH (Home) and CODEMAN_CASES_PATH at unrelated host paths, so the picker opened somewhere with no cases in sight. Worse: if CODEMAN_CASES_PATH is ever changed after cases already exist, the old cases directory lingers, still reachable, under Home — indistinguishable at a glance from the real one under the new Codeman Cases root. Prefer the Codeman Cases root in the fallback chain, ahead of the generic roots[0]. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
66eb01ba8f |
feat(docker): restore in-app self-update in the Compose deployment
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
|
||
|
|
1125f7c1c5 |
fix(test): strip the instance-selection env vars in test/setup.ts
`test/setup.ts` gives every test file a temp HOME so the suite cannot touch the real Codeman tree, and strips the env vars that would leak past it — but the list only covered auth and the gesture flag. The three vars `src/config/instance.ts` derives the data dir and tmux socket from were missing, and they reach past the temp HOME: - **`CODEMAN_DATA_DIR` is the one that matters.** It is an ABSOLUTE override read in `getDataDir()`, so it bypasses HOME entirely: a developer who exports it — or a shell left over from `codeman web -d` — has the suite reading and WRITING their real `state.json`, `users.json`, `intents.json` and `hook-secret`. - **`CODEMAN_INSTANCE`** moves the data dir to `~/.codeman-<name>` and the socket to `codeman-<name>`. Inside the temp HOME that is not data loss, but it silently changes the paths tests assert on — and `scripts/run-beta.sh` exports it, so any shell that has run a beta carries it. - **`CODEMAN_TMUX_SOCKET`** renames the socket `resolveTmuxSocketName()` returns. `TmuxManager` no-ops its shell commands under vitest, so this is assertion drift rather than a stray `tmux -L` against prod — same class of leak, same one-line fix. They are deleted in the setup file rather than in a hook because `CODEMAN_INSTANCE` is captured into a module-level const the first time `config/instance.ts` is imported; a `beforeEach` would already be too late. `test/test-env-isolation.test.ts` pins the whole list in two halves, because the obvious half is not enough: asserting the vars are unset passes trivially on a machine that never set them, so a removed `delete` line would sail through on almost every box and on CI. The static half reads `setup.ts` and asserts each name is deleted there, which fails everywhere. An anti-drift check catches the other direction — a var stripped in `setup.ts` but never given a reason in the list — and is scoped to the strip section so the teardown's restores are not mistaken for strips. Verified by demonstrating the leak: with the `CODEMAN_DATA_DIR` line removed and the var exported, the runtime assertion fails; with the line restored it passes. Full suite: no new failures against an upstream/master baseline. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
c5b84fb5f4 |
docs(cli-registry): annotate overlays.credStore as declared-for-later
Review item 4 named THREE live tables duplicating registry data. Two are now read from the entry (`defaultRemoteCommandForMode`, `defaultDockerCommandForMode`); the third, `resolveDockerCredentialArtifacts`, is not — and it was left neither wired nor annotated, which is the state that item explicitly rules out. It is not wired because the shape cannot express the live table: `credStore` is ONE store per CLI, and `CRED_STORES` needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek's entry declares none at all even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is at once the worst thing in that file to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest. It belongs in its own change, measured against a real container. So it is annotated instead, at the field, in the type's declared-for-later header, in docs/cli-registry.md, and in the pinned DECLARED_FOR_LATER list — the last of which means wiring it later makes a test fail rather than leaving a stale comment behind. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
6acf0dea0f |
fix(cron): scope the launch pre-flight to launcher CLIs, not every mode
CI caught three cron-service failures. Both are mine, from converting cron's per-mode ladders to capability reads without checking what each ladder's scope actually was. **The pre-flight.** cron only ever pre-flighted `deepseek` — dsh is a profile LAUNCHER, so "installed" is not "runnable" and a bare `dsh` can boot a profile that cannot drive a pane. I replaced that with an unscoped `resolveCliLaunchError(mode)`, which pre-flights EVERY mode, so a claude cron job on a box with no claude binary now failed with "Claude CLI not found" instead of reaching tmux-manager's own throw. Three tests assert the latter. It is now gated on `discovery.launcherProfile !== undefined`, which is byte-identical to the `mode === 'deepseek'` check it replaces and generalises to the next launcher. The equivalent HTTP-route conversion was already scoped (to `capabilities.external`, matching what that route has always pre-flighted); I simply failed to carry the same reasoning across. **The model.** cron's ladder was `mode !== 'shell' && mode !== 'deepseek'`, and I read it as `capabilities.model.source === 'claude-settings-file'` — which is the HTTP route's question, not cron's. There, every external CLI reads its model from its own config object earlier in the chain, so only claude reaches the global default; cron has no such config, so the same expression silently narrowed the default model from eight modes to one. Now `!== 'none'`, which is exactly the two entries the ladder excluded. Not caught by a test — found by re-deriving each ladder's scope after the first failure. Also names a fourth deliberate behaviour change in the changeset, found while tracing these: `session.ts` carried a hand-written list of modes with no direct-PTY fallback and OMP was missing from it, though CLAUDE.md's own text says "all eight require tmux". `requiresMux` comes off the entry now, so an omp session whose mux creation fails refuses instead of silently starting outside tmux. Verified by diffing failing tests BY NAME against an upstream/master baseline, rather than by file as before — which is how the regression slipped through: the three new failures landed inside a file already failing for unrelated Windows-path reasons, and the aggregate count happened to collide. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
4830e662f9 |
refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. Code that used to ask "which CLI is this?" reads the entry instead. Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every spawn command as a literal string, captured from the hand-written builders before they were deleted, and `test/location-overlay-commands.test.ts` does the same for all 20 remote and in-container pane commands. Config can never contain shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only in this release. OMP is included as a registry entry rather than a tenth hand-written builder, so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of `buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen and doctor ladders all drop out. Guard rails: - `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id branching reappears outside `stock.ts`, in any of its four shapes (`===`, `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the negated forms, which is how 36 of them survived an earlier pass. Every allowlisted branch carries its reason. - `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities; deriving one from another shipped the `until=stop`-hangs-on-shell bug. - `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config` wire field is separate, bridged only by `legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. - Registry data resolves AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks). A module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. - Six fields are annotated DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/ `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured. A test pins the list so it cannot quietly grow. Three user-visible changes, all deliberate and named: - `probeDockerCliVersion()` derives the in-container binary from the registry rather than assuming it equals the mode name (`antigravity` runs `agy`). - The remote CLI version probe now covers grok and deepseek, which the hardcoded map it replaces omitted while its own comment said the rule was "every mode except shell". - `codeman doctor`'s CLI rows are generated from the entries, so Claude's install hint is the install command rather than a docs URL, five CLIs gain hints they never had, and the row order follows the catalog. Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars (matching the `cliId` pattern) before its failure message quotes the value back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading the hand-editable `clis.json`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
26b4ffbb0f | fix(docker): install pnpm for DeepSeek profile | ||
|
|
e2179bd530 | chore(docker): remove local handover references | ||
|
|
b85f7659b7 | feat(docker): add Compose deployment support |