mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
master
450
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
48f30f3055 |
style: drop em-dashes from the text added in c2114615
House style, and these land in the changelog. Only the sentences added in the previous commit are touched; the em-dashes in contributor text and in the pre-existing COD-54/COD-115 comments are left alone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c211461500 |
fix: merge-time follow-ups for #409, #404 and #399
#409 (Claude truecolor). The changeset becomes the changelog, and its premise does not hold on tmux 3.2 or newer. Measured here on tmux 3.4: `default-terminal` sits at its compiled default of `tmux-256color`, a live claude pane reports `TERM=tmux-256color`, and supports-color reads that as 256 colors, where rgb(55,55,55) lands on ESC[48;5;237m — visible, just not the color the theme named. The invisible block the PR describes needs TERM to resolve to a 16-color entry: tmux older than 3.2, or a ~/.tmux.conf setting `default-terminal screen`, which Codeman's own tmux server does read (it passes no -f). Both the changeset and the invariants paragraph now say that, so the next report here gets paired with the reporter's tmux -V instead of being read as universal. The change itself stands on the simpler argument: claude was one of two entries not asking for truecolor while twelve do. Also reorders buildClaudeEnv(). It applied the registry's unset/exports AFTER the whole env was built, so a clis.json entry naming CODEMAN_HOOK_SECRET_FILE or PATH would strip it on the direct-PTY path while the tmux pane kept it — buildEnvExports() emits `...cliEnv` ahead of `export CODEMAN_MUX=1` and cannot. The block now runs first and Codeman's own keys are assigned on top, matching the pane. #404 (Ctrl+Z trap). Adds the missing changeset, and records what the trap does not cover: an agent CLI already holds its tty with ISIG off (verified on three live panes: `susp = ^Z -isig -icanon`), so this is defence for the startup window rather than a fix for the steady state, and two input paths still reach the PTY unfiltered — the mobile accessory bar's one-shot Ctrl and the CJK textarea. #399 (path picker sort). The server sorts by name and cuts at 500, so the client sorting those 500 by date gives "the newest of the first 500 by name", which is wrong in exactly the >500-entry folder the date sort exists for. The status line now says "(first 500 by name)" so the cut is legible, with the reasoning parked on _sortEntries. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e0ebbbdc91 |
Merge pull request #409 from irisitymichaelgrundberg/fix/claude-truecolor-in-panes
fix(terminal): let Claude use truecolor so its themed backgrounds render |
||
|
|
8c237223b0 |
Merge pull request #399 from shenlvkang-collab/pr/path-picker-sort-jump
feat(files): let the path picker jump to a typed path and sort by name or date |
||
|
|
dae2ac580f |
fix(terminal): read the colour env from the registry on every local spawn path
buildClaudeEnv(), the direct-PTY fallback taken when mux creation fails, now
reads getCli('claude').env and applies its unset and exports lists. It used to
delete COLORTERM and CLAUDECODE from a hand-maintained list of its own, which
left it contradicting the registry entry that the tmux pane and the attach
client both read. An engine value needing a mux name has nothing to resolve
against on this path, so it is skipped rather than guessed.
Claude no longer unsets NO_COLOR. The invisible-background bug does not need
it, and unsetting it overrides a preference the user set deliberately, so a
user who exports NO_COLOR globally keeps monochrome panes. The other seven
truecolor CLIs still unset it; that inconsistency is intentional and the
comment on the entry says so.
The invariants doc gains a Terminal colour env paragraph under Session launch
modes, where a reader looking up Claude will find it — the previous sentence
sat under a heading that lists only the non-Claude CLIs. It now says the lists
are the stock catalog and a clis.json override replaces them wholesale, and
that the declarations reach the tmux pane, its attach client and the direct
PTY but not a remote pane, whose command carries no env exports at all. Docker
hands COLORTERM=truecolor to every mode, including the two the registry says
must unset it.
The changeset named six peer CLIs and there are seven: deepseek also exports
truecolor. A test beside the existing OpenCode assertion pins the new
behaviour, so a future registry edit cannot make the backgrounds vanish again
in silence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7767b16d4f |
fix(terminal): let Claude use truecolor so its themed backgrounds render
Claude draws the user's own messages as a block of background color, and inside a Codeman pane that block was invisible. tmux hands each pane TERM=screen, which supports-color reads as 16 colors, and Claude's registry entry deleted COLORTERM on top of that. Claude therefore quantized every RGB color its theme asked for down to the basic palette, where rgb(55, 55, 55) and every other dark background becomes ESC[40m, the terminal's own black. Changing the color in a custom Claude theme moved nothing on screen. Claude now exports COLORTERM=truecolor and unsets NO_COLOR, matching codex, gemini, antigravity, pi, grok and omp. CLAUDECODE stays unset, because Claude reads it as a signal that it is running nested inside itself. Both the tmux session and the attach client read this one registry entry, so they cannot disagree. PR #3 introduced the unset in February, citing xterm.js#484 for the claim that xterm.js mishandles truecolor. xterm.js closed that issue in April 2019, Codeman now depends on @xterm/xterm 6, and TmuxManager already sets terminal-overrides ",*:Tc" on its own tmux server, so 24-bit color reaches the browser today for every CLI that asks for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a0628a40e8 |
fix(cli-registry): address maintainer review on #380
Rebased onto current master (the one real conflict was the import line
in docker-hosts.ts Ark0N flagged; kept both), then addressed every
point from the review:
**1. Rebase.** Done — this branch now sits on current upstream/master.
**2. Agent-image special cases are data now, not an id-keyed table
outside stock.ts.** `AGENT_IMAGE_SPECIAL_CASE_IDS`/`AGENT_IMAGE_SPECIAL_CASES`
are gone. `CliDiscovery.install.agentImageLayer?: { kind: 'dedicated';
reason: string }` is a field on the registry entry itself (pi,
deepseek), `reason` is required by schema.ts, both producers
(docker-hosts.ts and cli-catalog.mjs) filter on its presence instead
of an id, and the coverage test reads it from the generated catalogue.
Also added the npm-package-name validation to the TS producer, which
only the .mjs one had — same SAFE_PACKAGE regex, duplicated
(necessarily, one side can't import the other) and now pinned
byte-identical by a new parity test.
**3. Changeset said five, it's eight.** (Not nine — see the DeepSeek
point below, which changes the true count.) Reworded to state it
structurally rather than pin a number that will go stale again.
Then the four behavior-changing findings:
- **DeepSeek was offered as a normal install option but can't actually
drive a pane.** `npm install -g @deepseek-ai/dsh` installs the
launcher only; DeepSeek ships no profile that can run standalone.
The generator now emits an empty install command for any
`launcherProfile` entry, so install.sh's menu (which requires a
non-empty command) skips it and falls through to its docs URL hint
instead — matching what the old hand-written code did before this
PR replaced it.
- **wget-only hosts lost every automatic install, including the npm
ones that never needed curl.** The menu-building loop now filters
PER ENTRY (only a command starting with `curl ` is held back) rather
than wiping the whole menu when DOWNLOADER != curl.
- **The DISPLAY/TRUSTED split and the catalogue refresh didn't hold up
under review** (refresh's only real write was the label; it ran
before the Node existence check; its own eval-detection test was
tripped by the word "eval'd" in a comment). Dropped entirely per
your own recommendation — embedded catalogue only, no network
fetch, no second array. install-sh-invariants.test.ts now asserts
the refresh/DISPLAY machinery does not exist rather than testing its
internals.
The three take-or-leave items, applied:
- `dsh_banner_probe`'s bash 3.2 empty-array bug: `${runner[@]}` →
`${runner[@]+"${runner[@]}"}`. Verified live in a real `bash:3.2.57`
container with `timeout` removed from PATH — crashed before, clean
now, full `detect_all_clis` path exercised end to end.
- `docker-agent-image-coverage.test.ts` now anchors on each layer's
`<binary> --version` proof line instead of `Dockerfile.includes(binary)`,
which stayed true if a layer were deleted but its comment survived.
- Doc drift: docs/docker-cases.md (four → five, and now describes the
data field), docker/agent.Dockerfile's "other four CLIs" comment (no
longer a magic number — CLI_NPM_PACKAGES is generated and can grow),
CLAUDE.md's install.sh size (104KB → ~112KB) and its stale mention of
the now-dropped refresh.
Verified: tsc clean, prettier clean, the full targeted suite (142
tests across the 8 affected files) green, and the full `npm test` gate
diffed BY TEST NAME against a clean upstream/master baseline run on
this same machine — identical 201-name failure set both sides (168
tests / 67 files, all pre-existing Windows-environment noise: symlinks,
PTY spawning, POSIX permission bits — none of it touching anything
this PR changes), zero new failures either side of the diff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
|
||
|
|
c5c015d648 |
docs(cli-registry): document the catalogue's consumers and the trust boundary
Adds a "Consumers outside the server" section covering the two generated artifacts, why each exists (neither install.sh nor a .mjs can import TypeScript), what is deliberately NOT exported and why, the three-rule install command trust boundary, and the bash 3.2 constraint with the offset/length window shape it forces. The adding-a-CLI checklist gains the regenerate step, since forgetting it is how the installer would keep detecting the old set while the server offers the new one — the drift this change removes, one level out. docs/docker-cases.md gains how CLI_NPM_PACKAGES is derived, why it reads the stock catalogue and not the merged registry, and a table of the four documented Dockerfile special cases with their reasons. CLAUDE.md gains a command row and names the generated block, the bash 3.2 rule and the trust boundary in its install.sh paragraph. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12 |
||
|
|
61779745aa |
test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
|
||
|
|
41416566aa |
feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. Covers local hardware (llama.cpp, Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter). Off by default (customModelEndpointsEnabled, synced, default OFF). - Registry: capabilities.customModelInjection per CLI entry (env / configContentEnv / configDir / unsupported kinds) - Pure injection builder (custom-model-injection.ts) turning an endpoint + model id into the real env vars / config content per CLI - Endpoint store + CRUD routes (custom-model-hosts.ts, custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded - Session integration: Session.setCustomModel()/restartCli() (POST /api/sessions/:id/custom-model), reusing the existing respawn-pane -k primitive to restart the CLI process with new env - Multi-user hardening: every new redirect-capable env var added to its CLI's privilegedEnvKeys, closing a pre-existing gap where several were already reachable via the generic envOverrides field's prefix allowlist - Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries against a real endpoint outside the web UI, independent of tmux/sessions - Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying every CLI's injected values through a real HTTP shape Real end-to-end validation against a live llama-swap server (inside a codeman/agent:llm-test Docker image with all 9 CLI binaries) found and fixed three real bugs before they shipped: - Codex's config.toml schema was wrong ([model].default table instead of a top-level model string + [model_providers.custom]); fixing it then surfaced a genuine, documented protocol incompatibility (Codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement) - Claude Code's async session-title-generation call validates ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and hangs the whole -p invocation on an unrecognized name; documented for chunk 6, worked around in the standalone script only (--bare is NOT safe for a real interactive session, which needs hooks) - The discovery route's authStyle: 'both' option (send both Authorization and api-key headers) reliably hung a real server; removed the option entirely rather than just changing the default Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see PR.md and deployment_plan.md for the full chunk breakdown and confidence table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3 |
||
|
|
ae32daf135 |
fix(docker): address maintainer review on #377
Two real bugs the review caught, both verified live against a real build on the Unraid host: 1. entrypoint.sh's chown fired on ANY ownership mismatch, not just a directory the daemon itself created root-owned. A host tree legitimately owned by some other account - an existing CODEMAN_CASES_PATH the README already allows pointing at a normal projects directory, or appdata under a different PUID/PGID convention than the one in use - got silently recursively re-owned with one log line to explain it. Now gated on the target actually being root-owned; anything else is a clean refusal naming the directory, its owner, and PUID/PGID. Start-Codeman.sh also now pre-creates CODEMAN_CASES_PATH the same way it already did CODEMAN_APPDATA_PATH, so Compose never has to materialise a missing bind source as root in the first place - the in-container chown becomes a safety net, not the primary mechanism. 2. The CLI-update chown (chown -R .../node_modules /usr/local/bin) handed the runtime account write access to entrypoint.sh itself (root-owned, executed as root on every container start with CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node binary - owning the DIRECTORY is enough to rename it aside and drop a replacement, which would let a compromised session arrange for its own script to run as root at the next restart. The four CLIs now install into a dedicated /opt/codeman-cli prefix (NPM_CONFIG_PREFIX); only that directory is chowned, /usr/local stays root-owned throughout. Smaller fixes from the same review: - Start-Codeman.sh's volume-refresh label filter wasn't project-scoped: a second Compose stack on the same host sharing the `codeman-dist` volume KEY could have had ITS volume deleted. Added a com.docker.compose.project filter, resolved from this stack's own `compose config --format json`. - Override-file precedence was backwards (checked .yaml before .yml; Compose actually prefers .yml) - swapped, plus a warning when both exist. - entrypoint.sh's setpriv now also passes --bounding-set -all, so CapBnd actually clears post-drop rather than just CapPrm/CapEff. - A comment on git_head_commit() noting it returns nothing for a worktree checkout (.git as a file), consistent with the script's existing -d .git convention elsewhere. - Doc drift: CLAUDE.md's Docker Compose section still described the old pre-created-and-chowned-by-hand model and didn't mention the root-then-drop entrypoint; the state-files list was missing docker-build-source.json; docs/docker-compose.md and docker/.env.example still had the pre-rename `Coding/codeman` path in one place each. Verified end to end against a real build on the Unraid host: a root-owned bind source is corrected as before; a directory owned by neither root nor PUID:PGID is refused rather than silently rewritten; a correctly-owned directory is left alone entirely; the four CLIs resolve via PATH from /opt/codeman-cli while /usr/local/bin, /usr/local/lib/node_modules and entrypoint.sh itself stay root-owned; CapBnd is fully cleared post-drop. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
8fe3f34fc5 |
fix(docker): detect and refresh stale build-artefact volumes
codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded from the image only while empty, so a rebuilt image's fresh dist/ node_modules sat unused behind old volume content until something cleared it. The in-app self-updater never hit this (it rebuilds INSIDE the running container, into the very volume already in use), but a `docker compose build` triggered from outside it — Start-Codeman.sh, after a manual `git pull` — did: the container came back up looking unchanged, serving stale compiled routes against current source. Start-Codeman.sh now compares the checkout's HEAD commit and package-lock.json hash against a recorded marker (docker-build-source.json) and clears just the affected volume(s) before its own --build when either moved. The in-place self-update path writes that same marker after a successful build, so the two mechanisms agree on what the volumes currently reflect — without it, the next plain Start-Codeman.sh run would see the HEAD self-update just checked out, not recognise it as already accounted for, and wipe the volumes self-update just correctly rebuilt right back to the older baked image. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
65ddedd1d4 |
fix: act on the 1.27.0 pre-release review
A Fable 5.1 reviewer read the whole release diff against 1.26.2 and returned
SHIP WITH FIXES. These are its findings, verified before acting on each.
**The changelog advertised a feature the code refuses (major).** The #401
changeset and docs/web-tabs.md both listed `*.localhost` in the loopback set.
The follow-up in
|
||
|
|
02b0e27898 |
fix: merge-time follow-ups for #400, #401, #362 and #388
Each item is from the pre-merge review of the PR it names, applied on master rather than by pushing to a contributor branch. #400 (response viewer, shenlvkang-collab) - The brief view opened at `scrollTop = 0`, right when it was a single card holding the last row. Now that it renders the whole turn, the top is the turn's first narration line and the answer can be screens below it, while loadFullContext already scrolls to the bottom of the same turn. A multi-row turn now opens at its newest text; a single card still opens at the top. #401 (loopback links as web tabs, shenlvkang-collab) - Drop `*.localhost` from the auto-route set. Every other member is an address literal that can only mean this box; a `*.localhost` DNS name is not one, and a resolver with a search domain retries `evil.localhost` as `evil.localhost.<search domain>`. The link source is agent-written terminal output, so that set is the whole confinement on a tap that makes Codeman fetch a URL server-side and persist it. The page-side test stays broader (`isOnBoxHostname`), where a false positive only declines to proxy. - A link to the origin root navigated nothing: the path was flattened to '', which openWebview reads as "no deep link", leaving an open frame where it was. - `this.webviews` being set does not mean it is loaded. initWebviews() assigns a truthy empty map and only then awaits the list, so a tap during page load found nothing to reuse and POSTed a duplicate record. Join the in-flight refresh instead. - One dashboard per dev server rather than per host spelling, which is what the method's own comment already promised. - Toast on the auto-create: it writes webviews.json, broadcasts over SSE and adds a Run-dropdown row on every signed-in device, with a new tab as its only previous signal. #362 (remote omp continuation, timkjr) - Accept the allowlisted `mode === 'omp'` arm as-is; a blanket registry render would hand deepseek a locally-resolved --profile and bypass claude's own overlay. A registry-declared switch is the follow-up if a third mode needs it. - Revert the whole-file Prettier reformat of docs/remote-sessions.md (docs/ is hand-formatted and outside `npm run format`), keeping only the two new sections. - Correct three stale passages: architecture-invariants' `exec claude --dangerously-skip-permissions`, the `exec <cli>` paragraph (claude and omp now have their own arms, and the claude pane's PID is the login shell), and omp-integration's `-c 'omp'`. RemoteCommandMode gains deepseek and omp. - Add the missing `_maybeCaptureOmpSessionId` remote-guard test; the sibling guard in `_pinOmpRespawnId` had one and this path runs earlier, on the first idle turn. #388 (keyCode 229 recovery, aakhter) - Gate notifyCanonicalData on shouldSuppressTerminalQueryResponse and isTerminalFocusOrMouseReport. onData also carries the DA/DSR/CPR/OSC replies xterm answers during Ink redraws and its SGR mouse and focus reports; any of those landing between the keydown and the candidate's resolution was read as "xterm spoke for this keystroke", standing the recovery down and leaving the character dropped, worst on a busy agent pane. Reached through window.CodemanTerminalInput: the predicates live in a module IIFE that closes long before this call site, so bare references would throw into the surrounding try/catch and stop the notify from ever running. Every fix has a test that fails without it (verified by reverting each). Full gate green on the combined tree: 358 files, 6849 tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a28b04c368 |
Merge pull request #362 from timkjr/feat/omp-remote-continuation
fix(omp,remote): thread remote-omp resume/continue through respawn and reattach |
||
|
|
e35b68e253 |
Merge pull request #401 from shenlvkang-collab/pr/loopback-links-webtab
feat(webview): open localhost links through a proxied web tab from another device |
||
|
|
349a89ec3b |
fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads
A dashboard served through a web tab saw `/webview/<cap>/` as its
`location.pathname`, and no app has a route for that: a React Router, Vue
Router or Vite dev-server page painted its HTML and CSS and then replaced
them with its own "page not found" the moment its script ran (reproduced
with a minimal history-routed page).
The proxy's runtime shim now rewrites the history entry to the path the
page would see on its own origin, before any page script runs. The base
element still resolves relative URLs inside the prefix and every root-
absolute sink is rewritten back into it, so only what the page READS
changes. With the document URL masked the Referer-keyed 404 rescue can no
longer help a request the shim misses, so the remaining URL-taking entry
points (`Worker`, `SharedWorker`, `navigator.sendBeacon`, `window.open`)
are covered by the shim as well.
A navigation the page starts itself afterwards — `location.reload()`
(a dev server's full-reload HMR), a root-absolute `location.href` — lands
on Codeman's root with no capability anywhere: no prefix in the path, no
cookie in an opaque-origin frame, a Referer naming the masked page. It is
recognised by shape (a top-level iframe navigation asking for HTML, for a
path Codeman does not serve) and answered with a static page whose only
script posts `{type:'codeman:webview-lost', path}` to the parent; the tab
that owns the frame (matched by `event.source`, never by the payload)
remounts it inside the prefix at that path, bounded per frame. The
unauthenticated form is answered in the auth middleware before the
credential checks, so a dev server that reloads on every save cannot
rate-limit its own user out of Codeman; the authenticated form (Basic
auth, trusted mode) is answered by the 404 handler.
Verified end to end against a history-routed page: boots on `/`, its
API call succeeds, a reload inside the frame comes back routed on the
path it had pushed, `location.href = '/about'` comes back on `/about`,
and a deep link opens on its path.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
d9eeb039db |
feat(webview): open localhost links through a proxied web tab from another device
An agent prints `http://localhost:5173/` (a dev server, a preview it just served) and the user taps it on a phone. That address only exists on the Codeman box, so the link was a guaranteed connection error from any other device — while the web-tab proxy fetches from the server, where it works. A loopback link (`localhost`, `*.localhost`, 127/8, 0.0.0.0, ::1) activated in the terminal or clicked in the Response Viewer now opens as a proxied web tab whenever the Codeman page itself is not on that box. A saved proxied dashboard on the same origin is reused, with the link's own path, query and fragment opened inside it (a mounted frame is navigated, not torn down, so its state survives); otherwise one is saved under its host:port, sandboxed like any other web tab, so it is in the Run dropdown next time. Only loopback is routed this way. A LAN or tailnet address may well be reachable from the device (a VPN, the same Wi-Fi) and a direct open is the cheaper, richer path, so those keep opening in a new browser tab; on the box itself every link opens directly. The terminal link provider and the viewer's click handler consult one hook and fall through to their existing behaviour when it declines. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou |
||
|
|
bd61735393 |
fix(web): show the whole last turn in the Claude response viewer's brief view
The eye button rendered `data.text`, which is one row: the last assistant
row of the transcript. A Claude answer is a median of 3 model messages
(p90 11) split around tool calls, so the brief view usually showed the tail
of an answer ("Done.", "Let me look.") and the substance appeared only after
More. The full view was fine, which is why the brief one read as broken by
comparison.
The brief view now asks `?context=turn`. The reader answers with the
assistant messages of the last ANSWERED turn (`selectLastAnsweredTurn`: the
highest `turn` that has an assistant row, so a prompt queued after the
answer does not blank the view) and the frontend renders them exactly as
the full view renders that turn: one badge, then continuation segments,
gated on the numeric `turn` as before.
`data.text` is unchanged in every context — still the last assistant row,
never `messages.at(-1)` — because agent pollers hash it. Readers that emit
no turns (Codex, the pane parser, DeepSeek, an older server) return `text`
only for `context=turn`, and the brief view keeps its single card for them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
|
||
|
|
58b4cb06d8 |
feat(files): let the path picker jump to a typed path and sort by name or date
The picker's current-folder line was a read-only breadcrumb, so reaching a deep folder meant tapping through every level, and the listing was fixed to name order, so the file an agent had just written was somewhere in a 500-entry list. The current folder is now an editable field: Enter or Go jumps there, a full file path lands in its folder with that file selected, and a path that does not resolve keeps the listing you had and says so, instead of the reset to the root that a stale initialPath gets. A Sort control orders the listing by name or modified time in either direction, folders always first, and the choice is remembered per device like the hidden toggle. Each entry shows a compact modified time (time of day today, month-day this year, else the date). GET /api/filesystem/browse stamps every entry with mtimeMs to make that possible; the stat that already fetched a file's size now serves both, so it is still one stat per entry. Entries without an mtime (an older server, the in-container listing) sort after dated ones and then by name. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou |
||
|
|
890a1b0902 |
Merge pull request #395 from irisitymichaelgrundberg/fix/full-history-replay-row-alignment
fix(terminal): keep row alignment in the full-history pane replay |
||
|
|
a360763890 |
Merge pull request #394 from irisitymichaelgrundberg/fix/ctrl-v-pastes-twice
fix(paste): handle only the first paste event the Ctrl+V trap receives |
||
|
|
d4fe3afc9d |
feat(files): raise the download cap to 2GB and stream /api/download
The 50MB cap on file-raw, the attachment /raw route and /api/download was memory protection for a `readFile()` that no longer exists: file-raw and /raw were rewritten to stream through `sendFileBody()` and answer Range requests, so size costs a read stream rather than RSS (measured: a 600MB download moved peak RSS by ~37MB). All the cap still did was refuse legitimate downloads of build artifacts, videos and archives. It is now MAX_FILE_DOWNLOAD_BYTES in config/buffer-limits.ts, default 2GB, env CODEMAN_MAX_DOWNLOAD_BYTES, 0 = unlimited. `parseByteLimitEnv()` is separate from the `parseInt(...) || default` idiom used elsewhere in that file precisely because that idiom reads 0 as falsy and would silently restore the default for the one value that means "no limit". /api/download was the last route that really did buffer the whole file. It now shares sendFileBody() with the other two, so it streams, advertises Accept-Ranges, and is resumable. Its Content-Disposition also goes through buildContentDisposition() rather than raw interpolation. Refusals move from 400 to 413 across all three, which is the correct status for the case; with the cap at 2GB it is a path almost nothing reaches now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2b57c595df |
fix(terminal): gate the row-preserving skips on a capture, not the query flag
Review of the previous commit found the guard inverted: the three skips keyed on `?full=1`, which is only what the client asked for. When the capture comes back null — ENOBUFS, a timeout, a vanished pane, or a session with no mux at all — the reply falls back to the byte history, which IS a stream of successive frames and still needs stripping. Gating on the request returned it whole: measured at 82KB against 4KB for the same buffer without `full=1`. A direct-PTY session takes that path on every first selection, not only during an outage. The skips now key on `isFullCapture`, meaning a capture arrived. Three further defects the same review surfaced, all on this path: Keeping the trailing rows is only sound when a cursor move follows to count back up from them. On the two branches where the cursor query fails there is no move, so the caret was left at the bottom of the pane — worse than before. The cursor is now read first and settles both decisions together. The move is relative rather than absolute. `CUP` numbers rows from the top of the browser's screen, so it is only right while the browser's row count equals the pane's, and `resizeWindow` does not wait for tmux, so a capture can be taken before a requested resize applies. Measured against real tmux with a browser four rows shorter than the pane: the absolute move lands on a blank row, the relative one lands on the caret's row. An all-blank pane no longer reads as content. Retaining trailing rows and appending a move made it non-empty, and the caller treats non-empty as "replay this", so a blank screen would have replaced real history — the downgrade `_replayWouldShrinkBuffer` refuses, arriving from the server side where that guard cannot see it. The documentation claimed one line per screen row. `-J` joins a hard-wrapped row into its logical line, so that is false whenever any row wrapped: measured at 10 lines for a 12-row pane. Both entries now say what actually holds, and the stale "NOT repositioned" contract in the mux interface is updated too. Tests: the byte-history fallback is stripped, an empty capture leaves history intact, and the extracted helpers are unit-tested directly rather than through source-text matching. The slice window in the capture test is bounded at the next method, having overrun into its neighbours. |
||
|
|
323730a29d |
fix(terminal): keep row alignment in the full-history pane replay
Switching to a session left the caret one row below the composer's input line, on the box border, and every cursor-relative update the CLI sent afterwards was measured from the wrong row. Any fresh output repaired it, because the CLI then repainted the whole frame. Two things were wrong with the full-history replay, and they compound. The capture never restored the cursor. The visible-frame path ends with an absolute cursor move back to the pane's position; the linear path returned its text and left the caret wherever the last character landed, which for an agent CLI is the bottom-most row carrying text — the status line. The rows it addressed did not line up with the pane's rows either. Four transforms ran over the capture and each can delete a line: the trailing blank rows were stripped, redraw-bloat stripping ran, the trim that cuts everything above the Claude banner ran, and leading whitespace was removed. All four are right for a byte stream of successive frames. A capture is the rendered pane, one line per screen row, so each deletion shifted the frame out from under the restored cursor. The full-history path now appends the pane's own cursor position and keeps every row, so row N of the reply is row N of the pane. The visible-frame and tail paths are untouched. Restoring the cursor is what makes row alignment load-bearing here, and neither CLAUDE.md nor the architecture invariants said so — which is how four line-deleting transforms accumulated on the path. Both now record it. Verified against a live 315x59 pane: the reply carries 59 rows, its row 55 is the composer's input line matching tmux, and it ends with the cursor move that lands there. |
||
|
|
5130ca6633 |
fix(pr-bot): fail fast when the review model's budget is spent
Claude Code answers an exhausted model budget INSIDE the turn ("You've
reached your Fable limit. Run /usage-credits to continue or switch models
with /model.") and then sits there with nothing to write. The reviewer never
produces a report, so `runTurn` waited out its full 40-minute deadline and
reported a bare "timed out after 40 min without a report", which reads as a
hung reviewer rather than an account that needs attention.
Measured on 2026-09-08: #388, #393, #394 and #377 each lost 40 minutes this
way, and because every attempt counted, all four reached MAX_AUTO_RETRIES and
would NOT have been picked up again once the budget returned. One spent
afternoon quietly took the whole queue out of service.
`findModelLimitNotice()` reads the notice off the pane and `runTurn` returns
a new `limit` outcome instead of waiting. It is consulted in exactly two
places, both of which mean "the turn produced nothing": on a stop where
`isDone()` is still false, and on each timed-out wait slice. A review that
merely discusses usage limits in its own findings therefore cannot be
mistaken for one that hit the wall, and the pattern matches neither the model
name nor a straight apostrophe, since the pane renders a typographic one and
every model prints the same sentence.
A spent budget is an account condition, not a bad PR, so it no longer spends
the per-head retry budget: the queue resumes by itself when the budget does.
Telegram now names the cause and the file to change.
Tests use the pane captured verbatim off the run that lost the 40 minutes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
b87bc6871b |
fix(paste): handle only the first paste event the Ctrl+V trap receives
Ctrl+V in the terminal inserted the clipboard text twice. Right-click →
Paste inserted it once.
`_handleImagePaste()` appends a hidden contenteditable div, focuses it, and
reads the clipboard out of the paste event that lands there. Two separate
routes deliver that event for a single keypress. The function issues
`document.execCommand('paste')` itself, which in Firefox dispatches a
trusted paste event and then returns false, because the trap cancels the
event and the command never completes; Chromium and WebKit refuse that
command and dispatch nothing. The keydown's own default action delivers the
other, because xterm calls the custom key handler before its own `cancel()`,
so returning false never calls preventDefault. Firefox therefore ran the
trap's listener twice and both runs reached `terminal.paste()`. The
context-menu paste involves no keydown at all, which is why that path stayed
correct.
The trap now accepts the first paste event and cancels every later one, so
how many paste events a browser delivers no longer changes what the PTY
sees. Measured on a live install, one Ctrl+V each: Firefox two events and
two writes before this change, Chromium and WebKit one and one, and every
engine one write after it.
The `execCommand('paste')` call stays. Stripping it out also ends the
doubling, and all three engines still deliver one event without it, since
`trap.focus()` has already run when the key's default action resolves. It is
kept because the trap technique arrived in #84 for plain HTTP and for
mobile, and a desktop measurement says nothing about real iOS Safari or
Android Chrome: where a browser aims the default action at the element
focused when the keydown began, the command is the only route into the trap,
and the trap is the only place clipboard image blobs are read.
test/image-paste-trap.test.ts loads image-input.js into a `node:vm` context
with a fake document and fires two paste events at the trap. It covers text
and images, and fails on the old code with the text pasted twice and the
image uploaded twice.
Docs: the invariant goes into docs/architecture-invariants.md as a Terminal
paste section and into CLAUDE.md as a Frontend entry, both recording the
measured event counts and why the redundant call is still there. README.md
and the Keyboard Shortcuts and Input and Voice wiki pages gain a Ctrl+V row,
which all three tables were missing while listing every other clipboard
binding.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
797f0d387c |
fix(remote): address review feedback on omp/claude respawn continuity
- Remote omp command now renders through buildSpawnCommandFromRegistry (the mode-agnostic engine local/docker spawns use) instead of the buildOmpCommand() the CLI-registry refactor deleted. - Session._pinOmpRespawnId()/_maybeCaptureOmpSessionId() now skip host-local ~/.omp resolution entirely for a remote session and fall back to --continue: that resolver only ever reads THIS host's filesystem, which is meaningless (and could wrongly alias an unrelated local conversation) for a conversation that lives on the remote host. - Remote-claude launch now honors an explicit resumeSessionId distinct from sessionId (mirrors claudeDockerPaneCommand's shape), and validates sessionId the same way that sibling does before interpolating it into the remote shell command. - Add the still-missing header-cwd half of the trailing-slash test, and document respawn/reattach continuation + auto-reconnect-vs- clean-exit in docs/remote-sessions.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d5b75af628 |
fix(statusline): sticky telemetry collection, footer print-through, EOF fix
Responds to Ark0N's review round on the ephemeral-CLI-flag statusline injection rework: - Rebase-detail fixes: registry-gated telemetry eligibility via getCli(mode)?.capabilities.statusLineTelemetry instead of a hardcoded mode === 'claude' check, using the capability flag master's CLI-registry refactor already declares for exactly this purpose. - Design question settled: sticky (a). Rather than persisting the toggle as a new field and threading it through every session-creation path (cron, Ralph Loop API, quick-start), eliminated the per-session field entirely. readPlanUsageTelemetryEnabled() (hooks-config.ts) reads the existing showPlanUsageLimits setting fresh from settings.json at every claude create/respawn (TmuxManager.createSession/respawnPane) - no per-session state to survive a restart, and it applies uniformly to every creation path for free, since they all flow through the same TmuxManager methods. This required fixing a real bug found along the way: showPlanUsageLimits was not actually round-tripping through settings.json on save - settings-ui.js explicitly excluded it from the PUT body as a pure per-device display key. It now flows through normally (both true and false); the load-side per-device merge behavior is unchanged. Removed entirely as a result: the statusLineTelemetry field from CreateSessionSchema/SettingsUpdateSchema, CreateSessionOptions/ RespawnPaneOptions, Session._statusLineTelemetry (this is what makes the restart-persistence bug moot rather than patched), and the frontend send sites. - Footer print-through restored: the no-user-statusline branch of the exporter script now runs the telemetry POST in the foreground so its own stdout becomes the in-terminal footer, falling back to a plain "codeman" marker only on curl failure. - Background-subshell EOF fix: the wrap-a-real-statusline branch closes stdin too, not just stdout/stderr (`>/dev/null 2>&1 </dev/null &`) - the un-redirected subshell process itself, not curl, was what held a reader-to-EOF's pipe open for however long curl took to finish. Added curl --max-time 5 so a hung (not just refused) Codeman cannot wedge the render. Tests: real-shell-execution tests for the footer/EOF fixes (fake curl stand-in on PATH, real sh subprocess spawns, real elapsed-time measurements - verified non-vacuous against a hand-reconstructed old-style script), unit tests for readPlanUsageTelemetryEnabled. Adapted two existing tests whose payloads referenced the removed field. Fixed during independent code review: a stray indentation break and a test exercising the wrong (legacy) exporter code path. Docs synced: CLAUDE.md, docs/usage-limits-display-plan.md (old disk-based section marked superseded, kept for history), docs/architecture-invariants.md. Full suite green: 352 files, 6780 passed, 12 skipped, 0 failed. tsc/lint/format:check/frontend-syntax all clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4f2dfb4e6d |
fix(mobile): carry resumeId through the phone overview's past rows
#386 made Codex conversations resumable from Past Sessions, and resumeMobileOverviewSession() correctly passes row.resumeId on to resumeHistorySession(). The phone's own row projection never copied the field off the unified-list item though, so row.resumeId was always undefined there and a tapped Codex row started a FRESH session on a thread that was already on disk. The desktop path worked; only the phone was blind. The test fails without the projection line, and pins the other half too: a claude row must not grow a resumeId, since the field is what distinguishes "resume this conversation" from "start a new one". Docs: CLAUDE.md and architecture-invariants both still described the unified list as merging Claude transcript files. It has been three stores since this PR (Claude's ~/.claude/projects, omp's ~/.omp/agent/sessions, codex's ~/.codex/sessions), the alias field keeps its Claude-era name without being Claude-only, and the scanner-only rule behind resumeId was written down nowhere. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f1b7283393 |
fix(cli-registry): guard workDetect.workingLine like every other config regex
#385 made the composer glyph and the working status line per-CLI registry data, which is right, but `workingLine` arrived as a config-supplied regex validated with a bare `new RegExp()`. That skips `compileVersionRegex()`, the helper the registry uses for exactly this: a `~/.codeman/clis.json` override can set the field, the compiled pattern is run against every accumulated PTY chunk and every pane capture, and a nested quantifier there backtracks on the event loop for the whole server rather than one session. Route it through the helper in both places, which are not redundant: the schema refine rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` compiles through the same helper so the runtime cannot hold a pattern the schema would have refused. The helper returns null instead of throwing, so the Claude-pattern fallback stops being a try/catch and becomes structural. Both shipped patterns compile unchanged, and Claude's is behaviourally identical to CLAUDE_WORKING_LINE_PATTERN. Also match the Codex footer case-insensitively on the E. It was characterised against codex-cli 0.152.1, which prints a lowercase `esc`; a version capitalising it would make the whole fix silently inert, since the pane would simply never look like it was working. Docs: CLAUDE.md, architecture-invariants and cli-registry.md all still stated the Claude-mode-only rule this PR retires, and none of them named the new capability or the regex guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cfc8fe7e41 |
Merge pull request #381 from mtiller/feat/reverse-proxy-base-url
feat(web): support a reverse-proxy base URL |
||
|
|
bca1b764cc |
Merge pull request #367 from shenlvkang-collab/pr/claude-conversation-first-hand
fix(session): learn the live Claude conversation from the CLI's own hook |
||
|
|
1c1773278f |
Merge pull request #369 from shenlvkang-collab/pr/claude-response-viewer-per-message
fix(web): render one Claude response-viewer message per model message |
||
|
|
a2aaea3c0e |
docs(file-picker): state the Home/cases nesting the right way round, and document the new fallback chain
The two merge-time edits the #383 review asked for. The comment above the picker's fallback chain said Home is nested under Codeman Cases; on the native default it is the other way round (~/codeman-cases sits inside ~). And the "Filesystem path picker" paragraph in architecture-invariants still said the picker falls back to /mnt/d, which #383 changed to: the session's Current Folder, then the Codeman Cases root, then /mnt/d, then the first root. No code behaviour changes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
f33b37c008 |
feat(pr-bot): review open PRs in Codeman sessions and report over Telegram
Maintainer tooling in scripts/pr-bot/: a daemon (systemd user unit
codeman-pr-bot) that lists open PRs with gh, reviews each head commit once in
a Codeman claude session (`prbot-<n>`) running in a private `git clone
--shared`, and sends the verdict, ranked findings, checks and a recommendation
to Telegram with action buttons. Merge, close, post-comment and approve-CI
happen only from a Telegram command or button plus a confirmation tap; the
bot never writes to GitHub on its own. The Telegram token and chat id come
from the existing notifier bot's env file.
Verified live: three PRs reviewed end to end (383, 363, 368), reports
delivered with buttons, reviewer sessions on the pinned model. Findings
along the way, each fixed and documented: a linked worktree inherits the
main checkout's model pin (hence the shared clone), undici's 5-minute header
timeout cut off the first review, gh was missing from the service PATH, and
the periodic scan orphaned an in-flight review's record.
typecheck/lint/format now cover scripts/pr-bot; tests in
test/pr-bot-{report,state,commands}.test.ts; guide in docs/pr-bot.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
8ad2215118 |
fix(docker): close the three adoption gaps the negative guarantee missed
Review follow-ups to #357. Each is a path that still touched, or still hid, a container Codeman does not own. **Export still mutated it.** The four fail-closed layers cover create/start/ stop/remove, but `POST /api/docker-cases/:name/export` reaches the container twice through neither: a full export `docker commit`s it, and even a workspace-only export `docker pause`s it first for snapshot consistency. Pause freezes the owner's processes for as long as the tar takes, on a container we promised not to touch. Full export is refused for an adopted case (it packages someone else's container, with their logins, into a bundle Codeman hands out); workspace-only keeps working and no longer pauses, accepting a live filesystem the way `tar` does on any running host directory. **A freshly linked OWNED case became unusable.** The run menu now probes the container for its CLIs, and a failed probe hides every agent mode behind the reason. For an adopted case that is right. For an owned one the container does not exist until the first session launches it, so every newly linked Docker case answered `container "codeman-case-x" not found (adoption never creates a container — start it yourself first)` and offered nothing but Shell, for a container the launch chain was about to create itself. A failed probe is recorded only when the case is adopted; `CaseInfo.docker.owned` is on the wire so the frontend can tell them apart. Verified in a browser: owned-with-no- container offers all ten modes and no notice, adopted-but-stopped offers Shell and says why. **Multi-user gating.** Adoption is admin-only, unlike `docker-link` beside it. Linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has already confined to the caller's space; an adopted container's mounts are whatever its owner gave it, so one mounting `/` hands the adopter a shell over the whole host — exactly the workspace scoping multi-user mode exists to enforce. Listing the engine's containers and browsing directories inside an arbitrary one are machine-level reads and follow the docker-HOST policy for the same reason. The preflight is deliberately not admin-only: the run menu fires it for every docker case, so it admits a non-admin for a container already linked to a case they can access, and nothing else. Verified end to end against a real pre-existing root container (alpine + tmux, no bind mounts): adopt, claude session inside it, workspace export, session close and case unlink all left `StartedAt`, `RestartCount`, `Pid` and `Paused` untouched; the pane ran the CONTAINER's claude, without `--dangerously-skip-permissions`; a stopped container was refused at both preflight and launch and was never started. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1 |
||
|
|
3d8ffcb9a2 |
Merge pull request #357 from dignfei/feat/docker-adopt-existing-container
feat(docker): attach a case to an already-running container Conflicts came from work that landed after the PR was opened, and each is resolved onto the newer abstraction rather than by keeping the older code: - `defaultDockerCommandForMode` is registry-driven since #347, so the PR's `runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the refusal is visible only inside the container, so an adopted root container otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and `test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch. - The probe's mode list and its mode -> binary table both duplicated the registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is also what fixes the merge's silent regression: the hand-written list predates `omp`, and the run menu gates every docker case on this probe, so owned containers would have lost that mode. `shell` needs no arm — it declares no binary, so it is dropped from the lookup and reported available regardless. - The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one `missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto it. Its test now pins the single gate instead of counting seven arms. - The create arm keeps #349's swap-limit warning filter, which the adopted arm never reaches; the run-mode list gains `omp` from #353. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1 |
||
|
|
7e4914d991 |
feat(web): support a reverse-proxy base URL (--base-url / CODEMAN_BASE_URL)
Codeman can now be mounted under a sub-path behind a reverse proxy that forwards the prefix unchanged (e.g. https://host/codeman/). Default is `/` (root), which is byte-identical to the historical behavior. Design — few choke points, mirrored ingress/egress: - src/config/base-path.ts: pure single-source normalize/validate/join/strip. - Server ingress: stripBasePath() inside Fastify rewriteUrl, so routes stay declared prefix-agnostic; un-prefixed requests (hooks, health, docker bridge hitting the raw port) pass through unchanged. - Server egress: one onSend hook prepends the base to root-absolute Location headers (covers all redirects). - HTML: renderIndexHtml points <base href> at the mount and injects window.__CODEMAN_BASE__ — ONLY when a base is set (inert at root). - Frontend runtime URLs: CodemanBase.url() route builder in constants.js, applied transparently by a fetch wrapper and explicitly at the EventSource/WebSocket/window.open/<img|iframe|a>-src sites. - sw.js derives its base from self.location; manifest uses relative start_url/scope. - Web-tab proxy: proxyPrefixFor(cap, basePath) is the single base-aware root that cascades to the injected <base>, HTML/attr rewrites, runtimeUrlShim, Set-Cookie Path and Location; capabilityFromReferer strips the base off the browser Referer, while the ingress parsers stay base-agnostic (rewriteUrl already stripped it). --base-url rides the daemon relaunch (buildWebArgs) and the service unit (resolveServicePlan). constants.js is guarded against a missing `window` for isolated unit-test contexts. Tests: test/base-path.test.ts (pure helpers), base-path coverage in webview-proxy/render-index-html/daemon-control; CodemanBase stubbed in the vm-isolated panels-ui test contexts. Docs: Remote-Access.md (sub-path section + nginx example), security-architecture.md env table, CLAUDE.md pattern. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XUkPBxbumnct6qSrx4JDju |
||
|
|
eeb5f9d0b2 |
docs: web-tab egress guard, capability revocation and referrer policy
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE |
||
|
|
99ad9cb236 |
fix(docker): never exit the server unless something is known to restart it
#373 restarts the Compose container by exiting the server, which is right for the shipped deployment: `restart: unless-stopped` relaunches it. The updater verified that policy through the Docker socket and, when it could not (no socket mounted), failed open and exited anyway. Failing open is the correct choice for the GATE, where refusing would block every install without a socket, but not for the kill: a container the daemon does not restart goes down for good, with no UI left to recover it from. That is exactly the case a plain `docker run` of this image without `--restart` produces, and the image sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path. The decision now happens server-side, where both the socket and the Compose env are reachable, and rides down to the script as `--restart-by-exit 0|1`. It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there and only there, since that file is what sets the restart policy; the image ENV deliberately does not) or when the daemon confirmed an auto-restart policy. Otherwise the build still lands, the status becomes `completed-needs-manual-restart` with the `docker restart` hint, and the server keeps running. The shipped deployment is unchanged in effect: with the socket it was already confirmed, and without it the declaration now covers it. Also: a root-run `Start-Codeman.sh` (common on Unraid) created the fingerprint baseline's `.codeman` directory before the container's first start and left it root-owned, which the unprivileged server could then never write its own state into. It is chowned to PUID:PGID when running as root. Verified with a real image build of the merged tree (classic builder; this box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain present, the four CLIs at their pins, docker/.env absent, and `docker inspect $HOSTNAME` returns the restart policy through the mounted socket as that user. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu |
||
|
|
823f56a243 |
Merge pull request #373 from opticon454/feature/docker-self-update
feat(docker): restore in-app self-update in the Compose dep |
||
|
|
a81e87f440 |
fix(cli-registry): log why clis.json was ignored, and say 0600 when that is the rule
The loader refuses a `clis.json` with any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored. That is a defensible posture for a file that chooses the binaries Codeman spawns, but two things around it made the override feature look dead: the warning said "group/world-writable", which a 0644 file is not, and `LoadResult.warnings` was returned to a caller nobody wired up, so nothing anywhere printed it. A user following the docs got silence. The message now names the rule and the command that satisfies it, the loader logs every warning once on first load (the result is memoized, so once per process), the module header stops claiming that nothing ever writes (the quarantine rename of a malformed file is a write, on first use) and the registry doc gains a short section on the override file with the 0600 requirement in it. Whether the check should relax to writable bits only is a separate decision; this keeps the shipped behaviour and makes it visible. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu |
||
|
|
268e4819ff | feat: auto-name sessions from first prompt | ||
|
|
66eb01ba8f |
feat(docker): restore in-app self-update in the Compose deployment
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
|
||
|
|
c5b84fb5f4 |
docs(cli-registry): annotate overlays.credStore as declared-for-later
Review item 4 named THREE live tables duplicating registry data. Two are now read from the entry (`defaultRemoteCommandForMode`, `defaultDockerCommandForMode`); the third, `resolveDockerCredentialArtifacts`, is not — and it was left neither wired nor annotated, which is the state that item explicitly rules out. It is not wired because the shape cannot express the live table: `credStore` is ONE store per CLI, and `CRED_STORES` needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek's entry declares none at all even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is at once the worst thing in that file to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest. It belongs in its own change, measured against a real container. So it is annotated instead, at the field, in the type's declared-for-later header, in docs/cli-registry.md, and in the pinned DECLARED_FOR_LATER list — the last of which means wiring it later makes a test fail rather than leaving a stale comment behind. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
4830e662f9 |
refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. Code that used to ask "which CLI is this?" reads the entry instead. Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every spawn command as a literal string, captured from the hand-written builders before they were deleted, and `test/location-overlay-commands.test.ts` does the same for all 20 remote and in-container pane commands. Config can never contain shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only in this release. OMP is included as a registry entry rather than a tenth hand-written builder, so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of `buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen and doctor ladders all drop out. Guard rails: - `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id branching reappears outside `stock.ts`, in any of its four shapes (`===`, `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the negated forms, which is how 36 of them survived an earlier pass. Every allowlisted branch carries its reason. - `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities; deriving one from another shipped the `until=stop`-hangs-on-shell bug. - `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config` wire field is separate, bridged only by `legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. - Registry data resolves AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks). A module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. - Six fields are annotated DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/ `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured. A test pins the list so it cannot quietly grow. Three user-visible changes, all deliberate and named: - `probeDockerCliVersion()` derives the in-container binary from the registry rather than assuming it equals the mode name (`antigravity` runs `agy`). - The remote CLI version probe now covers grok and deepseek, which the hardcoded map it replaces omitted while its own comment said the rule was "every mode except shell". - `codeman doctor`'s CLI rows are generated from the entries, so Claude's install hint is the install command rather than a docs URL, five CLIs gain hints they never had, and the row order follows the catalog. Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars (matching the `cliId` pattern) before its failure message quotes the value back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading the hand-editable `clis.json`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
2a32b5064a |
Merge pull request #349 from opticon454/feature/docker-compose
Docker Compose deployment: Codeman runs in a container and spawns Docker cases as SIBLING containers through the mounted host socket (Docker-outside-of-Docker). Resolved the README conflict (master had grown to eight CLIs since the branch was cut) and moved the Compose blurb out of the feature bullets into Quick Start, next to the other ways of starting Codeman. Three review findings from the PR discussion are fixed here rather than left for a follow-up, because two of them are shipped-image problems: - `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against the whole context-relative path, so `docker/.env` — which the deployment's own README tells the user to fill with CODEMAN_PASSWORD and provider API keys — was picked up by `COPY . .` and baked into the image at /opt/codeman/docker/.env. Verified in both directions against a real build context: with a canary secret in docker/.env, the unfixed ignore file lets /ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for the checked-in .env.example) leaves only the example behind. - `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which still hardcoded ~/codeman-cases, so `codeman skill install --case <name>` reported "Case not found" on exactly the deployment the override exists for. Both now resolve through config/cases-dir.ts. state-store.ts keeps its own literal on purpose: that one migrates the historical ~/claudeman-cases directory by name and is about the old default, not the active location. - CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore joins the documented list of files that genuinely belong in the repo root. The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/ deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's lastClaudeSessionId was handed to every non-claude CLI. Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax, public assets, lockfile. |
||
|
|
ccfda623fe |
fix(session): learn the live Claude conversation from the CLI's own hook
Which conversation a pane is on was re-derived by correlating ~/.claude/history.jsonl against Session.lastSubmitAt — and lastSubmitAt is bumped only by input that flows through Codeman's own write path (Session.write / writeViaMux). A user who attaches to the pane's tmux session directly never set it, so resolveActiveClaudeSessionIdFromHistory() returned at its first line for that pane's whole life and the response viewer stayed pinned to the launch conversation, showing a pre-/clear transcript indefinitely. A UserPromptSubmit hook reports the live conversation id from inside the CLI process, delivered under the pane's own $CODEMAN_SESSION_ID. That binding is a fact rather than a correlation: it never consults workingDir, so it cannot be claimed by a sibling pane on the same folder, a closed tab, or a bare `claude` in the user's terminal. A pane holding such an id skips the correlation entirely, so the number of prompts eligible for cwd-based guessing goes DOWN, never up — the naive alternative (relax the guard, or synthesize an anchor from PTY activity) is the reverted bug the resolver's own comment describes. The hook also stamps lastSubmitAt, so it finally means "a prompt was submitted" rather than "typed into Codeman's web terminal". Conversations vouched for first-hand — and only those — extend a persisted claudeSessionChain, whose tail re-pins the conversation when a surviving tmux session is re-attached after a restart. ⚠️ start() resets the id at THREE points and the last one runs unconditionally after the mux branch, so the tail is applied there too; patching only the mux branch looks right and silently does nothing. ⚠️ The hook's stdout is discarded with curl's own -o /dev/null. Claude Code injects a UserPromptSubmit hook's stdout into the model's context ("Exit code 0 - stdout shown to Claude"), and a trailing >/dev/null does NOT work: curlCmd already ends `... 2>/dev/null || true`, and in `pipeline || true >/dev/null` the shell binds the redirection to `true`, which never runs on the success path. The discard is opt-in so the five SSE-fed events keep byte-identical command text and no workspace's settings file is rewritten for them. The staleness marker is quote-free for the matching reason: hooksJson is JSON.stringify'd, so a quoted needle never matches and the gate would rewrite every workspace on every spawn. Existing workspaces heal on their next Claude spawn through the staleness sweep. |
||
|
|
3eff1feb5d |
fix(web): render one Claude response-viewer message per model message
The Claude reader concatenated every assistant row between two human prompts
into one card, fusing up to 74 distinct model messages into a single card, and
it never read the attachment rows that hold a prompt typed while the agent was
working. Measured over 57 real transcripts on 2026-09-01, the viewer shows
1,806 messages instead of 356 and 353 user cards instead of 178, with the
assistant text sequence unchanged row for row and the response without
?context=full byte-identical on all 57 files.
One assistant row IS one whole model message: in that corpus no assistant row
carries more than one content block and no message id carries more than one
text block, so there was nothing to reassemble. Each row becomes its own
message carrying an additive {kind, label, turn}, and the frontend renders a
same-role run inside one turn as badge-less continuation segments — which is
what keeps a p90 of 11 messages per turn from reading as card spam. A numeric
turn gates that rendering, so Codex, the external-CLI pane parser and an older
server keep one badge per card.
A prompt typed while Claude is working is recorded ONLY as an
attachment/queued_command row. Taking it when origin.kind is 'human' and
commandMode is 'prompt' recovers 162 user cards from 163 such rows — one is a
verbatim repeat inside an unanswered user run and is collapsed by the existing
dedup guard — and restores the turn boundary whose absence let the assistant
runs fuse. The CLI's own queue entries are cleanly separable: of 322
queued_command rows, 159 are commandMode 'task-notification' and not one of
them carries an origin key.
This narrows #169 rather than reverting it: sidechain exclusion, the
restored-<uuid8> rebind, replayed-snapshot dedup and synthetic-row filtering
are all unchanged and still asserted.
|