Two follow-ups to #361's sticky telemetry switch.
GET /api/settings reconciled an absent showPlanUsageLimits by persisting
true, but readJsonConfig() answers {} for ANY read failure (a parse
error, EACCES, EMFILE, a read landing inside PUT's non-atomic write), not
only ENOENT, and every page load calls this route, so one unlucky read
replaced the whole settings file with a one-key file. The route is a
plain read again and the default moved into the reader:
readPlanUsageTelemetryEnabled() treats an absent key as ON, the same way
readWorkspaceHooksEnabled() does, which is what the desktop chip already
shows for an install that never touched the setting.
saveAppSettings() sent showPlanUsageLimits on every save. The chip
defaults OFF on handhelds, so a phone saving its font size persisted
false and switched collection off for every desktop, whose chip then
went stale with no error anywhere. The key is now stripped like the
other per-device display keys and re-added only when the save FLIPS the
chip relative to what the device had (planUsageCollectionFlip), so an
explicit toggle on any device still writes it in either direction.
Tests pin both: the GET route with a mocked filesystem (absent, missing,
EACCES, garbage, explicit), the reader default, and the flip helper plus
its wiring in saveAppSettings.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
planUsageChipEnabled() (settings-ui.js) shows the header chip and the App
Settings checkbox as already ON whenever showPlanUsageLimits has never been
set — a discoverability default from 1.9.3. readPlanUsageTelemetryEnabled()
(hooks-config.ts) deliberately treats an absent key as "no telemetry" — a
privacy default, pinned by its own unit tests (never POST usage data
without an explicit persisted yes). Nothing reconciled those two
independent guesses, so a fresh install showed a checked box that silently
collected nothing until the user opened Settings and hit Save at least
once.
Verified live: an install that had never touched this setting had no
showPlanUsageLimits key in settings.json at all, and its running Claude
process's argv carried no --settings flag — zero telemetry ever collected
despite the chip rendering as enabled.
GET /api/settings now persists the resolved default (true) the first time
the key is truly absent — not explicit false — so "chip visible" and
"telemetry collected" become the same fact. readPlanUsageTelemetryEnabled's
own absent-means-false contract is untouched; after this runs once the key
is never absent again, so that branch stays correct in isolation while
being unreachable in practice for any install that has ever called this
route. An explicit false set afterward is respected forever.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Responds to Ark0N's review round on the ephemeral-CLI-flag statusline
injection rework:
- Rebase-detail fixes: registry-gated telemetry eligibility via
getCli(mode)?.capabilities.statusLineTelemetry instead of a hardcoded
mode === 'claude' check, using the capability flag master's CLI-registry
refactor already declares for exactly this purpose.
- Design question settled: sticky (a). Rather than persisting the toggle
as a new field and threading it through every session-creation path
(cron, Ralph Loop API, quick-start), eliminated the per-session field
entirely. readPlanUsageTelemetryEnabled() (hooks-config.ts) reads the
existing showPlanUsageLimits setting fresh from settings.json at every
claude create/respawn (TmuxManager.createSession/respawnPane) - no
per-session state to survive a restart, and it applies uniformly to
every creation path for free, since they all flow through the same
TmuxManager methods.
This required fixing a real bug found along the way: showPlanUsageLimits
was not actually round-tripping through settings.json on save -
settings-ui.js explicitly excluded it from the PUT body as a pure
per-device display key. It now flows through normally (both true and
false); the load-side per-device merge behavior is unchanged.
Removed entirely as a result: the statusLineTelemetry field from
CreateSessionSchema/SettingsUpdateSchema, CreateSessionOptions/
RespawnPaneOptions, Session._statusLineTelemetry (this is what makes
the restart-persistence bug moot rather than patched), and the
frontend send sites.
- Footer print-through restored: the no-user-statusline branch of the
exporter script now runs the telemetry POST in the foreground so its
own stdout becomes the in-terminal footer, falling back to a plain
"codeman" marker only on curl failure.
- Background-subshell EOF fix: the wrap-a-real-statusline branch closes
stdin too, not just stdout/stderr (`>/dev/null 2>&1 </dev/null &`) -
the un-redirected subshell process itself, not curl, was what held a
reader-to-EOF's pipe open for however long curl took to finish. Added
curl --max-time 5 so a hung (not just refused) Codeman cannot wedge
the render.
Tests: real-shell-execution tests for the footer/EOF fixes (fake curl
stand-in on PATH, real sh subprocess spawns, real elapsed-time
measurements - verified non-vacuous against a hand-reconstructed
old-style script), unit tests for readPlanUsageTelemetryEnabled.
Adapted two existing tests whose payloads referenced the removed field.
Fixed during independent code review: a stray indentation break and a
test exercising the wrong (legacy) exporter code path.
Docs synced: CLAUDE.md, docs/usage-limits-display-plan.md (old
disk-based section marked superseded, kept for history),
docs/architecture-invariants.md.
Full suite green: 352 files, 6780 passed, 12 skipped, 0 failed.
tsc/lint/format:check/frontend-syntax all clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Codeman's plan-usage chip wrote a statusLine.command into the case's
.claude/settings.local.json to receive Claude Code's rate_limits blob.
That file-based statusLine took precedence over the user's own
global/project statusline for ANY `claude` run in that directory,
including entirely outside Codeman, with no disclosure in the App
Settings UI (labeled only as a header-display toggle) and no way to
remove it once written (the removal code path was unreachable dead
code — nothing ever called it with false).
Replace the disk write with an EPHEMERAL `claude --settings
'{"statusLine":{...}}'` CLI flag, resolved fresh at spawn time
(resolveStatusLineCliCommand in hooks-config.ts) and merged with
effort/ultracode into one --settings object (buildClaudeSettingsFlag
in tmux-manager.ts, since Claude Code accepts only one --settings
flag). Never touches disk, so a plain `claude` run outside Codeman is
untouched. Self-healing: any legacy disk-written exporter from an
older build is stripped the first time a session starts in that
workspace again. Still respects a user's own hand-authored statusLine
(skips the flag entirely rather than overriding it).
Mid-fix bug found and fixed: the exporter's command legitimately
depends on $CODEMAN_SESSION_ID/$CODEMAN_API_URL/$CODEMAN_HOOK_SECRET_FILE
and an internal $INPUT, all meant to be expanded only when Claude Code
itself executes the statusline, using the pane's tmux-setenv'd
environment. Passing that text through --settings routed it through
execSync's own implicit /bin/sh -c first (tmux respawn-pane's
`bash -c "..."` wrapper) — POSIX double quotes don't suppress $
expansion, so those vars got expanded prematurely against the
server's own environment (unset there), producing malformed JSON that
printed as literal error text in the statusline. Fixed by writing the
exporter as a real, shared script file (ensureStatusLineExporterScript,
marker-versioned so stale copies self-heal) and passing only its bare
path via --settings — nothing for any intermediate shell to mangle.
Verified against a real Claude CLI on an isolated tmux socket, and via
direct execSync reproduction of the exact nested wrapping
createSession/respawnPane use.
A hard "never inject, even ephemerally" kill-switch was added and then
removed in the same pass: with the disk-leak fixed, disabling
injection only cost the plan-usage telemetry the feature exists to
provide, for no remaining benefit.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015GyMnFWnUzc41TDeHg9juW
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.
Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.
Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.
OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.
Guard rails:
- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
branching reappears outside `stock.ts`, in any of its four shapes (`===`,
`!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
negated forms, which is how 36 of them survived an earlier pass. Every
allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
`privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
wire field is separate, bridged only by `legacyConfigAliases`. Getting
`privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
clamp's only handle on a CLI's privilege switch, and a wrong name clamps
nothing with no error and no failing test — so `schema.ts` rejects an entry
naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
`allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
thunks). A module-level const freezes at first import, so a CLI enabled while
the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
(`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
rather than measured. A test pins the list so it cannot quietly grow.
Three user-visible changes, all deliberate and named:
- `probeDockerCliVersion()` derives the in-container binary from the registry
rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
hardcoded map it replaces omitted while its own comment said the rule was
"every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
install hint is the install command rather than a docs URL, five CLIs gain
hints they never had, and the row order follows the catalog.
Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
Follow-up to #350, which fixed the actual blocker (issue #352): `dsh plugin` is
a thin forwarder that `spawnSync`s a literal `pnpm` with no npm fallback, so an
image without pnpm dies at exit 127 and takes the whole build with it.
That PR also pinned an allowlist of the two packages whose lifecycle scripts
pnpm blocked at the time. Replace it with a policy that cannot go stale: pnpm,
unlike npm, refuses dependency build scripts by default and FAILS the install
over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on pnpm 11.24), and the
names to allow move between rebuilds because `@deepseek-harness-tui/dsh-tui` is
resolved by dist-tag, not pinned: 0.9.3 pulled `@google/genai` (whose script is
a literal `preinstall: no-op`), 0.10.0-beta.x does not. An allowlist of two
names would have let the next tree break the build the same way. Allowing them
wholesale is also the exposure this image already accepts three layers up,
where `npm install -g` runs the install scripts of every transitive dep of the
five CLIs above with no gate at all.
Also correct a comment in the `/api/deepseek/install-profile` route that
asserted the opposite of what #352 proved ("dsh bundles its own package
manager, so no system pnpm is required"). The route's behavior is already
right: dsh's own "pnpm not found on PATH" stderr reaches the caller as the
OPERATION_FAILED detail, so the UI's "add a terminal profile" button names the
fix. Documented the prerequisite in docs/deepseek-integration.md, and taught
the docker-cases image smoke test about `dsh`/`omp` plus the profile check that
`dsh --version` does NOT cover.
Fifteen review findings on the dsh mode, the serious ones first:
- Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys.
_configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into
every dsh pane and applyEnvOverrides() lands after it, so a non-granted
owner who could redirect the base URL would have the operator's key sent
as a bearer credential to a host of their choosing.
- Wait registry: until=stop/blocked is refused on docker and remote-SSH
dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions).
The HERDR triple is set via LOCAL tmux setenv, which crosses neither
docker exec nor ssh, so such a session can never post a hook event and
the wait burned its whole timeout on every turn.
- Approvals: a dsh item is an ALERT, not an answerable card. The answer
route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the
option parser cannot read a third-party TUI's frames, so an answer was a
blind keystroke into a foreign composer), and the push notification
carries no Approve/Deny actions for dsh sessions.
- Status shim (v3): --seq is forwarded and the server drops stale retried
reports inside a 60s window (the TUI retries with backoff, so a retried
'working' could land after 'blocked' and resolve an approval whose
dialog was still on screen); 4xx responses exit 0 instead of retrying,
so one misconfigured session cannot feed the auth rate-limit bucket
until the hook endpoint 429s for the whole instance.
- Web-UI server: concurrent starts are serialized through a lock (two
racing POSTs used to pick the same port and orphan the winner), and the
readiness poll / timeout paths only clear or stop the singleton while it
is still theirs. First click actually opens the tab now
(refreshWebviews, not the nonexistent loadWebviews). DELETE
/api/deepseek/web requires the privileged grant in multi-user mode.
- Cron: deepseek jobs run the same two-part launch gate as the HTTP
create paths (impl moved into the resolver so all three share it) and no
longer stamp a Claude default model on the session.
- Parity sweeps: quick-start's docker branch rejects deepSeekConfig like
the remote branch; the Ralph auto-enable list gained deepseek;
HookEventType gained agent_working; the phone overview run menu filters
managed webview records like the desktop menu.
- install.sh: the dsh identity probe closes stdin (under curl|bash a
child that reads stdin eats the rest of the script), bounds the exec
with timeout where available, and is memoized to one scan per install.
- Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand
identity (it rendered as an unstyled UA-grey button); stale markup
comment about the web shortcut rewritten; clamp docs updated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a
shell tab running the server next to it. The shell was deliberate - the server
lived in an ordinary session so it was visible, scrollable, killable and died
with its tab, and nothing new had to supervise a long-lived HTTP server. That
reasoning was sound and the result was still wrong in use: opening a dashboard
should open one tab, and after the first launch the terminal is pure noise.
The server moves to a background child process owned by a new
`src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session
gave away for free is now explicit, which is most of the module:
- Exactly one server. A second click reuses the running one instead of racing
it for a port; the session flow could not do this at all, because two clicks
were simply two sessions.
- Restarted when the requested authority changes. `--trusted-host` fences dsh's
own /api against the browser authority, and a Codeman reachable at both
loopback and a tailnet name has two. Reusing a server fenced for the other
origin renders a page whose every call 403s, which reads as a broken
dashboard rather than a misconfigured one, so a mismatch restarts instead.
- Killed on shutdown. The child is detached so its whole plugin tree can be
signalled at once, which also means it would outlive Codeman and hold its
port against the next start - the exact EADDRINUSE this feature already got
wrong once.
- Boot output captured and returned. With no shell tab there is nowhere else
for a stack trace to land, so a failed spawn reports its own tail.
The endpoint is fenced at the same bar as the profile installer and for the
same reason: booting a dsh profile executes the plugin code in it, so this is a
privileged action even though it reads as "open a page". `authority` comes from
the client (`location.host`) because only the browser knows which origin is in
play, and it is regex-confined at the schema boundary - defence in depth behind
the argv-array spawn, admitting host:port in the shapes a browser authority can
take and nothing readable as a second argument.
`GET /api/deepseek/web-port` is gone; port selection moved into the supervisor,
which is the thing that knows whether a server is already running. The two
client-side probe helpers went with it, since the server now owns the wait.
Verified over the tailnet authority end to end: no session is created (session
count unchanged, one tab), the server runs on 3081 beside the user's own dsh
web on 3080, status reports the tailnet authority, and the proxied dashboard
renders with zero 4xx. Full gate green (6148 passed, +6).
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.
1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
precisely the port a DeepSeek user is most likely to be serving on already,
so the launch died with EADDRINUSE against the user's own server. The port
now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
free loopback port by BINDING it (a connect probe cannot tell "free" from
"listening but not answering yet").
2. It opened the tab unconditionally. The crashed server left a saved dashboard
pointing at nothing, with the failure only visible in a shell tab nobody had
a reason to look at. The launch now polls the existing webview probe until
the URL answers, and on timeout reports the error naming the shell tab
instead of persisting a dead dashboard.
3. The saved tab was untrusted, so the frame was sandboxed without
`allow-same-origin` and the dashboard was broken twice over: the dsh
client-runtime reads `localStorage` while loading its plugins and died there
("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
every `/api` call no matter which authority `--trusted-host` named. Passing
`location.host` only means anything once the frame actually carries that
origin, so `--trusted-host` had never once done its job. The managed tab is
now created `trusted: true`.
That trade is real and deliberate: a trusted proxied frame is same-origin
with Codeman and can reach Codeman's API. It is defensible only because this
dashboard is an agent harness Codeman just started itself, on loopback, which
can already run code as the user. It is not a precedent for trusting
third-party dashboards, which is why it is set at this one call site rather
than defaulted.
Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.
`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.
Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
Three review findings on the DeepSeek Harness mode, plus one the third exposed.
1. The multi-user clamp was bypassable by a sibling field on the same request.
clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but
DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER
_configureDeepSeek(), so a non-granted owner sending
envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated
instance: a session created with permissionMode "read-only" and that override
ran with DSH_PERMISSION_MODE=danger-full-access in its pane.
Every other CLI's bypass is a command-line flag reachable only through the
per-CLI config, which is why the config clamp alone is the whole gate for
them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted
owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to
what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on
that list because it aims the launcher at a profile tree whose plugin code
runs at boot, before any approval row can apply. Verified end to end in real
multi-user mode: a non-granted user sending both now gets workspace-write and
no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched.
2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout`
signals only the direct child, and a plugin install fans out into
package-manager children that keep the inherited stdio pipes open, so `close`
never fires and the held-open request leaks with no route-level deadline.
Reproduced: with a 1.5s built-in timeout the promise was still unsettled after
6s and both fan-out children were alive. Now detached: true plus negative-pid
SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a
last-resort reap for a grandchild that escaped the group. Same probe after the
change: close fires, direct child and both grandchildren dead.
3. hooksAvailableForMode() promised more than a dsh session can deliver.
deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that
triple is the only reason a dsh session posts hook events, so `until=stop` was
accepted and then blocked for the caller's whole timeout: the exact
infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now
takes HookCapabilityOptions and every call site passes sessionHookOptions(),
with the deepseek arm reading `!== false` so a forgotten one degrades to the
old behaviour. The refusal names the setting rather than saying "no Claude
Code hooks", which would send the caller hunting a bug that is really a
setting they chose. Profile conformance stays unknowable at request time and
is documented as such. The stale "True for `claude` and nothing else" docblock
is corrected.
4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a
claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent
capture read Claude's own transcript, and adding deepseek silently widened
both to a mode that has none. They compare mode === 'claude' directly now, and
a static check pins them there.
Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the
wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 --
bridge off plus explicit until=stop is a 400 naming the setting, bridge off with
no `until` still 200s on idle/exit, bridge on accepts stop.
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.
DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:
1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
$DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
`base` -- the interactive terminal front door is always a third-party
plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
button gates on the latter, because reporting only the binary would spawn a
pane that dies on arrival. When the binary is present but no profile is,
the run menu offers to install one (POST /api/deepseek/install-profile).
2. The permission switch is an env var, not a flag. The harness has no
command-line permission option; its sandbox/approval rows read
DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
own workspace-write, which still asks, so the multi-user clamp is the
only-if-sent branch and clamps to workspace-write, never read-only.
3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
earned that. The terminal front door reports idle/working/blocked to a
supervising process over a generic env-gated contract; a generated shim
(deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
report to /api/hook-event as stop / agent_working / permission_prompt. So a
dsh session gets definitive respawn triggers, real wait-endpoint signals and
real Approvals Inbox items instead of output-stabilization guesswork.
`agent_working` is new (157th SSE constant) and joins
APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
alert at once.
The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.
Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.
Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.
Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`GET /api/system/update/check` answers "is there a newer published release
tag?", which is the right question for an npm install but not for a git clone
that tracks a branch. Such an install can be many commits behind its own remote
while the latest tag says it is current, and nothing surfaces that.
Adds `GET /api/system/repo-status`: an informational companion that reports what
this CHECKOUT looks like against its own remotes — current branch and commit,
ahead/behind counts per remote, the remote's role (tracking / upstream / other),
and a bounded list of incoming commits.
Read-only and defensive: every git invocation is `execFileSync` with an argv
array and a timeout, a non-git or remote-less install reports a structured
`error` rather than throwing, and nothing here mutates the working tree or
touches the updater's own state.
Tests: 24 cases in test/repo-status.test.ts.
SessionMode gains 'pi', a first-class backend alongside Claude Code,
OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose
tab identity, welcome button, run-mode entry, cron agentType, Docker and
remote-SSH command defaults, and clone-repo Brain option.
Pi is a different shape of CLI from the other four, and three decisions
follow from that:
- It has NO permission prompts and no sandbox, so there is no
--dangerously-skip-permissions analog and none was invented. The
privilege-shaped knob is the tri-state approveProjectTrust, which makes
pi load and EXECUTE repo-local .pi/extensions TypeScript and install
missing project packages. clampExternalCliBypassForOwner() therefore
puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets
--no-approve even when no config was sent, because pi's own default is
a prompt the session user could answer themselves. That helper had zero
test coverage; it now has coverage for all four CLIs.
- Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars
share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode
context, so admitting them would widen the allowlist for every mode at
once. Auth goes through pi's /login or the server's own environment.
--api-key is deliberately never wired: it would put a provider secret on
the spawn command line.
- pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the
main screen with terminal-owned scrollback, and its 0.84.0 fullscreen
mode is runtime-switchable via /settings; that flip was measured to put
the pane into the alt screen, which the strip would have corrupted.
pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires
semver-shaped output, because `pi` is a short generic name a stray binary
can shadow; GET /api/pi/status surfaces path and version so a
misresolution is diagnosable rather than presenting as a broken mode.
Docker installs pi in its own --ignore-scripts step so that flag cannot
affect the other four CLIs, and seeds its credentials per-file rather than
whole-dir (~/.pi/agent also holds sessions, extensions and package trees).
Verified end to end against pi 0.84.1 on an isolated instance: resolver
search-dir fallback, flag construction, piConfig persistence across a full
server restart, the trust prompt and its --no-approve suppression, the
rose Run button on the default daylight-blue skin (the nested skin block
eats per-mode gradients unless the rule lives inside it), and the buffer
local-echo policy, which pi tolerates where codex did not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Extends the Gemini gating from bb7fb9e to the other welcome-screen
buttons that had the same problem: shown unconditionally even when the
underlying CLI isn't installed.
- Add isClaudeAvailable() (claude-cli-resolver.ts) and GET
/api/claude/status, mirroring the existing opencode/codex/gemini
resolvers and status endpoints.
- Opencode already had a working /api/opencode/status the welcome
screen just wasn't checking; wire it up the same way.
- Refactor loadGeminiAvailability() into a shared
_loadCliAvailability(buttonId, statusUrl) helper instead of
duplicating the fetch/try-catch three times.
Run-mode dropdown entries (Opencode/Codex) are intentionally left
unconditional here — follow-up PR.
Adds Antigravity as a sixth CLI backend alongside Claude Code, shell, OpenCode,
Codex and Gemini, following the existing pluggable-resolver pattern.
- `utils/antigravity-cli-resolver.ts` resolves the CLI, mirroring the other
resolvers; `GET /api/antigravity/status` reports availability and path.
- `ANTIGRAVITY_*` joins the `ALLOWED_ENV_PREFIXES` allowlist in schemas.ts, so
env overrides stay CLI-scoped rather than blanket-forwarded.
- Session, tmux-manager, mux-interface and types carry the new mode; secrets are
injected via socket-scoped `tmux setenv`, never on the spawn command line, so
the mode requires tmux with no direct PTY fallback like the other external CLIs.
- Frontend: Run-dropdown entry, agent-type option, `ag` tab badge and toolbar
colours. `runAntigravity()` routes remote/docker cases through
`POST /api/quick-start` and skips the local status probe for them.
Tests: test/antigravity-mode.test.ts, plus run-mode-ui and system-routes coverage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PUT /api/settings service toggles now resolve from `merged` (persisted +
incoming) instead of the raw request body, so a partial PUT no longer
starts the subagent watcher and stops the workflow + image watchers by
treating every omitted key as "apply the default". Pinned by a 4-case
regression test verified to fail against the old handler.
Also trims the links line from the Codeman callout in the
xterm-zerolag-input README.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The opt-in multi-user feature's only enforcement is web-layer scoping
(all sessions share one OS account). An adversarial review found 8 critical
+ 7 high cross-user holes that defeated it, plus mediums; all fixed here.
Single-user (flag-off) behavior stays byte-identical apart from documented
consistency deltas.
Ownership / confinement:
- DELETE /api/sessions (bulk) + /:id now owner-scope / findSessionOrFail
- quick-start, cron (create+fire), scheduled runs confine workingDir to the
owner's space; case link/docker-link/docker-import confine the host path
- resolveCasePath no longer resolves linked cases for non-admins; foreign
remote/docker cases are skipped (fall through to the caller's own local case)
- history, subagents/workflows, mux-sessions, orchestrator, cron run-history,
away-digest, and remote/docker host reads are owner- or admin-scoped
Permission policy (section 6.3):
- non-granted users are downgraded at every spawn site incl. legacy
/api/scheduled, PlanOrchestrator one-shots, remote launch, and the cron-fire
gemini/codex bypass switches; resolveClaudeModeForUsername now fails closed
Auth / store:
- verify-first login throttle (a correct password is never locked out),
/ws terminal subject to the change-password lockbox, cookie fast-path
re-validates identity live, role/grant changes revoke sessions, admin delete
runs the last-admin guard before any teardown
- users.json: distinguish missing (ENOENT) from corrupt/unreadable so a bad
read can't overwrite all accounts; unique per-process temp write path
Event streams:
- debounced session:updated + batched task:updated, clipboard, and push
notifications route by owner (fail closed); getLightState hides machine-wide
globalStats from non-admins
Tests: two suites updated to assert the fixed (secure) behavior. tsc, eslint,
and test:ci all green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Scopes real-time streams and the init snapshot so a multi-user client only
receives what it owns. No-op in single-user mode (identity-less clients).
- WS terminal (ws-routes): owner gate after the session lookup. A non-admin may
only attach to their own session (close 4003); the global auth hook already
ran on the upgrade and decorated req.authUser, so an unauthenticated upgrade
never reaches the handler.
- SSE (sse-stream-manager): per-client identity stored at addClient; broadcast()
and the terminal-batch flush both enforce a routing hint via canDeliver().
WebServer.broadcast auto-derives the hint (deriveSseHint): session-scoped event
families resolve the owner from the payload's session id (fail closed when the
owner can't be resolved), machine-level families (docker/tunnel/update/system/
cron) + host-plan telemetry are admin-only, everything else stays global. Raw
terminal bytes resolve the owner once and are withheld from non-owners.
- getLightState is filtered per connection AFTER the shared cache (sessions,
respawnStatus, subagents, workflowRuns by owner; scheduledRuns + planUsage
admin-only); applied to both the SSE init snapshot and GET /api/status.
- file-routes: getKnownSessionWorkingDir + getSessionAttachmentHistory (the
preview/thumbnail/history helpers that bypass findSessionOrFail) now owner-check
the session, closing a cross-user file-read path.
- GET /api/search: harvestSources is owner-scoped.
Deferred to a follow-up (documented in docs/multi-user-plan.md): away-digest +
subagent/workflow REST list scoping, push-subscription identity + routing,
per-user screenshot subdirs. The live-event versions of these are already routed
by the SSE hint; only the on-demand REST aggregates remain global for admins-only
follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Threads per-user ownership through sessions, cases, cron, and the permission
policy. All scoping is a no-op in single-user mode (isMultiUserMode() guards).
Sessions
- Session.owner stamped at every create path from req.authUser / job.owner:
POST /api/sessions, /api/run, /api/quick-start, ralph start, cron launch,
plan generation. Round-trips through recovery (MuxSession.owner mirror, read
muxSession.owner ?? savedState?.owner) and the mux layer.
- findSessionOrFail(ctx, id, req) now does a NOT_FOUND owner check (never 403, so
other users' session existence is not leaked); wired at ~50 call sites.
- List endpoints filtered by owner: GET /api/sessions, /api/sessions/unified
(live+persisted+lifecycle scoped, host-wide transcripts admin-only), cron jobs.
Permission policy (section 6.3)
- resolveClaudeModeForUsername wraps getClaudeModeConfig at every spawn site so a
non-granted user is forced to --permission-mode auto (bypass -> auto), including
recovery (or a reboot would un-downgrade). buildPromptArgs now respects the
session's claudeMode, closing the one-shot (runPrompt) bypass hole.
- Shell mode and cron launchCommand require canBypassPermissions: 403 at
POST /api/sessions, /api/quick-start create, cron job create, AND cron fire time
(re-checked against the owner's current grant).
Cases
- resolveCasesDir(user): per-user ~/codeman-users/<name>/cases in multi-user, the
shared ~/codeman-cases otherwise. All case CRUD + ralph + plan + quick-start
resolve through it. resolveCasePath is owner-aware.
- GET /api/cases scoped per user (own folders; legacy linked cases admin-only;
remote/docker cases owner-filtered). RemoteCase/DockerCase gain owner, stamped
at link/quickcreate/import.
- Remote + Docker host CRUD is admin-only.
- Non-admin workingDir confinement (the linchpin): realpath must resolve inside the
user's space, enforced at POST /api/sessions and /api/run BEFORE any disk write.
Limits
- sessionCapacityState / sessionCapacityMessage centralize the global + per-user
cap (CODEMAN_MAX_SESSIONS_PER_USER, default global/2), replacing the 6 copy-pasted
MAX_CONCURRENT_SESSIONS checks.
Tests: test/ownership-scoping.test.ts (case isolation, host-CRUD gate, workingDir +
shell gates, and the scoping helpers). Deferred to phase 4: WS owner gate, SSE
fan-out filtering, file-route preview/thumbnail helper scoping, push routing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a parallel multi-user auth branch (the single-user Basic-auth path is
left byte-identical). Off unless CODEMAN_MULTIUSER/--multiuser.
- middleware/auth.ts: mode-selecting registerAuthMiddleware. New async
multi-user hook verifies username:password against the user store (scrypt),
mints identity-carrying cookies, decorates req.authUser, enforces a per-IP
AND per-username failure bucket, and the mustChangePassword lockbox. The
hook-secret loopback bypass is now a single shared helper used by both
branches. FastifyRequest.authUser module augmentation.
- ports/auth-port.ts: AuthSessionRecord gains username/role/mustChangePassword.
- user-store.ts: verifyPassword (timing-equalized against user enumeration).
- route-helpers.ts: getAuthUser (synthetic admin fallback), canAccessOwned,
requireAdmin, revokeUserSessions; findSessionOrFail gains an optional req for
a NOT_FOUND owner check (dormant until phase 3 wires callers).
- routes/me-routes.ts: GET /api/me (synthetic admin in single-user) and
POST /api/me/password (verify current, min 8, clear mustChangePassword,
revoke other sessions).
- QR: QrTokenRecord + AuthSessionRecord carry a username; tunnel-manager
mintUserToken / consumeTokenWithIdentity / getQrSvgForCode; /q/:code binds
the cookie to the token's user (rejects identity-less tokens in multi-user);
GET /api/tunnel/qr mints a per-user token. Single-user keeps the rotating token.
- server.ts: bootstrap the initial admin from CODEMAN_USERNAME/PASSWORD on first
boot (refuse to start with no users); multi-user with >= 1 user satisfies the
non-loopback auth requirement and the tunnel-enable guard; userFailures bucket
disposal.
- types/api.ts: FORBIDDEN, PASSWORD_CHANGE_REQUIRED, USER_EXISTS, USER_NOT_FOUND,
LAST_ADMIN error codes (message + status wired).
- Session.owner field + getter/setter, SessionState.owner, MuxSession.owner,
CreateSessionOptions.owner (foundation for phase 3 ownership threading).
Tests: test/multiuser-auth.test.ts (10, live server on 3170/3171). Existing auth
suite (auth-security, qr-auth, cod54-hook-event, network-auth-policy) unchanged
and green; full test:ci sweep passes (3519 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Introduce src/config/terminal-history.ts: one place for terminal scrollback,
tmux history-limit, and PTY buffer byte caps, each overridable via env var or
the settings object and bounds-clamped via resolveTerminalHistoryConfig().
Defaults match the prior hardcoded values, so this is behavior-neutral. Wires
the resolver through buffer-limits, tmux-manager (incl. a setHistoryLimit so a
settings change applies live), session, server, system-routes, session-routes,
schemas, and the config port. Adds 4 optional settings keys (terminalScrollback
Lines, tmuxHistoryLimit, terminalBufferMaxBytes, terminalBufferTrimBytes) with
bounds + a trim<=max cross-check.
- Daylight Blue: Cloudflare Tunnel welcome button is now purple (was orange),
keeping Claude blue / Tunnel purple / OpenCode green distinct.
- Allow enabling the Cloudflare tunnel with no CODEMAN_PASSWORD via the UI: the
toggle now pops a security confirm dialog and, on confirm, sends an explicit
per-request acknowledgeUnauthTunnel:true (new action field, never persisted).
Server logs a loud warning whenever a passwordless public tunnel starts.
curl/API/CLI stay refused unless password/env/flag — no accidental exposure.
Tests: extend test/routes/system-routes-tunnel-guard.test.ts (ack allows + not
persisted; ack:false still refuses). Verified e2e on an isolated instance
(purple button, confirm dialog, retry carries the flag, no real tunnel opened).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Auto-popping draggable window per active ultracode/Workflow run, connected by a
glowing line to its originating session tab (resolved via claudeSessionId ===
sessionUuid). Mirrors the live agent grid; auto-closes after a run finishes;
dismissals are remembered. Additional to the existing docked panel.
New "Ultracode Floating Windows" setting (default OFF), independent of the
"Ultracode Agents" panel toggle; either toggle starts the workflow-run watcher.
Also bumps version to 1.1.3 and brings CLAUDE.md up to date for the ultracode
subsystem.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Opt-in (showUltracodeAgents, default OFF) panel that visualizes ultracode /
Workflow-tool runs like Claude Code's "working agents" TUI: LEFT = runs + phases
(selectable tasks), RIGHT = each run's agents with model, live state, tokens
burned, and tool calls.
Standalone — ZERO edits to subagent-watcher.ts. A new workflow-run-watcher.ts
singleton globs the run-state tree (~/.claude/projects/*/*/workflows/wf_*.json,
disjoint from the transcript tree), strips the heavy script/scriptPath/result/logs
fields (174KB -> ~25KB/run), and emits workflow:run_* SSE events. The LEFT list
ships lightweight summaries (getLightState replay + SSE); the RIGHT pane fetches
the full run (with agents[]) via GET /api/workflows/:runId on selection.
Backend: workflow-run-watcher.ts, types/workflow-run.ts, config/workflow-config.ts,
3 SSE events, getLightState workflowRuns replay, GET /api/workflows[/:runId],
showUltracodeAgents schema key + boot-gate (default OFF) + live toggleService.
Frontend: ultracode-panel.js (debounced master-detail render, run/phase select),
header launcher (btn-ultracode-agents--hidden marker -> mobile-guard-exempt),
App Settings toggle (SYNCED, deliberately not in displayKeys).
Agent states on disk are start|progress|done (start=queued; done has
durationMs/resultPreview). Tests: workflow-run-watcher (9), workflow-routes (3).
Verified: tsc/lint/prettier/frontend-syntax/public-assets/mobile-header-guard
clean; full test:ci green (2986 passed); live server + Playwright e2e against 25
real runs (28-agent grid, phase filter, OFF hides launcher).
Design: docs/ultracode-agent-viz-plan.md (rev. 3).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The plan-usage header chip (5h/7d %) was a SYNCED setting, so enabling
it on desktop turned it on for mobile too — even though the user never
enabled it there. Make the chip's DISPLAY purely per-device (default
OFF) like the response viewer / skin, while keeping telemetry COLLECTION
server-side.
Three leak sources fixed:
- server.ts renderIndexHtml force-revealed the chip from the synced
value (pre-paint), pushing the desktop choice onto every device.
Removed — the chip now ships hidden and the client reveals it
per-device via applyHeaderVisibilitySettings.
- settings-ui.js load-merge let the server value win, writing desktop's
`true` into the (separate) mobile settings blob. showPlanUsageLimits
is now a displayKey AND is dropped from the server payload on load, so
a stale server value is never seeded into a device that didn't enable
it. It's also stripped from the save payload so a mobile "off" can't
clobber the server.
- Collection was gated on the same synced flag. Decoupled via a new
`statusLineTelemetry` ACTION field (schema + system-routes): sent on
ENABLE only and never persisted, so the exporter is injected when a
device turns the chip on but is never yanked when another device has
it off (it's shared across sibling sessions). Session-create already
reads the per-device blob, so that path was already correct.
One-time migration clears a stale synced `true` from the mobile blob so
existing mobile installs default to OFF without a manual toggle.
Verified end-to-end on an isolated server: with showPlanUsageLimits=true
persisted, the rendered HTML ships the chip hidden; a fresh browser
context (mobile case) keeps it hidden while a context that explicitly
enabled it shows it; the PUT accepts statusLineTelemetry and does not
persist it. tsc + frontend-syntax + system-routes/index tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two changes so the feature works for any user the moment they enable it,
without manual steps or per-client state:
- Reconcile on settings change: PUT /api/settings now applies the statusLine
exporter across all ACTIVE Claude sessions' working dirs when
showPlanUsageLimits is toggled (inject on enable, remove on disable). This is
server-side and authoritative, so existing sessions get the footer + feed the
chip immediately — no need to create a new session, no dependency on a
browser's synced localStorage.
- Create is now ADD-ONLY: never remove the statusLine on session create.
Sessions in a repo share one settings.local.json, so a single create-with-false
(e.g. a client whose synced setting hadn't loaded) was yanking the statusLine
out from under all other live sessions in that repo, killing their footer and
the chip's data feed. Removal now happens only via the explicit settings toggle.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two hardening fixes for the public-tunnel exposure path (COD-54 / COD-55).
COD-54 — gate the /api/hook-event localhost bypass when a tunnel is up:
`cloudflared --url http://127.0.0.1:port` proxies internet traffic INTO the
loopback origin, so a tunneled hook request arrives with req.ip === 127.0.0.1
and the old bare-localhost bypass would pass it unauthenticated. Now:
- tunnel running → bypass requires a shared per-instance hook secret
(X-Codeman-Hook-Secret header; constant-time compare) + per-IP rate limiting
- tunnel not running (loopback-only, the normal case) → unchanged, so
already-deployed credential-less hooks keep working.
New src/config/hook-secret.ts; auth middleware takes a getTunnelRunning probe
(wired from server.ts via tunnelManager.isRunning()).
COD-55 — refuse starting the Cloudflare tunnel without auth:
enabling the tunnel publishes full terminal control to a public URL; with no
CODEMAN_PASSWORD the auth middleware is inactive and the bind guard never trips
(tunnel binds loopback). PUT /api/settings now refuses tunnelEnabled:true with a
403 (before persisting) unless CODEMAN_PASSWORD is set or
CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1 is acknowledged. New
isUnauthenticatedNetworkAcknowledged() in network-auth-policy; settings-ui
surfaces the refusal as an error toast and reverts the toggle.
Scope: the always-on CSRF/Origin guard, Host-header allowlist, and
network-auth-policy itself are already upstream (#113) and not re-proposed here.
Verification: tsc, eslint, prettier, check:frontend-syntax clean; full test:ci
green (2723 passed), incl. test/cod54-hook-event-auth and
test/routes/system-routes-tunnel-guard.
Point 1 of the v1.0 lock-in: commit to a stable HTTP API (the cleanest, fullest form).
Core (centralized):
- Every JSON /api response now uses ONE envelope via a Fastify preSerialization hook (src/web/server.ts): success -> { success:true, data:<payload> }; error -> { success:false, error, errorCode } with a conventional HTTP status. Non-JSON routes (file-raw, tail-file SSE, download, screenshots, /q redirect, WS) are skipped.
- Error-code -> HTTP status is a single source of truth (httpStatusForErrorCode in src/types/api.ts): 400/401/404/409/422/429/500. Expanded ApiErrorCode (added UNAUTHORIZED, CONFLICT, RATE_LIMITED). Errors are no longer HTTP 200.
- Versioned alias: /api/v1/* rewrites to /api/* (rewriteApiV1Url), so external clients pin to a stable surface while the bundled UI keeps using /api/*.
- Handlers stripped of manual 'success:true' (50 across 14 route files) so they return bare payloads the hook wraps uniformly; fixed the mux DELETE {success:<bool>} envelope collision (-> {killed}).
Frontend (48 call sites across 10 files):
- _apiJson() auto-unwraps { success:true, data } -> data (null on error), so most bare-shape readers are transparent. Raw-fetch sites relocate payload reads under .data; success/res.ok/error checks unchanged.
Docs: new docs/api-reference.md (envelope, status table, error codes, /api/v1, SSE); versioning-policy.md flipped — the HTTP/SSE API is now part of the stable, SemVer-covered surface.
Verification: full unit/route suite green (2680 passed) incl. ~166 updated assertions across 24 test files; typecheck/lint/format/frontend-syntax clean; a headless-chromium smoke loaded the migrated UI and drove the panels with 0 console/page errors; /api/status and /api/v1/status confirmed returning the uniform envelope live.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Update Codeman from the web UI: a "Check for updates" button queries GitHub
for the latest tagged release (git ls-remote fallback) and shows release
notes; "Update now" runs git checkout <tag> → npm install → npm run build →
restart, streaming live progress that survives the service restart.
- Release-tag channel; dirty trees auto-stashed (left for manual git stash pop)
- Cross-platform restart: systemd / launchd / manual, detected at runtime
- Updater runs detached (systemd-run --scope on Linux, setsid on macOS) so the
restart it triggers can't kill the build mid-flight
- Build-failure rollback to the pre-update commit; boot reconcile with an
update-id/freshness guard; 409 concurrency lock; runner staged outside the
repo; strict tag validation; CODEMAN_DISABLE_SELF_UPDATE kill-switch
- Endpoints: GET /api/system/update/check, POST /api/system/update,
GET /api/system/update/status
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Make the branch genuinely master-mergeable and fix several review findings:
- Defaults are now prod-safe: CODEMAN_INSTANCE defaults to '' (→ ~/.codeman,
-L codeman) and the web port back to 3000, so an existing install upgrades
cleanly. Port also honors a new CODEMAN_PORT env var. Run the beta isolated
alongside prod with scripts/run-beta.sh (CODEMAN_INSTANCE=beta + PORT 5000).
- .gitignore: anchor the root `public` symlink rule to `/public` (a bare
`public` also swallowed src/web/public, silently un-staging new web assets);
ignore the gesture wasm/model binaries explicitly instead.
- span-displays: add a macOS-only guard (400 elsewhere instead of spawning a
bash that fails invisibly); extract resolveSpanUrl() for unit testing.
- server.ts: memoize asset-version stat() calls (~1s TTL) so each index render
doesn't re-stat every script/link tag.
- styles.css: hide the multi-monitor button in solo (detached) windows.
- app.js: require two consecutive unanswered roll-calls before redocking, so a
timer-throttled background popup isn't wrongly un-marked.
- index.html: make the "skip to terminal" link base-href-safe (onclick scroll)
so it doesn't navigate to the dashboard from a /session/:id window.
- Tests: test/config/instance.test.ts, test/routes/system-span-displays.test.ts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replace the header notification bell (now hidden by default; still reachable
via Settings → Notifications and the drawer) with a multi-monitor button.
Clicking it POSTs /api/system/span-displays, which spawns the bundled
scripts/span-codeman.sh — a fresh, maximized browser --app window sized to the
union of all displays — so in-page floating session panels can be dragged
across the physical monitor seam. macOS only; needs the one-time "Displays
have separate Spaces" OFF prerequisite (documented in the script). The route
pins the spanned window to localhost with a digits-only port from the Host
header so nothing attacker-controllable reaches the launched browser.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Detach a session tab into its own browser window and back.
Detach/undock:
- GET /session/:id serves the SPA in "solo mode", reusing the existing
client (terminal, local-echo overlay, reconnect) so no terminal code is
duplicated. One PTY already fans out to N SSE/WS clients, so a detached
window is just another live client — no server fan-out work was needed.
- A pop-out icon per tab; detached tabs show a badge and focus the popup on
click; closing the popup re-docks. Cross-window state via BroadcastChannel
plus a WindowProxy poll, and survives a dashboard reload (roll-call).
app.detachSession(id) is a single idempotent entry point (future gesture
hook). <base href="/"> so relative assets resolve under /session/:id.
Beta-branch isolation (so it can run alongside a prod Codeman):
- Default port 3000 -> 5000.
- New src/config/instance.ts derives the data dir and tmux socket from
CODEMAN_INSTANCE (default "beta"): ~/.codeman-beta + tmux -L codeman-beta.
Every ~/.codeman path now goes through dataPath()/getDataDir() (state,
mux-sessions, settings, push keys, lifecycle log, screenshots, certs,
linked-cases, subagent window state). Overridable via CODEMAN_INSTANCE /
CODEMAN_DATA_DIR / CODEMAN_TMUX_SOCKET. Prevents a second instance from
discovering and attaching PTYs to the first instance's live tmux sessions.
Verified: tsc / eslint / prettier / lockfile clean; Playwright E2E (27 checks)
for detach/solo/redock; default isolation confirmed to see zero real sessions.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Extract readLinkedCases() helper and resolveCasePath() to eliminate 6x duplicated
linked-cases.json path construction and 5x duplicated file read/parse logic
- Replace O(n) .some() duplicate check with O(1) Set.has() in case listing
- Un-export isError() in types/api.ts (only used internally by getErrorMessage)
- Standardize reply.status() → reply.code() in system-routes (Fastify canonical API)
- Update CLAUDE.md: accurate frontend module listing, SSE event count (~106)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Green pulsing dot in the desktop header shows when Cloudflare tunnel is active.
Clicking opens a dropdown panel with tunnel URL, remote client count, auth
sessions, and start/stop/QR/revoke controls. Detects tunnel clients via
Cf-Connecting-Ip header to exclude local connections from the count.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>