Small cleanup items from upstream review (Ark0N/Codeman#353):
- OMP_SEARCH_DIRS now leads with ~/.local/bin, matching omp.sh's real
installer target (~/.omp/bin was an earlier unverified guess, confirmed
wrong against a real --no-cache Docker build).
- docs/omp-integration.md: fixed the dead GitHub URL (can1357/omp ->
can1357/oh-my-pi), corrected the CLI count (ninth backend, tenth
SessionMode incl. shell -- not eighth), matched the install-path guidance
to the resolver fix, updated the version example to the actually-tested
18.0.8, and added a Docker-section caveat: --resume pinning does not
currently reach an in-container omp process, since Docker panes never see
ompConfig.
- docs/architecture-invariants.md: fixed a heading missing ", OMP" (CLAUDE.md
already linked to the -omp anchor, so the link was dead) and added an OMP
specifics paragraph -- the one external CLI missing an entry in this doc.
- .changeset/omp-backend.md: corrected the sibling-CLI list (was missing Pi,
Grok, and DeepSeek Harness) and the backend count.
- Removed a stray orphaned comment fragment in the quick-start docker branch
and split two CSS lines that had two declarations jammed onto one line.
resumeHistorySession() creates the resumed row in its own mode via a
modeConfigKey map (opencode/pi/grok/omp -> continueSession, deepseek ->
resumeSession) and retires the old row afterward. codex, gemini and
antigravity were missing from that map, so resuming one of their rows
started a brand-new session with NO continuation while still deleting
the row it came from -- silent data loss dressed as the duplicate-row
fix. Gate row retirement on continuesSomething (true only for modes that
actually got a continuation config) instead of wiring an unverified
sessionId->native-conversation-id assumption for the three affected CLIs.
DELETE /api/sessions/:id reimplemented the ownership 404 check inline in
two places instead of going through findSessionOrFail, and its
persisted-only-session branch never broadcast session:deleted, so other
open tabs kept the retired row until their next unrelated fetch. Extract
the shared 404 into sessionNotFoundError(), add findPersistedSessionOrFail()
alongside findSessionOrFail() in route-helpers.ts (same ownership
contract, returns a SessionState instead of a live Session), and use both
from the route instead of inline checks. Add the missing broadcast.
The docs claimed omp "has no documented vendor-key namespace of its own"
and "the multi-user clamp has nothing to gate" for omp — both false. Per
omp's own docs/environment-variables.md, it reads ~40 provider keys from
env (pi's known 34-key problem in the same shape), and its own knobs are
mostly PI_* (already globally allowlisted): PI_CONFIG_DIR,
PI_CODING_AGENT_DIR, PI_CODING_AGENT_SESSION_DIR, PI_SUBPROCESS_CMD,
PI_SHELL_PREFIX. The first three also move the ~/.omp tree
omp-session-resolver.ts/omp-transcript.ts hardcode, silently degrading
pinning/history — a known gap shared with pi, documented but not fixed
here.
The OMP_* prefix this PR adds brings in OMP_AUTH_BROKER_URL/
OMP_AUTH_BROKER_TOKEN, where omp resolves credentials from — the same
shape DEEPSEEK_BASE_URL is already dropped for in
clampEnvOverridesForOwner(). Add both to OWNER_CLAMPED_ENV_KEYS so a
non-granted owner in multi-user mode can't redirect them, and correct the
false claims in CLAUDE.md, docs/omp-integration.md, and the stale
resolveOmpHome() comment. Also documents omp's default
tools.approvalMode: yolo, which was previously unstated.
findLatestOmpSessionId()'s newest-mtime pin ran eagerly inside
_buildRespawnPaneOptions(), which startInteractive() calls unconditionally
on every boot-recovery reattach — before anything checks whether the pane
is actually dead. With two omp tabs in the same case dir, this could pin
an ALIVE pane's session onto whichever sibling's file happened to be
newest on disk, purely as a side effect of building options that might
never lead to a respawn (reported in Ark0N/Codeman#353 review).
Move resolution out of the eager builder into _pinOmpRespawnId(), called
explicitly only where a respawn is actually confirmed: the dead-pane
branch in _setupOrAttachMuxSession() and reattachRemote(). Add
resolveAndClaimOmpSessionId(), which verifies each candidate's own file
header (cwd) rather than trusting the mangled-directory match alone, and
tracks claimed ids in a process-wide registry so two ambiguous resolutions
can't both pick the same sibling's conversation.
Found live 2026-08-27 by Tim: clicking Run OMP to start a brand-new session
in a case directory with prior omp history launched --resume <old-id>
instead of a clean `omp` invocation.
Root cause: Session._resolvedOmpRespawnConfig() resolves-and-pins the
newest on-disk omp conversation as a side effect on this._ompConfig. That
is correct when reattaching to an ALREADY-TRACKED mux session (a dead-pane
respawn, or a boot-recovery reattach - the constructor sets _muxSession
from persisted state before startInteractive() ever runs there), but it
ran unconditionally. startInteractive() computes
`respawnPaneOptions: this._buildRespawnPaneOptions()` eagerly in the same
object literal that builds `createSessionOptions.ompConfig: this._ompConfig`,
so for a genuinely brand-new session (no muxSession in its create config,
_muxSession still null) the resolve-and-pin side effect ran and poisoned
this._ompConfig before that field was even read.
Fix: gate the resolve-and-pin logic on `this._muxSession` already being
set. A fresh session has no muxSession yet and now passes through
untouched; a real reattach (muxSession present since construction) keeps
resolving and pinning exactly as before.
Verified live in production against the exact reported scenario (a fresh
omp session in a case dir with 8+ hours of prior omp history) - confirmed
both via the API (ompConfig stays empty, claudeSessionId equals the
session's own id) and visually in the GUI. Regression test constructs a
real Session + TmuxManager to exercise the actual private-method
interaction directly, since no existing test called startInteractive() at
all.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
OMP was the one external CLI mode with no dedicated user-guide doc, unlike
opencode/pi/grok/deepseek which each have one. Covers install, auth (omp
owns its own entirely - no Codeman-side login flow or bypass switch),
what Codeman wires up (OmpConfig), the exact-id pinning mechanism and the
directory-mangling bug behind it, kill-survival via transcript scanning,
terminal behavior, Docker/remote-SSH cases, and known gaps (no idle hook,
mid-turn kill data loss, unverified symlinked-$HOME behavior).
Cross-referenced from README.md's Multi-CLI doc list and docs/docker-cases.md's
credential-seeding summary (which now also documents OMP's sessions/-is-shared
exception to the seed-everything pattern the other CLIs use).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
OMP had full routing at the Docker layer (default pane command, schema) but
was never actually installed in docker/agent.Dockerfile, and had no
credential-isolation entry in docker-hosts.ts's CRED_STORES - a Docker-mode
OMP session would have failed with "omp: command not found", and even with
the binary present would have had no config/auth seeded, despite the README
already claiming OMP has "seamless auth, isolated credentials" in Docker.
- docker/agent.Dockerfile: install omp via its own installer (standalone
binary, same shape as grok/antigravity - not on npm). Verified against a
real --no-cache build: the installer actually targets ~/.local/bin, not
~/.omp/bin as the resolver's OMP_SEARCH_DIRS ordering would suggest -
confirmed omp/18.0.8 installs and runs correctly inside the image.
- src/docker-hosts.ts: add a .omp/agent CRED_STORES entry. Unlike every
sibling CLI in this family, sessions/ is SHARED (RW), not seeded: Codeman
reads ~/.omp/agent/sessions/**/*.jsonl host-side for history recovery and
--resume pinning (omp-transcript.ts, omp-session-resolver.ts), the same
reason codex's sessions/ is shared rather than seeded. Seeding it instead
would silently break the kill-survival feature for Docker cases. Only the
small config files (config.yml/mcp.json/models.yml/settings.yml) are
seeded; the SQLite caches and terminal-sessions/ stay container-local.
- test/docker-hosts.test.ts: pin the new CRED_STORES entry's behavior.
Found in passing (NOT fixed here, unrelated and pre-existing on master): the
agent image's DeepSeek (dsh) plugin-install step currently fails on a fresh
build ("pnpm not found on PATH"), confirmed via git diff against
origin/master that this line is untouched by this branch. Worth a separate
issue/PR.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Follow-up from a full-branch review pass (Opus) of the omp-mode integration:
- Add pinning tests for resolveOmpConfigForCreate() (session-routes.ts),
exported to make it testable: the exact "resume this OMP row from
history" pipeline that mangleOmpWorkingDir's earlier bug lived in had
zero coverage despite being the resolver module's whole reason to exist.
- Log a warning when findLatestOmpSessionId() finds nothing on disk and
continuation silently degrades to omp's own ambiguous --continue,
in both call sites (session create and respawn pinning) - previously
silent, making the degradation invisible to anyone debugging it.
- Require an absolute cwd before trusting a session file's working
directory in omp-transcript.ts's parser, so a corrupted/malformed
session file can't point a downstream resume at a relative or empty
path.
- Document (don't speculatively fix) an unverified symlinked-$HOME edge
case in mangleOmpWorkingDir(): the review's suggested realpath() fix
assumes omp itself resolves symlinks before mangling, which is
unconfirmed - guessing wrong there would trade one silent mismatch
for a different one.
- Incidental: fixed unrelated pre-existing prettier drift in
session-routes.ts (antigravity/opencode dynamic import line-wrapping)
that was blocking the pre-commit formatting gate on this file.
Confirmed as a non-issue: the model-name regex allowing "/" is
intentional (provider/model ids like "crof/glm-5.2" were used
successfully in live testing).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two bugs compounded to break continuation pinning on every real OMP
case (only /tmp-based manual testing happened to work by coincidence):
1. startInteractive() had a second, unconditional claudeSessionId
assignment after the mux branch that clobbered its correctly
resolved value back to the session's own id on every mux path.
2. mangleOmpWorkingDir() assumed omp mirrors Claude Code's directory
naming (home prefix kept), but omp actually strips $HOME first.
findLatestOmpSessionId() was silently returning null for every
case under ~/codeman-cases/, so resumeSessionId never resolved for
any real case dir - only /tmp paths (outside $HOME) worked, which
is every dir this feature was previously tested against.
Verified live: killed and relaunched the omp-verify server process
mid-session (plain reattach, pane stayed alive) and confirmed
claudeSessionId now resolves to the real omp transcript uuid instead
of the Codeman session's own id.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude conversations survive "Kill Tmux & Claude" because Codeman reads
them back independently from ~/.claude/projects, not from its own
session bookkeeping. omp conversations had no equivalent: kill the
Codeman session and the conversation vanished from Past Sessions
entirely, even though omp itself never forgot it on disk.
Adds omp-transcript.ts, a scanner over omp's own
~/.omp/agent/sessions/<mangled-cwd>/<uuid>.jsonl files (the same shape
as Claude Code's own transcript scanner, but simpler -- these files are
small enough to read whole instead of doing head/tail windows). Each
file's own "session" header line carries the real cwd and session id
directly, so unlike Claude's mangled-directory-name decoding this
never has to guess. Wired into gatherUnifiedInputs() as a second
history source alongside the Claude scan, and HistoryInput/
mergeUnifiedSessions() now carry an optional `mode` so a non-claude
history-only row still gets a real mode badge.
Also fixes the ambiguity behind the "continue picks the wrong
conversation" report from this session's testing: omp mints its OWN
session uuid, unrelated to Codeman's, so a live/persisted row and its
own history-scan row would otherwise show up as two separate entries
for the same conversation the moment the id gets resolved. Reuses the
existing claudeSessionId alias field (mergeUnifiedSessions' fold-into-
owner mechanism) to point at the resolved omp id, threading it through
every place `_claudeSessionId` gets (re)computed -- the constructor,
_resolvedOmpRespawnConfig, and a new _maybeCaptureOmpSessionId() that
opportunistically resolves it the first time a brand-new omp session
(one that has never gone through a respawn) goes idle.
Also closes a THIRD instance of the "ompConfig never got wired in
here" gap this session kept finding: restoreMuxSessions() in server.ts
restores every sibling CLI's config from persisted state on boot except
omp's, so a boot-recovered omp session always lost its resolved resume
id and fell back to guessing again.
Verified live end-to-end: told a session a secret, killed it fully
(Kill Tmux equivalent, killMux=true -- the Codeman session AND its tmux
pane both gone), and the conversation still showed up in the unified
list as a history-sourced row with the real first prompt as its title
and an omp mode badge, keyed by omp's own session id.
Known remaining gap, not fixed here: the claudeSessionId alias doesn't
yet resolve reliably on every boot-recovery path for a session that
was never respawned while alive (e.g. a plain re-attach to a pane that
was never dead) -- worth a follow-up, but doesn't affect the two things
that matter most: the conversation surviving a kill, and continuation
correctness once an id has been resolved (which happens on the very
next respawn either way).
Every non-claude "Resume" click creates a brand-new Codeman session
(there is no id to reattach to), but the old row was never cleaned up
-- click resume on the same conversation a few times and the session
list fills up with duplicate rows sharing one name. resumeHistorySession
now retires the row it resumed from after the new one starts.
That retirement needs DELETE to actually work on a row that was never
live in the first place (the normal case for anything showing up in
"Resume Conversation"): findSessionOrFail only checks the in-memory
live-session map, so DELETE 404s on a persisted-only entry today. Give
the route a fallback: when the id isn't live, look it up in persisted
state instead and demote/remove it there (respecting the existing
pinned-session protection). Verified live against a real persisted-only
row via the API, and added route-test coverage for both the success
and still-truly-unknown-id cases (which needed a demoteOrRemoveSession
mock the route harness didn't have).
Also includes an unrelated pre-existing prettier drift fix picked up
by npm run format (omp-cli-resolver.ts, antigravity/opencode import
wrapping in session-routes.ts).
respawnPane() -- the path used when a session's pane died (crash, idle
respawn, or the user's own /exit) but the Codeman session object is
still tracked -- never had ompConfig wired through at all, in either
its options destructure or its inner buildSpawnCommand() call. This is
a gap in the original OMP patch, distinct from the resumeHistorySession
fix (which only covers a session that has been fully closed and shows
up as a history row): reselecting a tab whose CLI process just exited
goes through this path instead, and always launched a bare, contextless
`omp` no matter what.
Beyond the wiring, respawning a dead pane is semantically different
from creating a brand-new session: the conversation is still "this
session" to the user, so _buildRespawnPaneOptions() now defaults
ompConfig to continueSession:true unless the session already carries
an explicit resumeSessionId (which still wins in buildOmpCommand).
Verified live: told a session a secret, exited OMP so the pane died
(session and tmux both left alone), forced the exact dead-pane-respawn
path, and the new process replied with the secret -- confirming
`omp --continue` fired instead of a blank omp.
resumeHistorySession() never sent mode when recreating a session from a
history/session-manager row, so the server default silently opened a
plain Claude session for every non-claude row -- reproduced live: OMP
rows spawned Claude sessions on click. Thread the row's mode through
every call site (welcome list, session manager, mobile overview) and
only send the Claude-specific resumeSessionId for claude rows.
Codeman has no live PTY-reattach outside server boot, and it's moot for
OMP anyway (exiting it kills the pane's only process), so route the
non-claude relaunch through each CLI's own continue-most-recent flag
instead of a context-free fresh start. OMP never got one: buildOmpCommand
only implemented --model/--resume despite omp --help documenting
-c/--continue. Added continueSession to OmpConfig end-to-end (type,
schema, builder) mirroring the existing opencode/pi/grok/deepseek
fields, and wired resumeHistorySession to use it.
Verified live: told a real omp session a secret, exited it, closed the
tab without killing tmux, relaunched with --continue in the same
directory, and had it recall the secret.
Two prose lists in skills/codeman/ named some but not all external CLI
modes after the omp-mode rebase, which is exactly the drift
test/agent-skill-mode-lists.test.ts exists to catch: SKILL.md's
no-hook-signals list was missing omp, and endpoints.md's version-probe
sentence named pi/grok/omp as a bare 3-mode run with no matching class.
- omp-cli-resolver.ts already uses createCliExecutableResolver; add dedicated
test/omp-cli-resolver.test.ts mirroring pi's (version-probe accept/reject,
negative-cache backoff, VITEST hermeticity gate)
- dependency-registry omp entry now requires OMP_VERSION_REGEX match like pi,
so codeman doctor and the run-mode resolver agree on what counts as installed
- system-routes /api/omp/status surfaces version
The rebase hand-repair dropped the closing brace of .welcome-btn-pi:hover
and .btn-toolbar.btn-run.mode-pi:hover before the inserted OMP rules.
The browser CSS parser drops every rule after an unclosed block, so the
deployed UI rendered as unstyled text bars (only ~456 of ~2583 rules
applied). Verified clean via esbuild --minify (no css-syntax-error) and
rebuilt dist.
The network-access prompt, where Tailscale serve is configured, runs AFTER
the build step. A build failure therefore exits before the question is ever
asked, and a user who then finishes the build by hand (rather than re-running
install.sh) ends up with a healthy loopback-only Codeman, a connected
Tailscale, and no serve mapping — with nothing anywhere pointing at
`install.sh tailscale`, the command that fixes it. Reported from a fresh
Ubuntu 24 install that died on the node-pty compile.
- maybe_offer_tailscale_repair(): on the update/re-run path, detect exactly
that state (loopback bind + tailscale Running + no serve mapping fronting
Codeman) and offer the retrofit. Silent for a deliberate non-loopback bind,
silent once a mapping exists, silent when tailscale is absent, and prints
the command instead of prompting when non-interactive. Returns 0 even when
setup fails so it can never abort an update.
- print_security_notice(): the loopback branch now names
`install.sh tailscale` when Tailscale is installed on the box, rather than
the generic "tailscale serve / cloudflared tunnel" advice.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
install.sh installs a build toolchain on Linux (node-pty has no Linux
prebuild, so a stock Ubuntu 24 server died inside node-gyp with
"not found: make"), plus review hardening for #339: the write-queue
reset paths now release the one-chunk-in-flight gate.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fifteen review findings on the dsh mode, the serious ones first:
- Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys.
_configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into
every dsh pane and applyEnvOverrides() lands after it, so a non-granted
owner who could redirect the base URL would have the operator's key sent
as a bearer credential to a host of their choosing.
- Wait registry: until=stop/blocked is refused on docker and remote-SSH
dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions).
The HERDR triple is set via LOCAL tmux setenv, which crosses neither
docker exec nor ssh, so such a session can never post a hook event and
the wait burned its whole timeout on every turn.
- Approvals: a dsh item is an ALERT, not an answerable card. The answer
route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the
option parser cannot read a third-party TUI's frames, so an answer was a
blind keystroke into a foreign composer), and the push notification
carries no Approve/Deny actions for dsh sessions.
- Status shim (v3): --seq is forwarded and the server drops stale retried
reports inside a 60s window (the TUI retries with backoff, so a retried
'working' could land after 'blocked' and resolve an approval whose
dialog was still on screen); 4xx responses exit 0 instead of retrying,
so one misconfigured session cannot feed the auth rate-limit bucket
until the hook endpoint 429s for the whole instance.
- Web-UI server: concurrent starts are serialized through a lock (two
racing POSTs used to pick the same port and orphan the winner), and the
readiness poll / timeout paths only clear or stop the singleton while it
is still theirs. First click actually opens the tab now
(refreshWebviews, not the nonexistent loadWebviews). DELETE
/api/deepseek/web requires the privileged grant in multi-user mode.
- Cron: deepseek jobs run the same two-part launch gate as the HTTP
create paths (impl moved into the resolver so all three share it) and no
longer stamp a Claude default model on the session.
- Parity sweeps: quick-start's docker branch rejects deepSeekConfig like
the remote branch; the Ralph auto-enable list gained deepseek;
HookEventType gained agent_working; the phone overview run menu filters
managed webview records like the desktop menu.
- install.sh: the dsh identity probe closes stdin (under curl|bash a
child that reads stdin eats the rest of the script), bounds the exec
with timeout where available, and is memoized to one scan per install.
- Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand
identity (it rendered as an unstyled UA-grey button); stale markup
comment about the web shortcut rewritten; clamp docs updated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Four review findings on the worker-transcript feature:
- Docker and remote-SSH dsh sessions now keep the pane segmenter: their
transcripts live in the container's / remote host's own ~/.dsh, which
the local reader can never see, so the transcript path returned
'nothing said yet' forever and an agent polling such a worker starved
on an answer that existed. Gated on !session.docker && !session.remote
(statically pinned) and documented in the integration guide.
- last-response reads are memoized on (path, mtime, size, blocks): the
skill's last_text polls once per second, and each poll decompressed and
reparsed the whole file on the event loop even when nothing had been
appended. An unchanged poll now costs one stat.
- The pairing ladder's comment claimed /new is served by step 2; in truth
the boot-window transcript wins for as long as it exists (deliberately:
preferring newest-eligible would hand a worker its busier sibling's
reply). The comment now states the real tradeoff instead of the
aspirational one. Same for decodeZstdFrames' 'skipped' wording — a
corrupt frame truncates the decode there, which is the safe behavior.
- stripReasoningPrefix no longer runs on user prompt text, so a prompt
containing a literal </think> renders whole in blocks view.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three review findings on the detailed-rows feature, all in its edge cases:
- The App Settings width select consulted the handheld defaults blob
(tabRailWidth: 256) BEFORE the rich-aware default, which the renderer
never reads — so a tablet's unsized rich rail rendered 320 while the
dialog said 256, and a routine Save persisted the 256 (below the 288px
tight threshold, permanently). The chain now mirrors
applyTabRailWidth()'s actual resolution.
- _setTabRailWidth() re-rendered on a compact flip but never re-ran
applyTabWrapSettings(), the one owner of the folder line, whose railRich
input reads the compact class this function just toggled. A rich rail
dragged below 240px kept emitting folder rows — persistently, for a
stored width < 240, since the boot wrap pass runs before the class is
first applied. The wrap pass now re-runs on the flip, with exactly one
render either way.
- Both reset affordances (handle dblclick, Enter on the handle) reset to
the hardcoded 256 even on a rich rail, landing it below the tight
threshold; both now resolve the rich-aware default (320), via a new
optional defaultWidth input on resolveTabRailKeyboardWidth().
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The agent skill could spawn a worker in any mode, but it could only
DRIVE a claude one: every other CLI has neither a real end-of-turn
signal nor an answer to read, so the recipes route them through output
markers.
dsh has both halves now -- its harness reports idle/working/blocked to
Codeman, and the previous commit reads its transcript -- so it joins
claude as a mode the four verbs work on unchanged. `spawn_workers alpha
beta:deepseek` is a mixed fleet in one call, and `sendwait` / `last_text`
/ `delete_session` need no per-mode variant.
Preamble 1.20.0 (SKILL.md's §0 heredoc regenerated from it):
- `spawn_worker` grows a deepseek branch that gates on the harness
composer. ⚠️ Readiness there is NOT the stop signal: the harness
reports idle at BOOT ~300 ms before its composer paints (measured
2.26 s vs 2.56 s after spawn), so a send-and-wait fired straight after
quick-start resolves on the boot edge, reports a turn that never ran,
and strands the prompt in a pane not yet taking input. Waiting for the
composer also spends that edge, since signals are edge-triggered.
- `spawn_workers` takes `name[:mode]`, so a mixed fleet stays one
concurrent call. Case names still have to be unique -- the mode never
disambiguates two workers that would share a directory.
- `sendwait` asks for `wait:"stop,exit"` instead of the `wait:true`
default set. That set also carries `idle`, which for an external CLI is
inferred from output stabilization: on a dsh worker whose TUI repaints
rarely, the re-wait resolved in 0 ms with `signal:"idle"` on a turn
with three minutes left to run. It also makes a wrong mode loud -- the
modes that cannot deliver `stop` answer 400 before writing anything,
instead of resolving on a flap.
- The self-heal resend carries `delivered:true` forward. The resend is a
tagged duplicate, so the server truthfully reports `delivered:false`
about a write it skipped, and §1's cleanup then read a completed turn
as an undelivered one and kept a finished worker forever.
- dsh workers spawn with the permission posture the Run button sends,
because the harness default still asks and a worker parked on an
approval row cannot finish a fan-out. The multi-user clamp still
applies.
Docs: a worked dsh flow in recipes.md, readiness and the signal rules in
verbs.md, and the corrections this makes necessary -- `stop`/`blocked`
are no longer claude-only, and `last-response` is no longer permanently
empty for deepseek. The integration guide gains a section on reading a
session back and driving one as a worker; its web-UI section was also
stale (that server moved out of a shell session).
The static guard that keeps those lists from naming some external CLIs but
not others is extended rather than exempted: it now knows the three real
classes inside that family (no transcript, no hook signals, and the
positive twin -- the modes whose answers can be read), with the hook class
derived from `hooksAvailableForMode()` so the predicate and the prose
cannot drift apart. Any other partial list still fails, and a new backend
belongs to none of the classes until someone says so.
`GET /api/sessions/:id/last-response` is how an agent (and the Response
Viewer) reads what a worker said. DeepSeek was falling through to the
pane segmenter with the other external CLIs, which for this mode is not
merely coarse but wrong: dsh-TUI paints a full-screen splash, so a
`last-response` call on a fresh dsh session answered with its ASCII-art
logo -- and anything polling for a worker's first reply reads that as a
reply.
dsh does not belong in that group. It writes a structured JSONL
transcript per session, so read it. Four things in that file shaped the
reader, all measured against real transcripts on disk:
1. dsh appends ONE ZSTD FRAME PER WRITE, and Node's zlib zstd decoder
(one-shot and streaming alike) stops at the first frame end: a real
56-line transcript decoded as 1 line / 158 bytes -- the session header
alone, i.e. a silent truncation that reads as "nothing said yet"
forever. `zstdFrameRanges()` walks frame and block headers to find
exact boundaries; splitting on the 4-byte magic would corrupt
everything after a magic sequence occurring inside compressed data.
zstd is resolved at RUNTIME because it landed in Node 22.15 while the
project floor is 22.0, so an older Node keeps the pane behaviour.
2. Every turn also records a plugin-sourced `user/message` (the runtime
context snapshot), which must not render as the user's own words.
3. A turn that ends in an error carries the provider's message; it is
surfaced as `Turn error: …` (and a non-error early stop as
`Turn ended: …`) rather than as an empty string, which an agent reads
as "still thinking" through fifteen polls.
4. Reply text is assembled per (turn, step): a finalized message wins and
the streamed deltas fill in only for a step that never finalized, so a
partial answer is readable mid-turn and never doubled. "Finalized" is
tracked as a set of steps rather than as non-empty text, because a
step whose whole reply was reasoning strips to '' at the `</think>`
boundary and would otherwise resurrect the raw deltas in its place.
Session-to-transcript pairing is by the transcript's own header `cwd`
plus a boot window against the session's createdAt, never by
reproducing dsh's directory mangling (already two forms on disk) and
never by newest-mtime alone -- mtime alone handed a freshly spawned
worker its predecessor's answer in the same case directory.
An empty result still wins over the pane; only a Node that cannot decode
zstd falls back to it.
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a
shell tab running the server next to it. The shell was deliberate - the server
lived in an ordinary session so it was visible, scrollable, killable and died
with its tab, and nothing new had to supervise a long-lived HTTP server. That
reasoning was sound and the result was still wrong in use: opening a dashboard
should open one tab, and after the first launch the terminal is pure noise.
The server moves to a background child process owned by a new
`src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session
gave away for free is now explicit, which is most of the module:
- Exactly one server. A second click reuses the running one instead of racing
it for a port; the session flow could not do this at all, because two clicks
were simply two sessions.
- Restarted when the requested authority changes. `--trusted-host` fences dsh's
own /api against the browser authority, and a Codeman reachable at both
loopback and a tailnet name has two. Reusing a server fenced for the other
origin renders a page whose every call 403s, which reads as a broken
dashboard rather than a misconfigured one, so a mismatch restarts instead.
- Killed on shutdown. The child is detached so its whole plugin tree can be
signalled at once, which also means it would outlive Codeman and hold its
port against the next start - the exact EADDRINUSE this feature already got
wrong once.
- Boot output captured and returned. With no shell tab there is nowhere else
for a stack trace to land, so a failed spawn reports its own tail.
The endpoint is fenced at the same bar as the profile installer and for the
same reason: booting a dsh profile executes the plugin code in it, so this is a
privileged action even though it reads as "open a page". `authority` comes from
the client (`location.host`) because only the browser knows which origin is in
play, and it is regex-confined at the schema boundary - defence in depth behind
the argv-array spawn, admitting host:port in the shapes a browser authority can
take and nothing readable as a second argument.
`GET /api/deepseek/web-port` is gone; port selection moved into the supervisor,
which is the thing that knows whether a server is already running. The two
client-side probe helpers went with it, since the server now owns the wait.
Verified over the tailnet authority end to end: no session is created (session
count unchanged, one tab), the server runs on 3081 beside the user's own dsh
web on 3080, status reports the tailnet authority, and the proxied dashboard
renders with zero 4xx. Full gate green (6148 passed, +6).
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.
1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
precisely the port a DeepSeek user is most likely to be serving on already,
so the launch died with EADDRINUSE against the user's own server. The port
now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
free loopback port by BINDING it (a connect probe cannot tell "free" from
"listening but not answering yet").
2. It opened the tab unconditionally. The crashed server left a saved dashboard
pointing at nothing, with the failure only visible in a shell tab nobody had
a reason to look at. The launch now polls the existing webview probe until
the URL answers, and on timeout reports the error naming the shell tab
instead of persisting a dead dashboard.
3. The saved tab was untrusted, so the frame was sandboxed without
`allow-same-origin` and the dashboard was broken twice over: the dsh
client-runtime reads `localStorage` while loading its plugins and died there
("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
every `/api` call no matter which authority `--trusted-host` named. Passing
`location.host` only means anything once the frame actually carries that
origin, so `--trusted-host` had never once done its job. The managed tab is
now created `trusted: true`.
That trade is real and deliberate: a trusted proxied frame is same-origin
with Codeman and can reach Codeman's API. It is defensible only because this
dashboard is an agent harness Codeman just started itself, on loopback, which
can already run code as the user. It is not a precedent for trusting
third-party dashboards, which is why it is set at this one call site rather
than defaulted.
Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.
`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.
Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
Two review nits on the vertical rail's detailed rows.
1. The tab-rail-tight rule (below 288px) hides `.tab-meta-created`, and its
comment claimed the value "survives in the row's title attribute either way".
It did not: the only title carrying it lived ON that element, and a
`display: none` element has no hover target, so the created stamp was not
shrunk but gone with no way to ask for it. Rather than just correcting the
comment, `_sidebarRichMetaHTML()` now puts BOTH absolute stamps on the
`.tab-meta` line itself, so the pill and the gaps around the stamps remain as
hover targets. An item's own title still wins where the item is visible.
2. applyTabOrientation() decided whether applyTabWrapSettings() had already
re-rendered by comparing `_tallTabsEnabled` before and after. That reads an
UNDEFINED previous value as "it rendered", but applyTabWrapSettings()
deliberately renders nothing on its first call ever (it only establishes the
baseline: `prevTallTabs !== undefined && prevTallTabs !== showFolder`). So on
a first call that also flips the folder row, neither function rendered and the
rows stayed stale. Reachable when the pre-paint script throws and leaves the
layout attributes on their catch-branch fallbacks for applyTabOrientation() to
correct. The guard now mirrors applyTabWrapSettings()'s own condition.
Both new tests were run against the unfixed code first and fail there, which is
the only thing that makes them regression tests. (The third, "does not render
twice", passes either way by design: it pins that fix 2 did not introduce a
double rebuild.)
Verified in a real browser against a live server with two sessions, driving the
narrowing through _setTabRailWidth() the way the resize drag does: at the 320
default the row reads "CREATED 2m ago · IDLE <1m" with the created element
displayed; at 256 the tight class is on, the created element computes to
display:none, the visible text drops to "IDLE <1m", and the meta line's title
still reads "First created: ...". At 220 the compact threshold drops rich rows
entirely. Screenshots confirm no truncation artifacts in either state.
Full gate green (6104 passed), typecheck, lint, format, frontend-syntax and
public-assets all clean.
The three smaller review nits, plus the first real test coverage for the status
shim (it had none: it is emitted as a STRING, so tsc never sees it).
1. The shim was written with a plain writeFileSync. The TUI can be exec'ing that
exact path while an upgraded Codeman refreshes it, and a reader catching a
half-written file gets a syntax error, exits non-zero, and is retried four
times per state change for a file that will never parse. Now temp + rename
(atomic within the directory), with the temp chmod'ed before the rename since
writeFileSync's mode only applies on create, and removed if the write throws.
SHIM_VERSION bumped to 2, because SHIM_SOURCE changed and an existing v1 shim
would otherwise keep matching the embedded marker and never be refreshed.
2. The pane-id comment claimed the ambient env "cannot be spoofed by an argument
the agent itself could influence". The agent runs IN that pane and can invoke
the shim with CODEMAN_SESSION_ID unset and any argv it likes. It buys nothing
it did not already have (the hook-secret file is readable from the same pane,
so it can POST /api/hook-event directly), but the comment read like a security
boundary. Rewritten to say what the preference actually buys: correct
attribution when a TUI mangles or re-uses the pane argument. Accidents, not
adversaries.
3. classifyProfile() folded the directory name into the same haystack as the
bundles, but only the TUI arm could match a bare name, so a stock profile
whose package.json has no dsh.profile.bundles (hand-edited, older layout,
mid-install) classified as `unknown` -> launchable -> eligible as the DEFAULT
pick, which is exactly the pane-dies-on-arrival failure the two-part
availability gate exists to prevent. The stock names are now a LAST-resort
fallback consulted after the bundle patterns, so real bundle evidence still
wins over a name the user chose. The loose `tui` arm gained word boundaries:
it decides which profile boots by default, and matching the middle of
`intuition` is not a rule anyone could predict.
New test/deepseek-status-shim.test.ts runs the generated script the way the
harness does -- real node process, real argv, real env, real listener -- and
covers the exit-code contract that makes the retry behaviour safe: mapped states
post and exit 0, an unknown verb or unmapped state exits 0 WITHOUT posting (a
non-zero there would be four HTTP requests per state change forever), a rejecting
server or an unreachable one exits non-zero so the caller retries, the hook secret
is read at execution time, and `node --check` parses the file (a template-literal
typo in SHIM_SOURCE is invisible to tsc).
Trap worth recording, hit while writing it: the tests must spawn the shim
ASYNCHRONOUSLY. The listener lives in the test process, so spawnSync blocks the
event loop that has to accept the connection, the shim waits out its own 1500ms
socket timeout and exits 1, and it reads exactly like a broken shim (measured:
Socket._onTimeout in its --trace-exit output, server logging nothing).
Verified: full gate green (6142 passed, +10), typecheck/lint/format clean.
Three review findings on the DeepSeek Harness mode, plus one the third exposed.
1. The multi-user clamp was bypassable by a sibling field on the same request.
clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but
DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER
_configureDeepSeek(), so a non-granted owner sending
envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated
instance: a session created with permissionMode "read-only" and that override
ran with DSH_PERMISSION_MODE=danger-full-access in its pane.
Every other CLI's bypass is a command-line flag reachable only through the
per-CLI config, which is why the config clamp alone is the whole gate for
them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted
owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to
what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on
that list because it aims the launcher at a profile tree whose plugin code
runs at boot, before any approval row can apply. Verified end to end in real
multi-user mode: a non-granted user sending both now gets workspace-write and
no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched.
2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout`
signals only the direct child, and a plugin install fans out into
package-manager children that keep the inherited stdio pipes open, so `close`
never fires and the held-open request leaks with no route-level deadline.
Reproduced: with a 1.5s built-in timeout the promise was still unsettled after
6s and both fan-out children were alive. Now detached: true plus negative-pid
SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a
last-resort reap for a grandchild that escaped the group. Same probe after the
change: close fires, direct child and both grandchildren dead.
3. hooksAvailableForMode() promised more than a dsh session can deliver.
deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that
triple is the only reason a dsh session posts hook events, so `until=stop` was
accepted and then blocked for the caller's whole timeout: the exact
infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now
takes HookCapabilityOptions and every call site passes sessionHookOptions(),
with the deepseek arm reading `!== false` so a forgotten one degrades to the
old behaviour. The refusal names the setting rather than saying "no Claude
Code hooks", which would send the caller hunting a bug that is really a
setting they chose. Profile conformance stays unknowable at request time and
is documented as such. The stale "True for `claude` and nothing else" docblock
is corrected.
4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a
claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent
capture read Claude's own transcript, and adding deepseek silently widened
both to a mode that has none. They compare mode === 'claude' directly now, and
a static check pins them there.
Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the
wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 --
bridge off plus explicit until=stop is a 400 naming the setting, bridge off with
no `until` still 200s on idle/exit, bridge on accepts stop.
The vertical tab rail (tabOrientation 'vertical') listed names and nothing
else, while the rich sidebar and both home screens already answered the
question a docked column exists to answer: which of these sessions wants me
next, and how long has it been like that. The rail is a docked column too, so
it now draws the same row.
- New per-device setting tabRailDetail ('rich' | 'simple', default rich),
App Settings -> Appearance -> Tabs, in SettingsUpdateSchema + displayKeys and
stamped as data-tab-rail-detail by the pre-paint script, so a detailed rail
does not flash through simple rows on every load.
- ONE gate for both vertical surfaces: isRichTabRows() =
isSessionSidebarRich() || isTabRailRich(). The row model, the markup and the
20s in-place clock are the existing rich-sidebar ones, classified by
_mobileOverviewState/_mobileOverviewSince, so the rail, the sidebar, the
desktop home rail and the phone overview cannot disagree about what
"working" means or which stamp measures it.
- Detail rides on its OWN attribute, exactly as the sidebar's does, so every
existing [data-tab-orientation='vertical'] rule keeps matching both variants
untouched. A flip of detail ALONE still forces a full render (the stamps line
is emitted by the row template, not toggled by CSS) and re-runs
applyTabWrapSettings(), which owns the folder line and is now rail-aware.
- CSS: every rich paint rule gains a rail twin as a COMMA-GROUPED selector,
never :is() - an :is() list takes its most specific argument, which would
lift the sidebar arm from (0,3,1) to the rail's (0,5,1) and let these rules
outrank things they never used to.
- Width is why there are thresholds. At 256px the stamps line ellipsizes
mid-word, the same reason the rich sidebar is 300px, so a rail that has never
been sized defaults to 320 (RICH_DEFAULT_WIDTH, the existing Wide preset,
which also keeps the settings select on a named choice). A width the user has
chosen is never overridden: below 288px the created stamp is dropped rather
than truncated (tab-rail-tight, CSS only) and below 240px the rows go back to
simple (tab-rail-compact, which re-renders).
- The rich clock is armed and disarmed by applyTabOrientation() as well as
applySessionListLayout(); a leaked interval would rewrite stamps in a list
that no longer has any.
Also fixes a data-loss bug in the inline tab rename that predates the rail and
reproduces in every layout, header strip included: Escape set the input to ''
and blurred it, and the blur handler commits - so cancelling a rename PUT an
empty name, and the tab fell back to its folder label (measured against a live
server: ["rail-alpha","","rail-gamma"]). Escape now calls cancelRename(), which
invalidates the edit so the blur that follows the input's removal is a no-op.
Tests: rail-detail gate, the three ways it turns back off (simple, compact,
horizontal), sidebar-wins, render-on-detail-flip and the plumbing/CSS guards in
test/session-list-layout.test.ts; the rename cancel in test/inline-rename.test.ts
(browser suite), pinned by running it against the old code first. Verified live
against a real server on an isolated instance: detailed/simple/compact/header/
sidebar variants, click-select, the ... menu, inline rename, Alt+N, the in-place
stamp tick and a full settings-picker round-trip including reload.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.
DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:
1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
$DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
`base` -- the interactive terminal front door is always a third-party
plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
button gates on the latter, because reporting only the binary would spawn a
pane that dies on arrival. When the binary is present but no profile is,
the run menu offers to install one (POST /api/deepseek/install-profile).
2. The permission switch is an env var, not a flag. The harness has no
command-line permission option; its sandbox/approval rows read
DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
own workspace-write, which still asks, so the multi-user clamp is the
only-if-sent branch and clamps to workspace-write, never read-only.
3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
earned that. The terminal front door reports idle/working/blocked to a
supervising process over a generic env-gated contract; a generated shim
(deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
report to /api/hook-event as stop / agent_working / permission_prompt. So a
dsh session gets definitive respawn triggers, real wait-endpoint signals and
real Approvals Inbox items instead of output-stabilization guesswork.
`agent_working` is new (157th SSE constant) and joins
APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
alert at once.
The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.
Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.
Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.
Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>