Compare commits

...
Author SHA1 Message Date
Codeman maintainer a0298cf2b1 fix(skill): keep the re-wait open while sendwait works the composer
The first shape of the Enter loop read the composer BETWEEN two short waits,
and tested `wait.ended` (the session exiting) where it meant `timedOut`. A
`stop` that fired while no wait was open was lost, since signals have no
history, and a re-wait that had already resolved on `stop` fell through into
another wait that could never see the edge again: measured twice, the answer
was on screen and sendwait ran its whole 580 s slice anyway.

The long re-wait (a tagged duplicate of the original frame) is now registered
first and kept open in the background for the rest of the call; the loop reads
the composer and re-sends Enter beside it, stops when the prompt has left or
the wait's response has landed, then returns that response. Measured: the
stranded prompt got one extra Enter and sendwait returned on `stop` at 36 s.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:49:02 +02:00
Codeman maintainer 19ffe9b7a8 fix(input): make sure a prompt sent through the API actually leaves the composer
Claude Code 2.1.277 takes typed text the moment its composer paints but
ignores Enter for the first 30 to 50 seconds after it (measured 2026-09-19
through the input route: an Enter at 28 s stranded the prompt, one at 51 s
submitted it). The text+Enter pair `sendInput` sends 50 ms apart therefore
left every programmatic prompt sitting unsent, and every waiter burned its
timeout on a turn that never started.

Server: `SubmitVerifier` (session-submit-verifier.ts), armed from
`writeViaMux` for every mux write that carried a carriage return, reads the
pane on a 2 s to 60 s schedule and re-sends Enter only while the last
composer line (the CLI's own prompt glyph) still holds the head of what was
sent. An empty composer, other text, or no composer line at all ends it; a
newer write replaces the schedule.

Skill: `sendwait` gets the same loop (`_composer_text`, no-break space
stripped by its bytes for BSD sed) for servers that predate this, and the
preamble version moves to 1.30.1 so seeded agents pick up the fresh copy.
SKILL.md's heredoc and the plugin mirror are regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:38:59 +02:00
Codeman maintainer 3cdb4bf42e docs(terminal): the merge-time notes promised on #436
The four edits the review said would be folded in at merge, none of them
code: the changeset becomes one user-facing paragraph, since it is what
CHANGELOG.md and the release notes print; the `_bufferLoadFinishOpts` comment
now names the second contributor to the duplicate window (`captureActivePaneBuffer`
is `execSync`, so anything painted into the pane before the server read it is
in the capture and is broadcast after the reply) and says why a `history`
payload keeps the pre-existing discard when its exposure is the same; the
`_finishBufferLoad` doc block moves from above `_beginBufferLoad` onto the
function it documents; and the test file's header describes both rules the
file now pins instead of only COD-144.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 21:38:40 +02:00
Ark0N 492f8d8ddf Merge pull request #436 from irisitymichaelgrundberg/fix/replay-output-that-arrived-after-the-capture
fix(terminal): keep the output a pane capture could not contain
2026-09-18 21:34:42 +02:00
Michael GrundbergandClaude Opus 5 3730bc7df5 docs(terminal): correct what selectSession does with the viewport
The JSDoc on `_syncStickyScrollBaseline` said `selectSession` deliberately ends
at the bottom, so the baseline the replay samples is already true there. It
does not. `selectSession` calls `scrollToBottom()` after the write and then
ends at `scrollToLastNonEmptyLine()` (app.js:6512), which targets
`lastNonEmptyLine - rows + 2` and therefore parks ABOVE `baseY` whenever the
replayed frame keeps trailing blank rows — which a full capture does on
purpose, since no transform that can delete a line may run over one.

Its baseline really is a stale true. What covers it is the sticky snap itself:
since de864e7d that snap fires only when the flush found the viewport already
at the bottom (`preserveViewportY === null`), which a parked selectSession
viewport is not. That commit landed on master after this branch was cut, so
the guard arrives with the merge rather than being present here.

`_onSessionClearTerminal` is unchanged in the comment and was correct: it
resets and rewrites with no scroll afterwards, so it does end at the bottom.

Comment only; no behaviour change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 18:56:15 +02:00
Michael GrundbergandClaude Opus 5 cfd771d1d8 test(terminal): pin all four buffer-load paths to the shared flush helper
The first version of this fix decided the flush policy in `selectSession`
alone, and a later pass found it still covering one path of four. Nothing in
the CI gate stops a fifth path, or an inlined `{ flushQueued: true }`, from
splitting that policy up again — the browser suite that would notice is
excluded from `npm test`.

A static scan over `selectSession`, `_onSessionNeedsRefresh`,
`_onSessionClearTerminal` and `_maybeRefetchFullHistory` asserts each one asks
`_bufferLoadFinishOpts`, reusing the `methodBody` slice the sticky-scroll guard
already needed. Verified by inlining the policy back into
`_onSessionClearTerminal`, which fails it by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 18:44:04 +02:00
Michael GrundbergandClaude Opus 5 75a028e825 fix(terminal): re-take the sticky-scroll baseline after a replay
A capture load now replays its queued tail, and that replay runs through
`batchTerminalWrite`, which samples `_wasAtBottomBeforeWrite` before it queues.
It runs inside `chunkedTerminalWrite`, before that promise resolves, with the
terminal freshly reset and rewritten — so the sample is always true. The caller
then restored the reader's position and the next `flushPendingWrites` scrolled
straight back to the bottom off the latched flag, undoing it. The only thing in
the way was `_hasRecentUserScrollUp()`, a 1500ms window a server-triggered
refresh is usually past.

`_syncStickyScrollBaseline()` re-takes the flag from wherever the viewport now
sits, and the two paths that restore a position call it right after doing so:
`_onSessionNeedsRefresh` and `_maybeRefetchFullHistory`. Those are the paths
#259 and #205 exist for, and they are also where a non-empty queue is most
likely, since a needsRefresh fires when output is flooding. Re-taking rather
than suppressing the sampling: suppressing leaves whatever stale value the flag
held from before the load, which on the full-history re-pull has no reason to
be false. `selectSession` and `_onSessionClearTerminal` deliberately end at the
bottom, so the sampled true is already the truth there and they do not call it.

`_bufferLoadFinishOpts` gains the coverage the CI gate can see: both mux
sources flush, `history` does not, and a payload naming no source does not.
Its only coverage was the browser suite, which CI does not run.

The JSDoc and the changeset now record the one duplicate window this cutoff
cannot close. The server appends output to the byte buffer in the same tick it
emits, but broadcasts on a batch timer — 8ms over WebSocket, 16 to 50ms over
SSE — so a batch pending when `capture-pane` ran leaves the server after the
reply and is replayed although the capture holds it. It is one batch interval
wide against a recovery window spanning the whole chunked write, and closing it
means flushing that batch server side before the capture.

The second browser test asserts its session was created, so a failed create
fails it instead of passing with zero hits.

docs/architecture-invariants.md no longer claims the replay leaves the
queued-event discard window alone. That clause now describes what decides how a
load ends, the baseline rule, the batch window, and the three covering tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:43:04 +02:00
github-actions[bot]Claude Opus 5github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
20fc7b3c3d chore: version packages (#447)
* chore: version packages

* chore: sync the CLAUDE.md version line to 1.30.0

The changesets bot does not touch this line, and pushing it to master
after merging the version PR starts a second Release run that has raced
the first before. Riding the bot's own branch keeps it to one push.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
2026-09-18 14:09:55 +02:00
Codeman maintainer 0e1191b774 chore(changeset): trim the contributor entries and add the 1.30.0 thanks
Changeset text becomes user-facing CHANGELOG, so the #429 entry is cut
from five bullets of internal bash-array detail down to what the change
does for someone running the installer, as promised on the PR. The #441
entry loses its em-dashes, which are not house style. Adds an entry for
the maintainer fixes applied while landing #442, and the Thanks section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:55:22 +02:00
Codeman maintainer bb8ada7e5f fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does.

1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred
   it from the name: a session the user renamed by hand to something shaped
   like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the
   next prompt overwrote their name. The route persists right after, so the
   loss went to disk. `restoreMuxSessions()` already passes it.

2. The already-live sets were snapshotted once before a loop that awaits a
   real `startInteractive()` per entry, so by the tenth entry the snapshot
   was tens of seconds old and a conversation resumed by hand from the
   Resume list in that window was invisible to it: two panes on one
   transcript, the exact thing the check exists to prevent. Both sets are
   now read per iteration, and the late case is spent rather than re-offered
   for the same reason the batch case is.

3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path.
   The stamp predates the reboot and the pane is new, so honouring it meant
   one click had every restored session type `continue` into itself about a
   minute later, unattended, against the route header's own promise that a
   restored session comes back idle and disarmed. The setting stays ENABLED,
   so it re-arms on the next real limit message. A Codeman restart still
   re-arms from the stamp, because the limit footer will not reprint on its
   own; the new option exists only to tell the two paths apart.

4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()`
   and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession`
   performs that it was missing. Cosmetic, but a run left open reads as
   still going in the away digest.

5. A restored claude session gets `seedAgentSessionPreamble()` like both
   create paths, so the agent skill's bootstrap stays a two-line loader.

6. The heuristic's container comment was wrong in one direction and quiet
   about the real gap: after a genuine host reboot a containerized Codeman
   sees the host's short uptime and the banner does appear. What it cannot
   see is a container-only restart, which is where this would help most.

7. The banner is hidden in a solo window, which shows one session and has
   no tab strip to put restored ones in.

Also reverts 17 of the 18 hunks in docs/api-reference.md, which were
Prettier reformatting of prose the PR does not otherwise touch (docs/ is
outside the format glob), keeping only the Reboot restore section and
repairing the two continuation lines that reformat de-indented; renumbers
reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which
loads after it; and gives the feature its CLAUDE.md entry plus a route
test for the multi-user workspace-forbidden branch, the only new rule that
had nothing behind it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:46:04 +02:00
Codeman maintainer ea5323d990 test(input): pin the batched commit-plus-Enter ordering #441 fixes
The unit harness proves WHICH candidate gets forwarded; the ordering is
the half that shipped the bug, and only a real xterm shows it. The new
browser case dispatches the character's keydown, its composed insertText
and Enter's keydown in ONE page task, the shape an Android soft keyboard
delivers through a single InputConnection transaction, and asserts what
reaches the send path.

Verified in both directions on this machine: with the drain in place the
wire is `o\r`; with the drain removed (master's behaviour) it is `\r` and
the character is gone entirely, because by the time the zero-delay timer
runs xterm has emitted the `\r` and bumped the canonical counter past the
candidate's snapshot, so the candidate stands down. The other four cases
pass in both states.

CLAUDE.md now names the decision point, what it costs (a keydown decides
with less evidence than the timer did) and why that is safe for Enter,
and says that the pin lives in a suite the CI gate does not run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:42:44 +02:00
Codeman maintainer dee674d3e2 fix(install): point the launcher-only caveat at the thing that resolves it
The caveat #429 added ends with "see the docs above", and "the docs
above" is CLI_DOCS[$i], which for DeepSeek is the upstream harness repo.
Per docs/deepseek-integration.md the harness ships only the web,
headless and base profiles, so following that link and running
`npm install -g @deepseek-ai/dsh` leaves the reader exactly where the
caveat is warning them about: a dsh that cannot drive a pane. What
actually resolves it is Codeman's own Run dropdown, which offers
"DeepSeek: add a terminal profile..." and installs one in a click.

The new wording stays generic for any future launcherProfile entry,
since Codeman is the thing being installed at all three call sites.

Also flips one word in the generator: the comment said "see
installCommandFor below" and that function is defined above it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:41:45 +02:00
Ark0N f32c4f60d5 Merge pull request #442 from irisitymichaelgrundberg/feat/restore-sessions-after-reboot
feat(sessions): offer to rebuild the sessions a host reboot destroyed
2026-09-18 13:41:24 +02:00
Ark0N 9a503872d9 Merge pull request #441 from shenlvkang-collab/fix/android-last-char
fix(input): deliver a recovered keystroke before the Enter that submits it
2026-09-18 13:41:21 +02:00
Ark0N 9d7b29d899 Merge pull request #429 from opticon454/chore/cli-catalog-followups
chore(cli-registry): clean up dead code and stale claims left after #380
2026-09-18 13:41:14 +02:00
Ark0N ff8dc92187 Merge pull request #424 from Ark0N/fix/terminal-history-anchor-after-parse
fix(terminal): restore the history anchor after xterm parses, not before
2026-09-18 13:41:08 +02:00
Codeman maintainer 1f61d21298 docs: correct six stale counts and claims in CLAUDE.md
Each of these was measurable and wrong: the CI note listed 5 excluded
Playwright tests where config/test-suites.ts has 9, never mentioned the
packages/xterm-zerolag-input run that follows the gate, and never
mentioned wiki-sync.yml at all; the format glob note omitted that lint
covers only src/**/*.ts; app.js is ~6.9K lines, not ~6.7K, and
voice-pcm-worklet.js is fetched from JS rather than sitting in the load
order; src/config/ holds 23 files plus the cli-registry/ subdir, not 21,
and nothing said that the repo-root config/ is a different directory;
the route count is ~232 with cases at 34, not ~228 with cases at 30.

Also adds the pointer to docs/wiki/ as the user-facing manual, which the
header describes every other doc surface but not that one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:41:02 +02:00
Michael GrundbergandClaude Opus 5 62ceb4e87b fix(sessions): correct what the missing-pid rule actually recognises
A second real reboot disproved the mechanism the previous commit was built
on. Typing `/exit` does not persist `pid: null`, and the session was
restored anyway.

The pid a session record carries is its `tmux attach-session` process, not
the agent. `/exit` ends the CLI inside the pane, `remain-on-exit` keeps the
pane, and the attach process stays alive throughout — so Codeman's PTY never
exits, no exit handler runs, and the record keeps both its pid and
`status: 'idle'`. The lifecycle log for the session that came back shows
created, started, stale_cleaned and recovered, with no exit event at all,
which is the proof: Codeman never learned the agent was gone.

So nothing durable distinguishes an exited agent from a session that was
idle when the power went, and this pass restores both. Ark0N/Codeman#446 is
about making Codeman notice the dead pane; contrary to what the previous
commit's message claimed, this genuinely does wait on that. Until a record
can say the agent is gone, the user dismisses or closes those sessions.

The rule itself is kept, because a record with no attach process does
describe a session that never started or whose pane died outright, and
refusing it is right. Only its documentation was wrong. The module header,
the branch comment and the test names now say what it recognises instead of
claiming the case it cannot see.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 19:49:06 +02:00
Michael GrundbergandClaude Opus 5 5108a24bf0 fix(sessions): never restore a session whose agent was already exited
Found by a real reboot, which is the first thing to catch it. Typing `/exit`
ends the CLI process and leaves the session record behind, and the
process-exit handler persists `pid: null` with `status: 'idle'` before
anything else runs. By status alone that is indistinguishable from a session
sitting idle when the power went, so the boot pass offered those sessions
back and a click spawned the agents the user had deliberately closed — the
exact case the eligibility rule exists to exclude.

The absent pid is what tells the two apart, and the plan step now refuses a
record without one, under its own `not-running` reason so the boot log says
why. On a healthy board every running session carries a pid; a record with
none describes an agent that is already gone.

Deliberately the conservative direction. A session that somehow persisted no
pid while genuinely running is not offered, and its conversation stays
reachable from the Resume list, which is where every session would be
without this feature. The opposite error spawns processes nobody asked for.

Ark0N/Codeman#446 covers the dead panes those exits leave behind, but this
does not wait on it: the rule belongs here whether or not the record's shape
changes later.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 18:46:15 +02:00
Michael GrundbergandClaude Opus 5 5f55f9cb65 fix(sessions): never let the reboot-restore plan fail recovery
The plan build runs inside the try that decides whether restoreMuxSessions()
succeeded, so a throw would be caught there, report restoration as failed,
and block the stale cleanup and layout reconciliation that follow. An
optional convenience would then break the recovery it exists to help. It is
guarded on its own now: the correct way for this to fail is an offer nobody
gets.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 14:11:58 +02:00
Michael GrundbergandClaude Opus 5 18ab2ab595 docs(sessions): correct what a failed rebuild is actually likely to be
Ran the feature against a real server for the first time, on an isolated
instance, and two claims in the code turned out to be wrong.

A rebuild that fails after the session is registered was documented as
commonly caused by a CLI binary missing from a freshly booted machine's
PATH. It is not: the resolver finds its binary by absolute path, so PATH
never enters into it, and a server started without claude on PATH restored
every session normally. Nor does an un-enterable workspace fail — tmux falls
back to another directory and the pane comes up there. Neither obvious cause
throws, so the discard path is defended rather than expected, and the
comments now say that instead of naming a cause that cannot happen.

The four review rounds that shaped this path all reasoned about a trigger
none of them could test. The path itself is still worth having, since a mux
failure would reach it, but its comments should not claim a likelihood the
machine disagrees with.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 08:26:29 +02:00
Michael GrundbergandClaude Opus 5 39976041e0 fix(sessions): let a dismiss reach the entries a restore is holding
Fourth review of the reboot-restore branch, and the third to find a defect
in the previous round's fix. This one is the same shape as its predecessor:
a counter keyed on one thing, compared against a set keyed on another.

The generation counter was indexed by the entry's owner, while the in-flight
set holds the caller doing the restoring. Those are the same person exactly
when a user restores their own sessions, which is every case the tests
covered. The route deliberately supports the other case: an admin may spend
another user's entries. So when an admin restored Bob's sessions and Bob
dismissed the banner, nothing matched, the entries came back, and a plan Bob
had explicitly dismissed was re-armed for another twenty-four hours.

Rather than reconcile the two key spaces, the counter is gone. `take()` now
parks the entries it hands out, remembering which caller is spending them,
and they stay parked until that restore ends. A dismiss filters the parked
entries by `canAccess(entry.owner)` — the same predicate it already applies
to the plan — so it reaches them wherever they are. `releaseFlight()` puts
back only what is still parked. Expiry and a fresh boot plan unpark
everything, for the same reason. There is one key space now, the entry's
owner, and the spender is only ever used to tell two concurrent flights
apart. That removes `generations`, `snapshotGenerations()`, `bump()`,
`bumpAll()` and the argument threaded through the route.

The discard grew the teardown it still lacked. A rebuild can fail after
startInteractive() resolved, and a restored workspace still carries
Codeman's hooks, so the CLI can post a hook event within milliseconds; the
transcript watcher that starts from it, the attachment registry, the wait
registry and the approvals inbox all outlive the listeners and would meet
the retry, which reuses the session id by design. Its steps also run in
reverse order now, so no live listener can reach a tracker that has already
stopped, and the mux kill has its own guard, because stop() kills the pane
in its last block after destroying four trackers.

Tests. The run-summary test named an interval and asserted a map entry, so
dropping stop() left it green; it now spies on stop(). Nothing pinned that
before-spawn must precede setupSessionListeners, which reads the flag that
phase restores, so swapping the two lines was silent; the ordering test now
includes the listener setup. The retry assertion was a tautology and now
asserts a different refs object. Both strengthened tests were verified by
reverting their fix. Two new tests cover the admin-restores-another-owner
cases this round was about. The server in the discard test is built once and
stopped, since its constructor registers handlers on module-level watchers,
and the workspace is removed through safeRmHomeTree.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 16:51:00 +02:00
Michael GrundbergandClaude Opus 5 71ed7b127c fix(sessions): make the discard a real inverse of the construction
Third review of the reboot-restore branch. The narrow discard the previous
commit introduced avoided everything cleanupSession() did wrongly, and in
dropping so much of it also dropped four things it had to keep.

The worst broke the retry the whole design rests on. setupSessionListeners()
returns early while sessionListenerRefs still holds the session id, and the
discard never cleared that entry. So the advertised flow — a rebuild fails
because the agent binary is missing, the user fixes their PATH and clicks
again — reused the same id, wired no listeners at all, and produced a tab
that never showed output, never updated its status and never persisted. That
is worse than the leak the discard was added to prevent. Three more
registrations leaked with it: a RunSummaryTracker and its interval, an image
watcher on the workspace, and the Ralph fix-plan watcher. The discard now
undoes each registration setupSessionListeners() makes, in its order, and
the per-session custom-model config directory, which holds the endpoint's
API key literally and which nothing else would ever remove.

The image-watcher flag was restored after the code that reads it, so a
session came back reporting the feature as on with nothing watching. It
moves to the before-spawn phase, and that phase now runs before the
listeners rather than after them.

The generation counter that lets a mid-restore dismiss win was global while
clear() is ownership-scoped, so one user's dismiss discarded another user's
unspent entries, permanently, because nothing rebuilds an in-memory plan. It
is now per owner. Bumping only the owners of entries the dismiss removed was
not enough either: take() has already emptied the plan by then, so a dismiss
landing mid-restore saw nothing of that owner's to remove and invalidated
nothing. The owners that matter are those with a restore in flight, filtered
by what the dismissing user may access, and that is what clear() now bumps.
Plan expiry bumps too, so a restore straddling the 24-hour boundary cannot
hand entries back and give an expired plan another full day.

Tests. discardPartiallyBuiltSession had no test at all: the only
implementation any test ran was the mock's one-line stub, which is why every
defect above was invisible. test/discard-partially-built-session.ts drives
the real WebServer, and the retry assertion fails if the listener refs are
left behind — verified by reverting the fix. The dismiss-race test drove the
registry by hand, so deleting the route's generation argument left it green;
it now goes through the route, and two further tests cover the multi-user
cases.

The mock context has now gone stale twice, because route tests pass it as
`ctx as never` and tsconfig.json includes only src, so nothing ever compares
it to the ports. A type-level guard is therefore inert — I wrote one and
confirmed it never fires. test/mocks/mock-route-context-completeness.ts
compares the mock's keys against WebServer.createRouteContext() at runtime
instead, and names what is missing.

Also: the API reference now says workspace-forbidden is judged against the
owner's grant, the banner's module header no longer claims Restore always
dismisses it, and the detail span gets the same min-width: 0 the phone rule
already needed.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:34:47 +02:00
Michael GrundbergandClaude Opus 5 fa52753e8b fix(sessions): undo a failed rebuild without deleting the user's data
A second review of the previous commit found that its own repair for the
session leak introduced three defects, all from reaching for
cleanupSession() to undo a half-built session. That function is the
user-initiated delete, not an undo.

It banked the session's historical token and cost totals into the lifetime
figures, and a reboot never runs cleanup, so those totals had never been
counted before; every failed rebuild added them again. It saw the pin that
had just been restored and demoted the record to `stopped`, which this pass
reads as the durable marker of a deliberate kill, so a pinned session whose
rebuild failed became permanently unrestorable. And it recursively removed
`.claude-images` from the working directory, which belongs to the workspace
rather than to the session, so a failed rebuild destroyed the pasted images
of any other live session in that repo.

discardPartiallyBuiltSession() now undoes only what the construction did:
the map entry, the tab-layout slot, the listeners and any pane the launch
created before throwing. The persisted record, the lifetime totals, the
Ralph state and the workspace's files are left alone.

Re-applying the persisted state also splits in two, which removes the first
two defects at the root rather than only at the call site. The half that
shapes the pane, the custom-model environment and the nice priority, still
runs before the spawn. The half that is the session's own history now runs
after it, so a session whose pane never started carries no totals and no pin
for anything downstream to misread.

The rest of that review. The multi-user workspace confinement re-check read
the requesting user's grant, and returns true for an admin, so the case its
own comment described was the one it missed; it now resolves the entry
owner's grant through isWorkingDirAllowedForUsername, the way cron does. A
forbidden workspace goes back on offer, matching both the registry's stated
contract and the API reference. The client re-reads the plan after a restore
instead of blanking the banner, so entries the server put back stay
reachable, and a 409 now says a restore is already running rather than
reporting a failure. A dismiss arriving mid-restore wins, through a
generation counter the route carries across its take. The re-application
also restores the tab colour, the image-watcher flag and the original
pinnedAt, via a new Session.restorePin that does not re-stamp the pin time.
The phone breakpoint gains min-width: 0, without which a nowrap flex item
never shrinks and the buttons still overflow, and it folds into the existing
phone block.

Ralph's loop configuration still does not survive a restore, because
toState() reads it off a live tracker and there is no way to keep it without
arming the loop. The method now says so rather than leaving it implied.

Tests. The capacity test could not fail on the property it existed for: it
filled the board past the cap before the loop, so a single pre-loop check
would have passed it. It now leaves one seat, so only a per-iteration check
restores exactly one entry. New tests cover the ordering around the spawn,
a throw before the loop returning the whole plan and releasing the flight,
the dismiss-during-restore race, and that the failure path calls the narrow
discard rather than the delete. The shared mock context gains the port
method it was missing, which is what made the first run of these tests fail
for the wrong reason.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:04:35 +02:00
Michael GrundbergandClaude Opus 5 fbede5cd2a fix(sessions): act on the dual review of the reboot-restore route
Fifteen findings from two independent reviews of #442, three of them
blocking. Every one is addressed here.

The three blockers all sat in the restore route. A rebuild that threw after
addSession left a registered session with no pane behind it, visible on the
board, holding a layout slot and written to state.json, with its plan entry
already spent; the catch now cleans the session up and puts the entry back.
The loop checked neither the global nor the per-user session cap, so one
click could take a board past a documented limit; capacity is now re-checked
per iteration, because the loop is itself creating the sessions it counts.
Worst of the three, a rebuilt session carried none of the state its
constructor has no parameter for and then persisted itself over the record
that held it, zeroing token and cost totals and dropping the pin. The pin
matters most: pruning keeps a record only while it is pinned, so discarding
it handed the record to the next stale sweep. A new
reapplyPersistedSessionState() on the session port restores the pin, the
token totals, auto-compact, auto-clear, auto-resume, nice priority, the
flicker filter and the custom-model selection, and it runs before both
startInteractive and the first persist.

The rest, in the order they bite a user. Every rebuild failure was reported
as workspace-missing, so the banner told users their repo was gone when the
agent had simply failed to start; there are now distinct reasons, and the
toast names each one. The client read restored and skipped off the outer
response object rather than through the uniform envelope, so every count
came back zero and neither toast ever fired. A board left open across the
reboot never learned an offer existed, because the banner was seeded only on
the page-load path; it now re-reads on every SSE init. The workspace check
was existence-only, skipping the multi-user confinement that the create
route applies, so a withdrawn grant would not be noticed. The banner had no
phone breakpoint while its text was nowrap and its buttons could not shrink.

Smaller: a missing workspace is now re-offered rather than dropped, while an
already-open conversation is dropped rather than re-offered forever; a throw
anywhere in the route returns the unspent entries instead of discarding the
plan; the single flight is keyed by owner, since take() already stops two
callers receiving one entry; the env clamp's header no longer claims a
protection it cannot provide on this path today, and names the check that
does bite; the three endpoints are documented in docs/api-reference.md; and
the module header now says that os.uptime() reads the host's clock, so the
feature is effectively off inside a container.

The review also explained why the tests missed all of this: they proved the
construction claim through their own copy of the construction rather than
through the route, and the route tests used workspaces that did not exist,
so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts
mocks the Session module to drive the route's real path, and covers the
cleanup, the reason reported, the re-application ordering, the broadcast and
the caps. The mock route context gains the port method and the mux call the
route needs.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 11:44:37 +02:00
Michael GrundbergandClaude Opus 5 da933d70be feat(sessions): offer to rebuild the sessions a host reboot destroyed
A host reboot takes the tmux server down with it, so every pane dies,
reconciliation finds nothing to attach to, and the board comes up empty.
Picking yesterday's work back up meant finding each conversation in history
and resuming it by hand, one at a time.

The boot pass now works out what the reboot killed and leaves it on offer.
It runs inside restoreMuxSessions(), in the window where reconciliation has
reported the dead sessions and cleanupStaleSessions() has not pruned their
records yet, which is the only place the records can still be read. The
board shows a banner, and nothing is created until the user clicks it.

A click rather than an automatic restore is what makes the reboot heuristic
acceptable. The heuristic cannot tell a reboot from a crash that took tmux
down inside the same window, so it decides whether to ASK, never whether to
act: a wrong yes costs a line of text the user dismisses instead of N CLI
processes nobody asked for.

Four things are re-checked when the click arrives rather than trusted from
boot, because hours can pass and the board moves on. The owner's privilege
grant re-resolves through the env clamp. The workspace must still be on
disk. A conversation the user already resumed by hand from the Resume list
is skipped, since two panes running --resume on one conversation would
fight over the same transcript. Entries leave the plan synchronously before
the first await, and the route is single-flighted, so a double-click or two
devices cannot both reach the same entry.

A restored session comes back attached, idle and disarmed. Respawn
controllers and Ralph loops are deliberately not re-armed: a machine that
just came up is the worst moment to turn an autonomous run loose. Its
workspace hooks are installed by the restore route itself, because the
boot-time sweep sits behind a gate that is false after a reboot and has
finished long before the click; without them a session goes silently blind,
with no stop or idle events for respawn, no Approvals Inbox item and no red
tab on a blocking dialog. Stats collection starts the same way.

The pane is new, so the conversation continues and the terminal scrollback
does not. The banner says so rather than letting an empty pane read as a
broken restore.

The plan lives in memory only. A server restart drops it, which costs the
convenience this adds and never the conversation: the conversation is the
transcript under ~/.claude/projects, which the Welcome screen's Resume list
and the Session Manager already read, so a dropped plan returns the user to
resuming by hand.

clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the
question it answers is about session privilege rather than about HTTP and
it now has a caller outside the route layer. Its test hook stays re-exported
from session-routes.ts.

Claude sessions only for this pass. The other CLIs name their thread in
their own config object, which this does not thread through yet. Remote and
docker sessions are skipped on purpose, because both need another host or a
container to be up and a freshly booted machine cannot promise either.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 08:05:55 +02:00
codeman-localandClaude Opus 5 a1c35da0d8 fix(input): deliver a recovered keystroke before the Enter that submits it
Every message typed on an Android phone lost its last character.

An Android soft keyboard commits the last typed character and sends the
Enter key in ONE InputConnection transaction, so the committed-text
`input` event and the Enter keydown are both processed before any
zero-delay timer runs. The orphaned-input recovery from #388 resolved
its candidate only on such a timer, and that lost the character twice
over:

  * ORDER — xterm emits `\r` synchronously from the Enter keydown, and
    the local-echo composer submits `pendingText` right there. The
    recovered character arrived one macrotask too late to be part of the
    prompt.
  * LOSS — that same `\r` bumps the canonical counter, so by the time
    the candidate resolved, `canonicalCount > snapshot` read as "xterm
    spoke for this keystroke" and stood the recovery down. The character
    was not merely late, it was dropped.

Drain pending candidates synchronously at the next keydown instead, from
xterm's custom key handler, which runs before xterm processes that key.
The counter then still holds the value it had while the candidate's own
keystroke was current, so the stand-down decision is made against the
right keystroke, and the recovered byte reaches the composer ahead of
whatever the new key emits. The timer stays as the fallback for a
keystroke with no key after it.

Physical keyboards are unaffected: there the timer has already resolved
the candidate long before the next key arrives.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 09:26:30 +08:00
Codeman maintainer bd286bf502 docs(wiki): catch the manual up to 1.29.0 and add the three run modes it never had
The wiki was written for seven run modes and never received Grok Build, DeepSeek
Harness or OMP. They now appear everywhere the others do: the modes table and
per-CLI notes, install commands, environment prefixes, the Quick Start table, the
requirements rows, the vocabulary, and every "seven modes" count.

The 1.27 to 1.29.0 changes land on the pages that own them: attaching a case to an
existing container, multi-case adoption and the copy-a-case picker (Docker Cases);
file reads over ssh in remote cases and what stays unavailable (Remote SSH Sessions,
Working With Files, Security); single-page app routing, frame recovery, localhost
links as tabs and the egress guard (Web Tabs); DeepSeek as the one non-Claude mode
with real stop/blocked signals and Approvals items, Codex's own work detection,
last-response, the model-endpoint routes and refreshed counts (HTTP API, Driving
From An Agent, Hooks, Notifications, Keeping Agents Running, Core Concepts);
Shift+drag, right-click copy, Auto Copy, the Ctrl+Z guard, font weight, the vertical
rail and its activity sort (Keyboard Shortcuts, Input And Voice, The Dashboard,
Settings Reference); the 600px phone cutoff, Codex shift arrows and iPhone Duo
(Mobile Guide); the Docker Compose route and its update rule (Installation, Running
As A Service); four new symptom entries and a "which CLIs" question (Troubleshooting,
FAQ).

Custom model endpoints are deliberately left to #430, which adds that page and edits
Agent CLIs, Settings Reference and the sidebar; these edits stay out of the regions
#430, #428 and #376 touch, and all three still merge cleanly on top.

Both READMEs: the web-tab menu entry is labelled "Add URL" in the UI, not
"Add dashboard".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 19:05:59 +02:00
github-actions[bot]Claude Fable 5.1github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
3248f35081 chore: version packages (#437)
* chore: version packages

* chore: sync the CLAUDE.md version line to 1.29.1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
2026-09-15 18:18:09 +02:00
Michael GrundbergandClaude Opus 5 c9515b1d4c fix(terminal): keep the output a pane capture could not contain
Live terminal events are queued while a buffer load runs, and the load discards
that queue when it ends. That is right when the loaded buffer is the server's
accumulated byte history. The route appends to that history right up to the
moment it serializes the response, so a queued event already appears in it and
replaying it would duplicate output, most visibly Ink's cursor-up redraws.

A tmux pane capture is a photograph, current only as of the instant
`capture-pane` ran. Output printed afterwards was queued and then dropped, and
nothing scheduled a re-fetch to recover it: `_onSessionNeedsRefresh` is wired
only to the 128KB overflow path. The CLI's next partial redraw then landed on a
frame the terminal never received.

How much went missing depended on which capture the route served. A `?full=1`
load returns the capture alone, with no history in front of it, so it lost
everything from the capture to the end of the chunked write. A `?tail=` load
returns history, a clear, and then the capture, and the route reads that history
after the capture, so it lost everything from the response to the end of that
write. The chunked write dominates either way. An agent CLI hides the loss on
its next full redraw; a shell session does not, because its output is linear and
nothing repaints it.

Queue entries now carry their arrival time, and `_finishBufferLoad` takes a
`since` cutoff, so a capture load replays exactly the tail that arrived after
the response headers. The earlier events stay dropped, because a payload that
carries history does hold those.

All four paths that fetch a terminal buffer and write it now decide this the
same way, through one `_bufferLoadFinishOpts` helper, so they cannot drift
apart: `selectSession`, `_onSessionNeedsRefresh`, `_onSessionClearTerminal` and
`_maybeRefetchFullHistory`. The second of those is the one that stings. It
exists to restore output the client already dropped once under backpressure, and
it was dropping more output while performing that recovery. The cache-hit write
inside `selectSession` stays on discard deliberately: it runs before the fetch,
so its queue holds only events the capture that follows already contains.

Two further things had to change for that tail to still exist when the load
ends, and a browser test is what found both. `chunkedTerminalWrite` is what ends
the load for every non-empty buffer, so the flush policy travels to its own
finish calls; the call in `selectSession` runs only when the write was skipped.
`_beginBufferLoad` no longer empties the queue when one load re-enters it, which
it does on every write, because that reset discarded the whole fetch window
before anything could replay it.

The response already distinguishes the sources. `source` reads `mux-visible` or
`mux-full-history` for a capture and `history` for the byte stream.

Follows #395, #396 and #397, which fixed the ways the replayed frame itself
could disagree with the terminal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-15 18:14:18 +02:00
Codeman maintainer 5b920cb43d feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream
and named tabs wrong five ways (every prompt, every write path, a bare Esc
eating the next prompt's first character, pasted newlines as Enter, any CSI
clearing the draft) and replaced the whole name, which dropped the case from
the tab and reset the w<n> counter. This lands the feature with each of those
closed:

- First prompt means the first: applyAutoName() flips a placeholder to
  `auto` whether or not the string changed. nameSource is now the tri-state
  placeholder | auto | manual; the name setter is the only manual path.
- Only user-originated input counts: write()/writeViaMux() take
  SessionWriteOptions.fromUser, set by the browser WS path and POST /input
  only, so Ralph, respawn, cron, approvals and the trust-dialog keys can
  never name a tab. A startMode 'shell' CLI never feeds the tracker (a
  capability, not an id check); the send-key route feeds trackUserInput()
  because its line feed bypasses the session.
- Prefix form `w3-case: title`: parseSessionPrefix() already renders it as
  the title with the prefix in the tooltip and the next-session counter
  still matches it. Composed within MAX_SESSION_NAME_LENGTH.
- Tracker rules per key: bare Esc resolves at chunk end; mouse/focus
  reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R
  taint the draft so Enter submits nothing rather than a fragment;
  bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space;
  the draft keeps its head past 8192 code points; an escape past 64 bytes
  is abandoned.
- Title: slash commands by shape (a path is a prompt), `!` escapes
  refused, first sentence only past 8 code points ("e.g." is not a title),
  72 code points on a word boundary.
- Synced `autoNameSessions` setting, default OFF (the prompt reaches
  mux-sessions.json, session:updated and /api/search), App Settings ->
  Appearance -> Tabs, read fresh per prompt after the eligibility check.

Tests: test/session-auto-name.test.ts (tracker, title, composition,
ownership, emit gating), the wiring test (once, prefix, setting off,
manual protected), test/routes/session-name-routes.test.ts (PUT /name
flips to manual and persists). Verified live on an isolated instance: API
and browser-typed prompts name the tab, a second prompt does not, shells
and renamed tabs are untouched, nameSource survives a restart.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 17:59:16 +02:00
Codeman maintainer c4322513d9 Merge pull request #376 from shenlvkang-collab/feat/auto-session-names-upstream 2026-09-15 17:19:57 +02:00
DevvynandClaude Sonnet 5 3f2928ae73 chore(cli-registry): clean up dead code and stale claims left after #380
Addresses the "left as they are"/"worth knowing" items Ark0N named when
merging #380 (the CLI-catalogue-driven install.sh + Docker agent image
PR), none of which were correctness-blocking but all of which were real:

- Removed install.sh's dead _cli_index/check_cli/get_cli_path helpers:
  the catalogue-driven menu and hints stopped calling them and nothing
  else ever did.
- The generator no longer emits CLI_KIND/CLI_NPM, two bash arrays
  install.sh never read (the .mjs/docker-hosts.ts producers already
  read the JSON catalogue's kind/npmPackage fields directly, so only
  the bash copies were dead).
- detect_all_clis now skips a disabled entry's probe entirely instead
  of running it and filtering the result downstream. No stock entry
  ships disabled today, so this closes a latent inefficiency before it
  is a latent bug rather than fixing an observed one.
- The install hint for a launcherProfile entry (DeepSeek today) now
  explains in one line why it's a docs link and not a command: its own
  docs page documents `npm install -g @deepseek-ai/dsh`, which installs
  the launcher only and can't drive a pane, the exact trap the menu
  already avoids by withholding the command. Driven by a new generated
  CLI_LAUNCHER_ONLY array (from discovery.launcherProfile), not an id
  check, so any future launcherProfile entry gets the same caveat free.
- Corrected the non-interactive-default comment: on a wget-only host,
  Claude's curl one-liner is filtered out of the offered list first, so
  the default becomes whichever npm-based entry sorts earliest instead
  (Codex today), not always Claude. Behaviour is unchanged — it was
  already printed, never silent — only the comment overclaimed.

Tests: extended test/install-sh-invariants.test.ts with a positive
guard for the new array and the trimmed array list, a negative guard
that CLI_KIND/CLI_NPM/the three dead helpers cannot come back, and two
real-bash tests (driven the same way the existing skip-menu tests are)
proving a disabled entry is genuinely never probed rather than merely
filtered after the fact.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011WzDjJnbK7zug8iQWnCc9z
2026-09-15 09:10:37 +08:00
Codeman maintainer 018f0c4160 docs(readme): catch both READMEs up to 1.29.0 and repair three merge-damaged lines
DeepSeek Harness joins every CLI list it was missing from (tagline, intro,
run-mode table, Multi-CLI bullet with its env prefixes, security allowlist,
architecture diagram), and the 1.27 to 1.29.0 features get their bullets:
custom model endpoints (HTTP API only, with the verified and gapped CLIs
named), web tabs, attaching a case to an existing container, remote SSH file
access, the plan-usage chip, the sidebar and activity-sorted rail, font
weight, skins and entrance animations, Approvals Inbox, Read My Mind,
Claude-login voice dictation, Shift+drag select and right-click copy. The
agent guide's rule 7 now counts deepseek among the hook-signalling modes and
the recipes read answers through last-response first; the API section carries
the new routes and current counts; the download cap reads 2 GB instead of the
retired 50 MB; the zerolag package test count is the measured 238.

The English file had three spots where the OMP merge of 2026-08-18 left two
copies of a line joined without a newline (the Docker credentials bullet, rule
7 of the agent guide, the CLI node of the mermaid diagram). All three are
single lines again.

The Chinese file was further behind: besides the above it had never received
the daemon and service block, the Tailscale install option, the Compose
paragraph, the Tab Alerts section, the codeman tui section, the agent-skill
walkthrough, the Community section or the closing star paragraph. Those are
translated in, so both files now share one section structure.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 00:56:15 +02:00
Codeman maintainer de864e7d63 fix(terminal): restore the history anchor after xterm parses, not before
flushPendingWrites() captured the viewport of a user who was reading
scrollback, called terminal.write(), and restored the anchor on the next
line. xterm parses on its own schedule, so at that point the buffer has not
moved: the guard `viewportY !== preserveViewportY` was false, scrollToLine
was never called at all, and the Codex redraw landed a tick later and took
the viewport to the live bottom with nothing left to pull it back. Scrolling
up during a stream still got dragged down, which is what #358 reports, and a
refresh was the only way back to a coherent view.

The restore moves inside xterm's write callback, the first moment the
redraw's effect exists, and runs before _scheduleTerminalWriteFlush() so a
deferred remainder re-captures the restored anchor rather than the bottom.

Two things follow from it running later:

- A live anchor now wins over the sticky scroll-to-bottom. The two are
  captured at different moments (_wasAtBottomBeforeWrite at the frame's
  first batchTerminalWrite, the anchor at flush time), so a scroll-up in
  between leaves both set, and running both would jump to the bottom and
  come back a frame later instead of staying put.
- The anchor is dropped if the active session changed or a buffer load
  started while the write was in flight. It indexes the buffer it was
  captured from, and selectSession() resets the terminal and chunk-loads a
  different scrollback.

The existing regression passed throughout, because its write mock moved the
viewport synchronously, which real xterm never does. The harness now models
an asynchronous parse (redraw lands, then the callback fires), and all five
of the anchor tests fail against the old code.

Fixes #358

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-15 00:10:41 +02:00
Codeman maintainer 88e3faa456 chore: version packages 2026-09-15 00:07:19 +02:00
Codeman maintainer 70fc6b32d5 docs: record the dup/last input ACK, Shift+drag and right-click copy, and multi-case adopted containers
Three behaviours landed from #375 without their doc entries: the
duplicate input ACK now carries `dup:true` and the server's watermark
(`docs/reliable-input-delivery.md` still described a bare ACK), Shift+drag
and right-click copy in the terminal (the shortcut list did not know
them), and one adopted container backing several cases at different
in-container directories (the Docker cases paragraph still implied one
case per container for adopted containers too).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:56:19 +02:00
Codeman maintainer 9591b973cf fix(docker): carry the owned flag on the wire the way master already does
The cherry-picked "copy an existing case" commit declared a second
`CaseInfo.docker.owned` and emitted `owned: true|false` on every docker
case, while master had meanwhile shipped the same field from the
adopted-container work with a narrower wire shape: `owned` is present
only when false, absent means owned. Two declarations failed typecheck,
and two emit styles on one response would have made the picker's answer
depend on which read path filled it.

Keep master's shape at both response sites (the case list and the
single-case lookup, which lacked the field entirely), fold the picker's
reason for the field into the existing doc comment, and repoint the test
that pinned "set on exactly two sites" at the surviving form, adding a
negative pin so the duplicate style cannot come back.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:56:19 +02:00
d fei 025f061383 fix(docker): pre-fill the copied case instead of blanking two fields
The previous version cleared the case name and the in-container directory on the
grounds that they must differ. That left a form with three fields mysteriously
filled and two empty, and turned the most common operation — changing
/srv/app/api to /srv/app/web — into retyping a long path.

Both are now pre-filled, with focus on the in-container directory and the caret
at the end, since the tail is what changes. What stops an unmodified submit is no
longer an empty field but a guard: the values applied are recorded, compared at
submit time, and if nothing changed the reason is stated next to the field and
focus moves to it, without sending a request that is certain to be refused.

The server refuses these anyway (a duplicate case name, a twin case on the same
container and directory) and its errors are clear; but making a round trip to be
told "you forgot to edit the field you are looking at" is worse than saying so on
the spot. The guard only applies when a source case was actually selected, so
filling the adopt form from scratch is unaffected.

⚠️ The status text is written into dockerLinkStatus. My first version referenced
an id that does not exist (dockerAdoptStatus), which made the explanation vanish
silently and left only a toast. The test now extracts that id from the code and
looks it up in index.html, pinning that it must really exist.

(cherry picked from commit ba21ae11f4)
2026-09-14 23:56:19 +02:00
d fei 7a5543da09 feat(docker): add "copy an existing case" to the adopt panel
The backend already lets one adopted container back several cases pointing at
different in-container directories, but using it meant retyping the container
name, host and workspace one by one — exactly the friction that leaves a
capability unused. Picking an existing case from a dropdown now carries those
three over, leaving only the two fields that must differ: the case name and the
in-container directory.

Clearing those two is the point of the feature, not a convenience: keeping the
old name is refused by the server as "case already exists", and keeping the old
directory is refused as "a twin case on the same container and directory". Both
errors are clear, but a form pre-filled with values that are guaranteed to be
rejected is a trap. Focus lands on the in-container directory — the thing the
user came here to change.

⚠️ Only adopted containers are listed (docker.owned === false). A Codeman-built
container's lifecycle belongs to its one case — a second case would be torn out
by that case's recreate or delete — so the server refuses it anyway, and listing
it here would only manufacture a baffling error. `owned` may be absent and absent
means owned, so the test is `!== false`, not truthiness.

CaseInfo.docker gains containerWorkdir and owned for this: the former is the
"which directory does this case use" half of the picker, without which the user
cannot tell what to change it to; the latter backs the filter above. ⚠️ Both
places that build a docker CaseInfo (the list endpoint and the single-case query)
must set them — filling in only one makes the picker work or not depending on
which read path was taken, and a test pins "exactly two".

(cherry picked from commit f1ed3a58e1)
2026-09-14 23:56:19 +02:00
d fei cbb7f635ff feat(docker): let one adopted container back several cases in different dirs
Once a container is adopted, it could not be adopted a second time. But a
container usually holds more than one project directory, and opening a case for
another one had no path forward except starting a second container — precisely
what adoption exists to avoid.

The original reason was in a comment: two cases sharing an adopted container
would make one case's teardown race the other's launch on the same tmux server.
That reason does not hold. The in-container tmux session name is
dockerTmuxSessionName(sessionId), i.e. codeman-dkr-<id8>, keyed by SESSION and
not by case, and buildDockerKillCommand tears down exactly that name, so killing
A never touches B — hosting multiple sessions is what a tmux server is for.

The other three routes into an adopted container's lifecycle do not pass through
here either, confirmed one by one: the stop and remove builders throw outright;
recreate refuses `owned === false` before it even resolves the container name;
and orphan reaping filters on `label=codeman.managed=1`, which a user-built
container does not carry — a structural exclusion.

That leaves exactly three cases worth refusing, none of them tmux-related, split
into the pure, unit-tested classifyAdoptContainerConflict:
- owned-case   the container belongs to a Codeman-created case, whose lifecycle
               Codeman manages: one recreate or delete there would pull the
               container out from under the adopting case.
               ⚠️ `owned` may be absent and absent means owned (cases predate
               the field), so the test is `!== false`, not truthiness.
- other-owner  already adopted by a different user. Adoption hands out a shell
               inside someone else's container.
- duplicate    same container, same directory. The second case would behave
               identically to the first, so name the existing one rather than
               silently minting a twin. A different in-container directory is
               the case this change exists to support and passes.

(cherry picked from commit 1cb6bde891)
2026-09-14 23:56:19 +02:00
d fei e5684d0bba fix(ui): don't create a compositing layer for a hidden full-screen overlay
`backdrop-filter` promotes an element to its own compositing layer. A
position:fixed full-screen layer that is created and then hidden was measured to
leave a stale hit-test region behind in Chrome: the page renders perfectly, but
pointer events across the viewport go nowhere.

The report came from a long-lived tab connected to a remote server, where a
connection blip shows and then hides #offlineOverlay. The symptoms were a
terminal that would not scroll and, at the same time, an unrelated
click-to-expand that also stopped responding, while a freshly opened tab was
fine; a read-only console command (getComputedStyle + elementFromPoint, both of
which force a hit-test recomputation) then cured it. Two unrelated features
dying together and one read-only command fixing both points at hit-testing
itself rather than at either feature.

So the `backdrop-filter` moves onto the actually-visible selector and the layer
is never created while hidden. Only the two persistent overlays change:
offline-overlay (toggled with [hidden]) and file-preview-overlay (toggled with
.visible). path-picker and path-preview are created and removed by JS, leave
nothing behind, and are untouched.

⚠️ This is an evidence-based inference, not a fix verified by reproduction:
reproducing it needs a long-lived page that has been through a connection blip,
which I could not manufacture in a controlled environment. The guard test pins
both halves — no such property while hidden, and a real blur while shown — so a
later cleanup cannot quietly delete the effect.

(cherry picked from commit 08442dfee1)
2026-09-14 23:56:19 +02:00
d fei c7cc8e28d5 fix(sse): stop reloading the whole terminal when a reconnect lands on the same session
handleInit() did not distinguish a first load from an SSE reconnect: it always
cleared the terminal caches in _resetAllAppState() and re-ran selectSession() for
the session that was already on screen. Every reconnect therefore refetched up to
1 MiB of buffer and reset+rewrote xterm. On a link that drops a connection about
once a minute (measured at ~57s intervals against a healthy server) that reads as
the page refreshing itself and throwing away your reading position.

A reconnect that lands back on the still-open session now keeps the terminal
caches and activeSessionId and resyncs through _onSessionNeedsRefresh(). That
path still reloads the buffer, so output produced during the outage is not lost,
but it preserves distance-from-bottom — the same rule #259 established for a
refresh the server triggered rather than the user. The WS is reconnected
explicitly when it is not already on that session, since skipping selectSession()
skips its _connectWs() call.

First load (gen === 1) takes exactly the path it took before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rv24Pk4qzrsDYdVyDyJQmT
(cherry picked from commit 435569c76e)
2026-09-14 23:56:19 +02:00
d fei 01da577053 fix(input): recover when the seq counter falls behind the server watermark
Browser input is delivered exactly once by (clientId, seq). The server records a
watermark per clientId and discards anything not above it as a duplicate — but
acknowledged it with an ACK indistinguishable from "applied". The client then
dropped the record from its queue, the UI looked perfectly normal, and the
terminal received nothing at all.

The counter is persisted to localStorage through a debounced write. Kill the page
between "sent" and "persisted" and the restored counter is below the server's
watermark, after which every keystroke lands under it, is discarded, and is
ACKed. Reloading does not help: the clientId is restored from localStorage
alongside that stale counter. Measured on a real session — typing into the same
session from a fresh browser (new clientId, no watermark on the server) worked
perfectly, which is what localised the fault to client state.

Three changes:
- on rejection the server replies {"t":"ia",seq,"dup":true,"last":<watermark>}.
  It still ACKs, so the client can drop the record from its queue, but it now
  says the input was not applied and supplies the number needed to climb out.
- on `dup` the client lifts its counter above the watermark and re-queues.
  ⚠️ Only records whose FIRST delivery is being retried are re-sent: a retry
  judged duplicate means the mechanism is working (the original did arrive), and
  re-sending would type the same text twice.
- the counter is now persisted synchronously. The queue payload can stay
  debounced, but the counter is the thing that has to survive a crash, and
  leaving it on the lossiest path cancels the only guarantee there is.

⚠️ Reading the watermark is defensive: the session arrives through a structured
port, and a port missing that method must not take the whole input path down —
a throw inside the handler means the ACK is never sent and the record is stuck in
the client queue forever, which is worse than the ambiguity being fixed. A mock
port's test timeout is what exposed this.

(cherry picked from commit 05bb7081cc)
2026-09-14 23:56:19 +02:00
d fei 631386d3f7 fix(cjk): forward Ctrl/Alt-modified navigation keys to the CLI
claude advertises "Jump to bottom (ctrl+End)", so that chord has to actually
reach it. But PASSTHROUGH_KEYS carried only the bare forms (End -> \x1b[F) and
CTRL_KEYS held just six letters (c/d/l/z/a/e), which cannot express End. Ctrl+End
therefore failed in both directions:

- with an empty composer it went out as a bare \x1b[F, the modifier silently
  dropped, so the CLI received a plain End;
- with text in the composer the forwarding branch requires empty, so nothing was
  forwarded and the browser default applied — the caret jumped to the end of the
  draft, which is the "the shortcut now edits my input box" the user saw.

Encode them as CSI 1;<mod><final> instead, and forward Ctrl/Alt-modified
navigation keys whether or not the composer is empty: they are commands for the
CLI, and the composer has no editing semantics for them worth preserving (bare
Home/End still use the old table and edit locally).

⚠️ Bare Shift is deliberately excluded: Shift+arrow selects text in the composer,
a real editing gesture that must stay local. Shift held together with Ctrl/Alt is
still encoded into the modifier mask.

(cherry picked from commit 3fbaadadfb)
2026-09-14 23:56:19 +02:00
d fei b3a6ba2eb6 feat(terminal): make Shift+drag select, and right-click copy the selection
In a native terminal running a TUI with mouse tracking on (claude, codex), Shift
is the "let me select text" modifier: it bypasses the application's mouse
reporting so the emulator selects locally. Users bring that habit here, where it
did nothing — measured, `hasSelection` was already false during a Shift+drag and
no clearSelection call ran at all, because there was never a selection to clear.

The mismatch is that the two Shifts mean different things. xterm reads Shift as
"force selection", but that path is only taken when the application really has
mouse tracking on. The server strips the mouse DECSETs for claude/codex/gemini
(isAltScreenStripMode), so xterm's mouseTrackingMode is permanently `none`, that
branch is unreachable, and Shift instead lands in _onIncrementalClick — which
EXTENDS an existing selection. Extension is a no-op while selectionStart is
empty, so the drag had no anchor.

So plant the anchor xterm is missing. The listener sits on the capture phase of
the `.xterm` root, an ancestor of the `.xterm-screen` that SelectionService binds
to, and therefore runs before xterm's own mousedown; xterm then extends from our
anchor and the drag behaves like any other. Length is 0 so a Shift+click without
a drag does not select a stray character. An existing selection is left alone —
that is a genuine extend gesture, and xterm handles it correctly.

Right-click copies the selection (the mintty/PuTTY convention), completing the
gesture: until now there was nowhere for a finished selection to go. With no
selection the native menu is not hijacked — taking it away while offering
nothing in return is a pure loss.

(cherry picked from commit 7ab5015737)
2026-09-14 23:56:19 +02:00
Codeman maintainer 897a63183f chore(typecheck): include the local-LLM harness smoke script
scripts/test-local-llm-harnesses.ts (#393) sits outside tsconfig.json's
include, so nothing type-checked it. config/tsconfig.scripts.json pulls
it in; npm run typecheck now runs both projects, the way the pr-bot
config used to be chained before the bot moved out of the repo.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:47:32 +02:00
Codeman maintainer 942bf37e48 fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a
custom OpenAI-compatible endpoint by injecting env vars or a config file
and restarting the CLI in place. Review of the apply path found four
things, two of them destructive. This lands all four plus the smaller
items from the same review.

1. Clearing a selection did not clear it. The injected vars reach the CLI
   via `tmux setenv`, which persists at the tmux-session level and is
   inherited by `respawn-pane` (measured: `setenv FOO bar` survived two
   successive `respawn-pane -k`), so deleting the keys from the session's
   envOverrides relaunched the CLI still pointed at the old endpoint, and
   for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just
   been deleted. `Session.setCustomModel()` now reports the removed keys,
   queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys`
   carries them into `applyEnvOverrides()`, which `setenv -u`s them before
   re-applying the live overrides, on the same path that already unsets
   the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket
   that `setenv -u HOME` hands the next respawn the global HOME back.

2. Applying a model to a local claude session killed the pane. The
   relaunch was `claude --session-id <id>` and Claude refuses an id that
   already has a transcript, and unlike the dead-pane respawn this one
   kills a working pane first. `restartCli()` now pins the live
   conversation id as the resume id for that respawn when the CLI's launch
   declares a `fallback` chain, which renders the same
   `--resume <id> || --session-id <id>` shape the docker and remote pane
   commands use. Gated on the registry shape, not the CLI id: an entry
   whose resume id is minted by the CLI itself never declares that chain.

3. pi, omp and grok wrote their config file and then launched without the
   `--model` that selects it, so the file was ignored. The registry entry
   now declares `customModelInjection.launchModel` (`custom/{modelId}` for
   pi and omp, grok's `[model.codeman-custom]` block name), the builder
   renders it, and `_withCustomModelLaunchModel()` applies it onto the
   respawn options through `legacyConfigField`, leaving the stored
   <Mode>Config untouched so a clear falls back to the user's own model.
   A model id the CLI's `model` token pattern cannot carry is refused
   with a 400 rather than silently dropped by the argv engine.

4. Remote (SSH) and Docker sessions reported `restarted: true` and changed
   nothing: their `restartCli()` reattaches the durable tmux rather than
   relaunching the agent, and the env lands on the local pane. Both are
   refused with a 400 until those paths are plumbed.

Smaller items from the same review:

- The selection survives a Codeman restart as the disk-only `__customModel`
  bookkeeping (endpoint, model, injected key NAMES, config dir, launch
  model; never the values, which carry the API key). Recovery re-derives
  the values from the endpoint store through the same apply path the route
  uses and keeps the bookkeeping even when the endpoint is gone, so a
  later clear still has keys to unset.
- Discovery goes through `webviewFetch()`, so the RESOLVED address is
  judged by the same egress guard the web-tab proxy uses, and `baseUrl`
  reuses `webviewUrlSchema` (http(s) only, no embedded credentials,
  link-local and cloud-metadata addresses refused). undici's `fetch failed`
  wrapper is unwrapped so the user sees the ECONNREFUSED underneath.
- `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session
  config dir 0700/0600 (pi and omp embed the key literally), and that dir
  is removed with the session.
- `PR.md` is gone from the repo root and the design doc moved to
  `docs/custom-model-endpoints-plan.md` with the LAN address and the
  personal name scrubbed; every reference follows. The guide's `authStyle`
  text matches the shipped schema (`bearer | api-key`, default `bearer`)
  and says that `customModelEndpointsEnabled` is read by nothing until
  the picker lands.
- `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts`
  (four real type errors fixed). It is not yet wired into `npm run typecheck`
  because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json`
  there is the one-line follow-up.

Tests: `test/session-custom-model-restart.test.ts` drives a real Session and
fails on the unfixed code for items 1 to 3; the route suite covers item 4
and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets
run before the overrides and that a shell-metachar key never reaches tmux.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:46:28 +02:00
Codeman maintainer 1e42cb4e2d Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses)
2026-09-14 23:46:27 +02:00
Codeman maintainer e49c48145b fix(files): fail closed on remote symlinks, guard PUT for remote cases, bound ssh fan-out
Follow-up to #421 (remote-case file reads over ssh), addressing the review.

Symlink escape on a host without `readlink -f` (blocker). The probe's
portable fallback canonicalized only the directory chain and returned the
final component unresolved, so on macOS < 12.3 `ws/notes.txt -> ~/.ssh/id_rsa`
came back as `.../ws/notes.txt` (with the target's size), passed every
containment and blocklist check that runs on `realPath`, and `cat` followed
the link. The fallback now walks the directory chain with `cd -P`/`pwd -P`
and follows the LAST component with plain `readlink` for a bounded number of
hops, and anything it cannot fully resolve (a loop, a readlink failure, the
hop cap) is reported with an `x` marker that parses as null, i.e. 404. It
never returns the unresolved string. Measured on a real /bin/sh with
`readlink -f` shadowed: the pre-fix script reports `/ws/notes.txt`, the fixed
one `/secret/id_rsa`; both branches (native and fallback) now agree.

`PUT /api/sessions/:id/file-content` never had the remote guard the PR
described. It sits ahead of `validateSessionFilePath`, which resolves against
the LOCAL filesystem, because with a same-named directory on the Codeman host
(an sshfs mount of the remote tree, the documented stop-gap) the write landed
on the local twin while the viewer believed it edited the remote file.

ssh fan-out is bounded. `src/remote-ssh-limiter.ts` is a
document-conversion-limiter-shaped semaphore (default 4, env
`CODEMAN_MAX_REMOTE_FILE_SSH`) around every probe and buffered read; the
attachment-history list resolves its whole history in ONE batched probe
(`probeRemoteAttachmentHistory`, threaded into
`registerExternalAttachment({remoteProbes})` so the guards run unchanged)
instead of one handshake per entry; and probes chunk at 40 paths because the
whole script is one argv string. Terminal output in a remote session is
written on the remote host, so a prompt-injected agent printing hundreds of
`codeman://attach` links forked one ssh per link, each holding a 20 s
timeout, and a 100-entry history re-listed on every attachment:detected
tripped OpenSSH's default MaxStartups. Streams are deliberately not counted
(one per browser request, held for a whole playback, and gated behind a
counted probe anyway).

Smaller items from the same review: probe records are NUL-terminated and
index-keyed after a leading NUL (a newline in a filename can no longer shift
the alignment, and the banner is fenced off without last-N-lines guessing);
size comes from `stat -c %s || stat -f %z`; the three IO functions refuse
under VITEST instead of opening a connection; an unreachable host now reads
as unknown (missing: false) for detected AND external history entries, where
external used to fold its 502 into missing; a client that aborted during the
guard probe has its body's ssh child reaped (`reply.raw.destroyed` is checked
before the close listener is attached); `describeExecError` never returns
Node's `Command failed: <ssh line>` message, which carried the identity path
and the probe script into a 502 body; and the docs note that
`isSensitivePath`'s three home-anchored entries resolve against the Codeman
host's home, not the remote one.

Tests: the probe script runs on a real /bin/sh with a `readlink` shim that
rejects `-f` (the escape, a relative chain through a symlinked directory, a
loop, a newline filename, banner chatter that itself looks like a record),
the limiter's cap and FIFO order, and route tests for the PUT guard (local
twin untouched, no connection), the single batched history probe, the
unreachable-host alignment and the aborted-client reap. All four route tests
fail against the pre-fix file-routes.ts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:42:06 +02:00
Codeman maintainer 792a251e35 Merge pull request #421 from Randalix/fix/remote-file-access
fix(files): read remote-case previews, downloads and attachments over ssh
2026-09-14 23:42:06 +02:00
Codeman maintainer 6dc27ae727 docs(webview): record the lost-frame page as the third unauthenticated 200, and the inline-style limit
The lost-frame recovery page is answered ahead of the credential checks in both
auth hooks, which makes it the third unauthenticated 200 beside the two hook
routes, and the only one decided by request headers alone. CLAUDE.md's security
table listed exactly two, and docs/web-tabs.md is not where anyone auditing that
looks, so it now has a row in the table and a fourth property in
docs/security-architecture.md section 10b, including the `/` carve-out and its
credential-free condition. Both state the property that comes with it: a
non-browser client can set those headers, so an unauthenticated caller can tell a
registered route (401) from a non-route (200) and enumerate the route table,
accepted because the routes are public in docs/api-reference.md.

docs/web-tabs.md gains the landing-page case in layer 6 and a Known limits entry:
masking trades away the Referer safety net, only HTML is rewritten server-side,
and a root-absolute url() inside an inline <style> block has the masked document
as its Referer, so it 404s where the Referer fallback used to rescue it. External
stylesheets are unaffected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:38:54 +02:00
Codeman maintainer 1306f731cf fix(webview): recover a proxied dashboard that reloads on its landing page
The runtime shim masks `/webview/<cap>/` off a proxied page's URL so its router
boots on the path it expects, and the landing page masks to exactly `/`. A
`location.reload()` there (a Vite dev server on a config change or a failed HMR
update, the likeliest case in the feature's own motivating scenario) therefore
asks for Codeman's root as an iframe navigation. `serveLostWebviewFrame()`
returned early for `/`, so on a passwordless install the frame received Codeman's
own app shell and rendered it inside the web tab, and with a password it got a
401 in the frame. Either way no `codeman:webview-lost` message was posted, and
because the document loaded fine the load handler cleared the failed-frame panel,
so the Reload / Open in new tab affordances never appeared. Before masking the
frame's URL was the prefixed one, so a reload worked; this was a regression.

`/` is the one lost-frame path a registered route also serves, so the route
table cannot tell that reload from a real navigation. Credentials can: nothing
in Codeman frames its own root, and a sandboxed frame is opaque-origin with no
cookie and no Authorization header. `carriesAuthCredentials()` (pure, in
webview-proxy.ts) makes that test, and `/` is now admitted by the auth hook only
when it fails; a framed `/` that does carry credentials still gets the shell.
Without a password no auth hook runs at all, so the index route applies the
same test itself (`isLostWebviewRootFrame`) before rendering the shell, and the
three places that emitted the recovery page share `sendLostWebviewFramePage()`.

Tests: the password form in webview-auth-exemption (recovery page for a
credential-free framed `/`, shell with valid Basic auth, 401 with a stale cookie
or a top-level navigation), the passwordless form against a real WebServer in
webview-lost-root-frame (port 3198), and the credential predicate in
webview-proxy. All three fail without the fix. Verified against a live isolated
instance as well: a framed `GET /` with no credentials answers the 470-byte
recovery page, a top-level `GET /` and a framed one carrying a cookie answer the
shell.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:38:54 +02:00
Codeman maintainer d9364f52e1 fix(webview): refuse a backslash or tab-led recovery path, which the URL parser reads as an origin
The lost-frame handler in webview-tabs.js remounts a web-tab frame at the path
the frame reports it lost. It promised "path only, never an origin" and collapsed
a leading run of slashes so `//host/x` could not jump the frame off the proxy,
but it left two spellings through that the WHATWG URL parser treats the same way:
a backslash, which is read as `/` for http(s) schemes, and an ASCII tab or
newline, which the parser deletes before it looks at anything, so `/\host/x` and
`/<tab>/host/x` both resolve to `https://host/x`. That mattered only in
direct-mode tabs, where `POST /api/webviews/:id/open` returns no embedUrl and the
recovered path is resolved with `new URL(path, src)` straight into the frame's
src; a page in such a tab could remount its own frame on a foreign origin.

Not an escalation (the page can already navigate itself anywhere, and the remount
carries no Codeman-origin access), but the comment did not hold and the existing
test only covered the form that already worked. The handler now strips tab, CR
and LF, collapses any leading run of `/` or `\` to one `/`, and refuses whatever
still opens a second separator. The new test drives the reachable direct-mode
branch with all four spellings and fails without the fix.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:38:54 +02:00
Codeman maintainer b0dddc9c57 Merge pull request #402 from shenlvkang-collab/pr/webview-route-masking
fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads
2026-09-14 23:38:54 +02:00
Codeman maintainer f5f399a8b7 test(docker): pin cap_add against the entrypoint, the PATH order and git_head_commit
The capability list is DERIVED from what the scripts do (chown => CHOWN +
DAC_OVERRIDE, a setpriv uid/gid drop => SETUID + SETGID, `init: true` next to a
uid drop => KILL) and compared to docker-compose.yaml's cap_add, the
entrypoint's own required_caps diagnosis, and the lists quoted in docker/README.md
and CLAUDE.md, so the drift that shipped the missing CAP_KILL fails here rather
than on someone's server. Also pinned: the CLI prefix is appended to PATH in
server.Dockerfile and entrypoint.sh pins its PATH before its first command;
Start-Codeman.sh derives PUID/PGID before creating the cases dir, builds before
`down`, writes the source marker only after a refresh, and never aborts on a
failed volume removal.

git_head_commit is run as the script defines it, extracted by its own
delimiters into a real bash, against temp repos made with real git: a symbolic
ref with a loose ref file, a detached HEAD, packed refs after `git pack-refs`,
a linked worktree (which must resolve nothing rather than something wrong) and
a directory that is not a checkout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:37:04 +02:00
Codeman maintainer 2bda191471 docs(docker): describe the root start and drop, and keep the override file out of the image
docker/README.md and docs/docker-compose.md now say that the container starts
as root, corrects a daemon-created bind source and drops to PUID:PGID with
setpriv, which capabilities that needs, and that a compose file written
elsewhere must carry them. The README's PowerShell example runs Compose from
inside docker/ so the override file is discovered, instead of the `-f
docker/docker-compose.yaml` form its own Local customisation section warns
silently drops it, and the reverse-proxy section no longer asks for an override
file now that docker-compose.yaml forwards CODEMAN_ALLOWED_HOSTS itself.

.dockerignore excludes docker-compose.override.* everywhere: it is the
documented home for host-specific settings and rode `COPY . .` into the image,
the same shape as the docker/.env exclusion above it (verified with a scratch
build context: the override files and docker/.env are absent, .env.example and
the compose file present).

CLAUDE.md's Compose paragraph carries the corrected cap list, the writability
probe, and the two traps behind it (KILL is for tini, the CLI prefix is
appended to PATH).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:37:04 +02:00
Codeman maintainer f92883704e fix(docker): create the cases dir with the runtime owner and record a refresh only when it happened
Start-Codeman.sh created CODEMAN_CASES_PATH with a plain `mkdir -p` BEFORE it
derived PUID/PGID from the appdata directory, so the new directory landed as
the invoking user's uid and primary gid. On a host set up the way the README
suggests (`chown -R 99:100 <appdata>`) that gid is not PGID, and the container
refused to start on a directory the script had just made. PUID/PGID are now
derived first and the directory is chowned to them right after creation, with
a clear host-side error when that is not possible. As root this always works,
which also retires the old "refusing to create as root" branch for this path.

The build-artefact volume refresh had three holes. The docker-build-source.json
marker was written whether or not a volume had actually been removed, and the
project name came from a sed over `docker compose config --format json` keyed
on two-space indentation: an empty name made the label filter match nothing,
nothing was removed, and the marker recorded the new HEAD, so the check never
fired again while the stale volume kept serving old code. The name is now
parsed indentation-agnostically, an empty result falls back to `down --volumes`
(the documented reset; both volumes re-seed from the image by a plain copy),
the marker is written only after a successful refresh, and a failed `docker
volume rm` warns and leaves the marker alone instead of aborting under set -e
with the stack down. The image is also built BEFORE `down`, so the deployment
is offline only for the recreate rather than for the whole rebuild.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:37:04 +02:00
Codeman maintainer 1851d80f3a fix(docker): keep SIGTERM reaching the server, pin root's PATH, probe writability
Three changes to how the Compose container starts as root and drops to
PUID:PGID, each reproduced on Docker 29.1.3 / Compose v5.5.0 with a minimal
image of the same shape as server.Dockerfile.

- cap_add gains KILL. `init: true` makes tini PID 1, and tini stays root while
  the entrypoint drops the server to PUID. Signalling a process of a different
  uid needs CAP_KILL, and `cap_drop: ALL` had removed it, so every `docker
  compose down`/`restart` ended in `[FATAL tini (1)] Unexpected error when
  forwarding signal: 'Operation not permitted'` and the server being SIGKILLed
  instead of running `server.stop()`. Measured: without KILL the trap never
  fires, with it the child logs `GOT SIGTERM`.
- /opt/codeman-cli/bin is appended to PATH, never prepended, and entrypoint.sh
  pins its own PATH to the system directories before its first command. The
  prefix is chowned to the runtime account so sessions can update the agent
  CLIs in place, and the root entrypoint resolved stat/chown/setpriv by bare
  name through it: a `setpriv` planted there by the unprivileged uid ran as
  uid 0 at the next start. The image's full PATH is handed back to the server
  at the exec (`env PATH=...`), since Codeman resolves the CLIs through it.
- The ownership gate becomes a writability probe. A directory owned by neither
  root nor PUID:PGID is no longer refused on ownership alone; it is tested with
  `setpriv --reuid PUID --regid PGID --groups <same groups> test -w`, the exact
  identity the server gets, so a group-writable tree, an ACL or a CIFS/NFS
  mount reporting some unrelated uid all pass, and the refusal names path,
  owner and PUID:PGID. Root-owned directories are still chowned first.

Also: a pre-flight runs the drop before touching anything and, when it fails,
prints the cap_add list the compose file needs, so an out-of-tree compose file
(Unraid's Compose Manager) gets a one-line diagnosis instead of a restart loop;
`--bounding-set -all` is gone, since it is a silent no-op without CAP_SETPCAP;
a root:root Docker socket now produces a warning that Docker cases will not
work rather than silently losing group 0 at the drop; and CODEMAN_ALLOWED_HOSTS
is forwarded from .env with an empty default (documented as a commented entry
in .env.example so the parity test and the updater's env gate both stay quiet).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:37:04 +02:00
Codeman maintainer a29e1f61ef Merge pull request #377 from opticon454/bugfix-docker-user-perms
fix(docker): bind-mount ownership, Compose override discovery, and the default runtime account
2026-09-14 23:37:04 +02:00
Codeman maintainer 653e3cdf96 Merge pull request #423 from Ark0N/fix/xterm6-selection-background
fix(terminal): name the selection colour the way xterm 6 does
2026-09-14 23:35:43 +02:00
Codeman maintainer e54a8b1189 Merge pull request #422 from Ark0N/test/install-dsh-probe-bash32
test(ci): exercise the dsh identity probe with timeout missing (bash 3.2)
2026-09-14 23:35:43 +02:00
Codeman maintainer 44a754ea73 test(ci): exercise the dsh identity probe with timeout missing
The bash 3.2 job added with #380 cannot reach dsh_banner_probe, which is
the function #382 was filed against: this image ships `timeout`, so the
optional-prefix array is never empty, and with no `dsh` binary anywhere on
PATH the probe is not called at all. The fix landed in 1.28.2 with nothing
guarding it, and the failure mode is a runtime abort under `set -u` that
`bash -n` cannot see, which is precisely why the reporter had to find it by
reading the source rather than by running anything.

So call the probe directly, with `timeout` hidden behind a narrowed PATH,
and refuse to pass if `timeout` is still reachable (a guard that silently
stops exercising its branch is worse than no guard). Both directions are
asserted: a real DeepSeek Harness banner is accepted, and Debian's unrelated
`dsh` is refused, so the check covers the identity half too.

Verified by reverting install.sh to the pre-fix expansion, where the step
fails with the exact error from the issue, `runner[@]: unbound variable`.

Refs #382

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 23:33:41 +02:00
Codeman maintainer 9acc5aad50 fix(terminal): name the selection colour the way xterm 6 does
Every per-skin xterm palette declared its selection layer as `selection`,
the key xterm.js renamed to `selectionBackground` in v5. An ITheme is a
plain object handed straight to the terminal, so an unknown key is not an
error, it is dropped: all seven skins have been drawing xterm's built-in
default, rgba(255,255,255,0.3), rather than the colour sitting next to it
in the palette.

Nobody saw it on the dark skins, where white at 30% is close to what those
palettes asked for. On the four light skins it is white over a near-white
background: blended, Paper Gray's selection differs from its own background
by 3/255. That is not a subtle highlight, it is no highlight, and it looks
exactly like a selection gesture that failed, which is part of what #360
reports on Android Chrome.

test/skin-themes.test.ts pins both halves: the key name, and that the
blended selection stays at least 16/255 from the background on every skin,
plus the light-skin fallback landing under that floor, which is what makes
this a fix rather than a rename.

Refs #360

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 23:33:36 +02:00
Codeman maintainer 7c3c5b8f72 fix(mobile): show the Codex shift-arrow keys only on codex sessions
The two keys #408 adds to the mobile keyboard accessory bar send
Shift+Left and Shift+Right, which are Codex bindings (edit the last
queued message, step back through the prompt stack). They shipped on
both agent layouts, so a claude, pi, grok, omp, deepseek or gemini
session got two keys that do nothing. That was not only cosmetic: a tap
goes through sendNavKey(), which adds the session to
_echoPassthroughSessions and hands editing to plain PTY echo until Enter
or Ctrl+C, so on a phone a dead key also switched off the local echo
that makes typing feel instant there.

The reveal now follows the shape the 🧠 key already uses. The buttons
stay in both templates, carry an accessory-btn-codex marker class, and
are display:none in styles.css until the bar element carries
codex-enabled. The class has to live on the bar rather than on the keys
because setMode() rebuilds the buttons' innerHTML on every layout
switch. syncCodexKeys() toggles it from the active session's mode
(the same lookup _isShellSession() uses) and is called at init and from
refreshForActiveSession(), which selectSession() already invokes on
every switch. A session's mode is readonly on the server and fixed at
create, so no other event can change the answer; the welcome screen
(no active session) reads as not codex and hides the keys.

The frontend id-branching guard (test/cli-registry-no-id-branching.test.ts)
scans only src/**/*.ts, so the mode comparison in a public JS file is
in bounds, the same as the existing shell check beside it.

Tests: the new describe block in test/mobile-shell-keyboard.test.ts pins
the marker class in both templates, the CSS pair, the class for a codex
session in both layouts, its absence for claude/shell/pi/omp/deepseek,
the re-sync in both directions on a session switch, the no-session case,
and the init + refresh wiring. All six positive assertions fail without
the source change. README and the changeset now say the keys are
Codex-only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:30:14 +02:00
Codeman maintainer 1e5a53830f Merge pull request #408 from shenlvkang-collab/feat/codex-shift-arrow-keys
feat(mobile): add Shift arrows for Codex queued input and prompt navigation
2026-09-14 23:30:14 +02:00
Randalix 63aafdf274 fix(files): serve remote-case attachments, the path a click takes outside the case
A clicked path that points OUTSIDE the case directory goes through the attachment
routes (the frontend's `_isExternalPreviewPath` sends every absolute path not under
`workingDir` to `POST /attachments`), and those had the same local-`fs` assumption
as file-raw: `realpathSync`/`fs.stat` on a path that only exists on the remote host,
so the file never opened — the case the #415 report was actually about.

- `registerExternalAttachment()` accepts `remote` and resolves through
  `remoteProbePaths` (canonical path, size/mtime, kind, plus the workspace root for
  the confinement check). Everything around it — blocklist, extension allowlist,
  workspace confinement, registry/dedupe — is now shared by both branches, so the
  remote path cannot drift from the local one.
- The by-id routes (`raw`, `preview`, `thumbnail`), the metadata poll and the
  attachment history list resolve over ssh too. `raw` streams with the same
  Range contract as file-raw; `preview` (office) and `thumbnail` answer 400 for a
  remote record; an unreachable host answers 502, a vanished file 404.
- Which host a record is read from follows the SESSION, never the path string: the
  same absolute path is a different file on each host, and a remote session never
  falls back to a local file with that name.
- Codex generated artifacts keep force-workspace confinement for a remote case: the
  well-known artifact directories are anchored at THIS host's home, so only a file
  inside the remote workspace is trusted.

Still local-only by design: writes, office conversion, thumbnails, the file
tree/picker and tail-file.
2026-09-14 17:06:42 +02:00
Randalix 013a5d9cc8 fix(files): read remote-case file previews and downloads over ssh
A remote case's workingDir is an absolute path on the remote host, but the
file read routes resolved it with local `fs`: `validateSessionFilePath`'s
realpathSync fails for a path that does not exist on the Codeman host, so
every preview of an agent-written file answered "File not found" (#415).

Add src/remote-files.ts as the single remote-read layer, built on the same
buildSshConnectionArgs() the launch uses:

- remoteProbePaths(): ONE round trip returning realpath + stat for the
  requested path AND the workspace root, so containment is checked against a
  remotely canonicalized root (a symlinked remotePath is ordinary).
- remoteCreateReadStream(): streams the body (cat, or tail -c +N | head -c L
  for a Range) with nothing buffered in memory, and reaps the ssh child when
  the response ends so an aborted download cannot orphan it.
- remoteReadFile(): bounded read for file-content.

file-raw, file-content, file-preview and file-thumbnail now share one local/
remote target resolution. Guards keep their local strength: lexical pre-check,
remote realpath, workspace containment, sensitive-path blocklist, and the size
cap applied to the remote size before any bytes are read. An unreachable host
answers 502 with the remote reason instead of a misleading 404. Nothing is ever
copied to the Codeman host and there is NO local fallback (an sshfs mount of
the same tree must not shadow the remote bytes).

Deliberately unchanged: writes (edit=1 / PUT now answer 400 explicitly while
the viewer hides its Edit affordance), office previews, thumbnails, file tree,
picker, external attachment registration and tail-file stay local-only.
2026-09-14 14:54:41 +02:00
DevvynandClaude Sonnet 5 b6f75b87f5 fix(custom-model): don't clamp DEEPSEEK_API_KEY as a privileged env key
CI caught a real regression: DEEPSEEK_API_KEY was added to deepseek's
privilegedEnvKeys alongside DEEPSEEK_BASE_URL on the theory that "the pair
travels together," but that contradicts the documented and tested design
(clampEnvOverridesForOwner()'s own docstring in session-routes.ts) — a
non-granted owner supplying their OWN DeepSeek key removes privilege
rather than granting it, since the exfiltration vector is the BASE URL
(which redirects the server's own forwarded key to a foreign host), not
the key itself. Removed it from the list; test/deepseek-mode.test.ts's
existing two clamp tests now pass again.

Also swapped that test's "unrelated override" example off CODEX_HOME,
which the earlier commit in this same PR legitimately made privileged
(closing a real pre-existing gap, documented in PR.md) — so it stopped
being a valid "unrelated" example the moment that fix landed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00
DevvynandClaude Sonnet 5 e18499aa67 docs(pr): drop the draft/WIP framing now that the PR is submitted
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00
DevvynandClaude Sonnet 5 61779745aa test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.

Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.

Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:

- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
  speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
  don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
  undocumented GATEWAY AuthType gemini-cli selects once
  GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
  tried
- deepseek: reaches the server (env vars are read) but gets a consistent
  HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)

Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).

deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00
DevvynandClaude Sonnet 5 41416566aa feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).

- Registry: capabilities.customModelInjection per CLI entry (env /
  configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
  model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
  custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
  (POST /api/sessions/:id/custom-model), reusing the existing
  respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
  CLI's privilegedEnvKeys, closing a pre-existing gap where several were
  already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
  against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
  every CLI's injected values through a real HTTP shape

Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
  a top-level model string + [model_providers.custom]); fixing it then
  surfaced a genuine, documented protocol incompatibility (Codex only
  speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
  don't implement)
- Claude Code's async session-title-generation call validates
  ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
  hangs the whole -p invocation on an unrecognized name; documented for
  chunk 6, worked around in the standalone script only (--bare is NOT
  safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
  and api-key headers) reliably hung a real server; removed the option
  entirely rather than just changing the default

Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00
DevvynandClaude Sonnet 5 c179daf869 fix(docker): re-assert /opt/codeman-cli ownership every start, not just at build
/opt/codeman-cli is chowned to PUID:PGID once, at image build time, from
the PUID/PGID build args. That bake only happens when the image is
actually rebuilt (`docker compose up --build`, which Start-Codeman.sh
always does) — a deployment that runs the compose file directly instead
(Unraid's Compose Manager, a native systemd unit, any plain
`docker compose up`/`restart`) can change PUID/PGID in .env and restart
without ever rebuilding. The container then runs as the NEW uid via
entrypoint's setpriv (Linux needs no /etc/passwd entry to setuid to an
arbitrary number) while the CLI directory is still owned by the OLD one
baked into the image layer — silently breaking the self-update-a-CLI-
in-place fix that directory exists for.

Unlike HOME/CODEMAN_CASES_PATH, this one is pure image content Codeman
itself populated, never host data that might legitimately belong to
someone else, so there is no ownership to be careful about — it is
always correct for it to be owned by whoever the container is about to
run as. Re-assert it unconditionally on every start.

Verified live: built an image with PUID=99/PGID=100, ran it with
PUID=1234/PGID=4321 (no rebuild, simulating a changed .env restarted
directly), confirmed /opt/codeman-cli ends up 1234:4321-owned and is
genuinely writable by the running process.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-13 17:41:31 +08:00
DevvynandClaude Sonnet 5 ae32daf135 fix(docker): address maintainer review on #377
Two real bugs the review caught, both verified live against a real
build on the Unraid host:

1. entrypoint.sh's chown fired on ANY ownership mismatch, not just a
   directory the daemon itself created root-owned. A host tree
   legitimately owned by some other account - an existing
   CODEMAN_CASES_PATH the README already allows pointing at a normal
   projects directory, or appdata under a different PUID/PGID
   convention than the one in use - got silently recursively re-owned
   with one log line to explain it. Now gated on the target actually
   being root-owned; anything else is a clean refusal naming the
   directory, its owner, and PUID/PGID. Start-Codeman.sh also now
   pre-creates CODEMAN_CASES_PATH the same way it already did
   CODEMAN_APPDATA_PATH, so Compose never has to materialise a missing
   bind source as root in the first place - the in-container chown
   becomes a safety net, not the primary mechanism.

2. The CLI-update chown (chown -R .../node_modules /usr/local/bin)
   handed the runtime account write access to entrypoint.sh itself
   (root-owned, executed as root on every container start with
   CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node binary - owning the
   DIRECTORY is enough to rename it aside and drop a replacement, which
   would let a compromised session arrange for its own script to run
   as root at the next restart. The four CLIs now install into a
   dedicated /opt/codeman-cli prefix (NPM_CONFIG_PREFIX); only that
   directory is chowned, /usr/local stays root-owned throughout.

Smaller fixes from the same review:

- Start-Codeman.sh's volume-refresh label filter wasn't project-scoped:
  a second Compose stack on the same host sharing the `codeman-dist`
  volume KEY could have had ITS volume deleted. Added a
  com.docker.compose.project filter, resolved from this stack's own
  `compose config --format json`.
- Override-file precedence was backwards (checked .yaml before .yml;
  Compose actually prefers .yml) - swapped, plus a warning when both
  exist.
- entrypoint.sh's setpriv now also passes --bounding-set -all, so
  CapBnd actually clears post-drop rather than just CapPrm/CapEff.
- A comment on git_head_commit() noting it returns nothing for a
  worktree checkout (.git as a file), consistent with the script's
  existing -d .git convention elsewhere.
- Doc drift: CLAUDE.md's Docker Compose section still described the
  old pre-created-and-chowned-by-hand model and didn't mention the
  root-then-drop entrypoint; the state-files list was missing
  docker-build-source.json; docs/docker-compose.md and
  docker/.env.example still had the pre-rename `Coding/codeman` path
  in one place each.

Verified end to end against a real build on the Unraid host: a
root-owned bind source is corrected as before; a directory owned by
neither root nor PUID:PGID is refused rather than silently rewritten;
a correctly-owned directory is left alone entirely; the four CLIs
resolve via PATH from /opt/codeman-cli while /usr/local/bin,
/usr/local/lib/node_modules and entrypoint.sh itself stay root-owned;
CapBnd is fully cleared post-drop.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-13 17:41:31 +08:00
DevvynandClaude Sonnet 5 8fe3f34fc5 fix(docker): detect and refresh stale build-artefact volumes
codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded
from the image only while empty, so a rebuilt image's fresh dist/
node_modules sat unused behind old volume content until something
cleared it. The in-app self-updater never hit this (it rebuilds INSIDE
the running container, into the very volume already in use), but a
`docker compose build` triggered from outside it — Start-Codeman.sh,
after a manual `git pull` — did: the container came back up looking
unchanged, serving stale compiled routes against current source.

Start-Codeman.sh now compares the checkout's HEAD commit and
package-lock.json hash against a recorded marker
(docker-build-source.json) and clears just the affected volume(s)
before its own --build when either moved.

The in-place self-update path writes that same marker after a
successful build, so the two mechanisms agree on what the volumes
currently reflect — without it, the next plain Start-Codeman.sh run
would see the HEAD self-update just checked out, not recognise it as
already accounted for, and wipe the volumes self-update just correctly
rebuilt right back to the older baked image.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-13 17:41:31 +08:00
DevvynandClaude Sonnet 5 89e2cb5814 fix(docker): let the runtime account update its own global CLIs
The four CLIs (claude, gemini, codex, opencode) are npm-installed
globally as root during the image build, before the unprivileged
runtime account exists. A session running as that account (e.g. a
codex-mode terminal) then hits EACCES the moment it tries to update
one in place, because npm renames the old package directory aside
before installing the new one, which needs write access to the
parent (/usr/local/lib/node_modules), not just the target package.

Chown that tree plus /usr/local/bin's CLI symlinks to PUID:PGID in
the same step that creates/renames the runtime account.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-13 17:41:31 +08:00
DevvynandClaude Sonnet 5 d38bf33a69 docs(docker): document the reverse-proxy host allowlist
CODEMAN_ALLOWED_HOSTS is a real, documented application setting (the Host-
header allowlist in network-auth-policy.ts), but docker-compose.yaml does not
forward it from .env into the container - Compose only passes through
variables explicitly listed under environment:, and this is not one of them.
Set without that passthrough, any request through a reverse proxy is rejected
with 403 Forbidden: host not allowed before it reaches any handler, and
nothing in the Docker deployment docs said why.

Document the variable and the override needed to forward it, using the
Local customisation mechanism already described above it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-13 17:41:31 +08:00
DevvynandClaude Opus 5 9702126046 chore(docker): name the default runtime account codeman
CODEMAN_RUNTIME_USER defaulted to `opencode`, which no longer matches the
project and is confusing in a deployment whose every other identifier is
codeman. Rename the default in .env.example and in the Dockerfile ARG that
mirrors it, and correct the example comment that referred to
/home/opencode/codeman-cases.

Also drop the `Coding/` component from the example application-data path.
CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH now suggest /mnt/user/appdata/codeman
and its codeman-cases child, matching the account name and removing a directory
level that meant nothing outside the original author's host. README.md is
updated to match, including the chown example.

The npm package `opencode-ai` and the references to the OpenCode CLI are
deliberately left alone: those name a different tool, not this account.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 17:41:31 +08:00
DevvynandClaude Opus 5 748bbf5423 fix(docker): honour docker-compose.override.yml in Start-Codeman.sh
Naming a Compose file with -f disables Compose's automatic discovery of the
override file, so Start-Codeman.sh silently ignored docker-compose.override.yml.
Any local customisation placed in the conventional override file was dropped
without warning, and the only way to notice was to inspect the running
container.

Collect the -f arguments into an array, append the override file when one is
present, and reuse that array for the final launch so the two cannot drift
apart again. Both .yml and .yaml are checked, in Compose's own precedence
order, and the chosen file is reported on startup.

Document the override file in docker/README.md, including the two things that
are easy to get wrong: it is ignored when -f is passed without naming it, and
it cannot remove a key such as ports, which Compose concatenates. Add the
override file to .gitignore.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 17:41:30 +08:00
DevvynandClaude Opus 5 10876aa440 fix(docker): correct bind-mount ownership before dropping privileges
Compose binds CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH from the host. When
either path does not exist yet - a first run, a cleared application-data
directory, a restored backup - the Docker daemon creates it owned by root. The
server runs unprivileged as CODEMAN_RUNTIME_USER, so it cannot create its own
state directory, and the container restarts forever on:

  Failed to start web server: EACCES: permission denied, mkdir '/home/<user>/.codeman'

Start-Codeman.sh already worked around this by preparing the directory on the
host, so the failure only appears when Compose is run directly, which the README
documents as a supported path.

Add docker/entrypoint.sh, which starts as root, corrects the ownership of both
bind mounts, then drops to PUID:PGID with setpriv. The Dockerfile's USER
instruction is replaced by that entrypoint and CMD is unchanged.
docker-compose.yaml adds back only the four capabilities the chown and the
privilege drop require, so cap_drop: ALL continues to remove everything else.

Two guards keep existing deployments working:

- A container started with an explicit `user:` is left alone. The entrypoint
  execs straight through, with no elevation and no chown.
- A chown that fails is a warning, not an error. Bind mounts backed by NFS,
  CIFS or a rootless daemon can refuse chown while remaining perfectly
  writable, and those deployments must keep starting.

PUID and PGID are also exported as runtime environment defaults so the image
behaves correctly when run without Compose, rather than depending on build args
alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 17:41:30 +08:00
codeman-local b357fe832e feat(mobile): add Shift arrow keys for Codex prompt navigation 2026-09-12 20:45:05 +08:00
shenlvkang-collabandClaude Fable 5.1 349a89ec3b fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads
A dashboard served through a web tab saw `/webview/<cap>/` as its
`location.pathname`, and no app has a route for that: a React Router, Vue
Router or Vite dev-server page painted its HTML and CSS and then replaced
them with its own "page not found" the moment its script ran (reproduced
with a minimal history-routed page).

The proxy's runtime shim now rewrites the history entry to the path the
page would see on its own origin, before any page script runs. The base
element still resolves relative URLs inside the prefix and every root-
absolute sink is rewritten back into it, so only what the page READS
changes. With the document URL masked the Referer-keyed 404 rescue can no
longer help a request the shim misses, so the remaining URL-taking entry
points (`Worker`, `SharedWorker`, `navigator.sendBeacon`, `window.open`)
are covered by the shim as well.

A navigation the page starts itself afterwards — `location.reload()`
(a dev server's full-reload HMR), a root-absolute `location.href` — lands
on Codeman's root with no capability anywhere: no prefix in the path, no
cookie in an opaque-origin frame, a Referer naming the masked page. It is
recognised by shape (a top-level iframe navigation asking for HTML, for a
path Codeman does not serve) and answered with a static page whose only
script posts `{type:'codeman:webview-lost', path}` to the parent; the tab
that owns the frame (matched by `event.source`, never by the payload)
remounts it inside the prefix at that path, bounded per frame. The
unauthenticated form is answered in the auth middleware before the
credential checks, so a dev server that reloads on every save cannot
rate-limit its own user out of Codeman; the authenticated form (Basic
auth, trusted mode) is answered by the 404 handler.

Verified end to end against a history-routed page: boots on `/`, its
API call succeeds, a reload inside the frame comes back routed on the
path it had pushed, `location.href = '/about'` comes back on `/about`,
and a deep link opens on its path.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-10 14:13:15 +08:00
codeman-local 268e4819ff feat: auto-name sessions from first prompt 2026-09-03 18:14:29 +08:00
163 changed files with 15391 additions and 698 deletions
@@ -0,0 +1,5 @@
---
"aicodeman": patch
---
fix(terminal): keep the output a pane capture could not contain. Opening a session, a backpressure refresh, a clear-terminal reload and a full-history re-pull all load the screen from a tmux pane capture, and anything the CLI printed between that capture and the end of the load used to be dropped, so its next partial redraw landed on a frame the terminal had never seen: missing or garbled output right after a tab switch or a refresh, plainest in a shell session. Each load now replays exactly the output that arrived after the capture, through one shared rule for all four paths, and a refresh that restores your scroll position no longer snaps back to the bottom afterwards.
+5
View File
@@ -0,0 +1,5 @@
---
"aicodeman": patch
---
fix(input): make sure a prompt sent through the API actually leaves the composer. Claude Code 2.1.277 started ignoring Enter for the first 30 to 50 seconds after the composer paints while still accepting the typed text, so a prompt sent right after a session came up sat unsent in the pane and every waiter (send-and-wait, the agent skill, cron, the maintainer bot) burned its whole timeout on a turn that never started. The server now reads the pane after every programmatic write that carried Enter and presses Enter again, on a 2 to 60 second schedule, only while the composer verifiably still holds the text it sent; an empty composer, other text, or a pane with no composer at all ends it. The agent skill's `sendwait` gets the same loop for servers that predate this, and its preamble version moves to 1.30.1 so an already-seeded agent picks up the fresh copy.
+1 -1
View File
@@ -10,7 +10,7 @@
"name": "codeman",
"source": "./plugins/codeman",
"description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.28.2",
"version": "1.30.0",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
+3
View File
@@ -9,6 +9,9 @@
**/.env
**/.env.*
!**/.env.example
# Same shape: docker/docker-compose.override.yml is the documented home for
# host-specific settings, so it must not ride COPY . . into the image either.
**/docker-compose.override.*
node_modules
dist
coverage
+26
View File
@@ -74,6 +74,32 @@ jobs:
offer_ai_cli_install >/dev/null 2>&1
echo "bash $BASH_VERSION: skipping the AI CLI install menu continues"
'
# Issue #382: the dsh identity probe builds an OPTIONAL `timeout` prefix as an
# array, and on stock macOS there is no `timeout`, so the array is empty and the
# expansion aborts the whole installer under `set -u`. The step above cannot
# reach that branch: this image HAS `timeout`, and with no `dsh` on PATH the
# probe is never called at all. So hide `timeout` and call it directly.
docker run --rm -v "$PWD":/w -w /w -e CODEMAN_INSTALL_SH_LIB=1 bash:3.2 bash -c '
set -euo pipefail
. /w/install.sh
printf "#!/bin/sh\necho \"DeepSeek Harness 0.1\"\n" > /tmp/dsh
printf "#!/bin/sh\necho \"dancer shell (Debian dsh)\"\n" > /tmp/not-dsh
chmod 755 /tmp/dsh /tmp/not-dsh
# A PATH the probe can still work on, minus the binary under test.
mkdir -p /tmp/nobin
for b in grep sh; do ln -sf "$(command -v $b)" "/tmp/nobin/$b"; done
export PATH=/tmp/nobin
if command -v timeout >/dev/null 2>&1; then
echo "timeout is still on PATH, so this is NOT exercising the empty-array branch" >&2
exit 1
fi
dsh_banner_probe /tmp/dsh
if dsh_banner_probe /tmp/not-dsh; then
echo "identity probe accepted a foreign dsh" >&2
exit 1
fi
echo "bash $BASH_VERSION: dsh identity probe survives a missing timeout"
'
- name: CLI catalogue artifacts are in sync with stock.ts
run: npm run generate:cli-catalog -- --check
+8
View File
@@ -48,6 +48,10 @@ Thumbs.db
.env.local
.env.*.local
# Local Compose customisation (host-specific, not part of the project)
docker-compose.override.yml
docker-compose.override.yaml
# State files (local to each machine)
.claude/ralph-loop.local.md
@@ -105,3 +109,7 @@ readme-preview.mjs
# Uploaded images land here under each session working dir (runtime artifact)
.claude-images/
# Local-LLM harness smoke-test config (real IPs/keys) — see the .example.json
# alongside it in scripts/, which IS tracked as the template.
scripts/local-llm-test.config.json
+104
View File
@@ -1,5 +1,109 @@
# aicodeman
## 1.30.0
### Minor Changes
- da933d7: Offer to rebuild the sessions a host reboot destroyed. A reboot takes the tmux server down with it, so every pane dies and the board comes up empty. Codeman now works out what was running, and the board offers to restore it behind a click. The conversations come back; the terminal scrollback does not, and the banner says so.
### Patch Changes
- a1c35da: Stop a phone keyboard losing the last character of every message it sends. Android soft keyboards commit the last typed character and send the Enter key in one InputConnection transaction, so the `input` event and the Enter keydown are both processed before any zero-delay timer runs. The orphaned-input recovery from #388 only resolved its candidate on such a timer, and lost it both ways: xterm emits `\r` synchronously from the Enter keydown, so the local-echo composer submitted the prompt before the recovered character existed, and that `\r` bumped the "did xterm speak for this keystroke" counter, so the candidate then stood itself down and dropped the character outright. Pending candidates are now drained synchronously at the next keydown, from xterm's custom key handler, which runs before xterm processes that key, so the counter still holds the value it had while the candidate's own keystroke was current, and the recovered byte reaches the composer ahead of the Enter. Typing on a physical keyboard is unaffected: there, the timer has already resolved the candidate before the next key arrives.
- 3f2928a: The installer's hint for a launcher-only CLI (DeepSeek today) now says why it is a docs link rather than a command you can run, and points at the thing that resolves it: the package installs a launcher that still needs a terminal profile, and Codeman's Run menu can add one in a click. Driven by a generated `CLI_LAUNCHER_ONLY` flag rather than an id check, so it covers any future entry of that shape. Also removes three dead lookup helpers and two never-read generated arrays from `install.sh`, skips a disabled entry's probe instead of filtering it afterwards, and corrects a comment that claimed the non-interactive default is always Claude Code (on a wget-only host its curl one-liner is filtered out first).
- 0e1191b: Maintainer fixes applied while landing the above. A session restored after a reboot keeps the name you gave it (the rebuild dropped the field that records who named a session, so a hand-renamed session came back looking auto-named and the next prompt overwrote it), and no longer types `continue` into itself on its own: a pending auto-resume stamp from before the reboot is dropped rather than re-armed, since the pane is new and one click could otherwise arm several unattended prompts at once. Auto-resume itself stays on and re-arms on the next real usage-limit message. The restore offer is also hidden in a detached single-session window, which has no tab strip to put restored sessions in, and a conversation that goes live while an earlier session in the same batch is starting is no longer restored a second time.
- 0e1191b: ### Thanks
- @irisitymichaelgrundberg for the reboot-restore banner (#442), and for the three real reboots behind it rather than a mocked one.
- @shenlvkang-collab for tracking down why Android keyboards lost the last character of every message (#441), including the half where the character was not late but gone.
- @opticon454 for going back and closing out the loose ends left as "worth knowing rather than fixing" after #380 (#429).
- de864e7: Keep the terminal anchored where you are reading while an agent streams (#358). Scrolling up during a Codex response could still be dragged back to the live bottom by the next redraw: the flush captured the viewport before writing and restored it immediately after, but xterm parses asynchronously, so at that moment the buffer had not moved yet, the restore compared the anchor against itself and did nothing, and the redraw landed a tick later with nothing left to pull the view back. The restore now runs inside xterm's own write callback, which is the first point at which the redraw's effect exists, and it holds across consecutive and chunked redraws. It is dropped if you switch sessions or a history replay starts before the write parses, since the anchor indexes the buffer it was captured from.
## 1.29.1
### Patch Changes
- 5b920cb: Auto-name sessions from the first prompt (#376, opt-in). With the new synced **Auto-name Sessions** setting on (App Settings → Appearance → Tabs, default off), a tab that still carries its generated name takes a title from the first real prompt you submit, keeping the case prefix: `w3-myapp` becomes `w3-myapp: fix the login redirect`. The strip shows the title with the prefix in the tooltip, and the next session in that case still counts up. It happens once per session, only for prompts you type or send through the input API (never a Ralph, respawn, cron or approval answer), never for shells, and a name you set yourself is never touched. Slash commands such as `/clear` do not become titles. The title is derived locally from the prompt's first sentence; no text leaves the machine. `nameSource` (`placeholder` / `auto` / `manual`) is a new additive field on session state.
Landed with the fixes the review of #376 asked for: first prompt only (not every prompt), a user-input gate so Ralph, respawn, cron and approval writes cannot name a tab, the prefix form so the case identity and `w<n>` counter survive, and a keystroke tracker that handles a bare Esc, bracketed pastes, wheel reports, Tab and history recall instead of mis-titling the tab.
### Thanks
- @shenlvkang-collab for #376, the auto-naming idea and the ownership plumbing (`nameSource`, the listener wiring, the restore path) it shipped with.
## 1.29.0
### Minor Changes
- **Custom model endpoints, HTTP API first** (#393). Any run mode that has a mechanism for it can be pointed at a custom OpenAI-compatible endpoint (a local llama.cpp, llama-swap, Ollama or vLLM, or a cloud gateway) instead of its native backend, per session. Endpoints are stored in `~/.codeman/custom-model-hosts.json` (`GET/POST/PUT/DELETE /api/model-endpoints`, admin-only in multi-user mode), their model lists are discovered from the endpoint's own `/v1/models`, and `POST /api/sessions/:id/custom-model` applies one to a session by restarting its CLI in place. The mechanism is per-CLI registry data (`capabilities.customModelInjection`): env vars for Claude, Gemini, Grok and DeepSeek, `OPENCODE_CONFIG_CONTENT` for opencode, an isolated config dir for Codex, Pi and OMP, unsupported for Antigravity. Verified live against a llama-swap server for claude, opencode, pi, grok and omp; gemini and deepseek reach the server and fail for reasons not yet understood, and codex only speaks the Responses API, so a plain chat-completions server cannot serve it. Those three are documented as gaps rather than shipped as working. The toolbar picker is a follow-up; until it lands the feature is HTTP-API only (`docs/custom-model-endpoints.md`), and the `customModelEndpointsEnabled` setting is declared but read by nothing yet. Merged with maintainer follow-ups: clearing a selection now actually clears it (the injected vars are delivered by `tmux setenv`, which `respawn-pane` inherits, so the relaunched CLI came back still pointed at the endpoint; retired keys are now `setenv -u`'d before the respawn), applying a model to a local claude session no longer kills the pane (the relaunch pins `--resume <id>` with the `--session-id` fallback, since Claude Code refuses a session id that already has a transcript), pi, omp and grok now select the generated model through a registry-declared `launchModel` (`custom/<id>`, `-m codeman-custom`) instead of writing a config the CLI then ignored, remote and Docker sessions are refused with a clear 400 until those paths are plumbed, the selection survives a Codeman restart, discovery goes through the egress-guarded `webviewFetch()`, key-bearing files are written 0600 and the per-session config dir is removed with the session, and the design plan moved from the repo root to `docs/custom-model-endpoints-plan.md`. Along the way the multi-user clamp learned about `GOOGLE_GEMINI_BASE_URL`, `GROK_BASE_URL`, `CODEX_HOME`, `PI_CONFIG_DIR` and `OPENCODE_CONFIG_CONTENT`, which were already reachable through `envOverrides` and now count as privileged keys.
**Single-page apps work as web tabs, and a frame that reloads comes back** (#402). A history-routed dashboard (React Router, Vue Router, a Vite dev server) read `/webview/<cap>/` as its `location.pathname` and rendered its own "page not found" the moment its script ran. The proxy's runtime shim now masks the prefix off the document URL before any page script runs, while every URL the page emits still goes through the rewrite layers (now including `Worker`, `SharedWorker`, `sendBeacon` and `window.open`). A navigation the page starts itself afterwards (a dev server's full reload, a root-absolute `location.href`) used to land on Codeman's root with no capability; it is now recognised by shape, answered with a static recovery page that posts the lost path to the owning tab, and the frame is remounted inside the prefix at that path, bounded to five recoveries a minute per frame. Merged with maintainer follow-ups: the recovery path is sanitised properly (a leading backslash, or a tab/newline the URL parser deletes before parsing, resolved `/\evil.com` to a foreign origin in a direct-mode tab); a reload on the dashboard's landing page is recovered too, on password-protected and passwordless installs alike (it used to render Codeman's own shell inside the web tab); and the recovery page is written down as the third unauthenticated 200 in the security table and `docs/security-architecture.md`, with the route-enumeration property it implies stated rather than left to be discovered.
**Shift arrows for Codex on the phone keyboard bar** (#408). Two keys, `⇧←` and `⇧→`, send the Shift-modified arrows Codex binds to editing the last queued message and walking the prompt stack (verified against Codex 0.154.0's `/keymap`). Merged with a maintainer follow-up: the keys are shown only on Codex sessions (a `codex-enabled` class on the bar, the same shape as the Read My Mind key), because tapping one in any other session did nothing except hand that session to plain PTY echo for the rest of the prompt.
**Remote (SSH) cases can finally show you their files** (#421, fixes #415). File previews, downloads, text reads and the out-of-workspace attachment path resolved every path against the Codeman host's own filesystem, so in a remote case every click ended in "File not found" while the file plainly existed on the other machine. A single new ssh read layer (`src/remote-files.ts`, built on the same `buildSshConnectionArgs()` the launch uses) probes realpath and stat for the file and the workspace root in one round trip, then streams the body with `cat` (or a `tail`/`head` slice for a `Range`), so the 200/206/416 contract holds and nothing is buffered on the server. Symlinks are resolved on the host that can resolve them, containment is checked against the resolved remote root, the size cap applies to the remote size before a byte is requested, an unreachable host is a 502 rather than a 404, and there is deliberately no local fallback: a same-named file on the Codeman host is never served under a remote name. Writes, Office previews and generated thumbnails answer 400 for a remote case instead of a misleading 404. Merged with maintainer follow-ups: the `readlink -f` fallback resolved only the directory chain, so on a host without it a symlink's final component was returned unresolved and `ws/notes.txt -> ~/.ssh/id_rsa` passed containment while `cat` served the key; it now follows the last component with plain `readlink` for a bounded number of hops and fails closed (404) on a loop or the cap; `PUT /api/sessions/:id/file-content` answers 400 for a remote case as the PR already claimed (it still validated against the local filesystem, so a same-named local directory took the write); ssh children are bounded by a small semaphore (`CODEMAN_MAX_REMOTE_FILE_SSH`, default 4) covering the attachment-history fan-out, which now probes the whole history in one batched call, and the fire-and-forget magic-link registrations an injected agent could use to fork hundreds of `ssh` processes; probe records are NUL-delimited and index-keyed so a newline in a filename cannot shift one path's result onto the next; and a 502 body never carries the ssh command line.
**Docker Compose: bind-mount ownership, override files, a `codeman` runtime account, and no more stale volumes** (#377). A missing bind source (first run, cleared appdata, restored backup) is created root-owned by the daemon, and the unprivileged server crash-looped on `EACCES` when Compose was run directly; the image now starts through an entrypoint that corrects a root-owned bind mount and drops to `PUID:PGID` with `setpriv`, and the compose file adds back only the capabilities that needs. `Start-Codeman.sh` honours `docker-compose.override.yml` (naming a Compose file with `-f` silently disables Compose's own discovery of it), pre-creates the cases directory like it already did for appdata, and detects when the checkout's HEAD or lockfile moved under the `codeman-node-modules`/`codeman-dist` volumes and refreshes them, which used to leave a `docker compose build` serving stale compiled routes. The default runtime account is named `codeman` (it was `opencode`), the four global agent CLIs live in their own `/opt/codeman-cli` prefix so the runtime account can update them in place without owning `/usr/local/bin`, and `CODEMAN_ALLOWED_HOSTS` is documented and forwarded. Merged with maintainer follow-ups: `cap_add` gains `KILL` (with `init: true` tini runs as root while the server runs as `PUID`, and without CAP_KILL its SIGTERM forward failed and the server was SIGKILLed on every `compose down`/`restart`); the CLI prefix is appended to `PATH` rather than prepended and the root entrypoint pins its own `PATH`, since a `PUID`-writable directory ahead of `/usr/bin` let the runtime account plant a `setpriv` that ran as root on the next start; the entrypoint decides with a real writability probe as the runtime identity instead of an owner comparison, so ACLs, group-writable trees and NFS/CIFS mounts work and only a genuinely unwritable directory is refused, by name; the cases directory is created with the runtime owner after `PUID`/`PGID` are known; the build-source marker is written only when a refresh actually happened, an empty Compose project name falls back to `down --volumes`, the build runs before the `down` so the stack is offline only for the recreate, `docker-compose.override.*` stays out of the image, and `test/docker-entrypoint.test.ts` pins `cap_add` against what the entrypoint needs. ⚠️ Compose users: run `Start-Codeman.sh` once for this release rather than a plain `docker compose up`, so the rebuilt image, the refreshed volumes and the new entrypoint arrive together.
**Selected text is visible again on the light skins** (#423, part of #360). Every skin palette named its selection layer `selection`, the key xterm renamed to `selectionBackground` in v5, so all seven skins had been painting xterm's default white at 30% instead of the colour next to it in the palette. Dark skins hid it; on the four light skins a selection was white on near-white. The key is renamed and `test/skin-themes.test.ts` pins it. CI additionally exercises `install.sh`'s dsh identity probe with `timeout` missing under bash 3.2 (#422), the guard #382's fix shipped without.
**Eight fixes salvaged from #375** (dignfei; landed with the author's commits preserved, the rest of that PR is covered below). Shift+drag starts a text selection in a pane whose mouse reports go to the CLI, and right-click copies the selection. Ctrl- and Alt-modified navigation keys typed through the CJK composer reach the CLI as the modified sequences instead of plain arrows. A browser whose reliable-input sequence counter fell behind the server's watermark (a restored tab, a cleared localStorage) now recovers: the duplicate ACK carries `dup: true` plus the watermark, the client lifts its counter and re-sends, so a session that had silently stopped accepting typed prompts accepts them again. An SSE reconnect that lands on the session you are already looking at keeps its terminal buffer and resyncs instead of resetting the whole terminal. The hidden offline overlay and the file-preview overlay only apply `backdrop-filter` while shown, which removes a stale compositing layer that swallowed clicks. One adopted Docker container can back several cases at different in-container directories, and the adopt panel gains a "copy an existing case" picker. Of the PR's 27 commits, 14 had already shipped through #357, the selection theme key rename shipped as #423, and foreign tmux adoption plus SSH password auth stay with the author.
### Thanks
- **@opticon454** for custom model endpoints (#393), including the part nobody enjoys: working out each CLI's real endpoint mechanism against real binaries and writing down which ones do not work yet instead of claiming they do; and for the Docker Compose deployment fixes (#377), rebased and reworked through three review rounds.
- **@shenlvkang-collab** for making single-page apps route inside web tabs and recovering a frame that reloads (#402), the best-engineered PR of this batch, and for the Codex Shift arrows on the phone keyboard bar (#408), verified against Codex's own keymap.
- **@dignfei** for the eight fixes salvaged from #375 (terminal selection and copy, CJK navigation keys, input recovery, SSE reconnect, overlay compositing, multi-case adopted containers), landed under their own name.
- **@Randalix** for reporting #415 and then fixing it themselves with the whole missing ssh read side for remote cases (#421), with a real-shell test for the probe script and a full route suite.
### Patch Changes
- 349a89e: fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads
A dashboard served through a web tab saw `/webview/<cap>/` as its `location.pathname`, and
no app has a route for that: a React Router, Vue Router or Vite dev-server page painted its
HTML and CSS and then replaced them with its own "page not found" the moment its script ran.
The proxy's runtime shim now rewrites the history entry to the path the page would see on its
own origin before any page script runs, while every URL the page emits still goes through
the existing rewrite layers (plus `Worker`, `sendBeacon` and `window.open`, which the masked
Referer can no longer rescue). A navigation the page starts itself afterwards — a dev
server's full-reload HMR, a root-absolute `location.href` — lands on Codeman's root with no
capability; it is recognised by shape (an iframe navigation asking for HTML for a path Codeman
does not serve), answered with a static page that tells the owning tab which path was lost,
and the tab remounts the frame inside the prefix at that path. That answer is served before
the credential checks, so it never counts as a failed login.
- 013a5d9: File previews, downloads and text reads now work in a **remote (SSH) case**.
A remote case's working directory is an absolute path on the _remote_ host, but the
file routes resolved it with local `fs` — so a clicked path (or the File Viewer) always
failed as "File not found" even though the file existed and the session was clearly
working in that directory. `GET /api/sessions/:id/file-raw`, `file-content`,
`file-preview` and `file-thumbnail` now resolve and read through the same
`buildSshConnectionArgs()` connection the launch uses (`src/remote-files.ts`, one
`realpath`+`stat` probe per request returning both the file and the workspace root).
Clicked paths that point OUTSIDE the case directory (a remote `/tmp` scratchpad capture,
a screenshot elsewhere in the remote home) go through the attachment routes, which had
the same local-`fs` assumption: registration, the by-id `raw` stream, the metadata poll
and the attachment history list now resolve over ssh as well, so the click-path works
whether the file sits inside or outside the case. Which host a record is read from
follows the SESSION, never the path string — the same absolute path means a different
file on each host, and a remote session never falls back to a local file.
The guards are unchanged in strength: the workspace boundary is still enforced (now
resolved on the host that can actually resolve it), the sensitive-path blocklist and
the size cap (`CODEMAN_MAX_DOWNLOAD_BYTES`) still apply before any bytes are read, and
`Range` requests keep working, so remote `<video>`/`<audio>` seeking behaves like a
local file. An unreachable host is reported as `502` with the remote reason instead of
a misleading 404. Nothing is ever copied to the Codeman host.
Still not available for remote cases, and now said explicitly instead of 404-ing:
editing a file (`edit=1` / `PUT` answer 400, the viewer hides its Edit affordance),
office-document previews and generated thumbnails (both need the bytes on the server's
disk), the file tree / path picker, and `tail-file`. Docker cases are unaffected (their
workspace is bind-mounted at the same absolute path).
- b357fe8: Add Shift+Left and Shift+Right buttons to the default and extended mobile agent keyboard bars, shown only on Codex sessions, enabling Codex queued-message editing and prompt-stack navigation. Flush locally buffered drafts before navigation and keep terminal focus after taps.
- 9acc5aa: Fix an invisible terminal text selection on the light skins (#360). Every xterm palette declared its selection colour under the key `selection`, which xterm.js renamed to `selectionBackground` in v5. An `ITheme` is a plain object, so the unknown key was dropped without an error and every skin fell back to xterm's own default of `rgba(255,255,255,0.3)`: unnoticeable on the dark skins, which wanted roughly that anyway, and effectively invisible on Paper Gray, Solarized Light, Catppuccin Latte and Rosé Pine Dawn, where white at 30% over a near-white background moves a channel by about 3/255. Selecting text on those skins now highlights it, with desktop drag-select and the mobile long-press both fixed by the same rename.
## 1.28.2
### Patch Changes
+22 -15
View File
File diff suppressed because one or more lines are too long
+58 -31
View File
@@ -5,7 +5,7 @@
<h2 align="center">Mission control for AI coding agents</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Grok &bull; OMP &bull; Terminal - One Dashboard &bull; Any Device</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Grok &bull; DeepSeek &bull; OMP &bull; Terminal - One Dashboard &bull; Any Device</em>
</p>
<p align="center">
@@ -27,7 +27,7 @@
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — parallel subagent visualization" width="900">
</p>
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, DeepSeek Harness, or OMP inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
Get started in one line (macOS & Linux, Windows via WSL):
@@ -42,7 +42,7 @@ codeman web
The installer asks before every system change, and re-running the same line updates in place. Full details: [Quick Start - Installation](#quick-start---installation).
- **One dashboard, eight CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **One dashboard, nine CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, DeepSeek, or OMP](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions), with your own dashboards open as [web tabs](#more-features) beside them
- **Truly phone-friendly** - a [touch-optimized terminal](#mobile-optimized-web-ui) with instant local echo, QR login, swipe navigation, and push notifications
- **Runs while you sleep** - [idle detection + respawn cycling](#respawn-controller) and auto-resume when a subscription limit resets, for 24+ hour unattended runs
- **See your agents think** - [live floating windows](#live-agent-visualization) for every subagent and teammate, with real-time transcripts
@@ -68,7 +68,7 @@ This installs Node.js, tmux and a build toolchain if missing (node-pty ships no
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the nine is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the nine is present; if none is found, it offers to install any of them from a menu (DeepSeek excepted, since its npm package installs only a launcher with no runnable profile), or you can skip and install one yourself later. After install:
```bash
codeman web
@@ -82,7 +82,7 @@ codeman users add alice --admin # create the first admin account
codeman web --multiuser # named logins + per-user case spaces
```
**Prefer Docker Compose?** A local-image Compose deployment ships in `docker/`: copy `docker/.env.example` to `docker/.env`, set `CODEMAN_PASSWORD`, then run `bash docker/Start-Codeman.sh` on Linux. Codeman runs in a container and spawns Docker cases as sibling containers through the host socket. See the [Docker deployment guide](docker/README.md) for direct Compose commands, storage and networking options.
**Prefer Docker Compose?** A local-image Compose deployment ships in `docker/`: copy `docker/.env.example` to `docker/.env`, set `CODEMAN_PASSWORD`, then run `bash docker/Start-Codeman.sh` on Linux. Codeman runs in a container and spawns Docker cases as sibling containers through the host socket. After updating, run the script again rather than a plain `docker compose up`, so the rebuilt image, refreshed volumes and entrypoint arrive together. See the [Docker deployment guide](docker/README.md) for direct Compose commands, storage and networking options.
Details in [Multi-User Mode](#multi-user-mode-opt-in) below.
@@ -209,10 +209,10 @@ The most responsive AI coding agent experience on any phone. Full xterm.js termi
<tr><td>Password typing on phone</td><td><b>QR code scan — instant auth</b></td></tr>
</table>
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accident
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accident; on Codex sessions the bar also shows `⇧←` / `⇧→` (Shift+Left / Shift+Right: edit the last queued message / return through the prompt stack)
- **Dedicated Enter button** — replays the keypress through the terminal, so text buffered by local echo is flushed first rather than stranded
- **Swipe navigation & smart keyboard handling** — swipe left/right to switch sessions; toolbar and terminal shift up when the keyboard opens (`visualViewport` API)
- **Built for phones** — safe-area insets for notch and home indicator, 44px touch targets, bottom-sheet case picker, native momentum scrolling
- **Built for phones** — safe-area insets for notch and home indicator, 44px touch targets, bottom-sheet case picker, native momentum scrolling; on a folding phone (iPhone Duo) dialogs stay clear of the hinge, and opening or closing the device is never mistaken for the keyboard
```bash
codeman web --https
@@ -255,7 +255,7 @@ Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in
| Field | What it does |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. **Add Case** creates one from scratch, links an existing folder, or clones a GitHub repo straight into one (**Clone Repo**). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, `Pi`, `Grok`, `OMP`, or `Terminal` (plain shell). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, `Pi`, `Grok`, `DeepSeek`, `OMP`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Models → New Claude sessions). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
@@ -263,7 +263,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
### 3. Read the dashboard
- **Tabs (top)** — one per session. `Alt+1`-`9` to jump, `Ctrl+Tab` for next, drag to reorder (tab order syncs across your devices).
- **Tabs (top)** — one per session. `Alt+1`-`9` to jump, `Ctrl+Tab` for next, drag to reorder (tab order syncs across your devices). Prefer a list? **App Settings → Appearance → Tabs** moves it into a left sidebar with a filter box (`Alt+B` collapses it) or a vertical rail whose rows sort by activity: blocked on you first, then longest running, then most recently quiet.
- **Terminal (center)** — a real `xterm.js` terminal; full TUIs render correctly. Type directly and press **Enter** to send. `Shift+Enter` inserts a newline.
- **Side panels** — Respawn, Orchestrator, Cron, Subagents, Settings (toggled from the toolbar).
@@ -271,8 +271,10 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
- **Type prompts** straight into the terminal — input is delivered exactly-once even across reconnects (a dropped link never loses or double-sends a prompt).
- **Paste or drag-and-drop images** directly into the session.
- **Voice input** — `Ctrl+Shift+V` (Deepgram Nova-3, with auto-silence stop).
- **Attachments** — register external files/docs and preview Office/PDF inline.
- **Voice input** — `Ctrl+Shift+V` (Deepgram Nova-3, or this machine's Claude Code login with no API key; auto-silence stop).
- **Attachments** — register external files/docs and preview Office/PDF inline; any file path an agent prints is clickable, in the terminal and in the chat view.
- **When it needs you** — the tab turns yellow (waiting for input) or red (a question is blocking). The **Approvals Inbox** _(opt-in)_ queues every pending prompt across sessions, answerable from the header bell or the phone home screen, and 🧠 **Read My Mind** _(opt-in)_ drafts your next prompt from the case's goals and recent work.
- **Copy what you see** — `Shift+drag` selects text even while the CLI owns the mouse, right-click copies it, and Auto Copy _(opt-in)_ copies a selection the moment you release it.
### 5. Make it autonomous
@@ -291,7 +293,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
### 7. Operate & maintain
- **App Settings** — model, effort, permission startup mode, theme/skin, notifications, display toggles, per-CLI options, a synced custom display name, and per-device English/Simplified Chinese UI language.
- **App Settings** — model, effort, permission startup mode, theme/skin, terminal font family and weight, entrance animations, notifications, display toggles, per-CLI options, a synced custom display name, and per-device English/Simplified Chinese UI language.
- **Run it in the background** — `codeman web -d` detaches from your shell (`--status`, `--stop`); `codeman service install` makes it a systemd user unit / macOS LaunchAgent that survives reboots. Both verify the server actually answers before reporting success, and both refuse to start a second server on one data dir. See [Keep it running in the background](#quick-start---installation).
- **Self-update** — git-clone installs update in place from **App Settings → System → Updates**.
- **Deploy your own changes** — see [Development](#development).
@@ -439,16 +441,21 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → System → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Clone a GitHub repo as a case** — paste a repository URL into **Add Case → Clone Repo** and Codeman clones it into `~/codeman-cases/<name>` and registers it as a normal case, ready to run an agent in. It preflights the URL while you type (tells you whether it can be cloned anonymously and offers the repo's real branches and tags for the optional branch/tag field), fills the case name in from the URL, and lets you pick which CLI the Run button should use. Public repositories over `https://`; Codeman never collects or stores credentials
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, **Pi**, **Grok**, or **OMP** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*` vs `GROK_*`/`XAI_*` vs `OMP_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md), [`docs/pi-integration.md`](docs/pi-integration.md), [`docs/grok-integration.md`](docs/grok-integration.md) and [`docs/omp-integration.md`](docs/omp-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, **Pi**, **Grok**, **DeepSeek Harness**, or **OMP** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*` vs `GROK_*`/`XAI_*` vs `DSH_*`/`DEEPSEEK_*` vs `OMP_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md), [`docs/pi-integration.md`](docs/pi-integration.md), [`docs/grok-integration.md`](docs/grok-integration.md), [`docs/deepseek-integration.md`](docs/deepseek-integration.md) and [`docs/omp-integration.md`](docs/omp-integration.md)
- **Custom model endpoints** _(new in 1.29.0, HTTP API for now)_ — point a session's CLI at any OpenAI-compatible endpoint instead of its native backend: a local llama.cpp, llama-swap, Ollama or vLLM box, or a cloud gateway such as Azure AI Foundry or OpenRouter. Save an endpoint once (`POST /api/model-endpoints`; its models are discovered from `/v1/models`), apply it to a session (`POST /api/sessions/:id/custom-model`), and the CLI restarts in place on that endpoint. Verified live for Claude, OpenCode, Pi, Grok and OMP; Codex, Gemini and DeepSeek have documented gaps, Antigravity has no mechanism. A toolbar picker is the follow-up. See [`docs/custom-model-endpoints.md`](docs/custom-model-endpoints.md)
- **Web tabs** — open Grafana, Uptime Kuma, a Vite dev server or any dashboard URL as a tab beside your sessions (Run dropdown → **Web / URL** → **Add URL**). Dashboards are proxied through Codeman's own origin, so an `http://` target works from a phone over HTTPS and through the tunnel, single-page apps route on their own paths, and a frame that reloads recovers itself. A `localhost` link an agent prints opens as a web tab automatically. See [`docs/web-tabs.md`](docs/web-tabs.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container, or attach a case to a container you already run; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host; file previews and downloads come over the same ssh connection. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
- **Voice input** — dictate prompts with Deepgram Nova-3 (Web Speech API fallback): toggle recording, auto-silence stop, live level meter (`Ctrl+Shift+V`)
- **Voice input** — dictate prompts with Deepgram Nova-3, or through this machine's Claude Code login with no API key at all (App Settings → Voice; Web Speech API fallback): toggle recording, auto-silence stop, live level meter (`Ctrl+Shift+V`)
- **Image input** — paste or drag-and-drop images straight into a session
- **Gesture control** _(opt-in)_ — a MediaPipe hand-tracking overlay to grab/drag session windows and pinch buttons, hands-free. Enable with `CODEMAN_GESTURE=1` + App Settings → Terminal & Input
- **Multi-monitor span** _(macOS)_ — one click opens a browser window maximized across all displays, so floating agent/gesture panels can cross the physical seam
- **File Viewer button** _(opt-in)_ — a header button that toggles the built-in file browser panel with one tap; enable under App Settings → Header & Panels → Header buttons
- **CJK / IME input** — full composition support for Chinese / Japanese / Korean
- **CJK / IME input** — full composition support for Chinese / Japanese / Korean, with Ctrl- and Alt-modified navigation keys passed through to the CLI
- **Plan usage in the header** — live Claude subscription usage (the 5-hour and weekly windows) from a statusline exporter Codeman hands to `claude` at spawn and never writes into your settings files, plus Codex limits from its own app-server; per device, on for desktops and off for phones
- **Session list, your way** — the header strip, a left sidebar with a filter box, or a vertical rail whose detailed rows carry created and state stamps and sort by activity; the phone home screen and the desktop home rail use the same order
- **Terminal looks** — seven skins, four of them light, per-device font family and weight (the bundled JetBrains Mono covers weights 100 to 800), and opt-in entrance animations for tabs, agent windows, the terminal pane and connection lines
- **OS notifications & hostname-aware titles** — desktop alerts and tab titles are prefixed `codeman:<host>` so multi-host setups stay unambiguous
---
@@ -461,8 +468,9 @@ Run a case inside its own hardened Docker container instead of directly on your
- **Resource templates** — expand the checkbox for a **Small / Medium / Large / GPU** preset (memory, CPUs, GPU), or set your own. **Disk is elastic** — storage grows as data flows in, no fixed cap.
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket; a **sealed** profile (no host credentials, network off) is one toggle away.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / Pi logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / OMP logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / Pi / Grok / OMP logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Attach to a container you already run** — tick **Attach to an existing container** on the Docker panel to link a case to it instead of creating one. Codeman only `exec`s into it and never starts, stops, restarts or removes it; one adopted container can back several cases at different directories, and **copy an existing case** pre-fills the form from a sibling. Admin-only in multi-user mode, since the container's mounts belong to whoever started it.
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
Prerequisite: just Docker (or Podman). The agent base image builds itself automatically on first use, with progress streamed to the UI (or pre-build it with `node scripts/build-agent-image.mjs`). Full guide: [`docs/docker-cases.md`](docs/docker-cases.md).
@@ -478,6 +486,7 @@ Point a case at another machine and run the agent **there**, over SSH, with the
- **Discover & attach**: list the `codeman-*` sessions already running on a host (started by that machine's own Codeman, or by another operator) and attach to one. Attached sessions you don't own **detach on tab close, never kill**.
- **Shared sessions**: several clients can attach the same remote session at different window sizes without clamping each other; discovery shows a "shared" badge with the client count.
- **Injection-safe**: every ssh command line flows through a single shell-escaping builder, and host/path/identity fields are schema-guarded.
- **Files too**: previews, downloads and text reads in a remote case go over the same ssh connection (one `realpath` + `stat` probe, then a streamed `cat`, `Range` seeking included), so a clicked path opens the file on the machine the agent is on. Nothing is copied to the Codeman host; editing and Office previews answer a clear 400 instead of a misleading 404.
Set it up under **New Case → Remote** (host, user, identity file, optional jump host). Full design: [`docs/remote-sessions.md`](docs/remote-sessions.md).
@@ -647,8 +656,8 @@ These run for **every** request — before auth, even on the default no-password
### Input, files & headers
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` / `PI_*` env-prefix allowlist gates which settings each CLI can receive
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` / `PI_*` / `GROK_*` / `XAI_*` / `DSH_*` / `DEEPSEEK_*` / `OMP_*` env-prefix allowlist gates which settings each CLI can receive, and the keys that could redirect a CLI's traffic (base URLs, config homes) are clamped for non-admin users
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 2 GB raw & download (`CODEMAN_MAX_DOWNLOAD_BYTES`; bodies stream and answer `Range` requests, so the cap is a sanity bound rather than memory protection); `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
- **Security headers** — `Content-Security-Policy` (`default-src 'self'`, every exception enumerated), `X-Content-Type-Options: nosniff`, `X-Frame-Options: SAMEORIGIN`, HSTS over HTTPS, and CORS reflected **only** for `localhost` / `127.0.0.1` / `::1`
### Supply chain & isolation
@@ -698,6 +707,10 @@ The web UI remains the primary surface; see **[docs/tui.md](docs/tui.md)** for t
| `Ctrl/Cmd +` / `-` | Font size |
| `Ctrl/Cmd+?` | Keyboard help |
| `Shift+Enter` | Insert newline (sent to terminal) |
| `Shift+drag` | Select text in a pane whose mouse events go to the CLI |
| Right-click | Copy the selection (the native menu stays when nothing is selected) |
| `Shift+Wheel` | Scroll the local scrollback while the wheel is forwarded to the CLI |
| `Ctrl+Z` | Swallowed in agent sessions so a running CLI cannot be suspended; normal job control in a shell |
| `Escape` | Close panels & modals |
---
@@ -762,7 +775,7 @@ Those `DONE_<task>_<random>` strings are the skill's **split marker** trick, and
| --------------------------------------------------------------------- | --------------------------------------------------------------------------------------------- |
| [`SKILL.md`](skills/codeman/SKILL.md) | Safety rules, the ready-made fast path (spawn N workers, task them, collect), and the verb index. Always loaded. |
| [`reference/verbs.md`](skills/codeman/reference/verbs.md) | The 14 verbs in detail: readiness, send-and-wait, markers, interrupts, cleanup. On demand. |
| [`reference/recipes.md`](skills/codeman/reference/recipes.md) | 6 worked multi-worker flows (fan-out, blocked-worker watch, messaging fan-out). On demand. |
| [`reference/recipes.md`](skills/codeman/reference/recipes.md) | 8 worked flows: claude, DeepSeek Harness and shell workers, fan-out, blocked-worker watch, messaging fan-out. On demand. |
| [`reference/endpoints.md`](skills/codeman/reference/endpoints.md) | Full endpoint tables, error codes, per-mode signal table, capacity limits. On demand. |
| [`reference/messaging.md`](skills/codeman/reference/messaging.md) | Talking to claude workers directly via Claude Code cross-session messaging. On demand. |
@@ -798,8 +811,8 @@ When a CLI runs in a Codeman-managed session, these environment variables are se
4. **Response envelope.** Most endpoints return `{ "success": true, "data": … }` (errors: `{ "success": false, "error", "errorCode" }`). A few legacy GETs return bare bodies — **handle both** (`body.data ?? body`).
5. **`/api/v1/*`** is a stable alias of `/api/*`.
6. **Wait instead of polling, and don't treat a timeout as an error.** The wait endpoints answer with HTTP `200` and `wait.timedOut: true` when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. `wait.timeoutMs` tells you the timeout the server actually applied after clamping (600s ceiling).
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/pi) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/omp) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
7. **Only `claude` and `deepseek` sessions emit `stop` and `blocked`.** Those two come from hooks (Claude Code's own, and the DeepSeek Harness status bridge); `shell` and the other external CLIs (opencode/codex/gemini/antigravity/pi/grok/omp) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
### Recipes
@@ -866,9 +879,20 @@ curl -sG "$API/api/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
# 5. Read the terminal back. ⚠️ Use terminal?tail=, NOT /output: the latter's
# textOutput is empty for every tmux-backed (i.e. every interactive) session.
# tail counts BYTES, and what comes back is terminal data, ANSI included.
# 5. Read the answer. claude / codex / deepseek sessions have last-response: it comes
# from the transcript, not the screen, so no TUI frames or repaint noise.
# ⚠️ Poll rather than read once: the transcript lands slightly after the stop
# signal, so a read right after send-and-wait returns often comes back empty.
for _ in $(seq 1 10); do
TXT=$(curl -s "$API/api/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
# 5b. Other modes (shell/opencode/gemini/antigravity/pi/grok/omp) have no transcript:
# read the terminal. ⚠️ Use terminal?tail=, NOT /output: the latter's textOutput
# is empty for every tmux-backed (i.e. every interactive) session. tail counts
# BYTES, and what comes back is terminal data, ANSI included.
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
# 6. Stream live events (session output, agent activity, status)
@@ -914,7 +938,7 @@ Codeman registers Claude Code hooks that `POST /api/hook-event` (`permission_pro
## API
REST over Fastify — **~200 handlers across 21 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
REST over Fastify — **~230 handlers across 25 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
### Sessions
@@ -925,11 +949,13 @@ REST over Fastify — **~200 handlers across 21 route modules**, plus an SSE str
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?, wait?, waitTimeout?}`: `clientId`+`seq` = exactly-once; `wait` blocks until the turn ends) |
| `GET` | `/api/sessions/:id/terminal` | Read terminal output (`?tail=<bytes>`, `?full=1`); the read path for interactive sessions |
| `GET` | `/api/sessions/:id/output` | Parsed one-shot output (`textOutput` is empty for tmux-backed sessions) |
| `GET` | `/api/sessions/:id/last-response` | The last answer as clean text, read from the transcript (claude, codex, deepseek) |
| `GET` | `/api/sessions/:id/wait` | Block until a signal fires (`?until=stop,idle,exit&timeout=&fresh=`); a timeout is a `200` |
| `GET` | `/api/sessions/:id/wait-output` | Block until a literal string appears (`?match=&nocase=&from=now\|buffer&timeout=`) |
| `GET` | `/api/sessions/unified` | Unified live + history list (Session Manager) — `?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | Pin/unpin in the Session Manager (`{pinned}`) |
| `PUT` | `/api/session-order` | Sync tab order across devices (`{order: [ids]}`) |
| `POST` | `/api/sessions/:id/custom-model` | Restart the session's CLI on a saved custom endpoint (`{endpointId, modelId}`; `{clear: true}` returns to the native backend) |
| `DELETE` | `/api/sessions/:id` | Delete session |
### Respawn
@@ -978,6 +1004,7 @@ REST over Fastify — **~200 handlers across 21 route modules**, plus an SSE str
| `GET` | `/api/system/update/check` | Check for a new release |
| `POST` | `/api/system/update` | Self-update (git-clone installs) |
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
| `GET` / `POST` | `/api/model-endpoints` | List / save custom OpenAI-compatible endpoints (`PUT` / `DELETE` `/:id`; admin-only in multi-user mode) |
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
> **Building something on top of Codeman?** [`docs/extending-codeman.md`](docs/extending-codeman.md) is the integration guide: render your own UI as a tab, subscribe to the SSE event stream to react when an agent needs you, drive Codeman from a script, and the traps worth knowing before you start. Codeman has no plugin runtime on purpose, so an integration is just your own process talking HTTP.
@@ -1014,8 +1041,8 @@ flowchart TB
end
subgraph External["External"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / OMP</small>"] BG["Background Agents<br/><small>(Task tool)</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi / Grok / DeepSeek / OMP</small>"]
BG["Background Agents<br/><small>(Task tool)</small>"]
end
end
@@ -1081,7 +1108,7 @@ Full details: [`docs/archive/code-structure-findings.md`](docs/archive/code-stru
[![npm](https://img.shields.io/npm/v/xterm-zerolag-input?style=flat-square&color=22c55e)](https://www.npmjs.com/package/xterm-zerolag-input)
Instant keystroke feedback overlay for xterm.js. Eliminates perceived input latency over high-RTT connections by rendering typed characters immediately as a pixel-perfect DOM overlay. Zero dependencies, 6.1 kB gzipped, configurable prompt detection, CJK/emoji wide-character support, full state machine with 175 tests.
Instant keystroke feedback overlay for xterm.js. Eliminates perceived input latency over high-RTT connections by rendering typed characters immediately as a pixel-perfect DOM overlay. Zero dependencies, 6.1 kB gzipped, configurable prompt detection, CJK/emoji wide-character support, full state machine with 238 tests.
```bash
npm install xterm-zerolag-input
+207 -51
View File
@@ -5,7 +5,7 @@
<h2 align="center">AI 编程智能体的任务控制中心</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Grok &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Grok &bull; DeepSeek &bull; OMP &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
</p>
<p align="center">
@@ -17,6 +17,8 @@
<a href="https://nodejs.org/"><img src="https://img.shields.io/badge/Node.js-22%2B-22c55e?style=flat-square&logo=node.js&logoColor=white" alt="Node.js 22+"></a>
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/TypeScript-5.9-3b82f6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript 5.9"></a>
<a href="https://fastify.dev/"><img src="https://img.shields.io/badge/Fastify-5.x-1e3a5f?style=flat-square&logo=fastify&logoColor=white" alt="Fastify"></a>
<a href="https://www.npmjs.com/package/aicodeman"><img src="https://img.shields.io/npm/v/aicodeman?style=flat-square&label=npm&color=22c55e" alt="npm version"></a>
<a href="https://github.com/Ark0N/Codeman/stargazers"><img src="https://img.shields.io/github/stars/Ark0N/Codeman?style=flat-square&color=eab308" alt="GitHub stars"></a>
<a href="https://github.com/Ark0N/Codeman/graphs/contributors"><img src="https://img.shields.io/github/contributors/Ark0N/Codeman?style=flat-square&color=3b82f6" alt="Contributors"></a>
<a href="https://github.com/Ark0N/Codeman/commits/master"><img src="https://img.shields.io/github/commit-activity/t/Ark0N/Codeman?style=flat-square&color=1e3a5f" alt="Total commits"></a>
</p>
@@ -25,12 +27,10 @@
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — 并行子智能体可视化" width="900">
</p>
<p align="center">
<img src="docs/images/codeman-tour-20260724.png" alt="Codeman 仪表盘导览:按项目分组的会话标签页、一键 Run 启动新智能体、页头实时用量" width="900">
</p>
> 本文档由英文版 [`README.md`](README.md) 翻译而来。如有出入,以英文版为准。
**Codeman** 是一个自托管的 AI 编程智能体任务控制中心。它在持久化的 tmux 会话里拉起 Claude Code、OpenCode、Codex、Antigravity、Gemini、Pi、Grok、DeepSeek Harness 或 OMP,把真实的终端流式传到任意浏览器,并在你离开之后让智能体继续干活:空闲时重新提示、用量限额重置后自动续跑、按计划执行任务,还能实时展示每一个后台智能体的工作。
一行命令即可安装(macOS 和 Linux,Windows 通过 WSL):
```bash
@@ -44,6 +44,17 @@ codeman web
安装器在每次系统改动前都会先询问;重跑同一条命令即可原地更新。详见[快速开始 — 安装](#快速开始--安装)。
- **一个仪表盘,九个 CLI**:每个会话可选 [Claude Code、OpenCode、Codex、Antigravity、Gemini、Pi、Grok、DeepSeek 或 OMP](#更多特性)(外加普通 shell),在本机、[Docker 容器](#隔离的-docker-会话)或 [SSH 远程主机](#远程-ssh-会话)上运行,你自己的仪表盘也能作为 [Web 标签页](#更多特性)并排打开
- **真正的手机友好**:[触控优化的终端](#移动端优化的-web-ui),即时本地回显、二维码登录、滑动导航与推送通知
- **睡觉时也在跑**:[空闲检测 + 重生循环](#重生控制器respawn-controller),订阅限额重置后自动续跑,支持 24 小时以上的无人值守运行
- **看见智能体在想什么**:每个子智能体和团队成员都有[实时浮动窗口](#实时智能体可视化),附带实时活动记录
- **什么都不会丢**:tmux 让会话挺过重启和断网,输入精确一次送达,完整的回滚缓冲区回放
- **自托管、私有**:默认仅环回、MIT 许可、无遥测,完全运行在你自己的机器上
<p align="center">
<img src="docs/images/codeman-tour-20260724.png" alt="Codeman 仪表盘导览:按项目分组的会话标签页、一键 Run 启动新智能体、页头实时用量" width="900">
</p>
---
## 快速开始 — 安装
@@ -52,13 +63,14 @@ codeman web
curl -fsSL https://getcodeman.com/install | bash
```
该脚本会在缺失时自动安装 Node.js 和 tmux,把 Codeman 克隆到 `~/.codeman/app` 并完成构建。几点须知:
该脚本会在缺失时自动安装 Node.js、tmux 和一套构建工具链(node-pty 没有 Linux 预编译包,需要从源码编译),把 Codeman 克隆到 `~/.codeman/app` 并完成构建。几点须知:
- **先询问,后改动。** 所有系统级改动(安装软件包、下载 AI CLI)都会先征求确认;结束时的菜单可选择:直接在本终端运行、安装为后台服务(systemd/launchd,开机自启),或暂不启动。不选就不会有任何后台进程。
- **怎么访问,由你决定。** 安装器提供三种到达仪表盘的方式:**Tailscale**(环回绑定,由 `tailscale serve` 代理,得到带真实证书的 `https://<机器名>.<tailnet>.ts.net`,用你的 tailnet 当登录,无需密码)、**局域网内任意设备**(`0.0.0.0`,会提示设置一个强烈推荐的密码),或**仅本机**(`127.0.0.1`,最安全)。绑定网络却跳过密码需要显式确认,并以醒目警告收尾。高亮的默认项反映机器上已有的状态(已在用 Tailscale 时默认 Tailscale,重跑时沿用现有绑定),直接回车绝不会引入新软件。手动运行的 `codeman web` 仍默认仅环回。
- **重跑即更新。** 再次运行同一条命令即可原地更新已完成的安装:`~/.codeman/app` 中的本地改动会被 stash(绝不丢弃),运行中的服务会自动重启并校验。若首次安装中途失败,重跑会继续完成完整的安装流程。也可以使用 `install.sh update` 与 `install.sh uninstall`。
- **CI / 无终端环境:** 没有终端时,涉及系统改动的步骤会带着说明中止,而不是静默执行;在自动化场景设置 `CODEMAN_NONINTERACTIVE=1` 即可批准这些步骤。
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli)、[Pi](https://pi.dev)、[Grok Build](https://github.com/xai-org/grok-build)、[DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) 或 [OMP](https://github.com/can1357/oh-my-pi)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这九个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli)、[Pi](https://pi.dev)、[Grok Build](https://github.com/xai-org/grok-build)、[DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) 或 [OMP](https://github.com/can1357/oh-my-pi)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这九个中已安装的任意一个;若一个都没有,会给出一个菜单让你安装其中任意一个(DeepSeek 除外,它的 npm 包只装一个启动器,没有可运行的 profile),也可以选择跳过、稍后自行安装。安装完成后:
```bash
codeman web
@@ -72,12 +84,34 @@ codeman users add alice --admin # 创建第一个管理员账号
codeman web --multiuser # 命名登录 + 按用户隔离的案例空间
```
**更喜欢 Docker Compose?** `docker/` 里附带一套本地镜像的 Compose 部署:把 `docker/.env.example` 复制为 `docker/.env`,设置 `CODEMAN_PASSWORD`,然后在 Linux 上运行 `bash docker/Start-Codeman.sh`。Codeman 自己跑在容器里,并通过宿主机的 socket 把 Docker 案例作为并列容器拉起。更新之后请再跑一次这个脚本,而不是直接 `docker compose up`,这样重建的镜像、刷新的卷和新的入口脚本会一起就位。直接的 Compose 命令、存储与网络选项见 [Docker 部署指南](docker/README.md)(英文)。
详见下文[多用户模式](#多用户模式可选启用)。
<details>
<summary><strong>作为后台服务运行</strong></summary>
<summary><strong>让它在后台一直运行</strong></summary>
安装器结尾的菜单(选项 2)可以帮你完成这一步,并在宣告成功前校验服务确实已启动。如需手动配置:
想让它活过你启动它的那个 shell,而且什么都不用配置:
```bash
codeman web -d # 脱离终端;日志写到 ~/.codeman/web.log
codeman web --status # 是否在运行,pid 是多少
codeman web --stop # 优雅的 SIGTERM;智能体继续留在 tmux 里运行
```
`-d` 会等到服务器真正应答后才报告成功,并且拒绝在同一个数据目录上启动第二个(两个服务器共用一个 tmux socket 会互相附着对方的会话)。
想让它在重启后自动回来,就装成服务。安装器结尾的菜单(选项 2)会替你完成;`codeman service` 是 `npm i -g aicodeman` 安装的等价物:
```bash
codeman service install # systemd 用户单元(Linux)或 LaunchAgent(macOS)
codeman service status
codeman service uninstall
```
`service install` 会把你当前的 PATH 写进单元文件,这比听起来重要得多:launchd 只给任务 `/usr/bin:/bin:/usr/sbin:/sbin`,所以手写的 plist 根本找不到 Homebrew 或 nvm 装的 `node`、`tmux` 或 `claude`。它绝不会把 `CODEMAN_PASSWORD` 复制进单元文件;服务需要认证的话请自行添加。
如需手动编写单元文件:
**Linux(systemd):**
@@ -177,17 +211,17 @@ Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.
<tr><td>在手机上手打密码</td><td><b>扫二维码 —— 即时认证</b></td></tr>
</table>
- **键盘配件栏** —— 在虚拟键盘上方提供 `/init`、`/clear`、`/compact` 快捷按钮;破坏性命令需双击确认,绝不误触
- **键盘配件栏** —— 在虚拟键盘上方提供 `/init`、`/clear`、`/compact` 快捷按钮;破坏性命令需双击确认,绝不误触;在 Codex 会话上还会显示 `⇧←` / `⇧→`(Shift+Left / Shift+Right:编辑上一条排队的消息 / 在提示栈里回退)
- **独立的 Enter 按钮** —— 以按键方式回放,先冲刷本地回显缓冲的文本,不会让内容滞留在屏幕上
- **滑动导航与智能键盘处理** —— 左右滑动切换会话;键盘弹出时工具栏与终端整体上移(`visualViewport` API)
- **为手机而生** —— 刘海与 Home 指示条的安全区适配、44px 触控目标、底部抽屉式 case 选择器、原生惯性滚动
- **为手机而生** —— 刘海与 Home 指示条的安全区适配、44px 触控目标、底部抽屉式 case 选择器、原生惯性滚动;折叠屏手机(iPhone Duo)上对话框会避开铰链,开合设备也绝不会被误判成键盘弹出
```bash
codeman web --https
# 在手机上打开:https://<你的IP>:3000
```
> `localhost` 走纯 HTTP 即可。从其他设备访问时请使用 `--https`,或使用 [Tailscale](https://tailscale.com/)(推荐)—— 它提供私有网络,让你无需 TLS 证书即可从手机访问 `http://<tailscale-ip>:3000`。
> `localhost` 走纯 HTTP 即可。从其他设备访问时请使用 `--https`,或使用 [Tailscale](https://tailscale.com/)(推荐):安装器可以替你配好(在网络访问提示处选择 **Tailscale**,或在已有安装上运行 `bash ~/.codeman/app/install.sh tailscale`)。这样你会得到带真实证书的 `https://<你的机器>.<tailnet>.ts.net`:只对你的 tailnet 可见、无需密码,手机上的 PWA 安装和推送通知也都能用。
### 安全的二维码认证
@@ -210,6 +244,8 @@ codeman web # localhost:3000(仅环回 —— 安全默
codeman web --port 8080 # 自定义端口(或设置 CODEMAN_PORT)
codeman web --https # 自签名 TLS(仅远程访问时需要)
codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_PASSWORD(见「安全」)
codeman web -d # 脱离终端:关掉 shell 也在跑(--status、--stop)
codeman service install # systemd/launchd 服务:重启后自动回来
```
打开打印出的 URL。整个页面是一个单一仪表盘;下面的一切都在这里完成。
@@ -220,16 +256,16 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
| 字段 | 作用 |
| ---------------------- | ------------------------------------------------------------------------------------------- |
| **工作目录 / case** | 智能体操作的文件夹。「case」就是一个 Codeman 记住的命名工作目录。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini`、`Pi`、`Grok` 或 `Terminal`(普通 shell)。 |
| **模型** | 每会话模型(App Settings → Claude Model)。软默认值 —— 会话内 `/model` 依然有效。 |
| **工作目录 / case** | 智能体操作的文件夹。「case」就是一个 Codeman 记住的命名工作目录。**Add Case** 可以从零创建、链接一个已有文件夹,或把一个 GitHub 仓库直接克隆成 case(**Clone Repo**)。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini`、`Pi`、`Grok`、`DeepSeek`、`OMP` 或 `Terminal`(普通 shell)。 |
| **模型** | 每会话模型(App Settings → Models → New Claude sessions)。软默认值 —— 会话内 `/model` 依然有效。 |
| **Effort / Ultracode** | 推理力度(`low`–`max`),或用 `ultracode` 开启动态多智能体工作流。随时可用 `/effort` 切换。 |
点击启动 —— Codeman 通过真实 PTY 拉起 CLI,并经 SSE 流式传输到你的浏览器。
### 3. 读懂仪表盘
- **标签(顶部)** —— 每个会话一个。`Alt+1`–`9` 跳转,`Ctrl+Tab` 下一个,拖拽排序(标签顺序会跨设备同步)。
- **标签(顶部)** —— 每个会话一个。`Alt+1`–`9` 跳转,`Ctrl+Tab` 下一个,拖拽排序(标签顺序会跨设备同步)。更喜欢列表?**App Settings → Appearance → Tabs** 可以把它挪进左侧边栏(带筛选框,`Alt+B` 折叠)或一条竖向导轨,导轨的行按活动状态排序:先是等你处理的,然后是跑得最久的,最后是刚刚安静下来的。
- **终端(中央)** —— 真实的 `xterm.js` 终端;完整 TUI 正常渲染。直接输入并按 **Enter** 发送。`Shift+Enter` 插入换行。
- **侧边面板** —— Respawn、Orchestrator、Cron、Subagents、Settings(从工具栏切换)。
@@ -237,8 +273,10 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
- **直接在终端输入提示** —— 即使跨越重连,输入也是精确一次送达(连接中断绝不会丢失或重复发送提示)。
- **粘贴或拖放图片**,直接进入会话。
- **语音输入** —— `Ctrl+Shift+V`(Deepgram Nova-3,自动静音停止)。
- **附件** —— 注册外部文件/文档,并内联预览 Office/PDF。
- **语音输入** —— `Ctrl+Shift+V`(Deepgram Nova-3,或者直接用这台机器的 Claude Code 登录、不需要任何 API key;自动静音停止)。
- **附件** —— 注册外部文件/文档,并内联预览 Office/PDF;智能体打印出的任何文件路径都可以点击,终端里和对话视图里都行。
- **需要你的时候** —— 标签会变黄(等待输入)或变红(有个问题挡住了它)。**审批收件箱(Approvals Inbox)**(可选启用)把所有会话里等着你的提示排成一个队列,可以从页头的铃铛或手机首页直接作答;🧠 **Read My Mind**(可选启用)会根据这个 case 的目标和最近的工作替你起草下一条提示。
- **看到什么就能复制什么** —— `Shift+拖动` 在 CLI 接管了鼠标时也能选中文本,右键复制选中内容,自动复制(Auto Copy,可选启用)在松开鼠标的瞬间就复制。
### 5. 让它自主运行
@@ -246,7 +284,7 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
| ---------------- | --------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ |
| **Respawn** | 长时间无人值守运行 —— 空闲/限额时自动重启 CLI,带自适应时序。预设:`solo-work`、`overnight-autonomous` 等 | Respawn 标签页 |
| **Orchestrator** | 把一个目标变成分阶段计划,并跨多个智能体推动完成。 | 编排器面板 |
| **Cron** | 已保存的、命名的定时任务(`once`/`interval`/`daily`/`weekly`),到期时拉起会话并发送提示。 | ⏰ Cron 按钮(可选启用:App Settings → Display → Header Displays) |
| **Cron** | 已保存的、命名的定时任务(`once`/`interval`/`daily`/`weekly`),到期时拉起会话并发送提示。 | ⏰ Cron 按钮(可选启用:App Settings → Header & Panels → Scheduling) |
| **Auto-resume** | 订阅限额重置后自动继续。 | Respawn 标签页(顶部) |
### 6. 随时随地访问
@@ -257,8 +295,9 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
### 7. 运维与维护
- **App Settings** —— 模型、effort、权限启动模式、主题/皮肤、通知、显示开关、各 CLI 的专属选项,以及跨设备同步的自定义显示名称和按设备保存的英文/简体中文界面语言。
- **自更新** —— git-clone 安装可在 **Settings → Updates** 中原地更新。
- **App Settings** —— 模型、effort、权限启动模式、主题/皮肤、终端字体与字重、入场动画、通知、显示开关、各 CLI 的专属选项,以及跨设备同步的自定义显示名称和按设备保存的英文/简体中文界面语言。
- **让它在后台运行** —— `codeman web -d` 脱离你的 shell(`--status`、`--stop`);`codeman service install` 把它装成 systemd 用户单元 / macOS LaunchAgent,重启后自动回来。两者都会先确认服务器真正应答再报告成功,也都拒绝在同一个数据目录上启动第二个服务器。见[让它在后台一直运行](#快速开始--安装)。
- **自更新** —— git-clone 安装可在 **App Settings → System → Updates** 中原地更新。
- **部署你自己的改动** —— 见[开发](#开发)。
> ⚠️ **安全提示:** 如果你正在 Codeman 受管会话*内部*工作(`echo $CODEMAN_MUX` → `1`),绝不要直接运行 `tmux kill-session` / `pkill claude` —— 请使用 Web UI 或 `./scripts/tmux-manager.sh`。
@@ -373,6 +412,14 @@ codeman web --title-hostname dev-box # codeman:dev-box(用于覆盖嘈
| **110k tokens** | 自动 `/compact` | 上下文被摘要,工作继续 |
| **140k tokens** | 自动 `/clear` | 以 `/init` 全新开始 |
### 标签提醒(Tab Alerts)
<p align="center">
<img src="docs/images/tab-alerts-glow-20260815.gif" alt="会话标签:一个普通的活动标签,旁边是黄色的等待输入标签和红色的需要决定标签,都带着呼吸式光晕" width="900">
</p>
每个标签一眼就能看出状态。运行中的会话保持绿色状态点。会话停下来等待输入时,标签变**黄**:稳定的描边、着色的背景、黄色的点,上面叠一层缓慢的呼吸光晕。当权限提示或提问**挡住**了智能体,标签变**红**,脉动更快。底色永远不会闪灭,所以哪怕只瞥一眼(或截一张图)也能读到真实状态;标签被选中时描边依然可见,页面刷新后会从服务端重新装载待处理的提醒,因此一个被挡住的会话绝不可能藏在一个看起来正常的标签后面。
### 通知
当会话需要关注时实时桌面提醒 —— `permission_prompt` 与 `elicitation_dialog` 触发关键的红色标签闪烁,`idle_prompt` 触发黄色闪烁。点击任意通知即可直接跳转到相关会话。Hook 按 case 目录自动配置。
@@ -393,17 +440,24 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
## 更多特性
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity**、**Gemini**、**Pi** 或 **Grok**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*`、`PI_*`、`GROK_*`/`XAI_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)、[`docs/pi-integration.md`](docs/pi-integration.md) 与 [`docs/grok-integration.md`](docs/grok-integration.md)
- **Docker 会话** —— 在隔离且加固的容器中运行案例。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一案例的多个会话共享一个容器;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
- **远程 SSH 会话**:把案例指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **后台守护进程与服务安装** —— `codeman web -d` 以脱离终端的方式运行服务器,带 pid 文件、`~/.codeman/web.log` 和经过校验的启动(它会轮询到服务器应答为止,所以端口冲突绝不会被当成成功);`codeman service install` 写入一个 systemd 用户单元(Linux)或 LaunchAgent(macOS),并把你 shell 的 PATH 一并写进去,这样 nvm 或 Homebrew 装的 `node`、`tmux` 和 `claude` 才真的找得到。机密永远不会写进单元文件
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → System → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
- **把 GitHub 仓库克隆成 case** —— 在 **Add Case → Clone Repo** 里粘贴一个仓库 URL,Codeman 会把它克隆到 `~/codeman-cases/<name>` 并注册为普通 case,随时可以跑智能体。输入时它会预检 URL(告诉你能否匿名克隆,并为可选的分支/标签字段提供仓库真实的分支与标签),从 URL 里填好 case 名,还让你选 Run 按钮该用哪个 CLI。支持 `https://` 的公开仓库;Codeman 绝不收集或保存凭据
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity**、**Gemini**、**Pi**、**Grok**、**DeepSeek Harness** 或 **OMP**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*`、`GEMINI_*`/`GOOGLE_*`、`PI_*`、`GROK_*`/`XAI_*`、`DSH_*`/`DEEPSEEK_*` 与 `OMP_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)、[`docs/pi-integration.md`](docs/pi-integration.md)、[`docs/grok-integration.md`](docs/grok-integration.md)、[`docs/deepseek-integration.md`](docs/deepseek-integration.md) 与 [`docs/omp-integration.md`](docs/omp-integration.md)
- **自定义模型端点**(1.29.0 新增,目前仅 HTTP API)—— 让某个会话的 CLI 指向任意 OpenAI 兼容端点,而不是它自己的官方后端:本地的 llama.cpp、llama-swap、Ollama 或 vLLM 机器,也可以是 Azure AI Foundry、OpenRouter 这类云端网关。端点只需保存一次(`POST /api/model-endpoints`,模型列表从它的 `/v1/models` 自动发现),再应用到会话(`POST /api/sessions/:id/custom-model`),CLI 就会在原地重启并接上该端点。Claude、OpenCode、Pi、Grok 与 OMP 已实测通过;Codex、Gemini 与 DeepSeek 存在已记录的缺口,Antigravity 没有可用机制。工具栏选择器是下一步。详见 [`docs/custom-model-endpoints.md`](docs/custom-model-endpoints.md)
- **Web 标签页** —— 把 Grafana、Uptime Kuma、一个 Vite 开发服务器或任何仪表盘 URL 作为标签页打开在会话旁边(Run 下拉菜单 → **Web / URL** → **Add URL**)。仪表盘通过 Codeman 自己的源代理,因此 `http://` 目标在手机上走 HTTPS 也能用、走隧道也能用;单页应用能在自己的路径上正常路由,页面自己重载后也能自行恢复。智能体打印出的 `localhost` 链接会自动以 Web 标签页打开。详见 [`docs/web-tabs.md`](docs/web-tabs.md)
- **Docker 会话** —— 在隔离且加固的容器中运行 case。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一 case 的多个会话共享一个容器,也可以把 case 挂到你已经在跑的容器上;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
- **远程 SSH 会话** —— 把 case 指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话;文件预览与下载走同一条 ssh 连接。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
- **语音输入** —— 用 Deepgram Nova-3 口述提示(带 Web Speech API 回退):切换录音、自动静音停止、实时音量表(`Ctrl+Shift+V`)
- **语音输入** —— 用 Deepgram Nova-3 口述提示,或者干脆用这台机器的 Claude Code 登录、不需要任何 API key(App Settings → Voice;带 Web Speech API 回退):切换录音、自动静音停止、实时音量表(`Ctrl+Shift+V`)
- **图像输入** —— 直接把图片粘贴或拖放进会话
- **手势控制** _(可选)_ —— 一个 MediaPipe 手部追踪叠加层,可徒手抓取/拖动会话窗口并捏合按钮。用 `CODEMAN_GESTURE=1` + App Settings → Display 启用
- **手势控制** _(可选)_ —— 一个 MediaPipe 手部追踪叠加层,可徒手抓取/拖动会话窗口并捏合按钮。用 `CODEMAN_GESTURE=1` + App Settings → Terminal & Input 启用
- **多显示器横跨** _(macOS)_ —— 一键打开一个横跨所有显示器最大化的浏览器窗口,让浮动的智能体/手势面板可以跨越物理拼接缝
- **文件查看器按钮** _(可选)_ —— 头部新增一个按钮,一键切换内置文件浏览器面板;在 App Settings → Display → Header Displays 中启用
- **CJK / 输入法支持** —— 完整支持中文 / 日文 / 韩文的组合输入
- **文件查看器按钮** _(可选)_ —— 页头新增一个按钮,一键切换内置文件浏览器面板;在 App Settings → Header & Panels → Header buttons 中启用
- **CJK / 输入法支持** —— 完整支持中文 / 日文 / 韩文的组合输入,Ctrl、Alt 修饰的导航键也会原样透传给 CLI
- **页头里的套餐用量** —— 页头实时显示 Claude 订阅用量(5 小时窗口与每周窗口),数据来自 Codeman 在拉起 `claude` 时临时交给它的 statusline 导出器,绝不会写进你的设置文件;Codex 的限额则来自它自己的 app-server。按设备生效:桌面默认开,手机默认关
- **会话列表,随你摆** —— 页头横条、带筛选框的左侧边栏,或一条竖向导轨,导轨的详细行带有创建时间与状态时长并按活动状态排序;手机首页和桌面首页导轨用的是同一套顺序
- **终端外观** —— 七套皮肤(其中四套浅色)、按设备保存的字体与字重(内置的 JetBrains Mono 覆盖 100 到 800 的字重),以及可选启用的入场动画,覆盖标签、智能体窗口、终端面板和连接线
- **操作系统通知与主机名感知标题** —— 桌面提醒与标签标题以 `codeman:<host>` 为前缀,使多主机配置不再含糊
---
@@ -416,7 +470,8 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
- **资源模板** —— 展开复选框可选 **Small / Medium / Large / GPU** 预设(内存、CPU、GPU),也可以完全自定义。**磁盘是弹性的** —— 存储随数据增长,没有固定上限。
- **按案例共享容器** —— 多个会话可以 `docker exec` 进同一个容器;结束某个会话绝不会影响其他会话所在的容器。
- **默认加固** —— 非 root、`--cap-drop ALL`、`no-new-privileges`、PID/内存上限,绝不使用 `--privileged` 或 docker socket;**密封(sealed)** 配置(不注入主机凭据、关闭网络)只需一个开关。
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Antigravity / Gemini / OpenCode / Pi 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Antigravity / Gemini / OpenCode / Pi / Grok / OMP 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
- **挂到你已经在跑的容器上** —— 在 Docker 面板勾选 **Attach to an existing container**,就能把 case 链接到一个现成容器,而不是新建一个。Codeman 只 `exec` 进去,绝不启动、停止、重启或删除它;一个被接管的容器可以在不同目录下支撑多个 case,**复制一个已有 case** 会用同一容器上的兄弟 case 预填表单。多用户模式下仅管理员可用,因为容器的挂载属于启动它的人。
- **迁移到另一台机器** —— 把容器的完整环境(工具链 + 工作区)导出为可移植的 `.tar.gz`,在另一台机器上导入到新案例即可继续。
- **持久耐用** —— Codeman 重启后重连会回到同一个存活的智能体;容器停止/重启后则从绑定挂载的转录恢复对话。
@@ -433,6 +488,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
- **发现与附着**:列出主机上已在运行的 `codeman-*` 会话(由那台机器自己的 Codeman 或其他操作者启动)并附着其一。非你所有的已附着会话在关闭标签时**只分离,绝不杀掉**。
- **共享会话**:多个客户端可以以不同窗口尺寸同时附着同一个远程会话而互不挤压;发现列表会显示带客户端计数的「shared」徽标。
- **注入安全**:所有 ssh 命令行都经由单一的 shell 转义构建器生成,主机/路径/身份文件字段均有模式校验。
- **文件也行**:远程 case 里的预览、下载和文本读取走同一条 ssh 连接(一次 `realpath` + `stat` 探测,然后流式 `cat`,支持 `Range` 拖动进度),所以点一个路径打开的就是智能体所在那台机器上的文件。什么都不会复制到 Codeman 主机;编辑和 Office 预览会明确返回 400,而不是一个误导性的 404。
在 **New Case → Remote** 中配置(主机、用户、身份文件、可选跳板机)。完整设计:[`docs/remote-sessions.md`](docs/remote-sessions.md)。
@@ -486,7 +542,7 @@ codeman users list
systemctl --user enable codeman-tunnel
loginctl enable-linger $USER
# 或通过 Codeman Web UI:Settings → Tunnel → 切换为开
# 或通过 Codeman Web UI:App Settings → System → Remote access → Cloudflare Tunnel
```
</details>
@@ -588,7 +644,7 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
- **默认仅环回** —— 绑定 `127.0.0.1`,仅可从本机访问,因此「无密码」默认配置开箱即安全。在未设置 `CODEMAN_PASSWORD` 的情况下绑定非环回主机会*启动但打印一条醒目警告*,并给出三个具体修复方案(设置密码、环回 + 一个带认证的隧道,或用 `--allow-unauthenticated-network` 显式确认)
- **可选认证,真实会话** —— 通过 `CODEMAN_USERNAME`(默认 `admin`)/ `CODEMAN_PASSWORD` 的 HTTP Basic 认证。成功后签发一个不透明的 256 位 `codeman_session` cookie(`randomBytes(32)`)—— 服务端校验,而非客户端签名,因此无法离线伪造(24h TTL、自动延长、设备上下文审计日志)
- **按 IP 速率限制** —— 失败 10 次 → `429` 并带 `Retry-After`(15 分钟衰减)。即便攻击者在同一 IP 上猛攻,有效 cookie 或正确密码也能*立即*恢复 —— 这很重要,因为所有隧道流量共享同一个环回 IP。二维码认证有自己独立的限制器
- **可配置的权限模式**:`--dangerously-skip-permissions` 只是默认值。**App Settings → Claude CLI → Startup Mode** 可以把新会话切换为 Anthropic 的分类器护栏 `auto` 模式(低打扰,需要 Claude Code 2.1.207+)、`normal` 提示模式,或一份显式的允许工具列表。多用户模式下,未获授权的用户会被强制为 `auto`,shell 会话与跳过权限需要按用户显式授权
- **可配置的权限模式**:`--dangerously-skip-permissions` 只是默认值。**App Settings → Agents & CLIs → Claude → Startup Mode** 可以把新会话切换为 Anthropic 的分类器护栏 `auto` 模式(低打扰,需要 Claude Code 2.1.207+)、`normal` 提示模式,或一份显式的允许工具列表。多用户模式下,未获授权的用户会被强制为 `auto`,shell 会话与跳过权限需要按用户显式授权
### 始终开启的浏览器加固(v0.9.5)
@@ -602,8 +658,8 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
### 输入、文件与响应头
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` / `PI_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **路径限定** —— 文件路由在边界检查前先 `realpath`(无 TOCTOU);`..`、绝对路径、以及解析到工作目录之外的符号链接都会被拒绝。上限:10 MB 文本预览 / 50 MB 原始与下载;`/api/download` 对敏感路径(`.env`、`*credentials*`、`~/.ssh/`、`.aws/credentials`)做黑名单。SVG/HTML 以 `octet-stream` + `nosniff` + attachment 提供,因此会被下载而非执行
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` / `PI_*` / `GROK_*` / `XAI_*` / `DSH_*` / `DEEPSEEK_*` / `OMP_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置,而那些能把 CLI 流量改道的键(base URL、配置目录)对非管理员用户会被钳制
- **路径限定** —— 文件路由在边界检查前先 `realpath`(无 TOCTOU);`..`、绝对路径、以及解析到工作目录之外的符号链接都会被拒绝。上限:10 MB 文本预览 / 2 GB 原始与下载(`CODEMAN_MAX_DOWNLOAD_BYTES`;响应体是流式的并支持 `Range` 请求,所以这个上限只是合理性边界,不是内存保护);`/api/download` 对敏感路径(`.env`、`*credentials*`、`~/.ssh/`、`.aws/credentials`)做黑名单。SVG/HTML 以 `octet-stream` + `nosniff` + attachment 提供,因此会被下载而非执行
- **安全响应头** —— `Content-Security-Policy`(`default-src 'self'`,每个例外都逐条列举)、`X-Content-Type-Options: nosniff`、`X-Frame-Options: SAMEORIGIN`、HTTPS 下的 HSTS,以及**仅**对 `localhost` / `127.0.0.1` / `::1` 反射的 CORS
### 供应链与隔离
@@ -615,6 +671,22 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
---
## 终端界面(`codeman tui`)
一个在终端里运行的全屏会话仪表盘。状态与 Web UI 完全一致,因为它就是同一个服务器的客户端:
```bash
codeman tui # 仪表盘
codeman tui --list # 带编号的会话列表,随即退出(可用于脚本)
codeman tui 2 # 直接附着到列表里的第 2 个会话
```
会话按 **NEEDS YOU → WORKING → IDLE → RECENT** 分组,等得最久的排最前。`↑↓`/`j`/`k` 选择,`1`-`9` 与 `[`/`]` 切换会话,`Enter` 附着进 tmux 面板(按 **`F1`** 回来)。在面板里,顶部的横条会一直显示会话条,`Alt+1`-`Alt+9` 不用离开就能切换。`y`/`n`/数字可以直接在列表里回答待处理的权限对话框,`p` 发送一行提示,`n` 新建会话并直接进入,`x` 杀掉一个(`y` 确认),`/` 搜索,`g` 显示离开摘要,`?` 是帮助,`q` 退出。窄于 72 列时它会去掉预览面板、变成单列列表,所以在手机上的 Termius 里依然好用。没有服务器在跑时,它仍会以仅附着的降级模式启动。
Web UI 仍是主要界面;完整指南见 **[docs/tui.md](docs/tui.md)**(英文)。
---
## 键盘快捷键
> Ctrl 绑定在 macOS 上也接受 Cmd。
@@ -626,15 +698,21 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
| `Ctrl/Cmd+Tab` | 下一个会话 |
| `Alt/Option+[` / `Alt/Option+]` | 上一个 / 下一个会话 |
| `Alt/Option+1`–`Alt/Option+9` | 切换到第 N 个标签(按物理键位,macOS Option 布局也适用) |
| `Alt/Option+B` | 折叠 / 展开会话侧边栏(仅侧边栏布局) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | 将当前标签左移 / 右移 |
| `Ctrl/Cmd+C` | 复制选中内容;未选中时中断代理 |
| `Ctrl+Shift+C` | 复制选中内容(永不中断) |
| `Ctrl/Cmd+V` | 粘贴,或上传剪贴板里的图片并粘贴其路径 |
| `Ctrl/Cmd+L` | 清屏 |
| `Ctrl+Shift+R` | 恢复终端尺寸 |
| `Ctrl+Shift+V` | 切换语音输入 |
| `Ctrl/Cmd +` / `-` | 字体大小 |
| `Ctrl/Cmd+?` | 键盘帮助 |
| `Shift+Enter` | 插入换行(发送到终端) |
| `Shift+拖动` | 在鼠标事件交给 CLI 的面板里选中文本 |
| 右键 | 复制选中内容(没有选中时保留原生菜单) |
| `Shift+滚轮` | 滚轮被转发给 CLI 时,滚动本地回滚缓冲区 |
| `Ctrl+Z` | 在智能体会话里被吞掉,运行中的 CLI 不会被挂起;shell 里照常是作业控制 |
| `Escape` | 关闭面板与模态框 |
---
@@ -643,16 +721,78 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
面向不经浏览器控制 Codeman 的 AI 智能体与自动化:一个拉起工作会话的智能体、一个 CI 机器人,或是**运行在 Codeman 会话*内部*、编排其他会话的 Claude Code**。UI 能做的一切都是 HTTP + CLI,因此智能体也能做。
> **捷径:装上打包好的智能体技能。** 下面这一整套(外加多工作会话的实战配方)已经作为 Claude Code 技能随仓库发布在 [`skills/codeman`](skills/codeman/SKILL.md),会话内部的智能体不必等你把文档粘进提示词就能驱动 Codeman。三种获取方式:
>
> - `npx skills add Ark0N/Codeman --skill codeman -g`:全局安装,任何支持技能的智能体都能用
> - Claude Code 插件:`/plugin marketplace add Ark0N/Codeman`,然后 `/plugin install codeman@codeman`:通过 Claude Code 自带的插件管理器全局安装,`/plugin update codeman` 跟随新版本;与 `codeman skill install` 二选一,两者都装会让技能出现两次(`codeman` 和 `codeman:codeman`)
> - `codeman skill install`(全局)或 `codeman skill install --case <name>`:给那些从 npm 安装、从未克隆过仓库的用户;`codeman skill uninstall` 可撤销
> - **App Settings → Agent Skill**(`agentSkillEnabled`,默认关闭):开启后,Codeman 会在每次于某个 case 中创建 Claude 会话时把技能注入该 case;case 里用户自己写的 `skills/codeman` 永远不会被覆盖
>
> 全局安装(`codeman skill install` 或 `npx skills add`)会被**本机每一个新建的 Claude Code 会话**读到,无论它在不在 Codeman 里。技能自带门禁:不在 Codeman 会话中(`CODEMAN_MUX` 未设置)时它拒绝动作,所以全局装上它对无关会话没有代价。
>
> ⚠️ 把 `agentSkillEnabled` 关回去**不会删掉已经注入的副本**(在创建时做清扫,会把技能从共用同一个 `.claude/` 目录的其他活动会话脚下抽走)。要删就按 case 删:`codeman skill uninstall --case <name>`。
### 智能体技能(从这里开始)
这一节的所有内容也打包成了一个 **Claude Code 技能**,位于 [`skills/codeman`](skills/codeman/SKILL.md)。装一次,就再也不用把 API 文档粘进提示词。你用大白话说想要什么,已经坐在 Codeman 会话里的智能体会自己加载配方并驱动 API。
#### 第 1 步:安装
| 方式 | 命令 | 范围 |
| ---------------- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
| Skills CLI | `npx skills add Ark0N/Codeman --skill codeman -g` | 全局,任何支持技能的智能体都能用 |
| Claude Code 插件 | `/plugin marketplace add Ark0N/Codeman`,然后 `/plugin install codeman@codeman` | 全局,通过 Claude Code 自带的插件管理器;`/plugin update codeman` 跟随新版本。与 `codeman skill install` 二选一:两者都装会让技能出现两次(`codeman` 和 `codeman:codeman`) |
| 内置 CLI | `codeman skill install` | 全局(`~/.claude/skills/codeman`),给那些从 npm 安装、从未克隆过仓库的用户 |
| 内置 CLI | `codeman skill install --case <name>` | 仅一个 case |
| Web UI | App Settings → Agents & CLIs → Claude → **Agent Skill** | 每次在某个 case 创建 Claude 会话时自动注入(`agentSkillEnabled`,跨设备同步,默认关闭) |
`codeman skill uninstall [--case <name>]` 可以撤销 CLI 安装,并且绝不会碰你自己写的 `skills/codeman`。
#### 第 2 步:开口要
整个界面就这么多。不用 curl,不用端点名,不用会话 id。下面这些提示照原样就能用:
| 你说 | 技能做的事 |
| ------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| _「现在有哪些会话在跑?」_ | 列出它们的名字、模式和状态。只读,随时可以问。 |
| _「在 `myapp` case 上起一个 shell 工作会话,跑测试套件,告诉我过没过。」_ | 拉起、等待一个拆开的完成标记、读回退出码、清理。 |
| _「起 3 个工作会话分别跑 lint、typecheck 和测试。并行跑,报告失败的。」_ | 扇出流程:每个任务一个会话,先全部启动,再逐个收集完成的。 |
| _「让一个 claude 工作会话在 `refactor-auth` 上总结 `src/session.ts`,然后关掉它。」_ | 拉起、走完就绪阶梯(包括首次运行的信任对话框)、发送并等待、读取干净的 transcript 答案、删除。 |
| _「盯着会话 w4,如果它卡在权限提示上就告诉我。」_ | 阻塞在 `blocked` 信号上,并把问题交给**你**。它绝不会替另一个会话回答提示。 |
#### 第 3 步:没有了
智能体会删掉它启动的每一个会话。你可以在仪表盘里看着标签出现又消失。
#### 一次真实的运行,从头到尾
> **你:** 起 3 个 shell 工作会话,并行跑 lint / typecheck / 前端语法检查,告诉我哪个失败了。
```text
lint -> 9f2d8e5f dispatched
typecheck -> aff9c691 dispatched 仪表盘里出现 3 个标签
syntax -> be9f1f15 dispatched
lint DONE_lint_17909 rc=0
typecheck DONE_typecheck_3409 rc=0 每完成一个就收集一个
syntax DONE_syntax_18501 rc=0
deleted 9f2d8e5f, aff9c691, be9f1f15 标签消失
```
那些 `DONE_<task>_<random>` 字符串就是技能的**拆分标记**技巧,也是扇出在没有 hook 的 `shell` 会话上依然可靠的原因:敲进去的那一行只含 `${M}_17909`,因此只有命令真正的*输出*里才会出现 `DONE_17909`。不拆开的标记会在命令还没跑之前就匹配到你自己按键的回显。
#### 盒子里有什么
| 文件 | 内容 |
| ------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- |
| [`SKILL.md`](skills/codeman/SKILL.md) | 安全规则、现成的快速路径(起 N 个工作会话、派任务、收集)和动词索引。始终加载。 |
| [`reference/verbs.md`](skills/codeman/reference/verbs.md) | 14 个动词的详细说明:就绪、发送并等待、标记、中断、清理。按需加载。 |
| [`reference/recipes.md`](skills/codeman/reference/recipes.md) | 8 个完整流程:claude、DeepSeek Harness 与 shell 工作会话、扇出、盯住被卡住的工作会话、消息扇出。按需加载。 |
| [`reference/endpoints.md`](skills/codeman/reference/endpoints.md) | 完整端点表、错误码、各模式的信号表、容量限制。按需加载。 |
| [`reference/messaging.md`](skills/codeman/reference/messaging.md) | 通过 Claude Code 跨会话消息直接和 claude 工作会话对话。按需加载。 |
里面的每一个配方都在真实服务器上验证过,注释记录的是实测出来而不是猜出来的失败模式。
#### 两件值得知道的事
- **它会自我门禁。** 不在 Codeman 会话里(`CODEMAN_MUX` 未设置)时,技能拒绝动作,也不去猜 API 地址,所以全局安装对无关的 Claude Code 会话没有任何代价。
- **它刻意保守。** 未经提示,它只会拉起会话、给它们发提示,并删除**它在同一段对话里自己创建的**会话(按精确 id,经由一个拒绝删除智能体自身会话的失败即关闭守卫)。删除 case(会抹掉一个真实的代码目录)、批量杀会话、改动 respawn/ralph/cron/orchestrator 以及写设置,都需要你开口并指名目标。
⚠️ 把 `agentSkillEnabled` 关回去**不会删掉已经注入的副本**(在创建时做清扫,会把技能从共用同一个 `.claude/` 目录的其他活动会话脚下抽走)。要删就按 case 删:`codeman skill uninstall --case <name>`。
---
**这一节余下的部分是手动路径**:同样的操作用裸 HTTP 来做,适合 CI 机器人、shell 脚本,或任何不支持技能的智能体。
### 检测自己身处 Codeman 内部
@@ -673,7 +813,7 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
4. **响应信封。** 多数端点返回 `{ "success": true, "data": … }`(错误:`{ "success": false, "error", "errorCode" }`)。少数遗留 GET 返回裸响应体 —— **两种都要处理**(`body.data ?? body`)。
5. **`/api/v1/*`** 是 `/api/*` 的稳定别名。
6. **用等待代替轮询,别把超时当成错误。** 等待类端点在没等到事情发生时也以 HTTP `200` 加 `wait.timedOut: true` 应答,所以要循环调用短等待(默认 60 秒),而不是发一个超长的调用:隧道会掐断空闲连接。`wait.timeoutMs` 告诉你服务端钳制之后真正采用的超时(上限 600 秒)。
7. **只有 `claude` 会话会发出 `stop` 与 `blocked`。** 这两个来自 Claude Code hook;`shell` 与外部 CLI(opencode/codex/gemini/antigravity/pi)只接受 `idle`、`working` 与 `exit`。在这些模式上显式索要 `stop` 会得到 `400`;不传 `until` 则永远安全。⚠️ `shell` 会话的 `idle` 只在启动时触发**一次**,此后再也不会,所以在那里用「发送并等待」只能等到超时:没有 hook 的会话请用 `wait-output` 标记来同步。
7. **只有 `claude` 与 `deepseek` 会话会发出 `stop` 与 `blocked`。** 这两个来自 hook(Claude Code 自己的,以及 DeepSeek Harness 的状态桥接);`shell` 与其他外部 CLI(opencode/codex/gemini/antigravity/pi/grok/omp)只接受 `idle`、`working` 与 `exit`。在这些模式上显式索要 `stop` 会得到 `400`;不传 `until` 则永远安全。⚠️ `shell` 会话的 `idle` 只在启动时触发**一次**,此后再也不会,所以在那里用「发送并等待」只能等到超时:没有 hook 的会话请用 `wait-output` 标记来同步。
8. **没有任何东西会报告「就绪」,得自己显式等。** 新会话在 PID 出现之前一律回答 `{"signal":"exit","immediate":true}`(意思是*还没启动*,不是*崩了*),而全新 case 里的 `claude` 工作会话接着会停在 CLI 的信任对话框上。此时给它发提示,等待会在约 2 秒后因 `idle` 解除,看上去和一个跑完的回合一模一样,而文本其实卡在对话框里。下面的配方 2b 就是避开它的顺序。
### 常用配方
@@ -738,7 +878,7 @@ curl -sG "$API/api/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
# 5. 读回答案。claude / codex 会话用 last-response:它取自 transcript 而不是屏幕,
# 5. 读回答案。claude / codex / deepseek 会话用 last-response:它取自 transcript 而不是屏幕,
# 因此不带 TUI 的画框与重画噪声。⚠️ 要轮询,别只读一次:transcript 落盘比 stop
# 信号稍晚,紧跟着「发送并等待」返回后立刻读,常常拿到空串。
for _ in $(seq 1 10); do
@@ -747,7 +887,7 @@ for _ in $(seq 1 10); do
done
printf '%s\n' "$TXT"
# 5b. 其他模式(shell/opencode/gemini/antigravity/pi)没有 transcript,读终端。
# 5b. 其他模式(shell/opencode/gemini/antigravity/pi/grok/omp)没有 transcript,读终端。
# ⚠️ 用 terminal?tail=,不要用 /output:后者的 textOutput 对每个由 tmux 承载的
# (也就是每个交互式)会话都是空的。tail 按字节计,返回的是含 ANSI 的终端数据。
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
@@ -780,7 +920,9 @@ codeman session start -d /path/to/repo # (s) 启动会话
codeman session list # 列出会话
codeman session logs <id> # 查看输出
codeman task add "fix the failing test" # (t) 排入任务
codeman attach <path> # 附着 Claude hook 上下文
codeman attach <path> # 为本地文件显示一张附件卡片
codeman tui --list # 带编号的会话列表(管道输出时为纯文本)
codeman tui 3 # 附着到该列表里的第 3 个会话
```
### Hook(事件*回流*到 Codeman)
@@ -793,7 +935,7 @@ Codeman 会注册 Claude Code hook,它们 `POST /api/hook-event`(`permission
## API
基于 Fastify 的 REST —— **21 个路由模块中约 200 个处理器**,外加一条 SSE 流和一条 WebSocket 终端通道。所有响应都使用 `ApiResponse<T>` 信封(`{success, data}` / `{success, error, errorCode}`);`/api/v1/*` 是稳定别名。以下是一个有代表性的子集:
基于 Fastify 的 REST —— **25 个路由模块中约 230 个处理器**,外加一条 SSE 流和一条 WebSocket 终端通道。所有响应都使用 `ApiResponse<T>` 信封(`{success, data}` / `{success, error, errorCode}`);`/api/v1/*` 是稳定别名。以下是一个有代表性的子集:
### 会话(Sessions)
@@ -804,11 +946,13 @@ Codeman 会注册 Claude Code hook,它们 `POST /api/hook-event`(`permission
| `POST` | `/api/sessions/:id/input` | 发送输入(`{input, useMux?, clientId?, seq?, wait?, waitTimeout?}`:`clientId`+`seq` = 精确一次;`wait` 阻塞到这一回合结束) |
| `GET` | `/api/sessions/:id/terminal` | 读取终端输出(`?tail=<bytes>`、`?full=1`):交互式会话的读取路径 |
| `GET` | `/api/sessions/:id/output` | 一次性的解析输出(tmux 承载的会话里 `textOutput` 为空) |
| `GET` | `/api/sessions/:id/last-response` | 从 transcript 读出的最后一条回答,纯文本(claude、codex、deepseek) |
| `GET` | `/api/sessions/:id/wait` | 阻塞到某个信号触发(`?until=stop,idle,exit&timeout=&fresh=`);超时是 `200` |
| `GET` | `/api/sessions/:id/wait-output` | 阻塞到某个字面串出现(`?match=&nocase=&from=now\|buffer&timeout=`) |
| `GET` | `/api/sessions/unified` | 统一的活动 + 历史清单(会话管理器):`?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | 在会话管理器中置顶 / 取消置顶(`{pinned}`) |
| `PUT` | `/api/session-order` | 跨设备同步标签顺序(`{order: [ids]}`) |
| `POST` | `/api/sessions/:id/custom-model` | 让会话的 CLI 在一个已保存的自定义端点上原地重启(`{endpointId, modelId}`;`{clear: true}` 回到官方后端) |
| `DELETE` | `/api/sessions/:id` | 删除会话 |
### 重生(Respawn)
@@ -857,6 +1001,7 @@ Codeman 会注册 Claude Code hook,它们 `POST /api/hook-event`(`permission
| `GET` | `/api/system/update/check` | 检查新发行版 |
| `POST` | `/api/system/update` | 自更新(git-clone 安装) |
| `POST` | `/api/clipboard` | 把文本推送到所有已连接浏览器(`{text}`) |
| `GET` / `POST` | `/api/model-endpoints` | 列出 / 保存自定义的 OpenAI 兼容端点(`PUT` / `DELETE` `/:id`;多用户模式下仅管理员) |
| `GET` | `/api/sessions/:id/run-summary` | 时间线 + 统计 |
> **想在 Codeman 之上做集成?**[`docs/extending-codeman.md`](docs/extending-codeman.md)(英文)是集成指南:把你自己的界面作为标签页嵌入、订阅 SSE 事件流以便在 agent 需要你时做出响应、用脚本驱动 Codeman,以及动手前值得先了解的那些坑。Codeman 刻意不提供插件运行时,所以一个集成就是你自己的进程在讲 HTTP。
@@ -893,7 +1038,7 @@ flowchart TB
end
subgraph External["外部"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi / Grok / DeepSeek / OMP</small>"]
BG["后台智能体<br/><small>(Task 工具)</small>"]
end
end
@@ -931,6 +1076,12 @@ npm test # 运行测试(与 CI 相同;浏览器/移动端
---
## 社区
提问、安装求助和想法都在 [GitHub Discussions](https://github.com/Ark0N/Codeman/discussions):[Q&A 板块](https://github.com/Ark0N/Codeman/discussions/categories/q-a)回答了最常见的那些(手机访问、通宵运行、更新),路线图则在 [Ideas](https://github.com/Ark0N/Codeman/discussions/categories/ideas) 里决定。Bug 请提到 [issues](https://github.com/Ark0N/Codeman/issues);报告通常一天内会得到回复,每个发行版都会点名感谢报告者和贡献者。想参与贡献?[CONTRIBUTING.md](.github/CONTRIBUTING.md) 是地图:皮肤、翻译和文档都是很好的第一个 PR,更大的特性先从一个 Discussion 开始。如果你对自己的配置很自豪,发到 [Show and tell](https://github.com/Ark0N/Codeman/discussions/300) 来。
---
## 代码库质量
本代码库经历了一次全面的 7 阶段重构,消除了上帝对象、集中了配置,并建立了模块化架构:
@@ -954,7 +1105,7 @@ npm test # 运行测试(与 CI 相同;浏览器/移动端
[![npm](https://img.shields.io/npm/v/xterm-zerolag-input?style=flat-square&color=22c55e)](https://www.npmjs.com/package/xterm-zerolag-input)
为 xterm.js 提供即时按键反馈的叠加层。通过把输入的字符立即渲染为像素级精准的 DOM 叠加层,消除高 RTT 连接下的感知输入延迟。零依赖、可配置的提示符检测、带 78 个测试的完整状态机。
为 xterm.js 提供即时按键反馈的叠加层。通过把输入的字符立即渲染为像素级精准的 DOM 叠加层,消除高 RTT 连接下的感知输入延迟。零依赖、gzip 后 6.1 kB、可配置的提示符检测、CJK/emoji 宽字符支持、带 238 个测试的完整状态机。
```bash
npm install xterm-zerolag-input
@@ -977,3 +1128,8 @@ MIT —— 见 [LICENSE](LICENSE)
<p align="center">
<strong>跟踪会话。可视化智能体。掌控重生。让它在你睡觉时持续运行。</strong>
</p>
<p align="center">
如果 Codeman 帮你省了时间,<a href="https://github.com/Ark0N/Codeman/stargazers">点个 star</a> 能让更多人找到它。<br>
欢迎到 <a href="https://github.com/Ark0N/Codeman/issues">Issues</a> 报告 bug 和提出特性想法。
</p>
+1
View File
@@ -27,6 +27,7 @@ export const BROWSER_TEST_GLOBS = [
'test/webgl-fallback.test.ts',
'test/terminal-copy-shortcut.test.ts',
'test/terminal-keycode229-recovery.browser.test.ts',
'test/capture-load-window.browser.test.ts',
'test/codex-predictive-echo.test.ts', // also needs a real codex binary
];
+11
View File
@@ -0,0 +1,11 @@
{
"extends": "../tsconfig.json",
"compilerOptions": {
"rootDir": "..",
"noEmit": true,
"declaration": false,
"declarationMap": false,
"sourceMap": false
},
"include": ["../scripts/test-local-llm-harnesses.ts"]
}
+11 -5
View File
@@ -13,24 +13,24 @@ TZ=Australia/Perth
# Name of the account that runs Codeman and all local CLI sessions. Changing
# this value rebuilds the image with a matching account.
CODEMAN_RUNTIME_USER=opencode
CODEMAN_RUNTIME_USER=codeman
# Required. Persistent Codeman application data, CLI credentials, and session
# state are stored here on the host and mounted at the runtime account's home
# directory in the container.
CODEMAN_APPDATA_PATH=/mnt/user/appdata/Coding/codeman
CODEMAN_APPDATA_PATH=/mnt/user/appdata/codeman
# Optional. Absolute host path of this Codeman checkout, mounted at
# /opt/codeman so App Settings -> Updates can update Codeman in place. The Bash
# start script detects it from the compose file's own location, so it only needs
# setting for direct `docker compose` use or a checkout kept elsewhere. Point it
# at a directory that is not a git checkout and in-app updates are unavailable.
# CODEMAN_REPO_PATH=/mnt/user/appdata/Coding/codeman/app
# CODEMAN_REPO_PATH=/mnt/user/appdata/codeman/app
# Required for Docker cases. This must be an absolute path on the Docker host.
# Codeman and each isolated case use this same path, so it cannot be a
# container-only path such as /home/opencode/codeman-cases.
CODEMAN_CASES_PATH=/mnt/user/appdata/Coding/codeman/codeman-cases
# container-only path such as /home/codeman/codeman-cases.
CODEMAN_CASES_PATH=/mnt/user/appdata/codeman/codeman-cases
# Required. Network bind address, host port, and local image tag.
CODEMAN_HOST=0.0.0.0
@@ -44,6 +44,12 @@ CODEMAN_PASSWORD=changeme
# Required. Username for Codeman HTTP Basic authentication.
CODEMAN_USERNAME=admin
# Optional. Extra Host-header allowlist entries for a reverse-proxied domain
# (comma-separated; a bare `.suffix` matches every subdomain). Without it a
# proxied request is rejected with `403 Forbidden: host not allowed`. See
# README.md, "Reverse-proxy host allowlist".
# CODEMAN_ALLOWED_HOSTS=codeman.example.com,.internal.example.com
# Optional: authenticate Gemini CLI without an interactive login.
GEMINI_API_KEY=
+46 -6
View File
@@ -11,18 +11,21 @@ cp docker/.env.example docker/.env
bash docker/Start-Codeman.sh
```
On PowerShell, use the following command instead.
On PowerShell, use the following commands instead. Running Compose from inside `docker/` with no `-f` lets it discover `docker-compose.override.yml` on its own (see [Local customisation](#local-customisation)); naming the file with `-f docker/docker-compose.yaml` from the repository root silently drops the override unless it is named too.
```powershell
Copy-Item docker/.env.example docker/.env
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
Set-Location docker
docker compose --env-file .env up --build -d
```
Every required value is defined and explained in `.env.example`. `GEMINI_API_KEY` is intentionally optional and may remain blank.
The container starts as root so `entrypoint.sh` can correct the ownership of a bind source the Docker daemon created (it creates a missing one as `root:root`), then drops to `PUID:PGID` with `setpriv` before the server starts, so Codeman itself never runs privileged. That drop needs `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the file's `cap_drop: ALL`; a compose file written elsewhere (Unraid's Compose Manager, a hand-written unit) must carry the same additions, and the entrypoint names them when they are missing. A directory owned by neither root nor `PUID:PGID` is never re-owned: it is probed for writability as the runtime account and refused with a clear message if that fails. Setting `user:` in Compose skips the whole step.
On Linux, `Start-Codeman.sh` stops with an error when required paths are missing. It creates the application-data directory when safe, detects its numeric owner as `PUID:PGID`, and detects `DOCKER_SOCKET_GID` from the configured Docker socket. It rejects a root-owned application-data directory because Codeman and its local CLI sessions must remain unprivileged.
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `opencode`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `codeman`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
To retain Docker-case support without root when running Compose directly, set `DOCKER_SOCKET_GID` to the numeric group ID of the host socket. On a standard Linux Docker host, obtain it with `stat -c '%g' /var/run/docker.sock`. The Bash start script detects it automatically.
@@ -38,6 +41,43 @@ Releases that change `server.Dockerfile`, `docker-compose.yaml`, or add a key to
changed, and asks you to run `Start-Codeman.sh` here on the host instead. Details:
[`../docs/docker-self-update.md`](../docs/docker-self-update.md).
## Local customisation
Compose merges `docker-compose.override.yml` on top of `docker-compose.yaml`. Keep host-specific changes there rather than editing `docker-compose.yaml`, so this repository can be updated without losing them. Both `docker-compose.override.yml` and `docker-compose.override.yaml` are ignored by Git.
`Start-Codeman.sh` names the Compose file explicitly, which disables Compose's automatic discovery of the override file, so the script adds it back when one is present and prints the file it used. Running `docker compose` from this folder without any `-f` option finds it automatically. When passing `-f docker/docker-compose.yaml` from the repository root, add `-f docker/docker-compose.override.yml` as well, or the override is silently ignored.
An override file adds to and replaces individual settings. It cannot delete a key from `docker-compose.yaml`, and Compose concatenates rather than replaces `ports`, so removing a published port still requires editing `docker-compose.yaml`. The example below replaces the restart policy and adds a mount, leaving every other setting in place:
```yaml
services:
codeman:
restart: always
volumes:
- /srv/projects:/srv/projects
```
### Reverse-proxy host allowlist
Codeman rejects any request whose `Host` header is not on its own allowlist - a
DNS-rebinding guard, not a Compose or Docker concern. Loopback, any IP literal,
the configured `--host`, and a few tunnel-provider suffixes are allowed by
default; a reverse-proxied domain is not, and is rejected with
`403 Forbidden: host not allowed` before the request reaches any handler.
Add the domain with `CODEMAN_ALLOWED_HOSTS` in `.env`:
```sh
CODEMAN_ALLOWED_HOSTS='codeman.example.com,.internal.example.com'
```
`docker-compose.yaml` forwards it into the container (Compose only passes
through the environment keys it explicitly lists, and this is one of them, with
an empty default so the line is optional in `.env`).
See the application's own `docs/wiki/Remote-Access.md` for the full allowlist
format and the tunnel providers it accepts by default.
## Application data storage
The default configuration uses a host-folder bind mount:
@@ -49,7 +89,7 @@ volumes:
target: /home/${CODEMAN_RUNTIME_USER}
```
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/Coding/codeman`.
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/codeman`.
`CODEMAN_CASES_PATH` is the separate host directory for managed case workspaces. It is mounted into Codeman at the same absolute path, allowing the host Docker daemon to bind it into an isolated case container. Set it to a child directory of `CODEMAN_APPDATA_PATH` unless you deliberately store workspaces elsewhere.
@@ -60,7 +100,7 @@ Set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` when `docker info` reports `SwapLimit=
For an existing installation created by a root-running image, change ownership of the application-data directory before upgrading so the configured `PUID` and `PGID` can read the saved credentials and state:
```sh
chown -R 99:100 /mnt/user/appdata/Coding/codeman
chown -R 99:100 /mnt/user/appdata/codeman
```
Replace `99:100` and the path with the values from your `.env` file.
@@ -69,7 +109,7 @@ Do not replace this bind mount with a Docker-managed named volume when Docker ca
## Static macvlan networking
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section and add the following to the `codeman` service:
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section from `docker-compose.yaml` and add the following to the `codeman` service. The service and network additions can instead be placed in `docker-compose.override.yml`, but the `ports:` removal cannot, as described under [Local customisation](#local-customisation):
```yaml
mac_address: ${CODEMAN_MAC_ADDRESS}
+187 -7
View File
@@ -12,11 +12,34 @@ if [[ ! -f "$env_file" ]]; then
exit 1
fi
compose_command=(docker compose --env-file "$env_file" -f "$compose_file")
# Naming a Compose file explicitly disables Compose's automatic discovery of
# the override file, so it has to be added back by hand. Without this, local
# customisation in docker-compose.override.yml is silently ignored. The
# candidates are checked in Compose's own precedence order - measured on
# Compose v5.5.0 with both present: it uses `.yml` and ignores `.yaml`.
override_yml="$script_dir/docker-compose.override.yml"
override_yaml="$script_dir/docker-compose.override.yaml"
if [[ -f "$override_yml" && -f "$override_yaml" ]]; then
printf 'Warning: both %s and %s exist; Compose uses .yml and ignores .yaml.\n' \
"$override_yml" "$override_yaml" >&2
fi
compose_files=(-f "$compose_file")
for override_file in "$override_yml" "$override_yaml"; do
if [[ -f "$override_file" ]]; then
compose_files+=(-f "$override_file")
printf 'Using Compose override file: %s\n' "$override_file"
break
fi
done
compose_command=(docker compose --env-file "$env_file" "${compose_files[@]}")
appdata_path=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "CODEMAN_APPDATA_PATH" { sub(/^[^=]*=/, ""); print; exit }'
)
cases_path=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "CODEMAN_CASES_PATH" { sub(/^[^=]*=/, ""); print; exit }'
)
docker_socket=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "DOCKER_SOCKET" { sub(/^[^=]*=/, ""); print; exit }'
@@ -36,11 +59,18 @@ if [[ ! -d "$appdata_path" ]]; then
mkdir -p -- "$appdata_path"
fi
if owner_ids=$(stat -c '%u:%g' -- "$appdata_path" 2>/dev/null); then
:
elif owner_ids=$(stat -f '%u:%g' "$appdata_path" 2>/dev/null); then
:
else
if [[ -z "$cases_path" ]]; then
printf 'Error: CODEMAN_CASES_PATH is not set in %s\n' "$env_file" >&2
exit 1
fi
# `stat -c` is GNU, `stat -f` is BSD/macOS; the bind sources live on the Docker
# host, so both need to work.
owner_of() {
stat -c '%u:%g' -- "$1" 2>/dev/null || stat -f '%u:%g' "$1" 2>/dev/null
}
if ! owner_ids=$(owner_of "$appdata_path"); then
printf 'Error: Cannot determine the owner of CODEMAN_APPDATA_PATH: %s\n' "$appdata_path" >&2
exit 1
fi
@@ -54,6 +84,34 @@ if [[ "$PUID" == '0' ]]; then
exit 1
fi
# Pre-creating this here, exactly like CODEMAN_APPDATA_PATH above, means Compose
# never has to materialise a missing bind source itself - which it does as
# root:root - so the in-container entrypoint's chown never has to run for this
# path at all. It happens AFTER PUID/PGID are known (they come from the appdata
# directory just above) so the new directory can be given that exact owner: a
# plain `mkdir -p` lands as the invoking user's uid and PRIMARY gid, and on a
# host set up the way the README suggests (`chown -R 99:100 <appdata>`) that gid
# is not PGID, which the container would then refuse to run on. Unlike appdata,
# an EXISTING cases directory is left exactly as it is: the README explicitly
# allows pointing this at a normal projects directory the host account already
# owns, and the container checks that it is WRITABLE as PUID:PGID rather than
# who owns it.
if [[ ! -d "$cases_path" ]]; then
mkdir -p -- "$cases_path"
if [[ "$(owner_of "$cases_path")" != "$PUID:$PGID" ]]; then
# As root this always succeeds; as a member of PGID a chgrp does; anyone
# else gets the clear error here, where the fix is obvious, rather than a
# restart loop from the container.
if ! chown -- "$PUID:$PGID" "$cases_path" 2>/dev/null; then
printf 'Error: created CODEMAN_CASES_PATH (%s) but could not make it %s:%s (the owner of CODEMAN_APPDATA_PATH).\n' \
"$cases_path" "$PUID" "$PGID" >&2
printf 'Run `chown %s:%s %s` as root, or create the directory as that account, then retry.\n' \
"$PUID" "$PGID" "$cases_path" >&2
exit 1
fi
fi
fi
if [[ -z "$docker_socket" || ! -S "$docker_socket" ]]; then
printf 'Error: DOCKER_SOCKET is not a Unix socket: %s\n' "${docker_socket:-<unset>}" >&2
exit 1
@@ -92,6 +150,31 @@ if [[ ! -d "$repo_path/.git" ]]; then
printf 'Note: %s is not a git checkout, so in-app updates are unavailable.\n' "$repo_path" >&2
fi
# Reads HEAD without requiring a `git` binary on the host — this script
# otherwise checks the checkout only by testing for `.git` as a directory, and
# resolving refs by hand keeps that the same "no host git needed" guarantee.
# ⚠️ A worktree checkout has `.git` as a FILE (`gitdir: <path>`), not a
# directory, so this returns nothing there and the volume-refresh check below
# silently no-ops — consistent with the `-d .git` test used everywhere else in
# this script, not a special case, but worth knowing if a worktree checkout
# stops picking up a stale-volume refresh it should have caught.
git_head_commit() {
local git_dir="$1/.git" head_ref ref_path
[[ -d "$git_dir" ]] || return 1
head_ref=$(cat -- "$git_dir/HEAD" 2>/dev/null) || return 1
if [[ "$head_ref" == ref:* ]]; then
ref_path="${head_ref#ref: }"
if [[ -f "$git_dir/$ref_path" ]]; then
cat -- "$git_dir/$ref_path"
else
# Packed after a `git gc`; the loose ref file above is gone.
awk -v ref="$ref_path" '$2 == ref { print $1; exit }' "$git_dir/packed-refs" 2>/dev/null
fi
else
printf '%s' "$head_ref"
fi
}
# Record what the container is about to be built and created FROM. The in-app
# updater compares these against the release it wants to apply: a release that
# changes either file cannot be applied by the container restarting itself (a
@@ -126,4 +209,101 @@ else
printf 'Warning: no sha256 tool found; in-app updates will not detect environment changes.\n' >&2
fi
exec docker compose --env-file "$env_file" -f "$compose_file" up --build -d
# codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded from
# the image only while EMPTY, so a rebuilt image's fresh output sits unused
# behind old volume content until something clears it. The in-app self-updater
# never hits this — it rebuilds INSIDE the running container, into the very
# volume already in use — but a `docker compose build` triggered from outside
# it (this script, after a `git pull`) does: the container comes back up
# looking unchanged. Detect that here and clear just the affected volume(s) so
# the build below actually takes effect. Best-effort: with no sha256 tool this
# quietly does nothing, same as the environment-gate block above.
volumes_to_refresh=()
if [[ -n "$dockerfile_sha" ]]; then
repo_head=$(git_head_commit "$repo_path" || true)
lockfile_sha=$(sha256_of "$repo_path/package-lock.json" 2>/dev/null || true)
source_state_file="$state_dir/docker-build-source.json"
prev_head=''
prev_lockfile_sha=''
if [[ -f "$source_state_file" ]]; then
prev_head=$(sed -n 's/.*"headCommit": *"\([^"]*\)".*/\1/p' "$source_state_file")
prev_lockfile_sha=$(sed -n 's/.*"lockfileSha256": *"\([^"]*\)".*/\1/p' "$source_state_file")
fi
[[ -n "$repo_head" && "$repo_head" != "$prev_head" ]] && volumes_to_refresh+=('codeman-dist')
[[ -n "$lockfile_sha" && "$lockfile_sha" != "$prev_lockfile_sha" ]] && volumes_to_refresh+=('codeman-node-modules')
fi
if [[ ${#volumes_to_refresh[@]} -eq 0 ]]; then
exec "${compose_command[@]}" up --build -d
fi
# Runs even on this script's very first invocation against an EXISTING
# deployment, deliberately: that deployment's volumes may already be stale
# (there was no earlier version of this check to have caught it), and clearing
# an already-empty or nonexistent volume is a harmless no-op, so there is no
# fresh-install case this needs to avoid.
printf 'Source changed since the last start; refreshing: %s\n' "${volumes_to_refresh[*]}"
# Build BEFORE taking the stack down: the image build is the slow part and needs
# no container stopped, so the deployment is offline only for the recreate.
"${compose_command[@]}" build
# `com.docker.compose.volume` is the volume KEY, not a project-qualified name -
# a second stack on the same host (a beta instance started with a different
# COMPOSE_PROJECT_NAME, say) that also declares a volume keyed `codeman-dist`
# shares that label, and `head -n1` would pick whichever the daemon happens to
# list first. Scope the lookup to THIS stack's own resolved project name so it
# can only ever match this stack's volume. The name is read from the resolved
# config's top-level `name` key, indentation-agnostic (the formatting is not a
# contract), and the FIRST `name` in the output is the project's: nested ones
# (a network's `name:`) come later. `--format json` needs Compose v2.3+.
project_name=$(
"${compose_command[@]}" config --format json 2>/dev/null |
sed -n 's/^[[:space:]]*"name":[[:space:]]*"\([^"]*\)".*$/\1/p' | head -n1
)
"${compose_command[@]}" down
# Track whether the volumes were actually cleared. The marker below is written
# ONLY on success: with an unresolvable project name the label filter would
# match nothing, nothing would be removed, and a marker recording the new HEAD
# would stop this check from ever firing again while the stale volume kept
# serving old code. A failed removal likewise leaves the marker alone, so the
# next start retries, and the stack is brought back up regardless rather than
# left down.
refreshed=1
if [[ -z "$project_name" ]]; then
# The documented reset (docs/docker-self-update.md): both volumes re-seed from
# the image by a plain copy, so clearing the extra one costs a copy, not data.
printf 'Warning: could not resolve the Compose project name; clearing both build-artefact volumes with `down --volumes` instead.\n' >&2
"${compose_command[@]}" down --volumes || refreshed=0
else
for key in "${volumes_to_refresh[@]}"; do
volume_name=$(
docker volume ls -q \
--filter "label=com.docker.compose.volume=$key" \
--filter "label=com.docker.compose.project=$project_name" |
head -n1
)
if [[ -n "$volume_name" ]] && ! docker volume rm -- "$volume_name"; then
printf 'Warning: could not remove volume %s; it will be retried on the next start.\n' "$volume_name" >&2
refreshed=0
fi
done
fi
if [[ "$refreshed" == '1' ]]; then
printf '{\n "headCommit": "%s",\n "lockfileSha256": "%s"\n}\n' \
"$repo_head" "$lockfile_sha" >"$source_state_file.tmp"
mv -- "$source_state_file.tmp" "$source_state_file"
if [[ "$EUID" == '0' ]]; then
chown -- "$PUID:$PGID" "$source_state_file"
fi
else
printf 'Warning: the build-artefact volumes were NOT refreshed; the container may serve stale code until the next successful start.\n' >&2
fi
# Already built above, so no --build here: a second build would only re-check
# the cache.
exec "${compose_command[@]}" up -d
+21
View File
@@ -32,6 +32,10 @@ services:
CODEMAN_DOCKER_HOST_HOME: ${CODEMAN_APPDATA_PATH}
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT: ${CODEMAN_DOCKER_DISABLE_SWAP_LIMIT}
CODEMAN_CASES_PATH: ${CODEMAN_CASES_PATH}
# Extra Host-header allowlist entries for a reverse-proxied deployment
# (docker/README.md, "Reverse-proxy host allowlist"). Optional, so it
# defaults to empty rather than requiring a line in every .env.
CODEMAN_ALLOWED_HOSTS: ${CODEMAN_ALLOWED_HOSTS:-}
CODEMAN_HOST: ${CODEMAN_HOST}
CODEMAN_PASSWORD: ${CODEMAN_PASSWORD}
CODEMAN_PORT: ${CODEMAN_PORT}
@@ -91,6 +95,23 @@ services:
- no-new-privileges:true
cap_drop:
- ALL
cap_add:
# The entrypoint corrects bind-mount ownership as root before dropping to
# PUID:PGID. Everything not listed here remains dropped by cap_drop above.
# test/docker-entrypoint.test.ts pins this list against what the
# entrypoint and `init: true` actually need, so a capability cannot go
# missing silently again.
- CHOWN
- DAC_OVERRIDE
# `init: true` makes tini PID 1, and tini stays ROOT while the entrypoint
# drops the server to PUID. Signalling a process of a different uid needs
# CAP_KILL; without it tini's SIGTERM forward fails ("Unexpected error
# when forwarding signal: 'Operation not permitted'"), tini dies, and the
# PID namespace teardown SIGKILLs the server instead of letting
# `server.stop()` flush state on every `docker compose down`/`restart`.
- KILL
- SETGID
- SETUID
healthcheck:
test:
- CMD-SHELL
+165
View File
@@ -0,0 +1,165 @@
#!/bin/sh
# Corrects ownership - host bind mounts, and the image-baked CLI prefix -
# then drops to PUID:PGID.
#
# Compose binds CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH from the host. When
# either path does not exist yet - a first run, a cleared application-data
# directory, a restored backup - the Docker daemon creates it owned by root,
# and an unprivileged server cannot then create its own state directory. The
# result is a container that restarts forever on:
#
# Failed to start web server: EACCES: permission denied, mkdir '/home/<user>/.codeman'
#
# Running this as root and dropping afterwards removes that failure mode without
# leaving the server privileged. The same root start also lets it re-assert
# /opt/codeman-cli's ownership on every start, not just at image build time -
# see the comment at that chown below for why that matters for anyone who
# runs the compose file directly rather than through Start-Codeman.sh.
#
# Capabilities this script needs against the compose file's `cap_drop: ALL`
# (test/docker-entrypoint.test.ts pins the list against docker-compose.yaml):
# CHOWN + DAC_OVERRIDE the chown of a root-owned bind source below
# SETUID + SETGID the setpriv drop itself
# KILL NOT used here, but required by the container: with
# `init: true` tini is PID 1 and runs as root while the
# server runs as PUID, and signalling a process of a
# different uid needs CAP_KILL. Without it every
# `docker compose down`/`restart` ends in tini dying with
# "Unexpected error when forwarding signal" and the
# server being SIGKILLed instead of stopping cleanly.
set -eu
# Honour an explicit `user:` in Compose: when the container was not started as
# root there is nothing to correct and no privilege to drop.
if [ "$(id -u)" -ne 0 ]; then
exec "$@"
fi
# Everything below runs as root and calls stat, chown, id, setpriv and friends
# by bare name, so the lookup path must not contain a directory the runtime
# account can write to. /opt/codeman-cli/bin is exactly that (it is chowned to
# PUID:PGID so sessions can update the agent CLIs in place), and the image
# appends it to PATH for the server's sake. Resolve root's commands through the
# system directories only, and hand the image's full PATH back to the server at
# the exec below, since Codeman resolves the agent CLIs through it.
runtime_path=$PATH
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
export PATH
: "${PUID:=1000}"
: "${PGID:=1000}"
# The capabilities the compose file must grant, named in the diagnosis below so
# an out-of-tree compose file (Unraid's Compose Manager, a hand-written unit)
# fails with a one-line fix instead of a restart loop.
required_caps='CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID'
# Pre-flight the drop itself before touching anything. A container started with
# `cap_drop: ALL` and none of the additions above fails here, and would otherwise
# die at the final exec with a bare "setpriv: setresuid failed: Operation not
# permitted" after chown had already failed, or worse, misreport a perfectly
# writable directory as unwritable because the probe below could not drop
# privileges to test it.
if ! setpriv --reuid "$PUID" --regid "$PGID" --clear-groups true 2>/dev/null; then
printf 'entrypoint: cannot drop privileges to PUID:PGID (%s:%s).\n' "$PUID" "$PGID" >&2
printf 'entrypoint: this image starts as root and drops with setpriv, which needs\n' >&2
printf 'entrypoint: cap_add: [%s]\n' "$required_caps" >&2
printf 'entrypoint: on top of cap_drop: ALL (see docker/docker-compose.yaml). Add them to the\n' >&2
printf 'entrypoint: compose file that started this container, or set `user:` to skip the drop entirely.\n' >&2
exit 1
fi
# Preserve the supplementary groups Compose granted through group_add - that is
# how the Docker socket stays reachable - while discarding root's own group.
supplementary=$(id -G | tr ' ' '\n' | grep -vx 0 | paste -sd, -)
[ -n "$supplementary" ] || supplementary="$PGID"
# Writable as the account the server is about to become? A real probe, run as
# exactly the identity the final exec below produces (PUID, PGID, the same
# supplementary groups, capabilities dropped), rather than a comparison of
# owners: ownership is not writability. A group-writable tree owned by another
# account, an ACL, or a CIFS/NFS mount that reports some unrelated uid are all
# fine to run on and would all fail an owner check.
writable_as_runtime() {
setpriv --reuid "$PUID" --regid "$PGID" --groups "$supplementary" test -w "$1" 2>/dev/null
}
for target in "${HOME:-}" "${CODEMAN_CASES_PATH:-}"; do
[ -n "$target" ] && [ -d "$target" ] || continue
owner=$(stat -c '%u:%g' "$target")
[ "$owner" = "${PUID}:${PGID}" ] && continue
# Only ever correct a directory the DAEMON created: root-owned, because
# neither PUID nor PGID existed yet when it materialised the missing bind
# source. Anything else - a host tree that legitimately belongs to some
# OTHER account, such as an existing CODEMAN_CASES_PATH the README already
# allows pointing at a normal project directory - is not this container's
# to reassign; recursively chowning it on every mismatch silently rewrote
# a credentials tree or a projects directory to PUID:PGID with one log
# line to explain it. Such a directory is left alone and only PROBED below.
#
# The chown is deliberately not fatal. A bind mount backed by NFS, CIFS or a
# rootless daemon can refuse chown while still being perfectly writable, and
# the probe below is what decides whether the server can run on it.
if [ "${owner%%:*}" = '0' ]; then
if chown -R "${PUID}:${PGID}" "$target" 2>/dev/null; then
printf 'entrypoint: corrected ownership of %s to %s:%s\n' "$target" "$PUID" "$PGID"
else
printf 'entrypoint: warning: cannot change ownership of %s to %s:%s; checking whether it is writable anyway\n' \
"$target" "$PUID" "$PGID" >&2
fi
fi
if writable_as_runtime "$target"; then
if [ "${owner%%:*}" != '0' ]; then
printf 'entrypoint: %s is owned by %s, not %s:%s, but is writable as the runtime account; leaving its ownership alone\n' \
"$target" "$owner" "$PUID" "$PGID"
fi
continue
fi
printf 'entrypoint: %s is not writable as PUID:PGID (%s:%s); it is owned by %s.\n' \
"$target" "$PUID" "$PGID" "$owner" >&2
printf 'entrypoint: refusing to change ownership of a directory this container did not create.\n' >&2
printf 'entrypoint: either chown it on the host, make it writable to %s:%s, or set PUID/PGID to match its owner.\n' \
"$PUID" "$PGID" >&2
exit 1
done
# /opt/codeman-cli (the four agent CLIs) is chowned to PUID:PGID once, at
# image BUILD time, from the PUID/PGID build args - server.Dockerfile's own
# comment on that RUN step explains why it lives in its own prefix rather than
# /usr/local. Unlike HOME/CODEMAN_CASES_PATH above, that bake happens only
# when the image is actually rebuilt (`docker compose up --build`, which
# Start-Codeman.sh always does) - a deployment that instead runs the compose
# file directly (Unraid's Compose Manager, a native Debian systemd unit, any
# `docker compose up`/`restart` with no --build) can change PUID/PGID in .env
# and restart without ever rebuilding, at which point the container runs as
# the NEW uid while the CLI directory is still owned by the OLD one baked into
# the image layer - silently breaking the very "self-update a CLI in place"
# fix this directory exists for. Re-assert it here, every start, unconditionally:
# unlike the host bind mounts above, this is pure image content Codeman itself
# populated, never host data that might legitimately belong to someone else,
# so there is no ownership to be careful about - it is always correct for it
# to be owned by whoever this container is about to run as.
if [ -d /opt/codeman-cli ] && [ "$(stat -c '%u:%g' /opt/codeman-cli)" != "${PUID}:${PGID}" ]; then
chown -R "${PUID}:${PGID}" /opt/codeman-cli
fi
# Discarding group 0 is right for root's own group, but it also discards a
# `group_add: 0` that was there to reach a Docker socket owned by root:root.
# The previous image ran as PUID with that group kept, so say so rather than
# letting Docker-case support vanish silently on such a host.
if [ -S /var/run/docker.sock ] && [ "$(stat -c '%g' /var/run/docker.sock)" = '0' ]; then
printf 'entrypoint: warning: /var/run/docker.sock is owned by group 0, which is dropped along with root;\n' >&2
printf 'entrypoint: warning: Docker cases will not work from this container. Give the socket a dedicated\n' >&2
printf 'entrypoint: warning: group on the host and set DOCKER_SOCKET_GID to it.\n' >&2
fi
# No `--bounding-set -all` here: it is a silent no-op without CAP_SETPCAP, which
# the compose file deliberately does not grant, and `no-new-privileges` already
# makes the bounding set moot. The reuid/regid drop leaves CapPrm/CapEff empty.
# The image's full PATH goes back to the server here; see the top of the file.
exec setpriv --reuid "$PUID" --regid "$PGID" --groups "$supplementary" \
env PATH="$runtime_path" "$@"
+47 -3
View File
@@ -24,7 +24,7 @@ RUN npm ci \
# docker/docker-compose.yaml. It does not run a Docker daemon in this container.
FROM node:22-bookworm-slim
ARG CODEMAN_RUNTIME_USER=opencode
ARG CODEMAN_RUNTIME_USER=codeman
ARG PUID=1000
ARG PGID=1000
@@ -71,6 +71,24 @@ COPY --from=docker:29-cli \
# Keep credentials out of the image. Users authenticate these CLIs at runtime
# through Codeman sessions, and the configured host bind mount retains state.
#
# Installed into a DEDICATED prefix, /opt/codeman-cli, not the base image's
# default /usr/local. A session needs write access to wherever these CLIs live
# so it can self-update one in place (observed via Codex's own
# `npm install -g @openai/codex`, which renames the old package directory
# aside before installing the new one — a rename needs write access to the
# PARENT directory, not just the target, so the runtime account needs that
# access at the directory level). Chowning /usr/local/bin and
# /usr/local/lib/node_modules directly to get it would ALSO hand away
# entrypoint.sh (COPY'd to /usr/local/bin below, root-owned, executed as root
# on every container start with CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node
# binary: owning the DIRECTORY is enough to rename it aside and drop a
# replacement, even though the file itself stays root-owned, which would let a
# compromised session arrange for its own script to run as root at the next
# restart — undoing the "the server itself never runs privileged" guarantee
# the entrypoint exists to provide. /opt/codeman-cli holds nothing else to
# escalate through, so owning it is exactly the CLI-update access it needs and
# no more.
#
# ⚠️ PINNED ON PURPOSE. Unpinned, the agent CLI versions a user ends up with are
# a function of WHEN their image was built, not of any commit — so a Codeman
# release that depends on newer CLI behaviour (the trust-dialog handling is
@@ -82,6 +100,15 @@ COPY --from=docker:29-cli \
#
# Bump these deliberately, in a release. `--no-cache` is still needed to rebuild
# this layer when only the pins change upstream.
# The prefix is APPENDED to PATH, never prepended: it is chowned to the runtime
# account below, and entrypoint.sh runs as root calling stat/chown/setpriv by
# bare name. A prefix ahead of /usr/bin would let a session drop a `setpriv`
# there and have it run as root at the next container start (measured with a
# minimal image of this exact shape). The four CLIs live only in this prefix,
# so they still resolve; entrypoint.sh additionally pins its own PATH to the
# system directories for the root part of the start.
ENV NPM_CONFIG_PREFIX=/opt/codeman-cli
ENV PATH=$PATH:/opt/codeman-cli/bin
RUN npm install --global \
@anthropic-ai/claude-code@2.1.258 \
@google/gemini-cli@0.58.0 \
@@ -93,6 +120,11 @@ RUN npm install --global \
# PGID match the host-owned application-data directory mounted by Compose. The
# requested GID may not exist in the base image, and a host UID such as 1000 may
# already belong to the baked `node` account, so handle both cases explicitly.
#
# The trailing chown hands the CLI prefix (/opt/codeman-cli, populated above)
# to that same account, so a session can self-update one of the CLIs in place.
# /usr/local stays root-owned throughout — see the comment on the npm install
# above for why that boundary matters.
RUN set -eux; \
case "${PUID}" in ''|*[!0-9]*) echo "PUID must be numeric" >&2; exit 1;; esac; \
case "${PGID}" in ''|*[!0-9]*) echo "PGID must be numeric" >&2; exit 1;; esac; \
@@ -120,7 +152,8 @@ RUN set -eux; \
--home-dir "/home/${CODEMAN_RUNTIME_USER}" \
--shell /bin/bash \
"${CODEMAN_RUNTIME_USER}"; \
fi
fi; \
chown -R "${PUID}:${PGID}" /opt/codeman-cli
WORKDIR /opt/codeman
@@ -135,8 +168,19 @@ ENV CODEMAN_IN_CONTAINER=1 \
HOME=/home/${CODEMAN_RUNTIME_USER} \
NODE_ENV=production
# Runtime defaults for the entrypoint, matching the account created above.
ENV PGID=${PGID} PUID=${PUID}
EXPOSE 3000
USER ${CODEMAN_RUNTIME_USER}
# The container starts as root so the entrypoint can correct the ownership of
# the host bind mounts, which the daemon creates as root whenever they do not
# already exist. The entrypoint then drops to PUID:PGID with setpriv, so the
# server itself never runs privileged. Setting `user:` in Compose bypasses both
# steps, leaving the caller in full control.
COPY docker/entrypoint.sh /usr/local/bin/entrypoint.sh
RUN chmod 0755 /usr/local/bin/entrypoint.sh
ENTRYPOINT ["/usr/local/bin/entrypoint.sh"]
CMD ["node", "dist/index.js", "web"]
+42
View File
@@ -479,6 +479,48 @@ re-captured, or the item acknowledged), `approval:resolved` (`{ id, sessionId, k
`resolution` one of `answered | resolved_in_terminal | superseded |
session_ended | dismissed | expired`).
## Reboot restore
A host reboot takes the tmux server down with it, so every pane dies and the
board comes up empty. At boot Codeman works out which sessions the reboot
destroyed and holds that plan in memory, and these endpoints let a client offer
it to the user. Nothing creates a pane until the user asks: the boot-time reboot
heuristic decides whether to ASK, never whether to act.
Claude-mode sessions only (others carry their conversation id in their own
config object); remote and docker sessions are never offered, because both need
another host or container to be up. The plan is in-memory, so a server restart
drops it and the offer is gone; the conversations themselves are unaffected,
since they live in the CLI's own transcript store and stay reachable from the
Resume list. A plan nobody spends expires after 24 hours.
- `GET /api/v1/reboot-restore` → `{ sessions: RestorableSession[],
scrollbackRestored: false }`, ownership-scoped in multi-user mode.
`RestorableSession`: `{ id, name?, workingDir, mode, owner? }`. The persisted
record itself is never sent. `scrollbackRestored` is always `false` and exists
so a client states it: a restored session is a NEW pane, so the conversation
continues and the terminal history does not.
- `POST /api/v1/reboot-restore/restore` with `{ sessionIds?: string[] }` (omit
to restore everything the caller can see) → `{ restored: RestorableSession[],
skipped: { sessionId, reason }[] }`. `reason` is one of `workspace-missing`
(the directory is gone), `workspace-forbidden` (in multi-user mode it is
outside the workspace of the user the session belongs to, re-checked against
that owner's current grant rather than the caller's), `already-live` (the conversation is already
open, typically resumed by hand from the Resume list), `capacity-reached`
(the global or per-user session cap), or `rebuild-failed` (the agent would not
start, most often a CLI binary missing from the server's PATH).
`409 CONFLICT` when that caller already has a restore running. Entries are
removed from the plan before any pane is built, so a double-click cannot put
two panes on one conversation; anything that never became a pane goes back on
offer, except `already-live`, which cannot stop being true. A restored session
comes back attached, idle and disarmed: respawn controllers and Ralph loops
are never re-armed automatically.
- `POST /api/v1/reboot-restore/dismiss` → `{ dismissed: n }`. Drops the offer
for everything the caller can see.
Each rebuilt session also emits the ordinary `session:created` SSE event, so
clients other than the one that clicked pick it up without refetching.
## Read My Mind intent profiles
Per-case profiles of what the user is trying to accomplish: user/agent-stated
File diff suppressed because one or more lines are too long
+363
View File
@@ -0,0 +1,363 @@
# Custom Model Endpoint Profiles (all harnesses, local or cloud)
## Context
The author pays for Claude Code but also runs a capable local model behind an
OpenAI-compatible server (llama.cpp) — and wants the same mechanism to work
against a **cloud** OpenAI-compatible endpoint too (e.g. Azure AI Foundry's
OpenAI-compatible inference endpoint, OpenRouter, a self-hosted gateway).
Right now every Codeman session mode defaults to its native cloud backend
with no way to redirect a session at any other endpoint from the UI — the
closest existing precedent is DeepSeek's server-env-sourced
`DEEPSEEK_BASE_URL`, which isn't user-facing.
**Scope note**: this plan originally said "local LLM." It now covers any
OpenAI-compatible endpoint the user configures — local (llama.cpp, Ollama,
vLLM) or cloud (Azure AI Foundry, OpenRouter, a company gateway). The
mechanism is identical (a base URL Codeman probes via `GET /v1/models`); the
only real differences are auth-header convention (cloud endpoints often want
an `api-key` header, e.g. Azure, rather than `Authorization: Bearer`) and
that a cloud "model" may actually be a deployment name distinct from the
underlying model family (Azure AI Foundry deployments) — both are called out
where they matter below. Naming throughout this plan is **"custom model
endpoint,"** not "local model," to keep that scope explicit.
### Additional use case: on-premises AI hardware
"Local" isn't limited to a desktop running llama.cpp — a growing category of
purpose-built, on-premises AI hardware exists specifically to run a serious
model on-site with an OpenAI-compatible server, and this feature is exactly
the on-ramp for pointing Codeman at one:
- **NVIDIA DGX Spark** (and the DGX Spark-class "Spark" mini-supercomputer
line) — a compact on-prem inference/training box aimed at running large
local models with an OpenAI-compatible API surface.
- **AMD "Strix Halo" (Ryzen AI Max)** on-prem AI mini-PCs — unified-memory
APU hardware marketed for local LLM inference, typically fronted by
llama.cpp/Ollama/vLLM the same way a home server would be.
Neither needs anything new from this design: both present a standard
`/v1/models` + `/v1/chat/completions` OpenAI-compatible surface once the
inference server is running, so they're just another `baseUrl` entry in the
custom-model-hosts store, same as llama.cpp or a cloud endpoint. The
justification for building this generically (rather than hardcoding "point
Claude at my llama.cpp box") is precisely this: **the same endpoint registry
and per-CLI injection mechanism should work unmodified for any current or
future OpenAI-compatible box or service** — a home GPU rig today, a Spark or
Strix Halo appliance tomorrow, a company's on-prem inference cluster after
that — without Codeman needing to know or care what's actually serving the
model on the other end of that URL.
A concrete example worth naming: **[Ark0N/Qwen5090](https://github.com/Ark0N/Qwen5090)**
(from the same GitHub account as this project's owner) is a one-click
Windows / one-command Linux installer that stands up Qwen3.8-27B locally on
an RTX 5090 (or another RTX 50-series card with ≥24GB) behind an
OpenAI-compatible API, served by any of vLLM, NInfer, or llama.cpp — MIT-
licensed tooling over Apache-2.0 Qwen weights. It's a direct, ready-made
target for this feature: point a custom-model-hosts entry at whichever
backend it's running, and it needs nothing further from Codeman's side. It's
also notable for already wiring up DeepSeek Harness and Claude Code as
coding agents against that local server itself, which is effectively the
same "point a Codeman-supported harness at a local endpoint" idea this
feature is generalizing — worth using as a real-world reference/test target
once chunk 5 (session integration) exists, alongside the author's own llama.cpp
box.
Each harness has its own (different-shaped) mechanism for pointing at a
custom OpenAI-compatible base URL + model — env vars for Claude, a JSON
config blob for opencode, a TOML file for Codex, etc. The author gave the
starting recipes for those three; the rest (Gemini, Pi, Grok, DeepSeek, OMP,
Antigravity) were researched for this plan and are flagged by confidence
below. A real end-to-end pass against the author's own llama-swap server
(`scripts/test-local-llm-harnesses.ts`, inside a `codeman/agent:llm-test`
Docker image with all 9 CLIs installed) then confirmed **claude and
opencode work end-to-end**, corrected a real Codex config.toml schema bug
the given recipe had (see the Codex row below), and surfaced that Codex's
_protocol_ — not just its config shape — does not work against a plain
OpenAI-Chat-Completions server like llama.cpp/llama-swap at all. Confidence
below reflects what was actually observed, not just what was planned.
The feature must be:
- **Off by default**, one settings toggle turns it on.
- Endpoint entry: user gives a base URL — a LAN address or a cloud URL —
plus an optional API key, and Codeman calls `GET <baseUrl>/v1/models` to
discover and store the available model (or deployment) list.
- A **new toolbar selector** (separate from the existing Run-mode menu, since
it's a modifier on top of whichever harness is already selected/running)
lets the user pick "Cloud (default)" — the harness's own native backend —
or a model discovered from one of the configured custom endpoints.
- Picking a custom-endpoint model for an **already-running session restarts
that session's CLI process** with the injected env/config pointed at that
endpoint (confirmed with the maintainer — these harnesses read endpoint config at
process start, not per-turn, so a live hot-swap isn't possible).
- **New sessions always default back to the harness's native cloud backend.**
A custom-endpoint selection is a per-session override, not a sticky global
default — starting a fresh CLI (any mode) always launches against its
native backend unless the user explicitly picks a custom endpoint for that
new session too. The toolbar selector is scoped to "this session," never
carried forward as the default for future sessions.
This follows the repo's existing data-driven CLI-registry philosophy
(`test/cli-registry-no-id-branching.test.ts`): per-CLI behavior is a
declared capability, never an `if (mode === 'claude')` branch.
## Per-CLI injection recipes (confidence-ranked)
| CLI | Mechanism | Confidence |
| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
Everything web-researched-but-unverified gets implemented but must be
smoke-tested against real installs of those CLIs before being called done —
call this out explicitly when implementing, don't just ship on faith.
**Cloud-endpoint specifics** to keep in mind per recipe above: an Azure AI
Foundry-style endpoint typically wants the API key in an `api-key` header
rather than (or in addition to) `Authorization: Bearer`, and its "model" is
often a deployment name rather than the underlying model family name — the
discovery step (`GET /v1/models`) still works the same way against Azure AI
Foundry's OpenAI-compatible endpoint shape, but a user may need to type the
deployment name manually if it isn't returned as expected.
## Architecture
### 1. Registry: new `capabilities.customModelInjection` field
Extend `src/config/cli-registry/types.ts` / `schema.ts` with a discriminated
union on each `CliEntry.capabilities`:
```ts
type CustomModelInjection =
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[] }
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json' }
| {
kind: 'configDir';
dirEnvVar: string;
fileName: string;
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml';
}
| { kind: 'unsupported' };
```
Declared per stock.ts entry per the table above. A pure function in a new
`src/custom-model-injection.ts` (`buildCustomModelInjection(entry, endpoint, modelId)`)
turns `(CliEntry, endpoint, modelId)` into either an `envOverrides` object
(kind `env`/`configContentEnv`) or a `{ dirEnvVar, files: [{path, content}] }`
descriptor (kind `configDir`) — unit-testable with no IO, mirroring how
`session-cli-builder.ts` is pure. The `configDir` kind additionally needs an
IO wrapper that writes those files under
`dataPath('custom-model-configs/<sessionId>/')` (new dir, cleaned up on
session delete — same lifecycle as other per-session generated state).
### 2. Endpoint registry: `src/custom-model-hosts.ts`
Same read-array/write-array shape as `src/remote-hosts.ts` /
`src/webview-store.ts`: `~/.codeman/custom-model-hosts.json` holding
`CustomModelEndpoint[] = { id, label, baseUrl, apiKey?, authStyle?: 'bearer'|'api-key'|'both', models?: string[], lastDiscoveredAt? }`.
`authStyle` defaults to `'both'` (send both header conventions on the
discovery probe, same approach the smoke-test script below uses) so one
endpoint entry works whether it's llama.cpp or Azure without the user having
to know which header their box wants in advance.
New route file `src/web/routes/custom-model-routes.ts` (registered in the
routes barrel), mirroring `case-routes.ts`'s remote/docker-host CRUD
(`GET/POST/PUT/DELETE /api/model-endpoints`, admin-gated in multi-user mode
the same way) plus:
- `POST /api/model-endpoints/:id/discover-models` — fetches
`${baseUrl}/v1/models`, stores the `data[].id` list, returns it. Bounded
timeout, and run the target through the **same SSRF egress guard already
used for web tabs** (`webview-egress-policy.ts` — reject link-local/cloud
metadata addresses) — this still matters for a cloud URL too, since the
guard is about preventing a redirect to internal infra, not about
local-vs-cloud.
**Why discovery rather than a free-text model field**: it removes the one
piece of configuration most likely to trip a user up — hand-typing the
exact model identifier a given inference server expects, which varies by
server and is an easy source of a silent "model not found" failure with no
useful error surfaced back through a CLI's own startup. Discovery also
means this design is not limited to a single-model box: a **multi-model
gateway** such as **[llama-swap](https://github.com/mostlygeek/llama-swap)**
(hot-swaps between several loaded llama.cpp model configs behind one
OpenAI-compatible endpoint) or a vLLM/LiteLLM/Ollama instance serving
several models advertises ALL of them through the same `/v1/models` call —
so one endpoint entry surfaces every model that gateway can serve, with no
extra per-model configuration on Codeman's side at all.
### 3. Settings
- New synced boolean `customModelEndpointsEnabled` in `SettingsUpdateSchema`
(`src/web/schemas.ts`), default `false`, documented inline like
`readMyMindEnabled`/`workspaceHooksEnabled`.
- New `.set-group` "Custom Model Endpoints" inside the **Agents & CLIs**
section (`settings-clis`, `index.html:2150+`) with the enable toggle plus
a list-editor (add/refresh-models/delete rows) for endpoints — closest
existing precedent is the respawn-presets array editor
(`schemas.ts:1285-1305`, `index.html:1243-1244`) for add/apply/delete-by-id
semantics, backed by the new CRUD routes above.
### 4. Toolbar UI
- New header/toolbar button (e.g. `#customModelBtn`, `btn-toolbar
btn-custom-model`), marker-hidden by default (`btn-custom-model--hidden`)
and revealed by `applyHeaderVisibilitySettings()` only when
`customModelEndpointsEnabled` is on — same pattern as the File
Viewer/Cron buttons.
- Clicking opens a dropdown (`#customModelMenu`, same `.run-mode-menu`-style
markup as the existing Run-mode gear menu) listing "Cloud (default)" plus
every discovered model, grouped by endpoint. An entry is disabled with a
tooltip when the active session's CLI has `customModelInjection.kind ===
'unsupported'` (Antigravity) or none declared.
- Selecting an entry calls a new route:
`POST /api/sessions/:id/custom-model { endpointId, modelId } | { clear: true }`.
Server: resolve the CLI entry for `session.mode`, build the injection via
§1, persist it as a new `session.customModel` state field (surfaced in
`toState()`/SSE so the tab can show a small badge, e.g. "🖥 qwen3 (local)"
or "☁ gpt-4o-mini (azure)", and the choice survives reload), merge into
the session's `envOverrides`, and **respawn the pane's CLI process**
through the same respawn/interactive-restart path
`session.ts`/`tmux-manager.ts` already use for effort/model changes
(`_configureCliEnv()` + `applyEnvOverrides()` at spawn time) — reuse,
don't reinvent, the existing kill-and-relaunch-in-pane machinery.
- New-session creation deliberately does **not** inherit a prior custom-
endpoint choice: `buildEnvOverrides()` (session-ui.js) never carries the
toolbar selection forward to the next `run()` call. Every new session
starts on its native backend; picking a custom endpoint in the toolbar for
a session applies only to that session (and, if done before Run is
clicked, to the one session about to be created — not to sessions created
afterward).
### 5. Multi-user security clamp
Every new env var this feature introduces that can redirect a session's
traffic (and thus wherever its credentials go) — `ANTHROPIC_BASE_URL`,
`GOOGLE_GEMINI_BASE_URL`, `GROK_BASE_URL`, the `CODEX_HOME`/`PI_CONFIG_DIR`
dir-redirects, plus the already-privileged `DEEPSEEK_BASE_URL` — must be
added to each CLI's `capabilities.privilegedEnvKeys` so
`clampEnvOverridesForOwner()` strips them for a non-granted multi-user
owner, exactly the precedent already documented for `DEEPSEEK_BASE_URL`/
`OMP_AUTH_BROKER_URL`. This matters _more_, not less, now that endpoints can
be cloud URLs: redirecting a non-granted user's session to an attacker's
cloud endpoint is a credential-exfiltration path, not just a mischief
redirect to a LAN box. Endpoint CRUD itself stays admin-only in multi-user
mode, same as remote/docker hosts.
## Files touched (representative, not exhaustive)
- `src/config/cli-registry/types.ts`, `schema.ts`, `stock.ts` — new capability + per-entry declarations
- `src/custom-model-injection.ts` (new) — pure per-CLI descriptor builder + unit tests
- `src/custom-model-hosts.ts` (new) — endpoint store
- `src/web/routes/custom-model-routes.ts` (new) — CRUD + discovery route
- `src/web/routes/session-routes.ts` — `POST /api/sessions/:id/custom-model`, clamp wiring
- `src/web/schemas.ts` — `customModelEndpointsEnabled`, endpoint/discover payload schemas, privileged-key updates
- `src/session.ts` — `customModel` state field, `toState()` surface
- `src/web/public/index.html`, `settings-ui.js`, `session-ui.js`, `styles.css` — settings group, toolbar button/menu, badge, accent CSS
- `src/web/sse-events.ts` + `constants.js` — if a dedicated SSE event is warranted for the badge (or just ride existing session-update broadcasts)
- `test/fixtures/mock-openai-server.ts` (new) + `test/custom-model-injection-contract.test.ts` (new) — see Mock-server validation below
- `scripts/test-local-llm-harnesses.ts` (already added, this branch; run via `npx tsx`) — the standalone real-CLI-and-real-endpoint smoke test, supporting any `--base-url` (local or cloud). Dynamic: derives its harness list and every env var/config it injects from the live CLI registry + `buildCustomModelInjection()` rather than a second hand-maintained copy — only the one-shot invocation flags (`ONE_SHOT` table) are CLI-specific info the registry doesn't model and stay hand-maintained
- `docs/custom-model-endpoints.md` (new) + a CLAUDE.md pointer bullet under External CLI modes / envOverrides
## Mock-server validation strategy (CI-runnable, no real CLI binaries needed)
Spawning nine real CLI binaries in CI isn't realistic, and neither the author's
llama.cpp box nor a real cloud subscription can be a CI dependency. So the
injection _logic_ gets a tier of automated coverage that sits between the
pure unit tests and the live manual checks in Verification:
1. **`test/fixtures/mock-openai-server.ts`** — a small in-process HTTP
server (plain `http.createServer`, no external deps, port picked per the
existing `const PORT = 3150+` convention) that:
- Serves `GET /v1/models` → a fixed fake model list (`{data:[{id:'qwen3'},...]}`),
for testing the discovery route.
- Serves `POST /v1/chat/completions` (OpenAI shape) **and**
`POST /v1/messages` (Anthropic Messages-API shape, since that's what
`ANTHROPIC_BASE_URL` traffic looks like) and records every request it
receives (headers, body, path) into an array the test can assert on —
including which auth header style it saw, so the `authStyle: 'both'`
default and Azure's `api-key` convention both get real coverage.
- Returns a minimal valid completion so a client library doesn't choke
on the response shape.
2. **`test/custom-model-injection-contract.test.ts`** — for every CLI with a
`customModelInjection` capability (i.e. every row in the table above
except `antigravity`):
- Point a fixture `CustomModelEndpoint` at the mock server's URL.
- Call `buildCustomModelInjection(entry, endpoint, modelId)` (the pure
function from §1) to get the real env vars / config-file content that
would be injected into that CLI's session.
- Replay those exact values through a minimal HTTP request shaped the
way that CLI is documented to send it (Anthropic Messages shape for
claude; OpenAI chat-completions shape for opencode/codex/pi/grok/omp;
`GOOGLE_GEMINI_BASE_URL`'s OpenAI-compat shape for gemini; dsh's
provider call for deepseek) against the mock server.
- Assert the mock server received the request **at the injected
`baseUrl`**, with **the injected API key** in the expected header, and
**the injected model id** in the body/path — i.e. prove the values
Codeman computes are internally consistent and would reach the right
place with the right identifiers, end to end, in CI, on every push.
- Also cover the `configDir` kind (codex/pi/omp): assert the written
`config.toml`/`models.json`/`models.yml` file parses and contains the
same base URL/key/model, and that it's written under the isolated
per-session dir rather than the user's real config path.
3. **Explicit, stated limitation** (goes in the test file's `@fileoverview`
and in this doc, not left implicit): this proves _"if the CLI honors its
documented env/config contract, it will hit the right endpoint with the
right model."_ It does **not** prove the real CLI binary actually reads
that env var / config file the way its docs say — that's still the job
of the live manual checks in Verification step 4-5 below, and is exactly
why the confidence table above did not stop at "researched" — every CLI
except antigravity (no mechanism at all) has since been run against a
real llama-swap server via `scripts/test-local-llm-harnesses.ts`:
claude/opencode/pi/grok/omp are confirmed PASS end-to-end, codex is
confirmed FAIL for a real documented protocol reason (Responses-API-only
since Feb 2026), and gemini/deepseek are confirmed reaching the server
but failing for reasons not yet root-caused (see their table rows). The
mock-server suite catches regressions in Codeman's own logic; it cannot
catch a CLI changing its env-var name in a future release, or a real
cloud endpoint behaving differently from a local llama.cpp box.
## Verification
1. `npm run typecheck && npm test` after each slice — this now includes the
mock-server contract suite from above, so injection-logic regressions
are caught automatically without touching real infrastructure.
2. Unit tests for `buildCustomModelInjection()` per CLI kind (pure, no IO).
3. Route tests (`app.inject`) for the new CRUD + discover-models endpoint
(mock `fetch` for `/v1/models`), and for the multi-user clamp on the new
privileged keys (mirror `test/routes/external-cli-bypass-clamp.test.ts`).
4. **Standalone real-binary smoke test**: `scripts/test-local-llm-harnesses.ts`
exercises every harness the CLI registry declares `customModelInjection`
support for against a real `--base-url` — local or cloud — outside of
Codeman's UI entirely, and is DYNAMIC (reads `enabledClis()` + calls the
real `buildCustomModelInjection()`, so a future registry change is picked
up automatically with zero edits to the script). Already run to
completion against the author's llama-swap server (a LAN address,
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected**
(Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED**
(reach the server, fail for undiagnosed reasons — see their table rows),
antigravity **SKIP** (no mechanism). Re-run this against a real cloud
endpoint (e.g. an Azure AI Foundry deployment) once one is available, to
prove the `authStyle`/deployment-name handling holds up outside llama.cpp.
5. Once the full feature (not just the standalone script) is built: add an
endpoint via the real UI, hit discover-models, confirm the returned model
list, pick Claude + the model on a real session, confirm via
`tmux -L codeman capture-pane`/`tmux showenv -t <pane>` that
`ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY`/`ANTHROPIC_DEFAULT_*_MODEL` are
set post-restart, and confirm the endpoint's own logs show the next
prompt actually landing there. Repeat for opencode and Codex at minimum
before considering this shippable; spot-check the web-researched CLIs
and correct the plan's confidence table with what's actually observed.
6. `npm run lint && npm run format:check`.
7. Update `CHANGELOG.md`/changeset per the COM workflow when shipping.
+147
View File
@@ -0,0 +1,147 @@
# Custom Model Endpoint Profiles
Point any Codeman-supported harness — Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, or OMP — at a custom OpenAI-compatible endpoint instead of
its native cloud backend, for a given session. "Custom endpoint" covers both
**local** hardware (llama.cpp, Ollama, vLLM, a home GPU rig, or purpose-built
boxes like NVIDIA DGX Spark or AMD Strix Halo mini-PCs) and **cloud**
services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
company gateway) — anything answering `GET /v1/models` and
`POST /v1/chat/completions` in the standard shape. Design doc, per-CLI
recipe confidence table, and security reasoning:
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
> **Status**: backend is implemented and tested (registry capability, the
> injection engine, the endpoint store + discovery route, the session
> restart route). The toolbar picker / settings UI described below as the
> intended surface is **not yet built** — until it lands, use the HTTP API
> directly (examples below). Antigravity has no known custom-endpoint
> mechanism and is not supported.
## Turning it on
App Settings → Agents & CLIs → **Custom Model Endpoints** (synced setting
`customModelEndpointsEnabled`, default **OFF**). Until the toolbar picker
lands, nothing reads this setting: the HTTP routes below work whether it is
on or off, and it exists now only so the picker has a switch to hang off
when it ships. The API equivalent:
```bash
curl -sk -X PUT https://localhost:3000/api/settings \
-H 'Content-Type: application/json' \
-d '{"customModelEndpointsEnabled": true}'
```
## Adding an endpoint
```bash
curl -sk -X POST https://localhost:3000/api/model-endpoints \
-H 'Content-Type: application/json' \
-d '{"id": "llama-box", "label": "Home llama.cpp", "baseUrl": "http://192.168.1.50:8080"}'
```
`apiKey` is optional (most local servers don't check it). `authStyle`
(`bearer` | `api-key`, default `bearer`) controls which auth header
convention discovery uses: `bearer` is `Authorization: Bearer <key>`
(llama.cpp, OpenAI-compatible servers, most gateways), `api-key` is the
`api-key: <key>` header Azure AI Foundry wants. There is deliberately no
"send both" option: measured against a real llama-swap server, a request
carrying both headers hung indefinitely. `baseUrl` must be `http(s)`, carry
no embedded credentials, and may not point at a link-local or cloud-metadata
address; discovery re-checks the address the name actually resolves to.
Discover its available models:
```bash
curl -sk -X POST https://localhost:3000/api/model-endpoints/llama-box/discover-models
```
This calls the endpoint's own `GET /v1/models` and stores the returned list
on the endpoint record; `GET /api/model-endpoints` lists everything
configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
Endpoint management is admin-only in multi-user mode, same as remote/docker
hosts — these are machine-level infra, not per-user settings.
## Applying a model to a session
```bash
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
-H 'Content-Type: application/json' \
-d '{"endpointId": "llama-box", "modelId": "qwen3"}'
```
This computes the CLI-specific env vars / config for that session's mode
(see the recipe table in `custom-model-endpoints-plan.md`) and **restarts the session's
CLI process in place** — same pane, same tmux session, fresh env. That
restart is necessary, not incidental: every supported harness reads its
endpoint config at process start, not per-turn, so there is no live
hot-swap. A Claude session is relaunched with `--resume <conversation> ||
--session-id <id>`, so it continues the conversation it was on; pi, omp and
grok are relaunched with the `--model` value that selects the injected
provider (`custom/<modelId>` for pi and omp, `codeman-custom` for grok),
since for those three the config file alone does not switch the model.
**Remote (SSH) and Docker sessions are refused** (400) for now: their restart
reattaches the durable remote/in-container tmux rather than relaunching the
agent, so the selection would report success and change nothing.
Clear back to the harness's native cloud default with:
```bash
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
-H 'Content-Type: application/json' -d '{"clear": true}'
```
Clearing also removes the env vars the selection injected from the tmux
session (they persist there and would otherwise be inherited by the
relaunched CLI) and deletes the per-session config directory
(`~/.codeman/custom-model-configs/<sessionId>`, written 0600 because pi and
omp embed the API key in it). That directory is also removed when the
session is deleted. The selection survives a Codeman restart: the endpoint
id, model and injected key NAMES are persisted, the values are re-derived
from the endpoint store on recovery, and the pane keeps running against the
endpoint in between because tmux retains its environment.
**New sessions always default back to the harness's native backend.** A
custom-endpoint selection is a per-session choice, never a sticky global
default — starting a fresh session doesn't inherit whatever the last one was
pointed at.
## Confidence per harness
Every harness except Antigravity has now been run end-to-end against a real
llama-swap server via `scripts/test-local-llm-harnesses.ts` (a dynamic
script that reads the live CLI registry, so a registry change is picked up
automatically). Results:
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
came back through the endpoint.
- **Codex** — the config is structurally correct, but Codex only speaks the
Responses API since Feb 2026, which llama.cpp/llama-swap don't implement.
This is a real protocol incompatibility, not a bug here; Codex support
needs a Responses-API-compatible endpoint.
- **Gemini** — fails with `Invalid auth method selected`, traced to an
undocumented `GATEWAY` auth path gemini-cli selects once
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
(several auth workarounds were tried and ruled out); do not rely on
Gemini support yet.
- **DeepSeek** — the request reaches the server (env vars are read) but
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind
each result. `scripts/test-local-llm-harnesses.ts` is the standalone script
used to check a harness against a real endpoint outside the web UI
entirely; see its own `--help` for usage.
## Security note
Every env var this feature can set that redirects a session's traffic
(`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, etc.) is
listed in that CLI's `privilegedEnvKeys` in the CLI registry, so a
non-granted multi-user owner cannot set one directly via the generic
`envOverrides` API field — only through this feature's own route, which
computes the value from an admin-configured, SSRF-guarded endpoint rather
than trusting arbitrary client input. See the "Multi-user security
hardening" section of `custom-model-endpoints-plan.md` for the full reasoning; several
of these were reachable via the generic `envOverrides` field even before
this feature existed, and building this surfaced and closed that gap.
+4 -2
View File
@@ -15,7 +15,7 @@ The application container mounts the Docker daemon socket so Codeman can create
## Start
Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/Coding/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes.
Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes.
```sh
cp docker/.env.example docker/.env
@@ -33,12 +33,14 @@ On Linux, run the stack with the start script. It determines `PUID` and `PGID` f
bash docker/Start-Codeman.sh
```
On other platforms, run Compose directly. `PUID` and `PGID` default to `1000:1000`; set them in `docker/.env` when the application-data directory has a different owner.
On other platforms, run Compose directly. `PUID` and `PGID` default to `1000:1000`; set them in `docker/.env` when the application-data directory has a different owner. Naming the file with `-f` disables Compose's own discovery of `docker/docker-compose.override.yml`, so add a second `-f` for it when you keep one (see `docker/README.md`, Local customisation).
```sh
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
```
The container starts as root, corrects the ownership of a bind source the daemon had to create, and drops to `PUID:PGID` with `setpriv` before Codeman starts; the capabilities that needs are declared in `docker/docker-compose.yaml` and named by the entrypoint when a compose file written elsewhere lacks them.
Open `http://localhost:3000` and sign in with the username and password from `docker/.env`.
## Operations
+15
View File
@@ -59,6 +59,7 @@ unchanged. The container path is a new `SupervisorKind`, not a new updater.
| `CODEMAN_RESTART_BY_EXIT=1` | The Compose file's declaration of that policy, so the updater may exit even with no Docker socket. |
| Toolchain + devDependencies in the image | Lets `npm install` and `npm run build` run inside the container. |
| `docker-env-applied.json` | Fingerprint baseline, written by `Start-Codeman.sh` on every start. |
| `docker-build-source.json` | What HEAD/`package-lock.json` the build artefact volumes currently reflect. Written by both `Start-Codeman.sh` and this in-place update, so the two agree on whether those volumes are stale. |
### Why build artefacts are in named volumes
@@ -72,6 +73,20 @@ Docker seeds an empty named volume from the image, so the first start inherits t
image's already-built `node_modules` and `dist` and pays no bootstrap cost.
`docker compose down -v` is the supported reset: the next start re-seeds them.
That seeding-only-while-empty behaviour has a second, less obvious edge: it also
means a plain `docker compose build` triggered from OUTSIDE the container (for
example `Start-Codeman.sh`, after a `git pull` done by hand rather than through
this in-app updater) produces a fresh image whose freshly-built `dist`/
`node_modules` then sit unused behind the volumes' OLD content — the container
comes back up looking unchanged. `Start-Codeman.sh` detects this by comparing the
checkout's current HEAD and `package-lock.json` hash against `docker-build-source.json`,
and clears just the affected volume(s) before its own `--build` if they moved.
This in-place update writes that same file after a successful build precisely so
that comparison does not fire on stale information: without it, the next plain
`Start-Codeman.sh` run would see the HEAD this update just checked out, not
recognise it as already accounted for, and wipe the volumes this update just
correctly rebuilt right back to the OLDER image.
### Why the runtime image carries a build toolchain
`npm run build` is `tsc` plus `esbuild`, both devDependencies, so the image no
+6 -3
View File
@@ -333,9 +333,12 @@ Out of scope per the issue, and the current behavior already degrades correctly:
- **Docker cases**: the workspace is a host directory bind-mounted at the same absolute path, so a host-side
write is visible in the container immediately. Edit mode works and needs nothing special. Worth one line
in the docs.
- **Remote SSH cases**: `workingDir` is a path on the remote host. `validateSessionFilePath` realpaths it
locally, which fails, so the write returns 404 exactly like the read routes do today. Confirm the viewer
shows a clean empty/error state rather than an unexplained failure, and do not attempt an SFTP path.
- **Remote SSH cases**: `workingDir` is a path on the remote host, and the READ routes now
resolve it over ssh (`src/remote-files.ts`, same `buildSshConnectionArgs` discipline as the
launch path — #415). What stays unsupported is the WRITE side: an `edit=1` / `PUT` answers
`400` "editing is not supported for files in a remote (SSH) case", `editable` is always
`false`, office previews and generated thumbnails answer `400`, and no remote file is ever
copied to the server's disk. Do not attempt an SFTP write path.
---
+11 -2
View File
@@ -50,8 +50,17 @@ each `(clientId, seq)` at most once, so a resend can't type the prompt twice.
last-applied is seen. A replayed/lower seq returns `false`. Bounded MRU map
(`MAX_INPUT_DEDUP_CLIENTS = 256`).
- **WS route** (`ws-routes.ts`) — parses optional `cid`/`seq` on `{t:'i'}`; applies
via `shouldApplyInput` (skips a duplicate, still ACKs with `{t:'ia',seq}` so the
client drops it). Untagged frames apply unconditionally (no behavior change).
via `shouldApplyInput`. An applied frame is ACKed with `{t:'ia',seq}`; a duplicate is
ACKed as `{t:'ia',seq,dup:true,last:<watermark>}`, where `last` is the server's
highest applied seq for that `clientId` (`Session.lastInputSeq`). The client drops
the record either way, and on `dup` it lifts its own counter to `last` first and
re-sends a FIRST-attempt record (a retry being called a duplicate is the mechanism
working: the original landed). Without `last`, a tab killed between a send and the
persisted counter write came back counting BELOW the server's watermark, and every
later keystroke was dropped-but-ACKed: a silently dead terminal a reload could not
fix, since the stale counter was restored from localStorage too. The client now
persists the counter synchronously on every send for the same reason. Untagged
frames apply unconditionally (no behavior change).
- **POST route** (`/api/sessions/:id/input`) — optional `seq`/`clientId` in
`SessionInputWithLimitSchema`; a deduped duplicate returns 200 without writing
(the 200 is the client's ACK). `curl`/legacy callers omit the fields and always
+101
View File
@@ -251,6 +251,107 @@ unreachable host answers "unknown", which also means do not revive. The answer
is cached per session and cleared whenever the pane is next seen alive, so a
stale `true` from one transport drop can never revive the NEXT clean exit.
## File access over SSH
A remote case's `workingDir` is an absolute path on the **remote** host
(`Session.workingDir = RemoteCase.remotePath`), so the file routes cannot use local
`fs`: a local `realpathSync` on a remote-only path fails by construction, which is why
previewing a file used to answer `404 File not found` for a case that was working
perfectly (#415). `src/remote-files.ts` is the one module that reads remote bytes,
and it follows the same rule as the launch path: every ssh command line comes from
`buildSshConnectionArgs()` — **never** a hand-built ssh line.
| Request | What happens |
|---------|--------------|
| `GET /api/sessions/:id/file-raw` | Streamed over `ssh` (`cat`, or `tail -c +N \| head -c L` for a `Range`); the same 200/206/416 contract as a local file, so `<video>`/`<audio>` seeking works |
| `GET /api/sessions/:id/file-content` | `cat` into memory, capped by the existing text limit; `edit=1` answers `400` (see below) and `editable` is always `false` |
| `PUT /api/sessions/:id/file-content` | `400` before any path is looked at: the guard sits AHEAD of the local path validation, because with a same-named directory on the Codeman host (an `sshfs` mount) the write would otherwise land on the local twin |
| `GET /api/sessions/:id/file-preview` | Non-office files redirect to `file-raw` (which works remotely); docx/pptx answer `400` |
| `GET /api/sessions/:id/file-thumbnail` | `400` for remote files |
| `POST /api/sessions/:id/attachments` | Registers an absolute path that lives on the **remote** host (a clicked link pointing outside the case directory) by probing it there |
| `GET /api/sessions/:id/attachments/:attachmentId/raw` | Streams the registered remote file over ssh, same 200/206/416 contract; `preview` (office) and `thumbnail` answer `400` |
| `GET /api/sessions/:id/attachments/:attachmentId`, `GET …/attachments` (history) | Size/mtime/existence resolved over ssh, so a remote entry is not reported `missing`; the history list resolves EVERY entry in one batched probe, never one connection per entry |
⚠️ The attachment route is the one a clicked path takes when it is **outside** the case
directory (a remote `/tmp` scratchpad capture, a screenshot elsewhere in the home dir):
the frontend's `_isExternalPreviewPath()` sends every absolute path that is not under
`workingDir` there, so fixing only `file-raw` would leave exactly that half broken.
Guard order is deliberately **the same as locally**, and the checks are not weakened
by the transport:
1. Ownership (`findSessionOrFail` / the scope helper) — unchanged.
2. Lexical containment of `workingDir + path` — a `../` escape is refused before any
connection is opened.
3. ONE ssh round trip that returns `realpath` **and** `stat` for the path **and** the
workspace root (`remoteProbePaths`). Resolving the root remotely is what keeps the
boundary honest for a symlinked `remotePath`. The probe uses `readlink -f` when
available; on a host without it (macOS before 12.3) a POSIX fallback canonicalizes
the directory chain with `cd -P`/`pwd -P` and then follows the LAST component with
plain `readlink` for a bounded number of hops. ⚠️ **The fallback fails closed**: a
path it cannot fully resolve (a loop, a `readlink` failure, the hop cap) is reported
as unresolvable and answers 404, never as its own unresolved string. An earlier
version resolved only the directory chain, so `ws/notes.txt -> ~/.ssh/id_rsa` passed
containment under the link's own path while `cat` followed it to the key.
Records come back NUL-separated and index-keyed (`<index>|kind|size|mtime|realPath`,
after a leading NUL that fences off any login banner), so a filename containing a
newline cannot shift the alignment.
4. Containment of the remote realpath against the remote root. The sensitive-path
blocklist then applies on whichever routes already apply it locally (`/api/download`,
attachment registration, edit mode — where resolving symlinks first is what makes it
meaningful); the remote branch neither drops a guard the local path has nor invents a
stricter one. One entry of that blocklist is host-bound by construction: the three
home-anchored members (`~/.claude.json`, `~/.claude/settings.json`,
`~/.claude/settings.local.json`) are compared against the **Codeman host's** home
directory, so they do not match a remote home at a different path. Everything else in
the list is depth-anchored (`/.ssh/`, `/.aws/credentials`, `/.claude/.credentials.json`,
`/etc/shadow`, ...) and applies to a remote path unchanged.
5. Size cap (`CODEMAN_MAX_DOWNLOAD_BYTES`) applied to the **remote** size, before the
body is requested.
The path arrives from the browser (`?path=`) and is interpolated as a single
`shellescape`-quoted token, in a command that is itself shellescaped into the ssh
line; `BatchMode=yes` means a host needing a passphrase fails fast instead of hanging.
A failed connection is reported as **502** with the remote reason — never a 404, which
used to make an unreachable host look like a typo in the agent's output. The reason is
the first stderr line, the timeout, or the exit code; never Node's `Command failed: …`
message, which would carry the identity-file path and the probe script into the body.
**Connections are bounded.** Every probe and buffered read runs through a small global
semaphore (`src/remote-ssh-limiter.ts`, default 4, `CODEMAN_MAX_REMOTE_FILE_SSH`), the
attachment-history list resolves its whole history in one batched probe instead of one
handshake per entry, and probes are chunked at 40 paths per round trip. Terminal output
in a remote session is written on the remote host, so a prompt-injected agent printing
hundreds of `codeman://attach` links used to make the server fork one `ssh` per link,
each holding a 20 s probe timeout, and a 100-entry history re-listed on every
`attachment:detected` event tripped OpenSSH's default `MaxStartups 10:30:100`. Streams
(`file-raw`, by-id `raw`) are not counted: one is held per browser request for the life
of a playback, and each is gated behind a counted probe anyway.
⚠️ **There is deliberately NO local fallback.** A remote case reads the remote bytes or
fails, even when a file with the same absolute name exists on the Codeman host — which
is the ordinary case for the documented stop-gap workaround, an `sshfs` mount of the
remote tree at the identical path. Serving the local twin instead would silently hand
back a DIFFERENT filesystem's bytes under a name the user believes is the remote file
(a stale mount, a different checkout, a leftover file), and the failure would be
invisible. An existing mount therefore stops being load-bearing for previews and
downloads but is harmless, and a missing remote file stays a 404 even if the mount
still has it.
**Not available over ssh (by choice, not by accident):** editing a file (writes would
need SFTP; `docs/file-viewer-edit-plan.md` §6), office-document previews and
generated thumbnails (both need the bytes on the server's disk — no remote file is ever
spilled onto the server), the file-tree/picker listings, and `tail-file`. Those routes
are still local-only, so with an `sshfs` mount in place they read the mounted copy —
the two views can only disagree when that mount is stale. Docker cases are unaffected:
their workspace is bind-mounted at the same absolute path, so local `fs` reads real bytes.
⚠️ A remote record stores the **remote** path, and the same absolute path STRING means a
different file on each host. What decides which host to read is therefore never the
path but the SESSION (`session.remote`): a remote session never falls back to local
`fs`, and a local session never opens an ssh connection — including for attachment
records, which are keyed to the session that registered them.
## API
Routes are registered in `src/web/routes/case-routes.ts`:
+5 -2
View File
@@ -125,7 +125,9 @@ loopback bind matters. The auth pipeline (`src/web/middleware/auth.ts`,
`onRequest` hook) runs in this order:
1. **Localhost‑only exemptions** (always first): `POST /api/hook-event` and the QR
`/q/` short‑code path are exempt when `req.ip` is loopback (see §3). While the
`/q/` short‑code path are exempt when `req.ip` is loopback (see §3). The three
web‑tab exemptions (§10b: the capability in the path, the `Referer` form, and
the lost‑frame recovery page) sit in this same slot, ahead of the credential checks. While the
**managed tunnel is running**, the hook‑event exemption additionally requires
the per‑instance `X-Codeman-Hook-Secret` header (COD‑54); failed presentations
are rate‑limited in a **dedicated bucket** (separate from Basic‑Auth failures)
@@ -514,9 +516,10 @@ Full feature guide: [`docker-cases.md`](docker-cases.md).
## 10b. Web tabs (dashboard proxy)
A saved dashboard URL renders as a tab, served through Codeman's own origin at `/webview/<capability>/`. User guide: [`web-tabs.md`](web-tabs.md). Three properties carry the security weight:
A saved dashboard URL renders as a tab, served through Codeman's own origin at `/webview/<capability>/`. User guide: [`web-tabs.md`](web-tabs.md). Four properties carry the security weight:
- **The proxy is exempt from cookie auth and the Origin/CSRF guard, and that is deliberate.** The iframe is sandboxed without `allow-same-origin`, so it is opaque‑origin: its requests are cross‑site, meaning the `SameSite=lax` session cookie is never attached and its writes and WS upgrades arrive with `Origin: null`. The credential is instead a 192‑bit capability in the path, minted only by an authenticated `POST /api/webviews/:id/open`, held in memory (a restart invalidates every one), rolling TTL, bound to the minting user, and granting nothing but "relay bytes to this one saved URL". ⚠️ **The Host allowlist is NOT bypassed**, so DNS‑rebinding protection is unaffected. A second `Referer`‑keyed form exists for root‑absolute assets and is the only exemption decided by a request‑supplied header, so it is fenced to safe methods on non‑`/api`, non‑`/ws`, non‑`/q` paths. Edges pinned by `test/webview-auth-exemption.test.ts`.
- **The lost‑frame recovery page is the third unauthenticated 200, and the only one decided by request headers alone.** The proxy's runtime shim masks `/webview/<cap>/` off the page's own URL so a single‑page app routes on the path it expects; a navigation the page then starts itself (`location.reload()`, a root‑absolute `location.href`) lands on Codeman's root with no capability anywhere, no cookie (opaque origin) and a Referer naming the masked page. `serveLostWebviewFrame()` in `middleware/auth.ts` recognises it by shape (`GET`/`HEAD`, `Sec-Fetch-Dest: iframe` or `frame`, `Accept: text/html`, `Sec-Fetch-Mode: navigate` or absent) and answers, BEFORE the credential checks and without counting an auth failure, with a static page whose only content is a `postMessage` of the lost path to the parent tab (`default-src 'none'` plus the hash of that one script, `no-store`, `referrer: no-referrer`, no reflected input). It is fenced to paths that are NOT registered routes and never `/api/`, `/ws/` or `/q/`, with one carve‑out: `/` itself, because the landing page masks to exactly `/` and its reload otherwise rendered Codeman's app shell inside the web tab. `/` is admitted only when the request carries neither the `codeman_session` cookie nor an `Authorization` header: nothing in Codeman frames its own root and a sandboxed frame has neither, while a framed `/` that does carry credentials still gets the shell. On a passwordless install no auth hook runs, so the index route applies the same test itself (`isLostWebviewRootFrame`). ⚠️ Known property, accepted rather than mitigated: those headers are trivially set by a non‑browser client, so an unauthenticated caller can distinguish a registered route (401) from a non‑route (200) and enumerate the route table; the routes are public in `docs/api-reference.md`, so nothing is learned. Pinned by `test/webview-auth-exemption.test.ts` (password) and `test/webview-lost-root-frame.test.ts` (passwordless).
- **Sandboxed by default; `allow-same-origin` is an explicit per‑dashboard opt‑in.** A proxied page is same‑origin with Codeman, so without the sandbox its JavaScript could read the Codeman document and call the agent‑spawning API. ⚠️ In BOTH modes the `Authorization` header and the `codeman_session` cookie are stripped before the upstream request, because a trusted (same‑origin) frame makes the browser attach Codeman's own Basic‑auth credentials to every proxied request; forwarding them would hand `CODEMAN_PASSWORD` to the dashboard.
- **Not an open relay, and not a privilege boundary.** `resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from). The one refused destination class is link‑local and cloud‑metadata addresses (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`, `168.63.129.16`, `100.100.100.200`, `metadata.google.internal`): `webview-egress-policy.ts` refuses them at save time, and `webview-egress.ts` re‑judges the RESOLVED address at connect time through a `lookup` hook on the proxy's undici Agent and on its WebSocket client, so a DNS name pointing into those ranges is refused as well. Loopback and RFC1918 stay allowed on purpose. Capabilities are revoked on logout, admin logout and user deletion, and proxied responses carry `Referrer-Policy: same-origin` so a dashboard cannot hand the capability‑bearing URL to a third‑party host it links.
+34 -4
View File
@@ -159,6 +159,24 @@ layers cooperate so a dashboard talking to its own backend just works:
using its `Referer` to identify the dashboard. This only fires for a request
that already missed every Codeman route, and never for one that resolves to a
real route, which is what keeps it from being an authentication bypass.
5. The same script **masks the proxy prefix off the page's own URL** before any
of the page's code runs (`history.replaceState` to the path the page would see
on its own origin). A single-page app routes on `location.pathname` at boot,
and `/webview/<cap>/` is a path no app has a route for: without this, a React
Router / Vue Router / Next dev server painted its HTML and CSS and then replaced
them with its own "page not found" the moment its script ran. The page only
*reads* the masked path; every URL it emits still goes through the layers above.
6. A navigation the page starts **itself** after that — `location.reload()` (a dev
server's full-reload HMR), a root-absolute `location.href = '/login'` — now
targets Codeman's root with no capability anywhere on it. Codeman recognises
that request by shape (a top-level `<iframe>` navigation asking for HTML, for a
path it does not serve) and answers a static page that does nothing but tell
the owning tab which path was lost; the tab remounts the frame inside the
prefix at that path. It never counts as a failed login, so a dev server that
reloads on every save cannot rate-limit its user out of Codeman. The landing
page is the one served path that gets the same answer: it masks to exactly
`/`, and a reload there is admitted as long as the request carries no Codeman
credentials, which a sandboxed frame never does.
On top of that, the proxy answers those requests with CORS headers. That sounds
wrong for same-host requests, but a sandboxed iframe has an *opaque* origin, so the
@@ -172,10 +190,22 @@ then every API call fails, which looks like the dashboard being broken.
EventSource, normal markup, the DOM sinks a page uses to build markup at runtime,
and `url()` inside stylesheets. Something that constructs requests by an unusual
route can still slip through. Symptom: the page renders but a panel stays empty.
- **Root-absolute `location` navigation.** A dashboard that navigates itself with
`location.href = '/login'` escapes the prefix, because `Location.href` is
unforgeable and cannot be patched the way the other sinks are. A relative
`location.href = 'login'` is fine (`<base>` covers it).
- **A root-absolute `url()` inside an inline `<style>` is not rescued.** Masking the
page's URL (layer 5) trades away the `Referer` safety net of layer 4 for
requests the shim cannot see, and only HTML is rewritten server-side. An
external stylesheet is fine: a `url()` it references is fetched with the
stylesheet's own URL as `Referer`, which is still inside the prefix. A
root-absolute `url(/img.png)` written directly into a `<style>` block in the
document has the masked document as its `Referer`, so it 404s where the
fallback used to rescue it. Symptom: one background image missing while
everything else renders. Narrow, and a `url()` the page sets from script is
still covered by layer 3.
- **Root-absolute `location` navigation is recovered, not prevented.** `Location`
is unforgeable, so `location.href = '/login'` or `location.reload()` really does
leave the prefix; the frame comes back through the recovery hop in layer 6 above,
which needs a browser that sends `Sec-Fetch-Dest` (every current one; iOS Safari
since 16.4). Older browsers show Codeman's 404 in the frame; the tab's **Reload**
button puts it back.
- **Cross-origin redirects are not followed.** If a dashboard bounces to a different
host (an external SSO provider, say), the proxy hands the redirect back unchanged
rather than relaying it, because relaying would make this an open proxy. Use
+85 -11
View File
@@ -1,9 +1,9 @@
# Agent CLIs
Codeman drives seven run modes: six agent CLIs plus a plain shell. This page covers picking
Codeman drives ten run modes: nine agent CLIs plus a plain shell. This page covers picking
one, setting it up, and the differences that actually change how you work.
## The seven modes
## The ten modes
| Mode | CLI | Get it |
| -------------------- | ---------------------------- | ---------------------------------------------------------------------- |
@@ -13,6 +13,9 @@ one, setting it up, and the differences that actually change how you work.
| **Gemini** | `gemini` | [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) |
| **Antigravity** | `agy` | [antigravity.google](https://antigravity.google) |
| **Pi** | `pi` | [pi.dev](https://pi.dev) |
| **Grok Build** | `grok` | [github.com/xai-org/grok-build](https://github.com/xai-org/grok-build) |
| **DeepSeek Harness** | `dsh` | [github.com/deepseek-ai/deepseek-harness](https://github.com/deepseek-ai/deepseek-harness) |
| **OMP** | `omp` | [github.com/can1357/oh-my-pi](https://github.com/can1357/oh-my-pi) |
| **Terminal / Shell** | your `$SHELL` | Already installed. |
Any combination works, including all of them. The run mode is chosen per session from the
@@ -47,8 +50,12 @@ If a CLI is installed but a Run button for it never appears:
precisely to avoid this; a hand-written plist or unit will not.
3. Restart the server after installing a new CLI.
`pi` is additionally version-probed rather than trusted by name, because `pi` is a generic
enough command that something else on your PATH may answer to it.
`pi`, `grok`, `omp` and `dsh` are additionally identity-probed rather than trusted by name:
`pi` and `omp` are generic enough that something else on your PATH may answer to them,
`grok` has npm squatters, and Debian ships an unrelated `dsh` (dancer's shell). Each has a
status endpoint (`/api/grok/status`, `/api/deepseek/status`, `/api/omp/status`) that reports
the path and version that actually resolved, so a misresolution is visible rather than
presenting as "the mode just does not work".
## Claude is the reference mode
@@ -62,15 +69,15 @@ output. The other CLIs expose no equivalent.
| Respawn cycling and unattended runs | Yes | Yes |
| Cron jobs | Yes | Yes |
| Docker cases, remote SSH cases | Yes | Yes |
| Precise idle detection (hook-driven) | Yes | Output-stabilization fallback, coarser |
| Precise idle detection | Yes | Codex: same screen check, via its own prompt and working line. DeepSeek: reports its state itself. Others: output stabilization, coarser |
| Auto-resume when a usage limit resets | Yes | No |
| Plan usage chip | Yes | No |
| Approvals Inbox | Yes | No |
| Approvals Inbox | Yes | DeepSeek yes; others no |
| Read My Mind | Yes | No |
| Ralph loop and its task tracker | Yes | No |
| Subagent and team windows | Yes | No |
| Model, effort, and ultracode controls | Yes | No |
| `stop` and `blocked` wait signals | Yes | 400 if you ask for them explicitly |
| `stop` and `blocked` wait signals | Yes | DeepSeek yes; elsewhere 400 if you ask for them explicitly |
| The bundled agent skill | Yes | No |
Everything that makes a session a session works everywhere. What is Claude-only is mostly
@@ -124,6 +131,11 @@ Two behaviours that are deliberate and worth knowing:
- **The wheel is not forwarded** into its transcript. Codex ignores the mouse reports
Codeman would send, so forwarding produced a dead wheel. Scrolling in a Codex session is
local scrollback.
- **Work detection is Codex's own.** Codex declares its `›` composer glyph and its
`esc to interrupt` working line, so it gets the same screen-checked idle detection Claude
does; before 1.26.1 every Codex session reported idle for its whole life. Codex
conversations also appear in Past Sessions and can be resumed, and on phones the keyboard
bar grows `⇧←` / `⇧→` for Codex's queued-message editing and prompt stack.
### Gemini
@@ -157,6 +169,60 @@ Pi needs the opposite instincts from every other CLI here.
Guide: [`docs/pi-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/pi-integration.md).
### Grok Build
xAI's `grok`, installed with `curl -fsSL https://x.ai/cli/install.sh | bash` into
`~/.grok/bin`. Codex-shaped on permissions and OpenCode-shaped on rendering:
- **Its bypass switch is `--always-approve`**, Grok's own `bypassPermissions` mode, and the
Run button sends it the way it sends Codex's. In multi-user mode a user without a grant
has it stripped.
- **Authentication is Grok's own**: browser OAuth on first run (a device-code screen inside
a Codeman pane), `grok login --device-auth` for headless hosts, or `XAI_API_KEY` as a
per-session environment override.
- It renders a full-screen TUI, so scrolling is local scrollback.
Guide: [`docs/grok-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/grok-integration.md).
### DeepSeek Harness
The mode wired least like the others, for two reasons worth knowing before you use it.
**`dsh` is a launcher, not an agent.** It boots a *profile*, and the three DeepSeek ships
(`web`, `headless`, `base`) cannot drive a terminal pane. So "installed" and "runnable" are
different questions: the Run menu offers **DeepSeek** only once a pane-capable profile
exists, and until then shows **DeepSeek — add a terminal profile…**, which installs the
community `dsh-tui` with one click (`pnpm` must be on PATH, because the launcher spawns it
directly).
**Permissions are an environment variable, not a flag.** The harness has no
skip-permissions switch. `DSH_PERMISSION_MODE` (`read-only`, `workspace-write`,
`danger-full-access`) is the whole control, and it is the one setting Codeman deliberately
carries as an environment variable, because the harness reads it as a soft boot-time
default. In multi-user mode a user without a grant is clamped to `workspace-write`.
The reward for the odd wiring: **DeepSeek is the one non-Claude mode with real signals.**
Its terminal front door reports idle, working and blocked to Codeman, so a DeepSeek
session gets precise idle detection, the `stop` and `blocked` wait signals, and Approvals
Inbox items. Answers are read from the harness's own transcript on disk rather than
scraped off the pane. The model is not a session setting; it is part of the profile.
Guide: [`docs/deepseek-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/deepseek-integration.md).
### OMP
Oh My Pi, installed with `curl -fsSL https://omp.sh/install | sh` into `~/.local/bin`.
OMP owns its auth, provider routing and approval mode entirely in `~/.omp`: there is no
Codeman-side login, key field, or bypass switch. Run `omp` once outside Codeman to finish
its own onboarding, and every session started through Codeman inherits that config. Its
documented default approval mode is `yolo`, so an OMP pane auto-approves tool use with no
flag from Codeman; change that in OMP's own config, not here.
OMP conversations appear in Past Sessions and can be resumed, and a respawn continues the
same conversation with `--continue`.
Guide: [`docs/omp-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/omp-integration.md).
### Terminal / Shell
A plain shell in a tmux session. No agent, no hooks, no idle detection.
@@ -180,9 +246,15 @@ respawns. Which variables are accepted depends on the mode:
| Gemini | `GEMINI_*`, `GOOGLE_*` |
| Antigravity | `ANTIGRAVITY_*` |
| Pi | `PI_*` |
| Grok | `GROK_*`, `XAI_*` |
| DeepSeek | `DSH_*`, `DEEPSEEK_*` |
| OMP | `OMP_*` |
Anything outside the allowlist is rejected at the schema. This is intentional: the allowlist
is one global list, so widening it for one CLI widens it for all of them.
is one global list, so widening it for one CLI widens it for all of them. In multi-user mode
the keys that could redirect a CLI's traffic or move its config home (`DSH_PERMISSION_MODE`,
`DSH_HOME`, `DEEPSEEK_BASE_URL`, `OMP_AUTH_BROKER_URL`, and the base URLs and config
directories of the others) are dropped for a user without the bypass grant.
Two things that deliberately do **not** travel as environment variables: **effort**, because
an environment variable hard-locks it and blocks `/effort`, and **model**, which is written
@@ -192,9 +264,11 @@ into the case's `.claude/settings.local.json` so that `/model` keeps working.
- **Claude Code** if you want every Codeman feature. Unattended overnight runs, usage-limit
auto-resume, the Approvals Inbox, and subagent visualization all assume it.
- **Codex, OpenCode, Gemini, Antigravity** when you prefer that agent or that model. You get
the session layer, respawn, cron, Docker, and remote SSH; you do not get the hook-driven
features.
- **Codex, OpenCode, Gemini, Antigravity, Grok, OMP** when you prefer that agent or that
model. You get the session layer, respawn, cron, Docker, and remote SSH; you do not get the
hook-driven features.
- **DeepSeek Harness** if you want DeepSeek's models with real status signals. It is the one
non-Claude mode that reports idle, working and blocked to Codeman itself.
- **Pi** if you want a fast, unsandboxed agent and you understand what project trust does.
- **Shell** for the times you want a terminal on your phone with no agent at all. It is a
genuinely useful mode, not a fallback.
+1 -1
View File
@@ -108,7 +108,7 @@ Conventions for wiki pages:
- Images are referenced from the main repository over raw URLs rather than being copied into
the wiki.
- Say what the default is, especially when it is off. Most of Codeman is opt-in.
- Label Claude-only behaviour every time it appears. Six of the seven run modes are not
- Label Claude-only behaviour every time it appears. Nine of the ten run modes are not
Claude.
## Conduct
+11 -9
View File
@@ -50,10 +50,10 @@ A session carries state the case does not:
## Run mode
The **run mode** is which CLI the session runs: `claude`, `opencode`, `codex`, `gemini`,
`antigravity`, `pi`, or `shell`. It is chosen at start and does not change afterwards; to
`antigravity`, `pi`, `grok`, `deepseek`, `omp`, or `shell`. It is chosen at start and does not change afterwards; to
switch, start another session.
Claude is the reference mode. Six of the seven are not Claude, and a number of Codeman
Claude is the reference mode. Nine of the ten are not Claude, and a number of Codeman
features are Claude-only for structural reasons rather than missing effort: they depend on
Claude Code's hook system or on parsing its terminal output. Every such feature is labelled
Claude-only where it appears, and [Agent CLIs](Agent-CLIs) lists them in one place.
@@ -68,8 +68,8 @@ Where a case runs is **separate from** which CLI it runs. There are three locati
| **Docker** | One long-lived container per case; sessions `docker exec` into it. See [Docker Cases](Docker-Cases). |
| **Remote SSH** | A durable tmux server on the remote host, fronted by a local pane running `ssh`. See [Remote SSH Sessions](Remote-SSH-Sessions). |
This matters because it is a common source of confusion: Docker is **not** an eighth run
mode. All seven run modes work in all three locations. A case is docker-backed or
This matters because it is a common source of confusion: Docker is **not** an eleventh run
mode. All ten run modes work in all three locations. A case is docker-backed or
ssh-backed; a session is claude or codex or shell.
**Web tabs** are the other thing that is not a session. A saved dashboard URL renders as a
@@ -155,9 +155,11 @@ report events back: a permission prompt appeared, the turn finished, the agent w
task completed. Those events drive tab alerts, the Approvals Inbox, notifications, and the
wait primitives.
This is why some features are Claude-only. The other CLIs have no equivalent hook system,
so for them Codeman falls back to watching terminal output, which is coarser: it can see
that something happened, not what it was.
This is why some features are Claude-only. The one partial exception is DeepSeek Harness,
whose terminal front door reports idle, working and blocked to Codeman over the harness's
own supervisor contract, so it gets the hook-driven signals without a hook file. The other
CLIs have no equivalent, so for them Codeman falls back to watching terminal output, which
is coarser: it can see that something happened, not what it was.
See [Hooks And Integrations](Hooks-And-Integrations).
@@ -167,7 +169,7 @@ See [Hooks And Integrations](Hooks-And-Integrations).
| --------------- | ---------------------------------------------------------------------------- |
| **Case** | Named working directory. |
| **Session** | One CLI in one tmux session. |
| **Run mode** | Which CLI: claude, opencode, codex, gemini, antigravity, pi, shell. |
| **Run mode** | Which CLI: claude, opencode, codex, gemini, antigravity, pi, grok, deepseek, omp, shell. |
| **Respawn** | Restarting the CLI on idle to keep an unattended run going. |
| **Ralph loop** | An autonomous single-session task loop. |
| **Orchestrator**| A phased plan driven across multiple agents. |
@@ -178,6 +180,6 @@ See [Hooks And Integrations](Hooks-And-Integrations).
## Read next
- [The Dashboard](The-Dashboard) - what the UI is showing you.
- [Agent CLIs](Agent-CLIs) - the seven run modes in detail.
- [Agent CLIs](Agent-CLIs) - the ten run modes in detail.
- [Keeping Agents Running](Keeping-Agents-Running) - respawn, idle detection, usage limits.
- [`docs/architecture-invariants.md`](https://github.com/Ark0N/Codeman/blob/master/docs/architecture-invariants.md) - the mechanisms behind all of this, for contributors.
+27 -6
View File
@@ -4,7 +4,7 @@ Run a case inside its own container instead of directly on your host: for isolat
reproducible toolchain, and for the ability to pick the whole environment up and move it to
another machine.
A docker case is a **location overlay**, not a run mode. All seven run modes work inside a
A docker case is a **location overlay**, not a run mode. All ten run modes work inside a
container. See [Core Concepts](Core-Concepts).
## One-time setup: the base image
@@ -26,7 +26,7 @@ A zero exit code proves the layers ran, not that the toolchain works. Verify:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy pi; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
'for c in claude codex gemini opencode agy pi grok dsh omp; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
The image is secret-free. Credentials are delivered at runtime, never baked in, so exports
@@ -79,6 +79,25 @@ Exactly one long-lived container per case, shared by every session in it.
conversation** from the bind-mounted transcript.
- Deleting the case removes the container. The workspace on the host survives.
## Attaching to a container you already run
Tick **Attach to an existing container** on **Add Case → Docker** to link a case to a
container that already exists instead of creating one. Codeman only `exec`s into it and
never creates, starts, stops, restarts or removes it, so a container that is missing or
stopped fails with a message rather than being fixed for you. Drift detection does not
apply (the container carries no Codeman configuration label). The full-image export is
refused, since it would `docker commit` someone else's container, and the workspace export
skips the pause that keeps an owned container consistent during the capture.
One adopted container can back several cases at different in-container directories, and
**copy an existing case** pre-fills the form from a sibling on the same container. An exact
twin (the same container and the same directory) is refused, as is a container another
user adopted.
Adoption is **admin-only in multi-user mode**. Linking creates Codeman's own container
with one bind mount that has already been checked; an adopted container's mounts belong to
whoever started it, and one that mounts `/` hands the adopter the host.
## Credentials
Your existing host logins work inside the container without logging in again. Credentials
@@ -92,10 +111,12 @@ the container instead.
Bind mounts are excluded from image capture, so exports stay secret-free.
One consequence worth knowing: Pi's credentials are seeded per file rather than as a whole
directory, because that directory also holds sessions, extensions, and installed packages,
which can be gigabytes. So in-container Pi sessions are invisible from the host, and `pi -c`
inside a docker case sees only that container's history.
One consequence worth knowing: Pi, Grok and OMP credentials are seeded per file rather than
as whole directories, because those directories also hold sessions, extensions, downloads and
installed packages, which can be gigabytes. So in-container Pi and Grok sessions are
invisible from the host (`pi -c` and `grok -c` inside a docker case see only that
container's history). OMP's `sessions/` is the exception and is shared read-write, because
Codeman reads it host-side for history and resume.
## Isolation
+18 -5
View File
@@ -54,7 +54,10 @@ create-time sweep would yank the skill out from under other live sessions sharin
directory. Remove them per case with `codeman skill uninstall --case <name>`.
The skill ships with the verb index always loaded, plus on-demand references for the verbs,
worked multi-worker recipes, endpoint tables, and cross-session messaging.
worked multi-worker recipes, endpoint tables, and cross-session messaging. It drives
DeepSeek Harness workers the same way it drives Claude ones (`spawn_workers alpha
beta:deepseek` is a mixed fleet in one call), since those are the two modes with real
completion signals.
## The manual path
@@ -92,8 +95,9 @@ Read these before writing any code. Each one has cost somebody an afternoon.
5. **Wait instead of polling, and a timeout is not an error.** The wait endpoints answer
`200` with `wait.timedOut: true`. Loop over short waits rather than one long call, because
tunnels cut idle connections.
6. **Only `claude` sessions emit `stop` and `blocked`.** They come from Claude Code hooks.
Shell and the external CLIs accept only `idle`, `working`, and `exit`; asking for `stop`
6. **Only `claude` and `deepseek` sessions emit `stop` and `blocked`.** Claude's come from
Claude Code hooks, DeepSeek's from the harness reporting its state to Codeman. Shell and
the other external CLIs accept only `idle`, `working`, and `exit`; asking for `stop`
explicitly there is a `400`, while omitting `until` is always safe. On a shell session
`idle` fires **once at startup and never again**, so synchronize hook-less sessions with an
output marker instead.
@@ -130,7 +134,10 @@ curl -s -X POST "$API/api/sessions/$ID/input" \
# Or wait for a marker in the output, which works on shell sessions too
curl -s "$API/api/sessions/$ID/wait-output?contains=DONE_17909&from=buffer" | jq
# Read the terminal back
# Read the last answer as clean text (claude, codex, deepseek sessions)
curl -s "$API/api/sessions/$ID/last-response" | jq -r '.data.text'
# Or read the terminal back
curl -s "$API/api/sessions/$ID/terminal?tail=4000" | jq -r '.data.output'
# Clean up, by exact id
@@ -157,7 +164,13 @@ Make it unique per call, because tmux repaints replay old screen text.
### Reading output
Use `terminal?tail=`, not `/output`. The latter's text field is empty for every tmux-backed
For `claude`, `codex` and `deepseek` sessions, read the answer from the transcript rather
than the screen: `GET /api/sessions/:id/last-response` returns the last reply as clean text
with no TUI frames or repaint noise. Poll it briefly rather than reading once, because the
transcript lands slightly after the `stop` signal, so a read immediately after send-and-wait
returns often comes back empty.
For everything else, use `terminal?tail=`, not `/output`. The latter's text field is empty for every tmux-backed
session, which is every interactive session. `tail` counts **bytes**, and what comes back is
terminal data with ANSI sequences included.
+6
View File
@@ -21,6 +21,12 @@ No. Codeman drives agent CLIs you have already installed and logged in yourself.
subscription or key that CLI uses is what pays for the tokens. Codeman never collects,
stores, or refreshes your credentials.
### Which agent CLIs does it support?
Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi, Grok Build, DeepSeek Harness and
OMP, plus a plain shell, chosen per session. Claude is the reference mode and a few features
are Claude-only; [Agent CLIs](Agent-CLIs) has the table.
### Does Codeman send my code or prompts anywhere?
No. There is no telemetry, no analytics, and no phone-home. The only network traffic
+18 -9
View File
@@ -66,14 +66,14 @@ self-signed certificate, add `-k`.
## Endpoint map
Roughly 200 handlers across 24 route modules. By domain:
Roughly 235 handlers across 26 route modules. By domain:
| Domain | Handlers | Covers |
| ------------------- | -------- | --------------------------------------------------- |
| System | 45 | Status, settings, search, digest, updates. |
| Sessions | 34 | Create, input, terminal, wait, kill. |
| Cases | 29 | Create, link, clone, remote and docker cases. |
| Files | 16 | Preview, edit, raw, attachments, path picker. |
| System | 56 | Status, settings, digest, updates, tunnel. |
| Sessions | 34 | Create, input, terminal, wait, last response, kill. |
| Cases | 34 | Create, link, clone, remote and docker cases. |
| Files | 17 | Preview, edit, raw, attachments, path picker. |
| Orchestrator | 10 | Plans and phases. |
| Ralph | 9 | Loop control and configuration. |
| Cron | 9 | Jobs and run history. |
@@ -82,10 +82,12 @@ Roughly 200 handlers across 24 route modules. By domain:
| Respawn | 7 | Respawn configuration and presets. |
| Webviews | 6 | Saved dashboards, plus the proxy. |
| Mux | 5 | tmux operations. |
| Custom model endpoints | 5 | Saved OpenAI-compatible endpoints, and applying one to a session. |
| Push | 4 | Web push subscriptions. |
| Read My Mind | 4 | Intent profiles and prediction. |
| Scheduled | 4 | The legacy scheduled-run concept. |
| Approvals | 3 | The inbox and answering. |
| Approvals | 4 | The inbox, answering, acknowledging. |
| Tab layout | 2 | Named tab groups per owner. |
| Teams, me, search, hooks, clipboard, telemetry, voice, ws | 1-2 each | |
Each route module documents its own endpoints in its file header.
@@ -114,12 +116,13 @@ Three semantics that break callers who assume otherwise:
`wait-output` matches a **literal substring, never a regex.** That is deliberate: no regex
means no catastrophic backtracking on attacker-influenced output.
Only `claude` sessions emit `stop` and `blocked`, because those come from Claude Code hooks.
Shell and external CLI sessions accept `idle`, `working`, and `exit`.
Only `claude` and `deepseek` sessions emit `stop` and `blocked`: Claude's come from Claude
Code hooks, DeepSeek's from the harness reporting its state to Codeman. Shell and the other
external CLI sessions accept `idle`, `working`, and `exit`.
## SSE
`GET /api/events` is the live event stream. 156 event names, kept in sync between server and
`GET /api/events` is the live event stream. 158 event names, kept in sync between server and
client with a test that fails on drift.
The heartbeat is a **named** `sse:heartbeat` event rather than an SSE comment, because
@@ -141,6 +144,12 @@ curl -s "$API/api/sessions" | jq '.data[].name' # live sessions
curl -s "$API/api/sessions/unified" | jq # live + historical, deduped
curl -s "$API/api/subagents" | jq # background agents
curl -s "$API/api/search?q=deploy" | jq # cross-session search
# with ID set to a session id:
curl -s "$API/api/sessions/$ID/last-response" | jq -r '.data.text' # last answer, from the transcript (claude, codex, deepseek)
curl -s "$API/api/model-endpoints" | jq # saved custom OpenAI-compatible endpoints
curl -s -X POST "$API/api/sessions/$ID/custom-model" -H 'Content-Type: application/json' \
-d '{"endpointId":"local-llama","modelId":"qwen3-27b"}' | jq # restart the CLI on that endpoint; {"clear":true} undoes it
```
## Limits
+4 -4
View File
@@ -5,8 +5,8 @@
<h3 align="center">Mission control for AI coding agents</h3>
Codeman runs your coding agents on your own machine and puts them behind one dashboard you
can open from any device. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, or
Pi inside persistent tmux sessions, streams the real terminal to the browser, and keeps
can open from any device. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi,
Grok, DeepSeek Harness, or OMP inside persistent tmux sessions, streams the real terminal to the browser, and keeps
working while you are away from the keyboard: it re-prompts idle agents, resumes when a
subscription limit resets, runs jobs on a schedule, and shows every background subagent
live.
@@ -33,7 +33,7 @@ codeman web # then open http://localhost:3000
**Already running it**
- [Agent CLIs](Agent-CLIs) - the seven run modes, their setup, and which features are Claude-only.
- [Agent CLIs](Agent-CLIs) - the ten run modes, their setup, and which features are Claude-only.
- [Mobile Guide](Mobile-Guide) - phone and tablet use, QR login, the touch keyboard bar.
- [Remote Access](Remote-Access) - Tailscale, Cloudflare tunnel, LAN plus password, QR login.
- [Keeping Agents Running](Keeping-Agents-Running) - idle detection, respawn cycling, auto-resume on usage limits.
@@ -122,7 +122,7 @@ codeman web # then open http://localhost:3000
| OS | macOS or Linux. Windows works through WSL2. |
| Node.js | 22 or newer. |
| tmux | Required. Sessions live in tmux, which is what makes them survive restarts. |
| An agent CLI | At least one of Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi. Plain shell sessions need none. |
| An agent CLI | At least one of Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi, Grok Build, DeepSeek Harness, OMP. Plain shell sessions need none. |
| Network | Binds to `127.0.0.1` by default. Reaching it from another device is a deliberate step: see [Remote Access](Remote-Access). |
Codeman is MIT licensed, self-hosted, and sends no telemetry. Everything runs on your
+6 -4
View File
@@ -19,9 +19,11 @@ terminal into something that can notify you.
| `teammate_idle` | An agent-team member goes idle. | Team surfaces. |
| `task_completed` | A task finishes. | Task tracking, run summary. |
This is why several Codeman features are Claude-only. The other CLIs have no hook system, so
for them Codeman watches terminal output, which reveals that something happened but not what
it was.
This is why several Codeman features are Claude-only. The one partial exception is DeepSeek
Harness, whose terminal front door reports idle, working and blocked to Codeman over the
harness's own supervisor contract, so it gets the hook-driven surfaces without any hook
file. The other CLIs have no equivalent, so for them Codeman watches terminal output, which
reveals that something happened but not what it was.
### How hooks get installed
@@ -67,7 +69,7 @@ sit beside the agents with no code at all. See [Web Tabs](Web-Tabs).
### 2. SSE events
`GET /api/events` streams everything Codeman knows: session lifecycle, output, agent
activity, approvals, cron runs. 155 named events, stable under semantic versioning.
activity, approvals, cron runs. 158 named events, stable under semantic versioning.
This is the seam for anything that reacts. A bot that pings your chat channel when an agent
needs a human is a short script over this stream.
+11
View File
@@ -27,6 +27,15 @@ The result is the property you want on a phone: a connection that drops mid-prom
loses the prompt and never delivers it twice. Two browser tabs on the same session coexist,
and only a reconnect from the *same* tab supersedes the old connection.
## Selecting and copying
Agent CLIs hold the mouse: clicks and drags are reported into the transcript rather than
selecting text. `Shift+drag` starts a selection anyway, right-click copies it (with nothing
selected the native context menu is left alone), and `Ctrl+Shift+C` copies without ever
interrupting. **Auto Copy Selection** in **App Settings → Terminal & Input**, off by
default, copies the moment you release the mouse. On phones, long-press selects; see
[Mobile Guide](Mobile-Guide).
## Zero-lag local echo
On touch devices, keystrokes are painted in the terminal immediately and sent when you press
@@ -55,6 +64,8 @@ reconcile against the real buffer and only apply while the cursor is on the comp
Chinese, Japanese, and Korean input needs an IME, and an IME needs a real text field.
Turning on CJK input in **App Settings → Terminal & Input** puts an always-visible textarea
below the terminal that owns composition, then delivers the composed text to the session.
Ctrl- and Alt-modified navigation keys typed through it reach the CLI as the modified
sequences, so word jumps and history keys keep working.
## Voice dictation
+26 -4
View File
@@ -9,7 +9,7 @@ Getting Codeman onto a machine, verifying it works, updating it, and removing it
| **macOS or Linux** | Windows works through WSL2. See [Windows](#windows-wsl) below. |
| **Node.js 22+** | The installer offers to install it if missing. |
| **tmux** | Not optional. Sessions live inside tmux, which is what makes them survive a server restart, a dropped connection, or a closed laptop. |
| **An agent CLI** | At least one of [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev). Plain shell sessions need none. See [Agent CLIs](Agent-CLIs). |
| **An agent CLI** | At least one of [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), [OMP](https://github.com/can1357/oh-my-pi). Plain shell sessions need none. See [Agent CLIs](Agent-CLIs). |
Codeman itself sends no telemetry and phones no home. The only network traffic is your
browser to your server, and whatever the agent CLI you chose does on its own.
@@ -20,13 +20,16 @@ browser to your server, and whatever the agent CLI you chose does on its own.
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if they are missing, clones Codeman into `~/.codeman/app`,
and builds it.
This installs Node.js, tmux and a build toolchain if they are missing (node-pty ships no
Linux prebuild, so it compiles from source), clones Codeman into `~/.codeman/app`, and
builds it.
What it asks you:
1. **Permission for every system change.** Package installs and agent CLI downloads are
prompted individually. Nothing is installed silently.
prompted individually. Nothing is installed silently. If no agent CLI is found, a menu
offers to install any of them (DeepSeek excepted: its npm package installs only a
launcher with no runnable profile), or you skip and install one yourself later.
2. **How the dashboard should be reachable.** Three choices:
- **Tailscale** (recommended for phone access): keeps the loopback bind and walks you
through `tailscale serve`, including the tailnet HTTPS toggle, then verifies the result
@@ -101,6 +104,21 @@ at server start, so markup changes need a restart.
See [Contributing](Contributing) for the rest of the development loop.
## Route D: Docker Compose
Codeman itself can run in a container and spawn Docker cases as sibling containers through
the host's Docker socket. Copy `docker/.env.example` to `docker/.env`, set
`CODEMAN_PASSWORD`, then:
```bash
bash docker/Start-Codeman.sh
```
Run the script again after updating rather than a plain `docker compose up`, so the rebuilt
image, the refreshed volumes and the entrypoint arrive together. The full guide, including
storage and networking options, is
[`docker/README.md`](https://github.com/Ark0N/Codeman/blob/master/docker/README.md).
## Installing an agent CLI
Codeman drives CLIs, it does not bundle them. Install at least one:
@@ -113,6 +131,9 @@ Codeman drives CLIs, it does not bundle them. Install at least one:
| **Antigravity** | See [antigravity.google](https://antigravity.google) | Google's successor to the consumer Gemini CLI. |
| **Gemini CLI** | See [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) | Enterprise only since Google's June 2026 consumer cutover. |
| **Pi** | See [pi.dev](https://pi.dev) | No permission prompts and no sandbox by design. Read [Agent CLIs](Agent-CLIs) before using it on a repo you care about. |
| **Grok Build** | `curl -fsSL https://x.ai/cli/install.sh \| bash` | xAI. Lands in `~/.grok/bin`; `grok login --device-auth` for headless hosts. |
| **DeepSeek Harness** | `npm i -g @deepseek-ai/dsh pnpm`, then a terminal profile | The npm package is only a launcher. Codeman's Run menu installs the community terminal profile for you. See [Agent CLIs](Agent-CLIs). |
| **OMP** | `curl -fsSL https://omp.sh/install \| sh` | Oh My Pi. Run it once by hand to finish its own onboarding. |
Log each CLI in once, by hand, before pointing Codeman at it. Codeman never collects or
stores your CLI credentials.
@@ -161,6 +182,7 @@ Full detail, including logs and the self-updater, is in
| Installer | Re-run the one-liner, or **App Settings → System → Updates** in the UI. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart. |
| Docker Compose | Re-run `Start-Codeman.sh`. The in-app updater works too, and refuses a release that changes the container definition until you re-run the script. |
The in-app updater covers git-clone installs supervised by systemd or launchd. It restarts
the process that is running it, so the actual work happens in a detached script and the
+16 -8
View File
@@ -27,9 +27,12 @@ keystroke echo. Idle now lands a few seconds after a turn genuinely ends.
There are several layers stacked on that: a completion message from the CLI, an AI check,
output silence, and token stability.
**For every other CLI**, there are no hooks to lean on, so detection is output
stabilization: the session is idle when output stops changing. Coarser, and it is why the
features further down this page are Claude-only.
**For the other CLIs** it depends on what the CLI tells Codeman. Codex declares its own
prompt glyph and working line, so it gets the same screen check Claude does (before 1.26.1
every Codex session reported idle for its whole life). DeepSeek Harness reports idle,
working and blocked to Codeman itself, which is as precise as hooks. Everything else is
output stabilization: the session is idle when output stops changing. Coarser, and it is
why the features further down this page are Claude-only.
## The Respawn Controller
@@ -101,13 +104,18 @@ subscription plan.
**Claude only.** A header chip showing live subscription usage, on by default on desktop and
off on phones.
It works by installing a status line exporter into Claude Code, which posts Claude's own
rate limit data back to Codeman. The exporter is marker-identified, so it only ever touches
a status line Codeman installed, never one you wrote yourself, and it prints your footer
through so the in-terminal status line still works.
It works through a status line exporter that Codeman hands to `claude` as an ephemeral
setting when it spawns the session, never written to disk, which posts Claude's own rate
limit data back to Codeman. Your own status line (project-local, project, then
`~/.claude/settings.json`) is wrapped and printed through, and a `claude` you run by hand
outside Codeman sees nothing of it. Workspaces an older Codeman wrote the exporter into are
cleaned up the first time a session starts there. Codex limits come from a read-only poll of
its own app-server. Known limit: sessions inside a Docker case do not feed the chip yet.
The chip and the exporter are the same setting. Turning the chip on without the exporter
would leave it showing a dash forever, so resolve it in one place: **App Settings**.
would leave it showing a dash forever, so resolve it in one place: **App Settings**. A
device writes the switch only when it flips the chip, so a phone (chip off by default)
saving its font size cannot switch collection off for your desktop.
## Circuit breakers
+3
View File
@@ -30,6 +30,9 @@ Press `Ctrl+?` in the app for the same list in a floating overlay.
| `Ctrl+Shift+R` | Restore terminal size. |
| `Ctrl` `+` / `Ctrl` `-` | Font size. |
| `Shift+Wheel` | Scroll the local buffer, even where the wheel is forwarded to the CLI. |
| `Shift+drag` | Start a selection in a pane whose mouse events go to the CLI. |
| Right-click | Copy the selection. With nothing selected the native menu is left alone. |
| `Ctrl+Z` | Swallowed in agent sessions so a running CLI cannot be suspended. Normal job control in a shell. |
## Everything else
+9 -3
View File
@@ -30,8 +30,12 @@ require a secure context.
| Toolbar | Bottom: Run, Stop, **Enter**, case picker, voice, settings. |
| Keyboard bar | Above the on-screen keyboard when it is open. |
Layout respects notch and home-indicator safe areas, touch targets are 44px, and the case
picker is a bottom sheet rather than a dropdown.
The phone layout applies up to 599px of viewport width, so the Plus and Pro Max iPhones,
the Pixel Pro and a folded Z Fold get it too; wider devices get the tablet layout. Layout
respects notch and home-indicator safe areas, touch targets are 44px, and the case picker is
a bottom sheet rather than a dropdown. On a folding phone (iPhone Duo) dialogs stay clear of
the hinge, and opening or closing the device is treated as the device changing shape, never
as the keyboard appearing.
**Swipe left and right** on the terminal to switch sessions.
@@ -58,7 +62,9 @@ A row of keys above the virtual keyboard, and what it contains depends on the se
**Agent sessions** get quick actions: `/init`, `/clear`, `/compact`, a clipboard key, `Esc`,
a path picker, an image key, and 🧠 when Read My Mind is on. Destructive commands need a
double press, so you cannot fire `/clear` with a stray thumb.
double press, so you cannot fire `/clear` with a stray thumb. On Codex sessions the bar also
shows `⇧←` and `⇧→`, the Shift-modified arrows Codex binds to editing the last queued
message and walking the prompt stack.
**Shell sessions** automatically swap it for terminal controls: `Ctrl`, `Esc`, `Tab`, four
arrows, paste, and dismiss. Your normal preference is remembered and restored when you
+6 -4
View File
@@ -33,8 +33,8 @@ reloading the dashboard while a permission dialog is blocking a session does not
with a normal-looking tab.
For Claude sessions, these come from Claude Code's hooks and are precise about *why* the
session stopped. For other CLIs there are no hooks, so you get the coarser output-based
signal.
session stopped; DeepSeek Harness sessions report the same states themselves. For the other
CLIs there are no hooks, so you get the coarser output-based signal.
## Window title and OS notifications
@@ -62,7 +62,8 @@ Once subscribed, a blocking prompt reaches your phone even from a locked screen.
## The Approvals Inbox
**Opt-in, off by default. Claude sessions only.**
**Opt-in, off by default. Claude sessions, plus DeepSeek Harness sessions, whose terminal
front door reports its prompts to Codeman.**
One queue of every prompt currently waiting on a human, across all your sessions, answerable
in place. When you have eight workers running, this is the difference between checking eight
@@ -136,7 +137,8 @@ from the lock screen.
- **No push over plain HTTP.** It is a browser requirement, not a Codeman one.
- **iOS needs the home screen install.** A Safari tab will never receive push.
- **The bell is invisible at zero.** That is deliberate, not a broken setting.
- **Approvals are Claude-only.** They are built on hook events the other CLIs do not emit.
- **Approvals need real signals.** They are built on hook events, which Claude emits and
DeepSeek Harness reports itself; the other CLIs do neither.
- **A stale menu answer is refused, not sent.** If you answer a card for a dialog that has
since gone away, Codeman declines rather than typing a digit into the composer.
+3
View File
@@ -67,6 +67,9 @@ one:
| **Gemini** | Enterprise only since Google's consumer cutover. |
| **Antigravity** | Google's successor to the consumer Gemini CLI. |
| **Pi** | No permission prompts and no sandbox by design. |
| **Grok Build** | xAI's CLI. |
| **DeepSeek Harness** | Needs a terminal profile; the menu offers to install one. |
| **OMP** | Oh My Pi, configured entirely through its own `~/.omp`. |
| **Terminal / Shell** | A plain shell, no agent. Also the **Run Shell** button. |
The dropdown also lists any saved dashboard URLs ([Web Tabs](Web-Tabs)) and your recent
+13 -2
View File
@@ -4,7 +4,7 @@ Point a case at another machine and the agent runs **there**, with the same dash
mobile UI, and autonomy features. Your laptop becomes a window onto a session living on the
remote host.
Like Docker, this is a **location overlay** on a case, not a run mode. All seven run modes
Like Docker, this is a **location overlay** on a case, not a run mode. All ten run modes
work remotely. See [Core Concepts](Core-Concepts).
## Why bother
@@ -52,7 +52,10 @@ A watcher with bounded backoff notices a dead SSH pane and quietly reattaches to
running remote session. On by default; the kill switch is in
**App Settings → Agents & CLIs → Remote auto-reconnect**.
Intentional kills are never revived. Closing a session means closing it.
Intentional kills are never revived. Closing a session means closing it. Neither is a clean
exit inside the pane (Ctrl-D, `exit`, Ctrl-C at the CLI's prompt): that tears the remote
tmux session down, and the watcher revives a session only when that durable session is
verifiably still alive. Only a transport drop is reconnected.
## Discover and attach
@@ -70,6 +73,14 @@ Attaching to someone else's session and closing your tab must not end their run,
not. Several clients can attach the same remote session at different window sizes without
clamping each other, and discovery shows a shared badge with the client count.
## Files
Previews, downloads and text reads in a remote case go over the same ssh connection the
session uses, so a clicked path opens the file on the machine the agent is on, `Range`
seeking included. Nothing is copied to the Codeman host. Editing, Office previews,
thumbnails, the file tree and the tail viewer are not available remotely and answer a clear
400 rather than a misleading 404. Details in [Working With Files](Working-With-Files).
## Security
Every SSH command line in Codeman flows through one builder that shell-escapes every
+13
View File
@@ -139,6 +139,7 @@ log stream --predicate 'process == "node"' # macOS, noisy
| Installer | Re-run the one-liner, or **App Settings → System → Updates**. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart. |
| Docker Compose | Re-run `Start-Codeman.sh`, or the in-app updater, which restarts the container in place. |
### The in-app updater
@@ -170,6 +171,18 @@ service without colliding with the main one. `CODEMAN_DATA_DIR` and `CODEMAN_TMU
exist for the rare case where they need to differ, but setting only one of them recreates
exactly the problem you were avoiding.
## Running Codeman itself in Docker
The Compose deployment in `docker/` runs the server in a container and spawns Docker cases
as sibling containers through the mounted host socket. Start it with
`bash docker/Start-Codeman.sh` rather than a bare `docker compose up`: the script pre-creates
the bind-mounted directories with the right owner, honours a `docker-compose.override.yml`,
and refreshes the build volumes when the checkout moved under them. The in-app updater
applies code only and restarts by letting the container exit, so it refuses a release that
changes the Dockerfile, the compose file, or adds a new `.env` key, until you re-run the
script. Guide:
[`docker/README.md`](https://github.com/Ark0N/Codeman/blob/master/docker/README.md).
## The tunnel as a service
```bash
+1
View File
@@ -67,6 +67,7 @@ be wrong for at least one of them:
| **File Viewer** | Real path resolution before boundary checks, so symlinks cannot escape. Sensitive trees blocked. Edit mode adds an extension allowlist, a size cap, `.git` denial, and optimistic concurrency. It never creates files. |
| **Attachments** | An id-based registry, so browser requests never carry absolute paths. The magic-link scanner is prompt-injectable by nature and is therefore force-confined to the session's workspace. Extension allowlist, not a blocklist. |
| **Path picker** | Its own root allowlist rather than the workspace confinement. In multi-user mode a non-admin gets only their own user space, because per-user spaces live inside the home directory. |
| **Remote cases** | Reads go over the session's own ssh connection and are resolved and contained on the remote host, with a bounded number of ssh children. Nothing is copied to the Codeman host; writes, Office previews and thumbnails are refused. |
Downloads block sensitive paths outright (`.env`, credentials files, `~/.ssh`, AWS
credentials), and SVG and HTML are served as downloads with `nosniff` so they cannot execute
+8 -1
View File
@@ -46,6 +46,7 @@ supervised by systemd or launchd; npm installs report as non-updatable. See
| Extended Keyboard Bar | Per device | Which accessory bar phones get. Shell sessions override it while they are active. |
| Wheel Scrolls Local History | Off | Keeps the wheel on the local buffer instead of forwarding it to the CLI. |
| Auto Copy Selection | Off | Copies highlighted terminal text to the clipboard the moment you finish selecting it. Ctrl+C still copies on demand. |
| Normal / Bold font weight | xterm defaults | Per device, each slot from 100 to 900. The bundled JetBrains Mono renders every step, so a lighter normal weight makes Claude's bold headings stand out. Applies live to the terminal, both echo overlays and open team panes. |
| WebGL Renderer | On | With a GPU-stall watchdog that falls back to DOM rendering. |
| Gesture Control | Off | Camera hand tracking. Also needs `CODEMAN_GESTURE=1` on the server. |
@@ -72,10 +73,13 @@ every session or only the active tab.
| Entrance Animations | Per-surface animation styles for tabs, terminals, windows, and lineage lines. All default to the legacy no-animation behaviour. |
| Display Name | Your name in the UI. Cosmetic only; it never renames the package, CLI, API, or storage. |
| Interface Language | English or Simplified Chinese. Per device. |
| Session List Layout | Header tab strip (default) or a collapsible left sidebar. See [The Dashboard](The-Dashboard#session-list-layout). |
| Session List Layout | Header tab strip (default), a collapsible left sidebar, or the sidebar with detailed rows. See [The Dashboard](The-Dashboard#session-list-layout). |
| Tab Orientation | Keeps the header list but turns the strip vertical beside the terminal, resizable, with detailed rows by default. Desktop and tablet only. |
| Vertical Rail Order | *By activity* (default) sorts the rail the way the home screens are sorted; *Manual* keeps your tab order and drag-reordering. |
| Tall Tabs | Taller tab strip. |
| Pop-out Button on Tabs | Adds the detach control to tabs, with a per-tab override. |
| Spawn Lineage Lines | Arcs from a parent tab to sessions it spawned. Desktop only, on by default. |
| Auto-name Sessions | Titles a new tab after its first prompt, keeping the case prefix (`w3-myapp: fix the login redirect`). Synced, off by default. See [The Dashboard](The-Dashboard#automatic-session-names). |
| Overview Home Screen | The phone home screen. On by default. |
### Models
@@ -155,6 +159,9 @@ Some things are configured before the server starts, not in the UI:
| `CODEMAN_DOCKER_BRIDGE_HOOKS` | Lets in-container hooks reach the host on a loopback bind. |
| `CODEMAN_FILE_PICKER_ROOTS` | Extra roots for the path picker. |
| `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK` | Acknowledges exposing the server with no password. |
| `CODEMAN_BASE_URL` | Mounts Codeman under a sub-path behind a reverse proxy that forwards the prefix unchanged. See [Remote Access](Remote-Access). |
| `CODEMAN_MAX_DOWNLOAD_BYTES` | Cap on raw file bodies and downloads. 2 GB by default, `0` for none. |
| `CODEMAN_MAX_REMOTE_FILE_SSH` | Concurrent ssh reads for files in remote cases. 4 by default. |
## Gotchas
+26 -5
View File
@@ -22,12 +22,14 @@ page says so and names the setting.
The session list lives in the header as a horizontal strip by default. With a lot of
sessions open that strip stops being scannable, so **App Settings → Appearance → Tabs →
Session List Layout** can move it into a vertical sidebar on the left instead.
Session List Layout** can move it into a vertical sidebar on the left instead, and
**Tab Orientation** can turn the strip itself into a vertical rail.
| Layout | Behaviour |
| -------------------- | --------------------------------------------------------------------------------- |
| **Header tab strip** | The default. Wraps to a second row on desktop, scrolls sideways on a phone. |
| **Left sidebar** | A vertical list with a filter box and a live session count. `Alt+B` collapses it to a narrow rail that keeps the status dots and task badges visible. On a phone it is an off-canvas drawer rather than a docked rail. |
| **Left sidebar** | A vertical list with a filter box and a live session count. `Alt+B` collapses it to a narrow rail that keeps the status dots and task badges visible. On a phone it is an off-canvas drawer rather than a docked rail. A detailed variant adds the home screen's per-session line (`created 3d ago · working 12m`) and a status pill. |
| **Vertical rail** | The strip turned vertical beside the terminal, resizable, with detailed rows by default. **Vertical Rail Order** sorts it by activity (blocked on you first, then longest running, then most recently quiet), the same order as the home screens; pick *Manual* to get your own order and drag-reordering back. Desktop and tablet only. |
It is the same list either way, just re-hosted: tab order, drag-to-reorder, the `Alt+1`
to `Alt+9` numbers and every status colour below behave identically in both. The setting is
@@ -66,6 +68,18 @@ reloading while a permission prompt is blocking does not lose the red tab.
Tabs can also be dragged to reorder.
### Automatic session names
Off by default. Turn on **Auto-name Sessions** (App Settings → Appearance → Tabs; synced
across devices) and a tab that still carries its generated name, such as `w3-myapp`, takes a
title from the first real prompt you submit, keeping the prefix: `w3-myapp: fix the login
redirect`. The strip shows the title and keeps the prefix in the tooltip, and the next
session in that case still counts up to `w4-myapp`. It happens once per session, only for
prompts you type or send through the input API (never a Ralph, respawn, cron or approval
answer), and never for shells. Slash commands such as `/clear` do not become titles; the
next prompt gets its turn. A name you set yourself, before or after, is never touched. The
title is derived locally from the prompt's first sentence; no text leaves the machine.
On phones the strip scrolls horizontally instead of wrapping, and the active tab is always
scrolled into view. It is not reordered to the front, so the `Alt+N` numbering stays stable.
@@ -142,6 +156,10 @@ Worth knowing:
always local scrollback. Other CLIs scroll locally.
- **Selection copy.** `Ctrl+C` copies when text is selected and interrupts when it is not.
`Ctrl+Shift+C` always copies.
- **Selecting where the CLI owns the mouse.** `Shift+drag` starts a selection even in a pane
whose mouse events are forwarded to the CLI, and right-click copies the selection (with
nothing selected the native menu is left alone). **Auto Copy Selection** in App Settings
copies the moment you release.
- **Zero-lag input.** On touch devices, keystrokes paint locally before the round trip. See
[Input And Voice](Input-And-Voice).
- **Renderer.** WebGL by default, with a watchdog that falls back to DOM rendering if the
@@ -155,8 +173,9 @@ which lists past sessions including Claude conversations started outside Codeman
Two extras depending on the device:
- **Desktop, wide windows**: your open tabs appear as a rail docked to the left edge, in tab
order, with created and last-active stamps. It needs at least 1180px of width; below that
- **Desktop, wide windows**: your open tabs appear as a rail docked to the left edge, in
overview order (blocked on you first, then longest running, then most recently quiet),
with created and state-duration stamps. It needs at least 1180px of width; below that
it is hidden so it cannot overlap the search panel.
- **Phones**: tapping the "C" logo gives a session overview instead: NEEDS YOU first, then
current sessions, then past ones. On by default.
@@ -192,7 +211,9 @@ so it is fast and cannot be turned into a traversal.
## Appearance
**App Settings → Appearance** carries the theme skins, including light ones. The choice is
applied before the first paint, so there is no flash of the wrong theme on load.
applied before the first paint, so there is no flash of the wrong theme on load. Terminal
font family and weight are per device too: a normal and a bold weight, each from 100 to
900, and the bundled JetBrains Mono renders every step.
The same section has the entrance animations for tabs, terminals, agent windows, and
lineage lines. All of them default to the legacy no-animation behaviour, so an untouched
+32 -2
View File
@@ -138,6 +138,12 @@ That is the PTY-exit circuit breaker. Repeated rapid PTY exits trip it, and it b
automatic restarts so a broken configuration does not spin forever. Reset it explicitly from
the session's controls. Reattaching does not clear it, deliberately.
### Typed prompts are silently ignored after restoring a tab
Update. A browser whose input sequence counter fell behind the server's (a restored tab,
cleared site data) used to have every prompt deduplicated away. Since 1.29.0 the duplicate
acknowledgement carries the watermark and the client re-sends.
### Sessions I did not create appeared, or my session resized itself
Two Codeman servers are running against the same data directory and tmux socket. The second
@@ -167,6 +173,16 @@ Things to try:
Codex ignores the mouse reports that forwarding would send, so Codeman does not forward
there. Scrolling is local, and `Shift+Wheel` behaves the same way.
### Selected text is invisible on a light skin
Update. Every skin named its selection colour under a key xterm renamed in v5, so the four
light skins painted white at 30% over near-white. Fixed in 1.29.0.
### `Ctrl+Z` suspended my agent
Update. Since 1.28.0 `Ctrl+Z` is swallowed in agent sessions, so a running CLI cannot be
stopped by job control. Shell sessions keep it.
### `Ctrl+C` copies when I wanted to interrupt
With a selection, `Ctrl+C` copies. With no selection, it interrupts. Clear the selection
@@ -252,11 +268,25 @@ node scripts/build-agent-image.mjs --no-cache
A plain rebuild reuses the cached `npm install -g` layer and keeps the CLIs frozen at their
original versions while reporting success.
### Every file in a remote case says "File not found"
Update. Before 1.29.0 the file routes resolved every path on the Codeman host, so in a
remote case every click failed while the file plainly existed on the other machine. Reads
now go over ssh; see [Working With Files](Working-With-Files). Editing and Office previews
stay unavailable remotely and say so with a 400.
### Compose: the server crash-loops with `EACCES` on first start
Start the stack with `bash docker/Start-Codeman.sh` rather than a plain `docker compose up`,
and update: since 1.29.0 the entrypoint corrects a root-owned bind mount before dropping
privileges. See [Running As A Service](Running-As-A-Service).
### A remote SSH session dropped and did not come back
A bounded-backoff watcher reattaches dropped sessions, and it is on by default. Intentional
kills are never revived. Check the host is reachable and that the remote tmux server is
still running.
kills are never revived, and neither is a clean exit inside the pane (Ctrl-D, `exit`): only
a transport drop is reconnected. Check the host is reachable and that the remote tmux server
is still running.
## Gathering diagnostics
+23
View File
@@ -24,6 +24,23 @@ Switching tabs does not reload a dashboard. Frames stay alive in the background,
took a while to authenticate is still there when you come back. Past six live frames, the
least recently viewed is dropped to bound memory.
## Single-page apps, reloads and links
A history-routed dashboard (React Router, Vue Router, a Vite dev server) sees the path it
would see on its own origin, not the proxy prefix, so it renders its real route instead of
its own "page not found". A navigation the page starts itself afterwards, a dev server's
full reload or a root-absolute `location.href`, would land outside the proxy with no
capability; Codeman recognises it, answers with a small recovery page, and remounts the
frame at the path that was lost, bounded to five recoveries a minute per frame. A reload on
the dashboard's landing page is recovered the same way.
A `localhost` or `127.0.0.1` link in agent output opens as a web tab automatically, reusing
a saved dashboard for the same server or saving one under its `host:port`. On a phone that
address only exists on the Codeman box, so the link would otherwise be a guaranteed
connection error. LAN and tailnet addresses still open directly. `*.localhost` names are
deliberately not auto-routed: they are DNS names rather than address literals, and the link
came from agent output. Add such a dashboard by hand instead.
## Why dashboards are proxied
A plain cross-origin iframe fails three ways at once in the setup Codeman actually ships in:
@@ -89,6 +106,12 @@ The proxy authenticates on an in-memory capability embedded in the path, which i
exempt from the cookie and Origin checks that every API route enforces. That exemption is
fenced to safe methods and non-API paths, and there is a test pinning it in place.
Saved URLs are refused when they point at a link-local or cloud-metadata address, at save
time and again against the address the name resolves to at connect time; loopback and
private ranges stay allowed, because a `localhost` Grafana is the feature. Capabilities are
revoked on logout, and proxied responses carry a same-origin referrer policy so a dashboard
cannot hand the capability-bearing URL to a third party.
Two failure modes that only appear inside a sandboxed frame, and that curl can never
reproduce, are handled: runtime-built root-absolute URLs escaping the injected base, and
same-host requests being CORS-checked with a null origin. Both present as the dashboard's own
+16
View File
@@ -110,6 +110,22 @@ it is written. Outside the workspace they open in the preview instead: the tail
Nothing is registered until you click. Opening a file this way does not add an attachment card.
## Remote (SSH) cases
In a remote case the workspace lives on the other machine, and so do the files. Previews,
downloads, text reads and the clicked-path route all go over the same ssh connection the
session uses: one `realpath` plus `stat` probe for the file and the workspace root, then a
streamed `cat` (or a slice of it, so video seeking works). Symlinks are resolved on the host
that can resolve them, the size cap applies to the remote size before a byte is requested,
and an unreachable host answers 502 rather than pretending the file is missing. Nothing is
ever copied onto the Codeman host, and a same-named local file is never served under a
remote name.
Not available over ssh, and said so with a 400 instead of a misleading 404: editing in
place, Office previews and generated thumbnails (both need the bytes on the server's disk),
the file tree and path picker, and the tail viewer. Docker cases are unaffected, because
their workspace is bind-mounted at the same path.
## The path picker
For choosing a path rather than typing one. It appears in two places:
+28 -37
View File
@@ -93,8 +93,7 @@ export PUPPETEER_SKIP_DOWNLOAD="${PUPPETEER_SKIP_DOWNLOAD:-1}"
CLI_IDS=('claude' 'shell' 'opencode' 'codex' 'gemini' 'antigravity' 'pi' 'grok' 'deepseek' 'omp')
CLI_LABELS=('Claude' 'Shell' 'OpenCode' 'Codex' 'Gemini' 'Antigravity' 'Pi' 'Grok' 'DeepSeek' 'OMP')
CLI_ENABLED=(1 1 1 1 1 1 1 1 1 1)
CLI_KIND=('agent' 'shell' 'agent' 'agent' 'agent' 'agent' 'agent' 'agent' 'agent' 'agent')
CLI_NPM=('@anthropic-ai/claude-code' '' 'opencode-ai' '@openai/codex' '@google/gemini-cli' '' '@earendil-works/pi-coding-agent' '' '@deepseek-ai/dsh' '')
CLI_LAUNCHER_ONLY=(0 0 0 0 0 0 0 0 1 0)
CLI_DOCS=('https://docs.claude.com/claude-code' '' 'https://opencode.ai/docs' 'https://developers.openai.com/codex/cli' 'https://github.com/google-gemini/gemini-cli' 'https://antigravity.google/cli' 'https://pi.dev' 'https://github.com/xai-org/grok-build' 'https://github.com/deepseek-ai/deepseek-harness' 'https://omp.sh')
CLI_CMD_LINUX=('curl -fsSL https://claude.ai/install.sh | bash' '' 'curl -fsSL https://opencode.ai/install | bash' 'npm install -g @openai/codex' 'npm install -g @google/gemini-cli' 'curl -fsSL https://antigravity.google/cli/install.sh | bash' 'npm install -g --ignore-scripts @earendil-works/pi-coding-agent' 'curl -fsSL https://x.ai/cli/install.sh | bash' '' 'curl -fsSL https://omp.sh/install | sh')
CLI_CMD_DARWIN=('curl -fsSL https://claude.ai/install.sh | bash' '' 'curl -fsSL https://opencode.ai/install | bash' 'npm install -g @openai/codex' 'npm install -g @google/gemini-cli' 'curl -fsSL https://antigravity.google/cli/install.sh | bash' 'npm install -g --ignore-scripts @earendil-works/pi-coding-agent' 'curl -fsSL https://x.ai/cli/install.sh | bash' '' 'brew install can1357/tap/omp')
@@ -410,22 +409,6 @@ check_build_tools() {
# test/install-sh-detection-parity.test.ts: the process PATH first (each declared
# binary name in turn), then each known install path, dir-major.
# Index of "$1" in CLI_IDS -> CLI_IDX, returning 1 with CLI_IDX=-1 when unknown.
# A global rather than an echo because this runs inside loops, and a subshell per
# lookup is a fork per CLI per call site.
CLI_IDX=-1
_cli_index() {
local want="$1" i
CLI_IDX=-1
for ((i = 0; i < ${#CLI_IDS[@]}; i++)); do
if [[ "${CLI_IDS[$i]}" == "$want" ]]; then
CLI_IDX=$i
return 0
fi
done
return 1
}
# `dsh` is the hardest name of the lot: Debian ships an unrelated `dsh`
# (dancer's shell). The server-side resolver settles it by demanding the
# harness's own help banner; detection here only feeds the "you have no AI CLI"
@@ -465,7 +448,8 @@ _cli_candidate_ok() {
# Resolve every CLI in ONE pass, memoized.
#
# CLI_FOUND_PATH is parallel to CLI_IDS ('' when not found). CLI_FOUND_COUNT
# CLI_FOUND_PATH is parallel to CLI_IDS ('' when not found, and also '' for a
# DISABLED entry — it is never probed at all, see below). CLI_FOUND_COUNT
# counts only ENABLED entries that have a binary to look for, which is what the
# "no AI CLI found" gate asks about — `shell` has no binary and must never make
# that gate think an agent is installed.
@@ -485,6 +469,16 @@ detect_all_clis() {
for ((i = 0; i < ${#CLI_IDS[@]}; i++)); do
found=""
# A disabled entry is never even probed: every consumer already filters
# on CLI_ENABLED before showing anything, so the command-v/stat calls
# below would be pure waste — and, unlike filtering downstream, skipping
# the probe here is what makes CLI_ENABLED mean "look for it" rather
# than just "offer it once found".
if [[ "${CLI_ENABLED[$i]}" != "1" ]]; then
CLI_FOUND_PATH[$i]=""
continue
fi
# 1. The process PATH, each declared binary name in turn.
bin_end=$((${CLI_BIN_OFF[$i]} + ${CLI_BIN_LEN[$i]}))
for ((j = ${CLI_BIN_OFF[$i]}; j < bin_end; j++)); do
@@ -520,20 +514,6 @@ detect_all_clis() {
return 0
}
# Is this CLI installed? Unknown id is "no", never an error.
check_cli() {
detect_all_clis
_cli_index "$1" || return 1
[[ -n "${CLI_FOUND_PATH[$CLI_IDX]}" ]]
}
# Where it was found, or nothing.
get_cli_path() {
detect_all_clis
_cli_index "$1" || return 1
printf '%s\n' "${CLI_FOUND_PATH[$CLI_IDX]}"
}
# ----------------------------------------------------------------------------
# Catalogue helpers
# ----------------------------------------------------------------------------
@@ -589,7 +569,12 @@ cli_catalog_names() {
# the registry but an empty one here: installing the launcher alone leaves
# nothing that can drive a pane, so the generator withholds the command for
# any launcherProfile entry (see installCommandFor in generate-cli-catalog.mts)
# and this hint falls through to the docs URL instead.
# and this hint falls through to the docs URL instead — CLI_LAUNCHER_ONLY adds
# one line explaining WHY it is a docs link and not a command, so a user who
# follows that link straight to `npm install -g @deepseek-ai/dsh` (which the
# docs page itself documents) does not land back in the same "installed but
# cannot drive a pane" trap the menu exists to avoid. Data-driven, not an id
# check: any future launcherProfile entry gets the same caveat for free.
cli_catalog_print_install_hints() {
detect_all_clis
local i
@@ -601,6 +586,9 @@ cli_catalog_print_install_hints() {
echo -e " ${CYAN}${CLI_INSTALL_CMD_TRUSTED[$i]}${NC} # ${CLI_LABELS[$i]}"
elif [[ -n "${CLI_DOCS[$i]}" ]]; then
echo -e " ${CLI_LABELS[$i]}: see ${CYAN}${CLI_DOCS[$i]}${NC}"
if [[ "${CLI_LAUNCHER_ONLY[$i]}" == "1" ]]; then
echo -e " (installs a launcher only: it still needs a terminal profile, and Codeman's Run menu can add one)"
fi
fi
done
}
@@ -679,9 +667,12 @@ offer_ai_cli_install() {
local cli_choice=""
if [[ "$NONINTERACTIVE" == "1" ]] || ! has_tty; then
# Explicit automation opt-in: default to the first offered entry,
# which is registry order, which is Claude Code (order 0) — the
# same default this prompt has always taken non-interactively.
# Explicit automation opt-in: default to the first OFFERED entry.
# That is registry order, which is Claude Code (order 0), UNLESS
# this is a wget-only host and Claude's curl one-liner was just
# filtered out of offer_idx above — there, the first survivor is
# whichever npm-based entry sorts earliest (Codex today), not
# Claude. Printed either way so the choice is never silent.
cli_choice="1"
info "CODEMAN_NONINTERACTIVE=1: defaulting to ${CLI_LABELS[${offer_idx[0]}]}"
else
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.28.2",
"version": "1.30.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.28.2",
"version": "1.30.0",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
+2 -2
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.28.2",
"version": "1.30.0",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -29,7 +29,7 @@
"test:mobile": "vitest run --config test/mobile/vitest.config.ts",
"check:frontend-syntax": "node scripts/check-frontend-syntax.mjs",
"fix:node-pty": "node scripts/fix-node-pty.mjs",
"typecheck": "tsc --noEmit",
"typecheck": "tsc --noEmit && tsc -p config/tsconfig.scripts.json",
"lint": "eslint --config config/eslint.config.js 'src/**/*.ts'",
"lint:fix": "eslint --config config/eslint.config.js 'src/**/*.ts' --fix",
"format": "prettier --write 'src/**/*.ts' 'src/web/public/**/*.{js,css,html,json}'",
+1 -1
View File
@@ -1,7 +1,7 @@
{
"name": "codeman",
"description": "Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.28.2",
"version": "1.30.0",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
+75 -26
View File
@@ -47,7 +47,7 @@ later call opens with, and your first REAL call performs them anyway:
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
```
⚠️ **Never spend a Bash call on this check alone.** §1's block opens with this same
@@ -75,8 +75,8 @@ PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh"
mkdir -p "$(dirname "$PRE")"
# Rewrite unless the file already ends with THIS version's stamp, so a stale or a
# half-written file self-heals here instead of costing you a round trip to rm it.
grep -qs '^CODEMAN_PREAMBLE=1.22.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
grep -qs '^CODEMAN_PREAMBLE=1.30.1$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -150,6 +150,27 @@ _trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
# ---- the composer: is the prompt still sitting there, unsent? ----
# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the
# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that
# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt
# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through
# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS
# the composer and keeps pressing Enter while the prompt is still there.
_composer_text() { # <sid> -> the composer's text with ALL whitespace removed: "" once
# the prompt was taken, "?" when the pane shows no composer at all. The composer is
# the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher
# up in the transcript, so only the last one says whether the text was taken.
local t
t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1)
[ -n "$t" ] || { printf '?'; return 0; }
# Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does
# not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH).
printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g"
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
@@ -262,12 +283,16 @@ spawn_workers() {
# worker a silent no-op that still "succeeds" and reports the previous turn's state.
# Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a
# deliberate duplicate, at the SAME number (§5.3).
# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the
# typed prompt stranded on the composer while a long wait runs its whole timeout
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude
# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints),
# leaving the typed prompt stranded on the composer while a long wait runs its whole
# timeout (observed live, twelve reviews in a row). So the first wait is short; on its
# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate:
# the server re-waits without retyping, §5.3) and kept open in the background, while
# the composer is READ (_composer_text) and, as long as the prompt is still sitting
# there, a bare \r goes out about every ten seconds, up to twelve times. An empty
# composer ends the loop, so a prompt that was taken is never nudged again, and the
# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
@@ -275,7 +300,7 @@ spawn_workers() {
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
@@ -290,16 +315,38 @@ sendwait() {
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
# ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call,
# in the background, while the Enter loop below works the composer. Signals have
# no history: a `stop` that fires while no wait is open (during a composer read
# between two short waits, measured) is lost, and the next wait then runs its
# whole timeout on a turn that already ended. The resend is a tagged DUPLICATE,
# so the server skips the write and re-waits without retyping (§5.3).
tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" &
bg=$!
# The prompt's head with whitespace removed, matched literally (the "$head"
# quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob).
head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24)
while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended
c=$(_composer_text "$sid")
if [ "$c" = '?' ]; then
[ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it
elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then
break # composer empty (taken) or holding other text
fi
n=$((n+1))
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done
done
wait "$bg"
# The duplicate reports `delivered:false` -- truthfully, but about the wrong send.
# The first one delivered, so carry that forward, or §1's cleanup reads a completed
# turn as an undelivered one and keeps a finished worker forever.
r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp")
rm -f "$tmp"
fi
printf '%s\n' "$r"
}
@@ -325,10 +372,10 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.22.0
CODEMAN_PREAMBLE=1.30.1
PREAMBLE
)
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
```
Every later Bash call that touches the API starts with the same two loader lines from
@@ -379,7 +426,7 @@ and no per-call body to hand-build.
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null # §0 loader
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
N=(alpha beta) # INVENT one fresh case name per worker; never list cases first
# (a name may carry a mode: `beta:deepseek`, see below)
T=('reply with one line: the absolute path of your working directory'
@@ -440,9 +487,11 @@ Four things this block leans on, each one link away, no detour needed to run it:
skill: §5.1. Those workspaces do get hooks now, unless the operator disabled it.
- `sendwait` supplies the `\r`, picks a fresh `seq`, and self-heals a stranded Enter.
A prompt without the `\r` is never submitted (§3), a reused `seq` is silently
swallowed as an already-applied duplicate, and an Enter eaten by an Ink repaint
strands the prompt on the composer until a bare `\r` follows: all three are reasons
to let `sendwait` build the call rather than hand-rolling it.
swallowed as an already-applied duplicate, and a lost Enter strands the prompt on the
composer until a bare `\r` follows: Claude Code 2.1.277 and later ignore Enter for the
first 30 to 50 seconds after the composer paints while still taking the text, so
`sendwait` reads the composer and keeps pressing Enter until the prompt has left it.
All three are reasons to let `sendwait` build the call rather than hand-rolling it.
- Each `sendwait` costs that worker one billed turn, as does every prompt you send it.
- Deleting the sessions does **not** remove the case directories. They are marked as
agent-created, so `GET /api/v1/cases/agent-created` lists them for cleanup: §5.14.
+66 -19
View File
@@ -1,4 +1,4 @@
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -72,6 +72,27 @@ _trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
# ---- the composer: is the prompt still sitting there, unsent? ----
# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the
# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that
# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt
# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through
# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS
# the composer and keeps pressing Enter while the prompt is still there.
_composer_text() { # <sid> -> the composer's text with ALL whitespace removed: "" once
# the prompt was taken, "?" when the pane shows no composer at all. The composer is
# the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher
# up in the transcript, so only the last one says whether the text was taken.
local t
t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1)
[ -n "$t" ] || { printf '?'; return 0; }
# Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does
# not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH).
printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g"
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
@@ -184,12 +205,16 @@ spawn_workers() {
# worker a silent no-op that still "succeeds" and reports the previous turn's state.
# Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a
# deliberate duplicate, at the SAME number (§5.3).
# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the
# typed prompt stranded on the composer while a long wait runs its whole timeout
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude
# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints),
# leaving the typed prompt stranded on the composer while a long wait runs its whole
# timeout (observed live, twelve reviews in a row). So the first wait is short; on its
# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate:
# the server re-waits without retyping, §5.3) and kept open in the background, while
# the composer is READ (_composer_text) and, as long as the prompt is still sitting
# there, a bare \r goes out about every ten seconds, up to twelve times. An empty
# composer ends the loop, so a prompt that was taken is never nudged again, and the
# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
@@ -197,7 +222,7 @@ spawn_workers() {
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
@@ -212,16 +237,38 @@ sendwait() {
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
# ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call,
# in the background, while the Enter loop below works the composer. Signals have
# no history: a `stop` that fires while no wait is open (during a composer read
# between two short waits, measured) is lost, and the next wait then runs its
# whole timeout on a turn that already ended. The resend is a tagged DUPLICATE,
# so the server skips the write and re-waits without retyping (§5.3).
tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" &
bg=$!
# The prompt's head with whitespace removed, matched literally (the "$head"
# quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob).
head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24)
while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended
c=$(_composer_text "$sid")
if [ "$c" = '?' ]; then
[ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it
elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then
break # composer empty (taken) or holding other text
fi
n=$((n+1))
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done
done
wait "$bg"
# The duplicate reports `delivered:false` -- truthfully, but about the wrong send.
# The first one delivered, so carry that forward, or §1's cleanup reads a completed
# turn as an undelivered one and keeps a finished worker forever.
r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp")
rm -f "$tmp"
fi
printf '%s\n' "$r"
}
@@ -247,4 +294,4 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.22.0
CODEMAN_PREAMBLE=1.30.1
@@ -21,7 +21,7 @@ by sourcing the preamble file the §0 bootstrap wrote, and checking its version
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
```
Do **not** re-paste the preamble body into each call. Sourcing it is what retires the
+7 -6
View File
@@ -135,8 +135,7 @@ export function renderInstallShBlock(entries: CliEntry[] = STOCK_CLIS): string {
const ids: string[] = [];
const labels: string[] = [];
const enabled: string[] = [];
const kinds: string[] = [];
const npm: string[] = [];
const launcherOnly: string[] = [];
const docs: string[] = [];
const cmdLinux: string[] = [];
const cmdDarwin: string[] = [];
@@ -151,8 +150,11 @@ export function renderInstallShBlock(entries: CliEntry[] = STOCK_CLIS): string {
ids.push(shQuote(entry.id as string));
labels.push(shQuote(entry.label));
enabled.push(entry.enabled ? '1' : '0');
kinds.push(shQuote(entry.kind));
npm.push(shQuote(entry.discovery.install.npmPackage ?? ''));
// Parallel to CLI_IDS: 1 when this entry's install command installs a launcher rather
// than something that can drive a pane on its own (see installCommandFor above). Purely
// derived from discovery.launcherProfile — install.sh's hint printer reads this to add a
// caveat instead of hardcoding which id it means.
launcherOnly.push(entry.discovery.launcherProfile ? '1' : '0');
docs.push(shQuote(entry.discovery.install.docsUrl ?? ''));
cmdLinux.push(shQuote(installCommandFor(entry, 'linux')));
cmdDarwin.push(shQuote(installCommandFor(entry, 'darwin')));
@@ -195,8 +197,7 @@ export function renderInstallShBlock(entries: CliEntry[] = STOCK_CLIS): string {
arr('CLI_IDS', ids),
arr('CLI_LABELS', labels),
arr('CLI_ENABLED', enabled),
arr('CLI_KIND', kinds),
arr('CLI_NPM', npm),
arr('CLI_LAUNCHER_ONLY', launcherOnly),
arr('CLI_DOCS', docs),
arr('CLI_CMD_LINUX', cmdLinux),
arr('CLI_CMD_DARWIN', cmdDarwin),
@@ -0,0 +1,9 @@
{
"_comment": "Copy this file to local-llm-test.config.json (gitignored) and fill in your own values. CLI flags on scripts/test-local-llm-harnesses.mjs always override these. Any field can be omitted. apiKey is OPTIONAL — omit it entirely (or delete this line) for an endpoint like llama.cpp that doesn't check one; it defaults to a harmless placeholder either way.",
"baseUrl": "http://192.168.1.50:8080",
"model": "qwen3",
"apiKey": "",
"prompt": "Reply with exactly: hello world",
"timeout": 30000,
"only": []
}
+23
View File
@@ -193,6 +193,29 @@ run_step "installing" "Installing dependencies" npm install --no-fund --no-audit
# 5) Build (gate the restart on success — never restart into a torn dist/).
run_step "building" "Building" npm run build || rollback_and_fail "Build failed"
# Docker Compose only: record what HEAD/package-lock.json the freshly-built
# codeman-dist/codeman-node-modules volumes now reflect. `Start-Codeman.sh`
# reads this same file (`$appdata_path/.codeman/…`, i.e. this container's own
# $HOME/.codeman since that path IS the appdata bind mount) to detect source
# changes an EXTERNAL `docker compose build` made and refresh those volumes —
# without this, the next plain `Start-Codeman.sh` run would see the HEAD this
# update just checked out, not recognise it as already accounted for, and wipe
# the volumes this update just correctly rebuilt right back to the OLDER image.
if [[ "$SUPERVISOR" == "docker-compose" ]]; then
build_source_file="$HOME/.codeman/docker-build-source.json"
mkdir -p -- "$HOME/.codeman"
build_head=$(git rev-parse HEAD 2>/dev/null || true)
build_lockfile_sha=''
if command -v sha256sum >/dev/null 2>&1; then
build_lockfile_sha=$(sha256sum -- package-lock.json 2>/dev/null | cut -d' ' -f1)
elif command -v shasum >/dev/null 2>&1; then
build_lockfile_sha=$(shasum -a 256 package-lock.json 2>/dev/null | cut -d' ' -f1)
fi
printf '{\n "headCommit": "%s",\n "lockfileSha256": "%s"\n}\n' \
"$build_head" "$build_lockfile_sha" >"$build_source_file.tmp" \
&& mv -- "$build_source_file.tmp" "$build_source_file"
fi
# 6) Restart the service so the new code loads. Write the terminal pre-restart
# marker FIRST so the freshly-booted server can reconcile it deterministically.
write_status "restarting" "Restarting Codeman…"
+699
View File
@@ -0,0 +1,699 @@
#!/usr/bin/env -S npx tsx
/**
* Standalone smoke-test for pointing each Codeman-supported harness CLI at a
* custom OpenAI-compatible endpoint — local (llama.cpp, Ollama, vLLM, ...) or
* cloud (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
* self-hosted gateway, ...). Anything that answers GET /v1/models and POST
* /v1/chat/completions in the standard shape qualifies; --base-url is not
* assumed to be a LAN address.
*
* This is intentionally OUTSIDE the npm test suite and outside Codeman's own
* session/tmux machinery: it spawns each real CLI binary directly, one-shot,
* with the env vars / config files that CLI's own docs say redirect it to a
* custom endpoint, and checks it can answer "hello world".
*
* DYNAMIC BY DESIGN: this file imports the SAME `enabledClis()` registry and
* `buildCustomModelInjection()` builder the production feature uses (see
* ../src/config/cli-registry/, ../src/custom-model-injection.ts,
* ../src/custom-model-injection-apply.ts) rather than keeping a second,
* hand-maintained copy of each CLI's env vars/config shape. A registry
* change (a new CLI, an edited env var name, a fixed config template) is
* picked up here automatically with zero edits to this file. Only the
* ONE-SHOT INVOCATION FLAGS (how to make each CLI answer one prompt and
* exit — information the registry doesn't model at all, since it only knows
* how to launch the interactive TUI) stay in the small ONE_SHOT table below;
* a CLI newly added to the registry with no ONE_SHOT entry is reported
* UNKNOWN rather than silently skipped or guessed at.
*
* Cloud endpoints often differ from a bare llama.cpp box in two ways this
* script accounts for: (1) auth may be an `api-key` header (Azure's
* convention) rather than `Authorization: Bearer` — see --auth-style below.
* (2) a cloud endpoint's "model" may actually be a deployment name distinct
* from the model family (Azure AI Foundry deployments) — always pass
* --model explicitly for those rather than relying on GET /v1/models
* discovery.
*
* IMPORTANT CONFIDENCE NOTE: claude and opencode are verified end-to-end
* against a real llama-swap server. codex's config STRUCTURE is verified,
* but it only speaks the Responses API (dropped Chat-Completions support
* Feb 2026) — expect it to fail against a plain OpenAI-compatible server,
* that's a real protocol gap, not a bug here. gemini/pi/grok/omp have their
* ONE-SHOT INVOCATION flags confirmed against real installed binaries'
* `--help` output, but their custom-endpoint env/config conventions remain
* web-researched, unverified. deepseek (dsh) is a profile launcher with no
* documented one-shot prompt flag at all — best-effort only. antigravity
* has no known CLI/env/config mechanism (GUI-only per public docs) — its
* registry entry declares `customModelInjection: { kind: 'unsupported' }`,
* which this script picks up dynamically and always skips.
*
* Usage:
* npx tsx scripts/test-local-llm-harnesses.ts --base-url http://192.168.1.50:8080 [options]
* npx tsx scripts/test-local-llm-harnesses.ts --base-url https://<resource>.services.ai.azure.com/openai/v1 --model <deployment-name> --api-key $AZURE_AI_KEY
*
* Options:
* --base-url <url> Required. Root URL of the OpenAI-compatible endpoint (local or cloud).
* --model <name> Model/deployment id to request. Default: first from GET /v1/models.
* --api-key <key> API key to send. Default: local-dummy-key (fine for llama.cpp; required for most cloud endpoints).
* --auth-style <style> "bearer" (default, Authorization: Bearer) or "api-key" (the
* `api-key` header some cloud gateways, e.g. Azure, want).
* NEVER send both — live-tested against a real server, doing
* so reliably HANGS the request indefinitely.
* --prompt <text> Prompt to send. Default: "Reply with exactly: hello world".
* --only <id,id,...> Restrict to these harness ids (comma-separated).
* --timeout <ms> Per-harness spawn timeout. Default: 30000.
* --probe-help Instead of testing, resolve each installed binary and print --help.
* --keep-temp Don't delete generated per-harness config dirs afterward.
* --list Dry run: print the resolved plan per harness, execute nothing.
* -h, --help Show this help.
*/
import { execFileSync, spawn } from 'node:child_process';
import { mkdtempSync, rmSync, readFileSync, existsSync } from 'node:fs';
import { tmpdir, homedir } from 'node:os';
import { join, delimiter, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
import { enabledClis } from '../src/config/cli-registry/index.js';
import type { CliEntry } from '../src/config/cli-registry/types.js';
import {
buildCustomModelInjection,
GROK_CUSTOM_MODEL_NAME,
type CustomModelEndpoint,
} from '../src/custom-model-injection.js';
import { applyConfigDirInjection } from '../src/custom-model-injection-apply.js';
const TAG = '[test-local-llm-harnesses]';
const SCRIPT_DIR = dirname(fileURLToPath(import.meta.url));
const CONFIG_PATH = join(SCRIPT_DIR, 'local-llm-test.config.json');
const CONFIG_EXAMPLE_PATH = join(SCRIPT_DIR, 'local-llm-test.config.example.json');
type AuthStyle = 'bearer' | 'api-key';
interface ConfigDefaults {
baseUrl?: string | null;
model?: string | null;
apiKey?: string;
authStyle?: AuthStyle;
prompt?: string;
only?: string[] | null;
timeout?: number;
}
/**
* Loads scripts/local-llm-test.config.json (gitignored — real IP/model/key,
* per-machine) if present, so you don't have to retype --base-url every run.
* See local-llm-test.config.example.json (tracked) for the shape. CLI flags
* always override whatever this file sets; this only supplies defaults.
*/
function loadConfigFile(): ConfigDefaults {
if (!existsSync(CONFIG_PATH)) return {};
try {
const raw = JSON.parse(readFileSync(CONFIG_PATH, 'utf8'));
return {
baseUrl: raw.baseUrl ?? null,
model: raw.model ?? null,
apiKey: raw.apiKey || undefined, // empty string counts as "not set", not a real key
authStyle: raw.authStyle === 'api-key' ? 'api-key' : undefined, // never 'both'
prompt: raw.prompt ?? undefined,
only: Array.isArray(raw.only) && raw.only.length ? raw.only : null,
timeout: typeof raw.timeout === 'number' ? raw.timeout : undefined,
};
} catch (err) {
console.error(`${TAG} failed to parse ${CONFIG_PATH}: ${(err as Error).message} (ignoring it)`);
return {};
}
}
interface Opts {
baseUrl: string | null;
model: string | null;
apiKey: string;
authStyle: AuthStyle;
prompt: string;
only: string[] | null;
timeout: number;
probeHelp: boolean;
keepTemp: boolean;
list: boolean;
help: boolean;
}
function parseArgs(argv: string[], configDefaults: ConfigDefaults): Opts {
const opts: Opts = {
baseUrl: configDefaults.baseUrl ?? null,
model: configDefaults.model ?? null,
apiKey: configDefaults.apiKey ?? 'local-dummy-key',
authStyle: configDefaults.authStyle ?? 'bearer',
prompt: configDefaults.prompt ?? 'Reply with exactly: hello world',
only: configDefaults.only ?? null,
timeout: configDefaults.timeout ?? 30000,
probeHelp: false,
keepTemp: false,
list: false,
help: false,
};
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
switch (a) {
case '--base-url':
opts.baseUrl = argv[++i];
break;
case '--model':
opts.model = argv[++i];
break;
case '--api-key':
opts.apiKey = argv[++i];
break;
case '--auth-style':
opts.authStyle = argv[++i] as AuthStyle;
if (opts.authStyle !== 'bearer' && opts.authStyle !== 'api-key') {
console.error(`${TAG} --auth-style must be "bearer" or "api-key"`);
opts.help = true;
}
break;
case '--prompt':
opts.prompt = argv[++i];
break;
case '--only':
opts.only = argv[++i]
.split(',')
.map((s) => s.trim())
.filter(Boolean);
break;
case '--timeout':
opts.timeout = Number(argv[++i]);
break;
case '--probe-help':
opts.probeHelp = true;
break;
case '--keep-temp':
opts.keepTemp = true;
break;
case '--list':
opts.list = true;
break;
case '-h':
case '--help':
opts.help = true;
break;
default:
console.error(`${TAG} unknown argument: ${a}`);
opts.help = true;
}
}
return opts;
}
function printUsage(): void {
console.log(`Usage: npx tsx scripts/test-local-llm-harnesses.ts [--base-url <url>] [options]
Reads defaults from scripts/local-llm-test.config.json if it exists (copy
scripts/local-llm-test.config.example.json to create it — gitignored, since
it holds a real IP/model/key). CLI flags always override the config file.
--base-url becomes optional once that file supplies one.
Works against any custom OpenAI-compatible endpoint, local or cloud
(llama.cpp, Ollama, vLLM, Azure AI Foundry, OpenRouter, a self-hosted
gateway, ...) — anything answering GET /v1/models and POST
/v1/chat/completions in the standard shape.
Options:
--base-url <url> Required. Root URL of the OpenAI-compatible endpoint.
--model <name> Model/deployment id to request. Default: first from GET /v1/models.
--api-key <key> API key to send. Default: local-dummy-key (required for most cloud endpoints).
--auth-style <style> "bearer" (default) or "api-key" (Azure-style). Never both — sending
both headers together reliably hangs some real servers.
--prompt <text> Prompt to send. Default: "Reply with exactly: hello world".
--only <id,id,...> Restrict to these harness ids.
--timeout <ms> Per-harness spawn timeout. Default: 30000.
--probe-help Print each installed binary's --help instead of testing.
--keep-temp Keep generated per-harness config dirs afterward.
--list Dry run: print the resolved plan, execute nothing.
-h, --help Show this help.
Harness ids are read from the CLI registry at run time — pass an unknown
one and the error message lists what's actually enabled right now.
Examples:
npx tsx scripts/test-local-llm-harnesses.ts --base-url http://192.168.1.50:8080
npx tsx scripts/test-local-llm-harnesses.ts --base-url https://<resource>.services.ai.azure.com/openai/v1 --model <deployment-name> --api-key $AZURE_AI_KEY`);
}
const HOME = homedir();
/** Expands a leading `~` the way the CLI registry's own search dirs are written. */
function expandHome(p: string): string {
if (p === '~') return HOME;
if (p.startsWith('~/')) return join(HOME, p.slice(2));
return p;
}
function pathWithExtraDirs(extraDirs: string[]): string {
return [...extraDirs.map(expandHome), '/usr/local/bin', process.env.PATH ?? ''].join(delimiter);
}
/** Resolve a binary by trying `<bin> --version` with the CLI's own registry search dirs prefixed onto PATH. */
function resolveBinary(bin: string, searchDirs: string[]): string | null {
try {
execFileSync(bin, ['--version'], {
timeout: 5000,
stdio: 'pipe',
env: { ...process.env, PATH: pathWithExtraDirs(searchDirs) },
});
return bin;
} catch (err) {
// Some CLIs (e.g. dsh) don't support --version cleanly for identity but
// still exist on PATH; a non-ENOENT failure still counts as "found".
if (err && (err as NodeJS.ErrnoException).code === 'ENOENT') return null;
return bin;
}
}
function printHelp(bin: string, searchDirs: string[]): void {
try {
const out = execFileSync(bin, ['--help'], {
timeout: 5000,
stdio: 'pipe',
env: { ...process.env, PATH: pathWithExtraDirs(searchDirs) },
});
console.log(out.toString());
} catch (err) {
const e = err as { stdout?: Buffer; message?: string };
console.log((e.stdout ?? e.message ?? String(err)).toString());
}
}
// --- one-shot invocation table (NOT in the registry — genuinely separate info) ---
type Confidence = 'verified' | 'researched' | 'unknown';
interface OneShot {
/** `modelId` is the RAW model/deployment id (e.g. "qwen3.5-0.8b-...") — CLIs whose
* config wraps it under a provider/block name (pi/omp's "custom/<id>", grok's fixed
* block name) build the full `--model` value here, not in the injection layer. */
argv: (prompt: string, modelId: string) => string[];
confidence: Confidence;
note?: string;
}
/**
* How to make each CLI answer ONE prompt and exit. The registry has no concept
* of this (it only knows the interactive TUI launch line), so this table is
* necessarily hand-maintained — but it is the ONLY hand-maintained part left;
* everything about WHERE the prompt goes (env vars, config files) comes from
* the real registry + `buildCustomModelInjection()` above.
*
* A CLI enabled in the registry with no entry here reports UNKNOWN rather
* than being silently skipped or guessed at — see `resolveOneShot()`.
*/
const ONE_SHOT: Record<string, OneShot> = {
claude: {
confidence: 'verified',
// Claude Code's async session-title-generation call also uses
// ANTHROPIC_DEFAULT_HAIKU_MODEL and validates it against Claude's OWN internal
// recognized-model list, printing [claude-code:unrecognized_model] to stderr for
// a local model name. Confirmed live: `--settings '{"autoTitle":false}'` does NOT
// stop it (still hung the whole run); `--bare` does — the warning still prints,
// but the actual prompt now runs and returns the real answer. Confirmed against
// a real llama-swap server. ⚠️ `--bare` also disables hooks/LSP/plugin sync/
// CLAUDE.md auto-discovery — fine for this ISOLATED one-shot test, never safe to
// apply to a real interactive Codeman session (which needs hooks).
argv: (prompt) => ['--dangerously-skip-permissions', '--bare', '-p', prompt],
},
opencode: { confidence: 'verified', argv: (prompt) => ['run', prompt] },
codex: {
confidence: 'verified',
note: 'config STRUCTURE verified; codex only speaks the Responses API (dropped Chat-Completions Feb 2026) — expect FAIL against a plain OpenAI-compatible server, that is a protocol gap, not a bug here.',
argv: (prompt) => ['exec', '--dangerously-bypass-approvals-and-sandbox', prompt],
},
gemini: {
confidence: 'researched',
// --skip-trust: without it, an untrusted-folder check silently overrides
// --approval-mode yolo back to 'default' (confirmed live: "Approval mode
// overridden to 'default' because the current folder is not trusted").
argv: (prompt) => ['-p', prompt, '--approval-mode', 'yolo', '--skip-trust'],
},
pi: {
confidence: 'verified',
// --model custom/<id>: without an explicit --model, pi uses its own default
// provider (not our injected "custom" one) and fails with "No API key found
// for the selected model" — confirmed live. "custom" matches the provider name
// pi-models-json writes in custom-model-injection.ts. Verified end-to-end
// against a real llama-swap server after two real bugs were found and fixed:
// pi's `models` field must be an ARRAY of `{id}` objects (an object keyed by
// id silently loaded zero models), and PI_CONFIG_DIR does nothing for pi at
// all (grepped pi's own bundled source — not present anywhere); the actual
// working redirect is the CHILD PROCESS's `HOME` itself, since pi hardcodes
// `~/.pi/agent/models.json` with no dedicated override.
argv: (prompt, modelId) => ['--approve', '--model', `custom/${modelId}`, '-p', prompt],
},
grok: {
confidence: 'verified',
// -m <block name>: grok's config.toml (grok-toml template) declares the custom
// model under a fixed [model.<name>] block; GROK_CUSTOM_MODEL_NAME is that same
// name, imported from custom-model-injection.ts so the two can never drift apart.
// Verified end-to-end against a real llama-swap server after correcting the
// ORIGINAL recipe, which was wrong (env vars, not a config file — see the
// customModelInjection comment on grok's registry entry).
argv: (prompt) => ['--always-approve', '-m', GROK_CUSTOM_MODEL_NAME, '-p', prompt],
},
deepseek: {
confidence: 'unknown',
note: 'dsh is a profile launcher, not a documented one-shot prompt flag. Best-effort only.',
argv: (prompt) => ['--profile', 'headless', prompt],
},
omp: {
confidence: 'verified',
// --model custom/<id>: same reasoning as pi — omp's own default model has no
// credential, so without an explicit --model it never reaches our injected
// provider at all. Verified end-to-end against a real llama-swap server after
// the same two fixes as pi (array-shaped `models`, HOME-redirect instead of
// PI_CONFIG_DIR — omp hardcodes `~/.omp/agent/models.yml`).
argv: (prompt, modelId) => ['--model', `custom/${modelId}`, '-p', prompt],
},
};
// --- baseline server check ---------------------------------------------------
async function baselineCheck(
baseUrl: string,
apiKey: string,
authStyle: AuthStyle,
model: string | null,
prompt: string,
timeoutMs: number
): Promise<string> {
console.log(`\n=== Step 0: baseline check against ${baseUrl} (auth: ${authStyle}) ===`);
// Exactly ONE header, never both. An earlier version sent both auth conventions
// (Bearer + api-key) on the theory that an unused header is harmless — live-
// tested against a real llama-swap server, sending both reliably HUNG the
// request indefinitely (reproduced 3x: Bearer alone ~500ms, api-key alone
// ~600ms, both together no response inside a 15s timeout). Use --auth-style
// api-key for endpoints that specifically want that header (e.g. Azure AI
// Foundry); default 'bearer' covers everything else.
const authHeaders: Record<string, string> =
authStyle === 'api-key' ? { 'api-key': apiKey } : { Authorization: `Bearer ${apiKey}` };
let discoveredModel = model;
try {
const res = await fetch(`${baseUrl}/v1/models`, {
headers: authHeaders,
signal: AbortSignal.timeout(timeoutMs),
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = (await res.json()) as { data?: Array<{ id: string }> };
const ids: string[] = (body.data ?? []).map((m) => m.id);
console.log(`GET /v1/models -> ${ids.length ? ids.join(', ') : '(empty list)'}`);
if (!discoveredModel && ids.length) discoveredModel = ids[0];
} catch (err) {
console.error(`${TAG} GET /v1/models failed: ${(err as Error).message}`);
console.error(`${TAG} Is the server actually running at ${baseUrl}? Aborting.`);
process.exit(1);
}
if (!discoveredModel) {
console.error(`${TAG} No --model given and none discovered from /v1/models. Aborting.`);
process.exit(1);
}
// Live-tested against a real llama-swap server: a POST issued right after a GET on
// the same Node process reliably HANGS indefinitely (reproduced repeatedly — GET
// alone ~30ms, POST alone ~1-2s, GET-then-immediate-POST times out completely; a
// 2s pause between them fixed it every time). This looks like Node's fetch (undici)
// reusing a pooled keep-alive connection the server doesn't handle cleanly for a
// second request right behind a first. A short pause is the simplest portable fix
// (no extra deps, no need for undici's Agent/dispatcher API).
await new Promise((resolve) => setTimeout(resolve, 2000));
try {
const res = await fetch(`${baseUrl}/v1/chat/completions`, {
method: 'POST',
headers: { 'Content-Type': 'application/json', ...authHeaders },
body: JSON.stringify({
model: discoveredModel,
messages: [{ role: 'user', content: prompt }],
}),
signal: AbortSignal.timeout(timeoutMs),
});
if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
const body = (await res.json()) as { choices?: Array<{ message?: { content?: string } }> };
const reply: string = body.choices?.[0]?.message?.content ?? '';
if (!reply.trim()) throw new Error('empty reply');
console.log(`POST /v1/chat/completions -> "${reply.trim().slice(0, 200)}"`);
console.log('Server baseline: PASS\n');
} catch (err) {
console.error(`${TAG} POST /v1/chat/completions failed: ${(err as Error).message}`);
console.error(`${TAG} Server responded to /v1/models but not to a chat request. Aborting.`);
process.exit(1);
}
return discoveredModel;
}
// --- per-harness run ----------------------------------------------------------
interface ChildResult {
code: number | null;
stdout: string;
stderr: string;
timedOut: boolean;
}
function runChild(bin: string, argv: string[], env: Record<string, string>, searchDirs: string[], timeoutMs: number) {
return new Promise<ChildResult>((resolve) => {
let stdout = '';
let stderr = '';
let settled = false;
const child = spawn(bin, argv, {
env: { ...process.env, ...env, PATH: pathWithExtraDirs(searchDirs) },
stdio: ['ignore', 'pipe', 'pipe'],
});
const timer = setTimeout(() => {
if (settled) return;
settled = true;
child.kill('SIGKILL');
resolve({ code: null, stdout, stderr, timedOut: true });
}, timeoutMs);
child.stdout.on('data', (d) => (stdout += d.toString()));
child.stderr.on('data', (d) => (stderr += d.toString()));
child.on('error', (err) => {
if (settled) return;
settled = true;
clearTimeout(timer);
resolve({ code: null, stdout, stderr: `${stderr}\n${err.message}`, timedOut: false });
});
child.on('close', (code) => {
if (settled) return;
settled = true;
clearTimeout(timer);
resolve({ code, stdout, stderr, timedOut: false });
});
});
}
interface HarnessResult {
id: string;
confidence: Confidence | 'unsupported' | 'no-one-shot-recipe';
status: 'PASS' | 'FAIL' | 'UNCONFIRMED' | 'SKIP' | 'LIST';
detail: string;
}
async function runHarness(
entry: CliEntry,
opts: Opts,
model: string,
endpoint: CustomModelEndpoint
): Promise<HarnessResult> {
const id = entry.id;
const injectionCap = entry.capabilities.customModelInjection;
// Dynamic: driven by the REGISTRY's own capability, not a hardcoded id check.
// A future CLI declared unsupported is skipped automatically, same as antigravity today.
if (injectionCap.kind === 'unsupported') {
return {
id,
confidence: 'unsupported',
status: 'SKIP',
detail: 'no known custom-model mechanism (registry: unsupported)',
};
}
const oneShot = ONE_SHOT[id];
if (!oneShot) {
return {
id,
confidence: 'no-one-shot-recipe',
status: 'SKIP',
detail:
'registry supports custom-model injection for this CLI, but this script has no ONE_SHOT invocation entry yet — add one to test it',
};
}
const binary = entry.discovery.binaries[0] ?? id;
const searchDirs = entry.discovery.searchDirs;
const resolved = resolveBinary(binary, searchDirs);
if (!resolved) {
return {
id,
confidence: oneShot.confidence,
status: 'SKIP',
detail: `binary "${binary}" not found on PATH or search dirs`,
};
}
// The REAL injection logic — same function the production route calls.
const injection = buildCustomModelInjection(entry, endpoint, model);
let env: Record<string, string> = {};
let tempDir: string | null = null;
if (injection.kind === 'env') {
env = injection.envOverrides;
} else if (injection.kind === 'configDir') {
tempDir = mkdtempSync(join(tmpdir(), `codeman-local-llm-test-${id}-`));
env = applyConfigDirInjection(tempDir, injection);
}
// injection.kind === 'unsupported' already handled via injectionCap above.
const argv = oneShot.argv(opts.prompt, model);
if (opts.list) {
const detail = `${binary} ${argv.join(' ')} | env: ${Object.keys(env).join(', ')}${tempDir ? ` | configDir: ${tempDir}` : ''}`;
if (tempDir && !opts.keepTemp) rmSync(tempDir, { recursive: true, force: true });
return { id, confidence: oneShot.confidence, status: 'LIST', detail };
}
const { code, stdout, stderr, timedOut } = await runChild(binary, argv, env, searchDirs, opts.timeout);
let detailSuffix = '';
if (tempDir && !opts.keepTemp) rmSync(tempDir, { recursive: true, force: true });
else if (tempDir) detailSuffix = ` [config kept at ${tempDir}]`;
if (timedOut) {
return {
id,
confidence: oneShot.confidence,
status: 'FAIL',
detail: `timed out after ${opts.timeout}ms. stderr: ${stderr.slice(-300)}${detailSuffix}`,
};
}
const reply = stdout.trim();
const matched = /hello/i.test(reply) && /world/i.test(reply);
const softStatus: HarnessResult['status'] = oneShot.confidence === 'verified' ? 'FAIL' : 'UNCONFIRMED';
if (code !== 0) {
return {
id,
confidence: oneShot.confidence,
status: softStatus,
detail: `exit ${code}. stderr: ${stderr.trim().slice(-300) || '(empty)'}${detailSuffix}`,
};
}
if (!reply) {
return { id, confidence: oneShot.confidence, status: softStatus, detail: `exit 0 but empty stdout${detailSuffix}` };
}
if (matched) {
return { id, confidence: oneShot.confidence, status: 'PASS', detail: `${reply.slice(0, 200)}${detailSuffix}` };
}
return {
id,
confidence: oneShot.confidence,
status: 'UNCONFIRMED',
detail: `reply didn't match heuristic, judge by eye: "${reply.slice(0, 300)}"${detailSuffix}`,
};
}
// --- main ---------------------------------------------------------------------
async function main(): Promise<void> {
const configDefaults = loadConfigFile();
const opts = parseArgs(process.argv.slice(2), configDefaults);
if (opts.help) {
printUsage();
process.exit(0);
}
// Dynamic: pulled from the live registry, not a hardcoded id list. `kind === 'agent'`
// excludes 'shell' (no model/endpoint concept). Antigravity stays in this list (it IS
// an enabled agent CLI) — it's the `unsupported` capability check in runHarness that
// skips it, not an exclusion here.
const allEntries = enabledClis().filter((e) => e.kind === 'agent');
const byId = new Map<string, CliEntry>(allEntries.map((e) => [e.id as string, e]));
const ids: string[] = opts.only ?? [...byId.keys()];
const unknownIds = ids.filter((id) => !byId.has(id));
if (unknownIds.length) {
console.error(`${TAG} unknown harness id(s): ${unknownIds.join(', ')}`);
console.error(`${TAG} known ids (from the live CLI registry): ${[...byId.keys()].join(', ')}`);
process.exit(1);
}
const entries = ids.map((id) => byId.get(id)!);
// --probe-help never touches the network — no --base-url needed for it.
if (opts.probeHelp) {
for (const entry of entries) {
const binary = entry.discovery.binaries[0] ?? entry.id;
const resolved = resolveBinary(binary, entry.discovery.searchDirs);
console.log(`\n=== ${entry.id} (${binary}) ===`);
if (!resolved) {
console.log('(not found on PATH or search dirs)');
continue;
}
printHelp(binary, entry.discovery.searchDirs);
}
process.exit(0);
}
if (!opts.baseUrl) {
console.error(`${TAG} --base-url is required (pass it, or set "baseUrl" in ${CONFIG_PATH}).`);
console.error(`${TAG} See ${CONFIG_EXAMPLE_PATH} for the config file shape.\n`);
printUsage();
process.exit(1);
}
opts.baseUrl = opts.baseUrl.replace(/\/+$/, '');
const endpoint: CustomModelEndpoint = {
id: 'standalone-test',
label: 'standalone test',
baseUrl: opts.baseUrl,
apiKey: opts.apiKey,
};
// --list is a pure dry run: never touch the network, even if --model was given.
let model: string;
if (opts.list) {
model = opts.model ?? 'local-model';
console.log(`\n=== Step 0 skipped (--list never hits the network; using placeholder "${model}") ===\n`);
} else {
model = await baselineCheck(opts.baseUrl, opts.apiKey, opts.authStyle, opts.model, opts.prompt, opts.timeout);
}
console.log(`=== Testing ${entries.length} harness(es) ===`);
const results: HarnessResult[] = [];
for (const entry of entries) {
process.stdout.write(`\n--- ${entry.id} ---\n`);
const result = await runHarness(entry, opts, model, endpoint);
results.push(result);
console.log(`${result.status}: ${result.detail}`);
}
console.log('\n=== Summary ===');
const width = Math.max(...results.map((r) => r.id.length)) + 2;
for (const r of results) {
console.log(`${r.id.padEnd(width)} [${r.confidence.padEnd(20)}] ${r.status.padEnd(11)} ${r.detail.slice(0, 100)}`);
}
const hardFail = results.some((r) => r.status === 'FAIL' && r.confidence === 'verified');
if (hardFail) {
console.error(
`\n${TAG} at least one VERIFIED harness FAILed — that's a real regression, not just an unconfirmed guess.`
);
process.exit(1);
}
process.exit(0);
}
main().catch((err) => {
console.error(`${TAG} unexpected error:`, err);
process.exit(1);
});
+75 -26
View File
@@ -47,7 +47,7 @@ later call opens with, and your first REAL call performs them anyway:
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
```
⚠️ **Never spend a Bash call on this check alone.** §1's block opens with this same
@@ -75,8 +75,8 @@ PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh"
mkdir -p "$(dirname "$PRE")"
# Rewrite unless the file already ends with THIS version's stamp, so a stale or a
# half-written file self-heals here instead of costing you a round trip to rm it.
grep -qs '^CODEMAN_PREAMBLE=1.22.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
grep -qs '^CODEMAN_PREAMBLE=1.30.1$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -150,6 +150,27 @@ _trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
# ---- the composer: is the prompt still sitting there, unsent? ----
# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the
# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that
# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt
# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through
# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS
# the composer and keeps pressing Enter while the prompt is still there.
_composer_text() { # <sid> -> the composer's text with ALL whitespace removed: "" once
# the prompt was taken, "?" when the pane shows no composer at all. The composer is
# the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher
# up in the transcript, so only the last one says whether the text was taken.
local t
t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1)
[ -n "$t" ] || { printf '?'; return 0; }
# Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does
# not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH).
printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g"
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
@@ -262,12 +283,16 @@ spawn_workers() {
# worker a silent no-op that still "succeeds" and reports the previous turn's state.
# Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a
# deliberate duplicate, at the SAME number (§5.3).
# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the
# typed prompt stranded on the composer while a long wait runs its whole timeout
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude
# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints),
# leaving the typed prompt stranded on the composer while a long wait runs its whole
# timeout (observed live, twelve reviews in a row). So the first wait is short; on its
# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate:
# the server re-waits without retyping, §5.3) and kept open in the background, while
# the composer is READ (_composer_text) and, as long as the prompt is still sitting
# there, a bare \r goes out about every ten seconds, up to twelve times. An empty
# composer ends the loop, so a prompt that was taken is never nudged again, and the
# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
@@ -275,7 +300,7 @@ spawn_workers() {
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
@@ -290,16 +315,38 @@ sendwait() {
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
# ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call,
# in the background, while the Enter loop below works the composer. Signals have
# no history: a `stop` that fires while no wait is open (during a composer read
# between two short waits, measured) is lost, and the next wait then runs its
# whole timeout on a turn that already ended. The resend is a tagged DUPLICATE,
# so the server skips the write and re-waits without retyping (§5.3).
tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" &
bg=$!
# The prompt's head with whitespace removed, matched literally (the "$head"
# quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob).
head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24)
while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended
c=$(_composer_text "$sid")
if [ "$c" = '?' ]; then
[ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it
elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then
break # composer empty (taken) or holding other text
fi
n=$((n+1))
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done
done
wait "$bg"
# The duplicate reports `delivered:false` -- truthfully, but about the wrong send.
# The first one delivered, so carry that forward, or §1's cleanup reads a completed
# turn as an undelivered one and keeps a finished worker forever.
r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp")
rm -f "$tmp"
fi
printf '%s\n' "$r"
}
@@ -325,10 +372,10 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.22.0
CODEMAN_PREAMBLE=1.30.1
PREAMBLE
)
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
```
Every later Bash call that touches the API starts with the same two loader lines from
@@ -379,7 +426,7 @@ and no per-call body to hand-build.
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null # §0 loader
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
N=(alpha beta) # INVENT one fresh case name per worker; never list cases first
# (a name may carry a mode: `beta:deepseek`, see below)
T=('reply with one line: the absolute path of your working directory'
@@ -440,9 +487,11 @@ Four things this block leans on, each one link away, no detour needed to run it:
skill: §5.1. Those workspaces do get hooks now, unless the operator disabled it.
- `sendwait` supplies the `\r`, picks a fresh `seq`, and self-heals a stranded Enter.
A prompt without the `\r` is never submitted (§3), a reused `seq` is silently
swallowed as an already-applied duplicate, and an Enter eaten by an Ink repaint
strands the prompt on the composer until a bare `\r` follows: all three are reasons
to let `sendwait` build the call rather than hand-rolling it.
swallowed as an already-applied duplicate, and a lost Enter strands the prompt on the
composer until a bare `\r` follows: Claude Code 2.1.277 and later ignore Enter for the
first 30 to 50 seconds after the composer paints while still taking the text, so
`sendwait` reads the composer and keeps pressing Enter until the prompt has left it.
All three are reasons to let `sendwait` build the call rather than hand-rolling it.
- Each `sendwait` costs that worker one billed turn, as does every prompt you send it.
- Deleting the sessions does **not** remove the case directories. They are marked as
agent-created, so `GET /api/v1/cases/agent-created` lists them for cleanup: §5.14.
+66 -19
View File
@@ -1,4 +1,4 @@
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -72,6 +72,27 @@ _trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
# ---- the composer: is the prompt still sitting there, unsent? ----
# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the
# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that
# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt
# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through
# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS
# the composer and keeps pressing Enter while the prompt is still there.
_composer_text() { # <sid> -> the composer's text with ALL whitespace removed: "" once
# the prompt was taken, "?" when the pane shows no composer at all. The composer is
# the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher
# up in the transcript, so only the last one says whether the text was taken.
local t
t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1)
[ -n "$t" ] || { printf '?'; return 0; }
# Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does
# not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH).
printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g"
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
@@ -184,12 +205,16 @@ spawn_workers() {
# worker a silent no-op that still "succeeds" and reports the previous turn's state.
# Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a
# deliberate duplicate, at the SAME number (§5.3).
# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the
# typed prompt stranded on the composer while a long wait runs its whole timeout
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude
# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints),
# leaving the typed prompt stranded on the composer while a long wait runs its whole
# timeout (observed live, twelve reviews in a row). So the first wait is short; on its
# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate:
# the server re-waits without retyping, §5.3) and kept open in the background, while
# the composer is READ (_composer_text) and, as long as the prompt is still sitting
# there, a bare \r goes out about every ten seconds, up to twelve times. An empty
# composer ends the loop, so a prompt that was taken is never nudged again, and the
# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
@@ -197,7 +222,7 @@ spawn_workers() {
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
@@ -212,16 +237,38 @@ sendwait() {
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
# ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call,
# in the background, while the Enter loop below works the composer. Signals have
# no history: a `stop` that fires while no wait is open (during a composer read
# between two short waits, measured) is lost, and the next wait then runs its
# whole timeout on a turn that already ended. The resend is a tagged DUPLICATE,
# so the server skips the write and re-waits without retyping (§5.3).
tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" &
bg=$!
# The prompt's head with whitespace removed, matched literally (the "$head"
# quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob).
head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24)
while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended
c=$(_composer_text "$sid")
if [ "$c" = '?' ]; then
[ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it
elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then
break # composer empty (taken) or holding other text
fi
n=$((n+1))
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done
done
wait "$bg"
# The duplicate reports `delivered:false` -- truthfully, but about the wrong send.
# The first one delivered, so carry that forward, or §1's cleanup reads a completed
# turn as an undelivered one and keeps a finished worker forever.
r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp")
rm -f "$tmp"
fi
printf '%s\n' "$r"
}
@@ -247,4 +294,4 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.22.0
CODEMAN_PREAMBLE=1.30.1
+1 -1
View File
@@ -21,7 +21,7 @@ by sourcing the preamble file the §0 bootstrap wrote, and checking its version
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
```
Do **not** re-paste the preamble body into each call. Sourcing it is what retires the
+124 -12
View File
@@ -10,10 +10,12 @@ import { randomUUID } from 'node:crypto';
import { realpathSync } from 'node:fs';
import fs from 'node:fs/promises';
import { basename, extname, isAbsolute } from 'node:path';
import { isBlockedAttachmentPath, loadAttachmentGuardConfig } from './config/attachment-guard.js';
import { isBlockedAttachmentPath, isUnderTree, loadAttachmentGuardConfig } from './config/attachment-guard.js';
import { EDITABLE_EXTENSIONS } from './config/file-editing.js';
import { validateSessionFilePath } from './web/route-helpers.js';
import { remoteProbePaths, RemoteFileAccessError, type RemoteProbe } from './remote-files.js';
import type { AttachmentDetectedEvent, AttachmentDetectedType } from './types.js';
import type { SessionRemote } from './types/session.js';
/**
* Playable media extensions, single-sourced here because the WORKSPACE preview
@@ -215,6 +217,106 @@ export interface RegisterExternalAttachmentOptions {
* `codeman attach` CLI (which POSTs directly when a session id is known).
*/
forceWorkspaceConfinement?: boolean;
/**
* Remote (SSH) case: the path exists on the REMOTE host, so it is resolved and
* stat'ed there (`remoteProbePaths`) instead of with local `realpathSync`/`fs.stat`,
* which cannot see it at all (#415). A file outside the case directory is
* unreachable exactly like a file inside it.
*
* `sessionWorkingDir` must then be the REMOTE path too, and the workspace
* confinement check (when active) compares against the remotely canonicalized root,
* so a symlinked `remotePath` does not refuse every registration.
*/
remote?: SessionRemote;
/**
* Remote only: `[file, workspaceRoot]` probes a caller already resolved in a BATCHED
* `remoteProbePaths` call (the attachment-history list does one round trip for the
* whole history). Skips this registration's own ssh probe; every guard below still
* runs on the same resolved path it would have produced itself.
*/
remoteProbes?: readonly [RemoteProbe | null, RemoteProbe | null];
}
/**
* A path an attachment request resolved to, on whichever host it lives — the local
* filesystem or the remote host of a remote-SSH case. The rest of
* {@link registerExternalAttachment} (guards, extension allowlist, registry) is then
* host-agnostic: it only ever sees canonical absolute paths and numbers.
*/
interface ResolvedAttachmentFile {
resolvedPath: string;
size: number;
mtimeMs: number;
isFile: boolean;
extension: string;
/** Remote only: the workspace root, with symlinks resolved on the remote host. */
workspaceRoot?: string;
}
/** `extension` the way the attachment registry defines it (no dot, lowercased). */
function attachmentExtensionOf(path: string): string {
return extname(path).toLowerCase().replace(/^\./, '');
}
/** Local resolution: the historical realpath + stat. */
async function resolveLocalAttachment(requestedPath: string): Promise<ResolvedAttachmentFile> {
let resolvedPath: string;
try {
resolvedPath = realpathSync(requestedPath);
} catch {
throw new AttachmentRegistrationError('Attachment file not found', 404);
}
const stat = await fs.stat(resolvedPath);
return {
resolvedPath,
size: stat.size,
mtimeMs: stat.mtimeMs ?? 0,
isFile: typeof stat.isFile === 'function' ? stat.isFile() : true,
extension: attachmentExtensionOf(resolvedPath),
};
}
/**
* Remote resolution for a remote-SSH case: ONE ssh round trip returns the
* symlink-resolved path, the size/mtime and the kind, for the file AND (when a
* workspace is known) its root, which the confinement check compares against.
*/
async function resolveRemoteAttachment(
requestedPath: string,
remote: SessionRemote,
sessionWorkingDir?: string,
preResolved?: readonly [RemoteProbe | null, RemoteProbe | null]
): Promise<ResolvedAttachmentFile> {
const paths = sessionWorkingDir ? [requestedPath, sessionWorkingDir] : [requestedPath];
let probes: ReadonlyArray<RemoteProbe | null>;
if (preResolved) {
probes = preResolved;
} else {
try {
probes = await remoteProbePaths(remote, paths);
} catch (err) {
// 502 marks the TRANSPORT as the failure, distinct from the file's own 404/403,
// so a history listing can report the entry as unknown rather than missing.
throw new AttachmentRegistrationError(
err instanceof RemoteFileAccessError ? err.message : 'remote host unreachable',
502
);
}
}
const [probe, rootProbe] = probes;
if (!probe) {
throw new AttachmentRegistrationError('Attachment file not found', 404);
}
return {
resolvedPath: probe.realPath,
size: probe.size,
mtimeMs: probe.mtimeMs,
isFile: probe.kind === 'file',
extension: attachmentExtensionOf(probe.realPath),
workspaceRoot: rootProbe?.realPath,
};
}
export async function registerExternalAttachment(
@@ -226,12 +328,9 @@ export async function registerExternalAttachment(
throw new AttachmentRegistrationError('Attachment path must be an absolute local path');
}
let resolvedPath: string;
try {
resolvedPath = realpathSync(requestedPath);
} catch {
throw new AttachmentRegistrationError('Attachment file not found', 404);
}
const resolved = await (options.remote
? resolveRemoteAttachment(requestedPath, options.remote, options.sessionWorkingDir, options.remoteProbes)
: resolveLocalAttachment(requestedPath));
// COD-53: enforce the active attachment-guard policy on the symlink-resolved
// path before doing anything else.
@@ -243,7 +342,10 @@ export async function registerExternalAttachment(
// the caller forces it for this registration (the magic-link scanner — see
// forceWorkspaceConfinement). Strictly more restrictive than the blocklist.
const workingDir = options.sessionWorkingDir;
if (!workingDir || !validateSessionFilePath(workingDir, resolvedPath)) {
const confined = options.remote
? !!workingDir && isUnderTree(resolved.resolvedPath, resolved.workspaceRoot ?? workingDir)
: !!workingDir && !!validateSessionFilePath(workingDir, resolved.resolvedPath);
if (!confined) {
throw new AttachmentRegistrationError('Access to this file is blocked', 403);
}
}
@@ -253,20 +355,30 @@ export async function registerExternalAttachment(
// operator-configured extra trees. Symlinks are already resolved above.
// Cross-workspace attachment of non-blocked files stays allowed, so
// codeman-publish and the ~/.codeman review loop keep working.
if (isBlockedAttachmentPath(resolvedPath, guard.blockedTrees)) {
//
// The list is a pattern list over ABSOLUTE paths, so it is host-agnostic and holds
// for a remote path exactly as it does for a local one, with ONE exception worth
// knowing: `isSensitivePath`'s three home-anchored members (`~/.claude.json`,
// `~/.claude/settings.json`, `~/.claude/settings.local.json`) resolve against THIS
// host's `homedir()`, so on a remote host with a different home they do not match.
// Everything else in that list is depth-anchored (`/.ssh/`, `/.aws/credentials`,
// `/.claude/.credentials.json`, ...) and applies unchanged.
if (isBlockedAttachmentPath(resolved.resolvedPath, guard.blockedTrees)) {
throw new AttachmentRegistrationError('Access to this file is blocked', 403);
}
const extension = extname(resolvedPath).toLowerCase().replace(/^\./, '');
const resolvedPath = resolved.resolvedPath;
const extension = resolved.extension;
if (!isSupportedAttachmentExtension(extension)) {
throw new AttachmentRegistrationError('Unsupported attachment type');
}
const stat = await fs.stat(resolvedPath);
if (typeof stat.isFile === 'function' && !stat.isFile()) {
if (!resolved.isFile) {
throw new AttachmentRegistrationError('Attachment path is not a file');
}
const stat = { size: resolved.size, mtimeMs: resolved.mtimeMs };
const existing = attachmentRegistry.findByFilePath(sessionId, resolvedPath);
if (existing) {
existing.size = stat.size;
+44
View File
@@ -261,6 +261,19 @@ const echoSchema = z
})
.strict();
/**
* `capabilities.customModelInjection.launchModel`: the `model` launch-param value that
* selects the injected provider, with `{modelId}` standing for the chosen id. Bounded to
* the characters the `model`/`model-pi` token patterns accept plus the placeholder braces,
* so a template can never smuggle a token the argv engine would have to quote.
*/
const launchModelTemplate = z
.string()
.min(1)
.max(120)
.regex(/^[a-zA-Z0-9._\-/:{}]+$/)
.optional();
const capabilitiesSchema = z
.object({
external: z.boolean(),
@@ -317,6 +330,37 @@ const capabilitiesSchema = z
privilegedEnvKeys: z.array(envName).max(8),
gates: z.record(z.string(), z.object({ minVersion: z.string().max(20), failClosed: z.boolean() }).strict()),
maxFrameBytes: z.number().int().positive().optional(),
customModelInjection: z.discriminatedUnion('kind', [
z
.object({
kind: z.literal('env'),
baseUrlVar: envName,
apiKeyVar: envName,
// Empty is valid: deepseek's model routing is a profile-composition concern, not
// an env var, so it declares baseUrl/apiKey injection with no model var at all.
modelVars: z.array(envName).max(8),
launchModel: launchModelTemplate,
})
.strict(),
z
.object({
kind: z.literal('configContentEnv'),
envVar: envName,
template: z.literal('opencode-json'),
launchModel: launchModelTemplate,
})
.strict(),
z
.object({
kind: z.literal('configDir'),
dirEnvVar: envName,
fileName: z.string().min(1).max(80),
template: z.enum(['codex-toml', 'pi-models-json', 'omp-models-yml', 'grok-toml']),
launchModel: launchModelTemplate,
})
.strict(),
z.object({ kind: z.literal('unsupported') }).strict(),
]),
})
.strict();
+157 -2
View File
@@ -189,6 +189,12 @@ const CLAUDE: CliEntry = {
unset: ['CLAUDECODE'],
tmuxSetenvKeys: [],
dockerExecEnvNames: [],
// Deliberately excludes ANTHROPIC_* (base URL / API key / default-model overrides):
// custom-model-injection.ts's claude recipe uses those names, but they must reach a
// session ONLY through the admin-configured, SSRF-guarded custom-model route, never
// through a plain client-supplied envOverrides field. Widening this prefix would let
// any session-create caller redirect a session's Anthropic traffic and credentials to
// an arbitrary, unvalidated URL.
allowedPrefixes: ['CLAUDE_CODE_'],
allowedKeys: ['CLAUDE_CONFIG_DIR'],
},
@@ -220,8 +226,28 @@ const CLAUDE: CliEntry = {
statusLineTelemetry: true,
model: { source: 'claude-settings-file' },
privilegedParams: [],
privilegedEnvKeys: [],
// ANTHROPIC_* is NOT in allowedPrefixes/allowedKeys above (deliberately — see the
// allowedPrefixes comment nearby), so these are unreachable via plain envOverrides
// today; listed here only so the dedicated custom-model route (docs/custom-model-endpoints-plan.md
// chunk 5) clamps them for a non-granted multi-user owner the same way every other
// CLI's injection vars are clamped, the day that route widens who can set them.
privilegedEnvKeys: [
'ANTHROPIC_BASE_URL',
'ANTHROPIC_API_KEY',
'ANTHROPIC_DEFAULT_SONNET_MODEL',
'ANTHROPIC_DEFAULT_HAIKU_MODEL',
'ANTHROPIC_DEFAULT_OPUS_MODEL',
],
gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } },
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — verified by hand against a real
// llama.cpp server. Claude reads these at process start only, so switching requires a
// respawn, never a live hot-swap.
customModelInjection: {
kind: 'env',
baseUrlVar: 'ANTHROPIC_BASE_URL',
apiKeyVar: 'ANTHROPIC_API_KEY',
modelVars: ['ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL'],
},
},
overlays: {
// Mirrors the local default so the remote/in-container agent runs non-interactively
@@ -287,6 +313,7 @@ const SHELL: CliEntry = {
privilegedParams: [],
privilegedEnvKeys: [],
gates: {},
customModelInjection: { kind: 'unsupported' }, // a raw shell has no "model" concept
},
overlays: {
// No `remote` entry: defaultRemoteCommandForMode special-cases kind==='shell' directly
@@ -366,6 +393,15 @@ const OPENCODE: CliEntry = {
...agentDefaults(),
altScreen: 'strip-mux-only',
echo: { policy: 'buffer', anchor: { kind: 'cursor' }, predictProfile: undefined },
// Verified by hand against a real llama.cpp server. Reuses the SAME env var opencode's
// own `env.configContentVar` already declares — the builder in custom-model-injection.ts
// must merge into whatever opencode config Codeman would otherwise send, not clobber it.
customModelInjection: { kind: 'configContentEnv', envVar: 'OPENCODE_CONFIG_CONTENT', template: 'opencode-json' },
// OPENCODE_CONFIG_CONTENT already matches the OPENCODE_ allowedPrefix above, so it was
// ALREADY reachable via plain envOverrides before this feature existed — it replaces
// opencode's whole config, provider api keys included, so a non-granted multi-user owner
// sending it is a pre-existing credential-redirection gap, not one this feature opens.
privilegedEnvKeys: ['OPENCODE_CONFIG_CONTENT'],
},
overlays: {
credStore: { rel: '.config/opencode', seedWhole: true },
@@ -455,6 +491,23 @@ const CODEX: CliEntry = {
// `dangerouslyBypassApprovals` on the wire), so it is the one that would have caught a
// regression; `schema.ts` now rejects a name that is not a declared param.
privilegedParams: [{ param: 'bypassApprovals', clampTo: false }],
// Verified by hand against a real llama.cpp server. Written to an isolated CODEX_HOME
// so the user's real ~/.codex/config.toml is never touched.
customModelInjection: {
kind: 'configDir',
dirEnvVar: 'CODEX_HOME',
fileName: 'config.toml',
template: 'codex-toml',
},
// CODEX_HOME already matches the CODEX_ allowedPrefix above, so it was ALREADY
// reachable via plain envOverrides before this feature existed. It is arguably
// MORE sensitive than a bare base-url var: a redirected CODEX_HOME points codex at a
// config.toml a non-granted owner fully controls, which can restate sandbox/approval
// policy INSIDE that file — a path the argv-level `bypassApprovals` clamp above
// cannot see or stop.
// CODEMAN_CUSTOM_MODEL_API_KEY: the credential config.toml's env_key references
// (see custom-model-injection.ts) — same reasoning as CODEX_HOME above.
privilegedEnvKeys: ['CODEX_HOME', 'CODEMAN_CUSTOM_MODEL_API_KEY'],
},
overlays: {
credStore: {
@@ -538,6 +591,20 @@ const GEMINI: CliEntry = {
// MATERIALIZE a config (not just touch an already-sent one) or a non-granted owner who
// sends no geminiConfig at all would still get yolo for free.
privilegedParams: [{ param: 'approvalMode', clampTo: 'auto_edit', materializeWhenAbsent: true }],
// Web-researched, unverified — needs a restart to pick up (CLI reads these at process
// start). Confirm the exact model-override env var name against the installed
// gemini-cli version before shipping.
customModelInjection: {
kind: 'env',
baseUrlVar: 'GOOGLE_GEMINI_BASE_URL',
apiKeyVar: 'GEMINI_API_KEY',
modelVars: ['GEMINI_MODEL'],
},
// All three already match the GEMINI_/GOOGLE_ allowedPrefixes above, so they were
// ALREADY reachable via plain envOverrides before this feature existed — a non-granted
// multi-user owner redirecting a gemini session's endpoint/credentials is a
// pre-existing gap this feature's analysis surfaced, not one it opens.
privilegedEnvKeys: ['GOOGLE_GEMINI_BASE_URL', 'GEMINI_API_KEY', 'GEMINI_MODEL'],
},
overlays: {
credStore: { rel: '.gemini', seedWhole: true }, // also covers antigravity — see its own entry
@@ -603,6 +670,10 @@ const ANTIGRAVITY: CliEntry = {
// Like codex: an ABSENT config already defaults safe (no bypass flag), so only a
// SENT config needs the flag forced off — nothing is materialized.
privilegedParams: [{ param: 'dangerouslySkipPermissions', clampTo: false }],
// No known CLI/env/config mechanism — Antigravity's own docs describe a GUI-only
// custom-endpoint setting and explicitly say it "cannot currently" become the core
// reasoning model. Toolbar entry stays disabled for this mode.
customModelInjection: { kind: 'unsupported' },
},
overlays: {
// No credStore of its own: agy nests its whole state under ~/.gemini/antigravity-cli/,
@@ -693,6 +764,34 @@ const PI: CliEntry = {
// just answer "yes" to, so omitting --approve is not itself a clamp — MATERIALIZE
// approveProjectTrust:false so buildPiCommand emits --no-approve outright.
privilegedParams: [{ param: 'approveProjectTrust', clampTo: false, materializeWhenAbsent: true }],
// CORRECTED after live-testing: `PI_CONFIG_DIR` does NOT exist anywhere in pi's own
// bundled source (grepped the installed package directly) — it does nothing for pi
// itself, despite being a real Codeman env var that OTHER things (omp) read. The
// confirmed working redirect is `HOME` itself: pi hardcodes `~/.pi/agent/models.json`
// with no dedicated override, so redirecting the CHILD PROCESS's HOME is what
// actually relocates it (verified: a model written under an isolated HOME's
// `.pi/agent/models.json` shows up in `pi --list-models` and answers a real prompt
// against a real llama-swap server; PI_CONFIG_DIR alone left it silently unable to
// see any provider). ⚠️ This is a bigger blast radius than a dedicated config-dir
// var: it also redirects pi's real sessions/auth/extensions for the DURATION of a
// custom-model session, not just its provider config — document this trade-off
// wherever this capability is surfaced.
customModelInjection: {
kind: 'configDir',
dirEnvVar: 'HOME',
fileName: '.pi/agent/models.json',
template: 'pi-models-json',
// Writing models.json is not enough: without `--model custom/<id>` pi stays on its
// own default provider and fails with "No API key found for the selected model"
// (confirmed live). `custom` is the provider name pi-models-json declares.
launchModel: 'custom/{modelId}',
},
// HOME is not `PI_`-prefixed, so unlike the old (wrong) PI_CONFIG_DIR guess this was
// never reachable via the generic envOverrides allowlist at all — listed here anyway,
// matching the documented pattern for every other CLI's dir-redirect var, since a
// redirected HOME is at least as sensitive as CODEX_HOME/GROK_HOME (pi executes
// repo-local .pi/extensions TypeScript — see the External CLI modes note in CLAUDE.md).
privilegedEnvKeys: ['HOME'],
},
overlays: {
credStore: {
@@ -788,6 +887,28 @@ const GROK: CliEntry = {
// already its safe interactive ask-mode, so the multi-user clamp only needs to force an
// EXPLICITLY-SENT bypass flag back off — nothing is materialized when config is absent.
privilegedParams: [{ param: 'alwaysApprove', clampTo: false }],
// CORRECTED after live-testing against a real grok binary: the original `env` kind
// (GROK_BASE_URL/GROK_MODEL/XAI_API_KEY) produced "Not signed in" — those env vars
// are NOT grok's real custom-endpoint mechanism. The real one (verified against
// xAI's own docs) is a `[model.<name>]` block in a config.toml under GROK_HOME,
// the same configDir shape as codex/pi/omp. `api_backend = "chat_completions"` is
// explicitly supported (unlike codex, which dropped it) — grok CAN talk to a plain
// OpenAI Chat-Completions server directly.
customModelInjection: {
kind: 'configDir',
dirEnvVar: 'GROK_HOME',
fileName: 'config.toml',
template: 'grok-toml',
// The `[model.<name>]` block the grok-toml template writes; `--model <name>` is what
// selects it (GROK_CUSTOM_MODEL_NAME in custom-model-injection.ts, pinned equal by
// test/custom-model-injection.test.ts so the two cannot drift).
launchModel: 'codeman-custom',
},
// GROK_HOME already matches the GROK_ allowedPrefix above, so it was ALREADY
// reachable via plain envOverrides before this feature existed — same reasoning
// as CODEX_HOME: a redirected config dir can restate policy the argv-level
// `alwaysApprove` clamp above cannot see.
privilegedEnvKeys: ['GROK_HOME'],
},
overlays: {
// ~/.grok also holds sessions/, memory/, downloads/ (the ~160MB binary), completions/,
@@ -943,7 +1064,23 @@ const DEEPSEEK: CliEntry = {
// The half no other CLI needs. `DSH_*` is an allowlisted envOverrides prefix and
// applyEnvOverrides() runs LAST, so without this a non-granted owner could send
// DSH_PERMISSION_MODE on the same request and land after the config clamp.
// ⚠️ DEEPSEEK_API_KEY deliberately stays OUT of this list (see the docstring on
// clampEnvOverridesForOwner() in session-routes.ts): _configureCliEnv() forwards the
// SERVER's own key into every dsh pane, so DEEPSEEK_BASE_URL is the exfiltration
// vector, not the key itself — a non-granted owner supplying THEIR OWN key removes
// privilege rather than granting it, and clamping it here was a real regression
// (test/deepseek-mode.test.ts) fixed before this shipped.
privilegedEnvKeys: ['DSH_PERMISSION_MODE', 'DSH_HOME', 'DEEPSEEK_BASE_URL'],
// Web-researched, unverified, partial: reuses the already-existing DEEPSEEK_BASE_URL/
// DEEPSEEK_API_KEY keys above. No modelVars — dsh's model is a profile-composition
// entry (see `model: { source: 'none' }` above), not an env var, so forcing a specific
// model name may not fully work; verify against a real profile before shipping.
customModelInjection: {
kind: 'env',
baseUrlVar: 'DEEPSEEK_BASE_URL',
apiKeyVar: 'DEEPSEEK_API_KEY',
modelVars: [],
},
},
overlays: {
// No credStore: dsh keeps everything under $DSH_HOME (default ~/.dsh), which is
@@ -1045,7 +1182,25 @@ const OMP: CliEntry = {
// Where omp resolves its auth from. No known concrete exfiltration path today (omp
// forwards no operator-held key into a pane), but a non-granted owner redirecting where
// a shared multi-tenant deployment resolves auth is not something to allow silently.
privilegedEnvKeys: ['OMP_AUTH_BROKER_URL', 'OMP_AUTH_BROKER_TOKEN'],
// HOME added for custom-model-injection.ts's omp recipe (see below). Unlike pi,
// PI_CONFIG_DIR genuinely IS one of the env vars omp reads (per the DeepSeek/OMP
// note in CLAUDE.md) — but live-testing this feature found it did NOT relocate
// omp's model config the way expected, while redirecting HOME itself (like pi)
// worked immediately (verified end-to-end: a real "hello world" reply came back).
privilegedEnvKeys: ['OMP_AUTH_BROKER_URL', 'OMP_AUTH_BROKER_TOKEN', 'HOME'],
// Verified end-to-end against a real llama-swap server (live-tested, not just
// researched — a real "hello world" reply came back). Same HOME-redirect mechanism
// as pi (see its customModelInjection comment for the full reasoning) — omp hardcodes
// `~/.omp/agent/models.yml` with no dedicated config-dir override either.
customModelInjection: {
kind: 'configDir',
dirEnvVar: 'HOME',
fileName: '.omp/agent/models.yml',
template: 'omp-models-yml',
// Same as pi: omp's own default model has no credential, so without an explicit
// `--model custom/<id>` it never reaches the injected provider at all.
launchModel: 'custom/{modelId}',
},
},
overlays: {
// `~/.omp/agent` also holds agent.db/history.db/models.db (SQLite caches) and
+51
View File
@@ -457,6 +457,57 @@ export interface CliCapabilities {
gates: Record<string, { minVersion: string; failClosed: boolean }>;
/** Cap on a single terminal frame, when this CLI needs a tighter one than the default. */
maxFrameBytes?: number;
/**
* How this CLI is pointed at a user-supplied custom OpenAI-compatible
* endpoint (local, e.g. llama.cpp, or cloud, e.g. Azure AI Foundry) — the
* Custom Model Endpoint Profiles feature (`docs/custom-model-endpoints-plan.md`). Declared
* per entry, never branched on id, same as every other capability here.
*
* `env`: plain env vars (claude's `ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY`/
* `ANTHROPIC_DEFAULT_*_MODEL`). `configContentEnv`: a full config blob
* carried in one env var (opencode's `OPENCODE_CONFIG_CONTENT`).
* `configDir`: a generated config file under an isolated, dir-redirect-env-
* pointed directory so the user's real CLI config is never touched
* (codex's `CODEX_HOME`/`config.toml`, pi/omp's `PI_CONFIG_DIR`, grok's
* `GROK_HOME`/`config.toml`). `unsupported`: no known mechanism
* (antigravity) — the toolbar entry stays disabled for this CLI.
*
* ⚠️ grok was ORIGINALLY declared as `env` kind (`GROK_BASE_URL`/
* `GROK_MODEL`/`XAI_API_KEY`) — that recipe was WRONG, not just unverified:
* live-tested against a real grok binary, it produced "Not signed in",
* because those env vars are not grok's real custom-endpoint mechanism at
* all. The real one is a `[model.<name>]` block in a `config.toml` under
* `GROK_HOME` (verified against xAI's own docs), same shape as codex/pi/
* omp — this is why the confidence table in docs/custom-model-endpoints-plan.md exists:
* "researched" web docs can still be plausible-sounding and wrong.
*
* Every env var name this introduces that can redirect a session's
* traffic MUST also appear in `privilegedEnvKeys` above, exactly like
* `DEEPSEEK_BASE_URL` — a non-granted multi-user owner redirecting a
* session to their own endpoint is a credential-exfiltration path, not
* just a mischief redirect.
*
* `launchModel` is the value the entry's own `model` launch param must carry
* for the CLI to SELECT the injected provider, as a template where
* `{modelId}` is the chosen model id. Writing the config file is not enough
* for pi and omp (`--model custom/<id>`, or the CLI stays on its own default
* provider and reports "No API key found for the selected model") or for
* grok (`--model codeman-custom`, the `[model.<name>]` block the config
* declares). Absent = the config alone selects the model (claude's env vars,
* opencode's blob, codex's top-level `model` key). Applied by the session's
* respawn options through the entry's `legacyConfigField`, never by id.
*/
customModelInjection:
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[]; launchModel?: string }
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string }
| {
kind: 'configDir';
dirEnvVar: string;
fileName: string;
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' | 'grok-toml';
launchModel?: string;
}
| { kind: 'unsupported' };
}
// ---------------------------------------------------------------------------
+65
View File
@@ -0,0 +1,65 @@
/**
* @fileoverview Read/write-array store for user-configured custom OpenAI-compatible
* model endpoints (local or cloud — docs/custom-model-endpoints-plan.md). Same
* shape as `remote-hosts.ts` / `webview-store.ts`: `~/.codeman/custom-model-hosts.json`
* holding a plain array, read/written whole. The file can hold API keys, so it is
* written 0600 via tmp+rename like `intents.json` (`mode` on `writeFile` applies only
* to a file being created; the rename is what keeps an existing file's bytes and
* mode from ever being observable half-written or world-readable).
*/
import { existsSync, mkdirSync } from 'node:fs';
import fs from 'node:fs/promises';
import { join } from 'node:path';
const CUSTOM_MODEL_HOSTS_FILE = 'custom-model-hosts.json';
export type CustomModelAuthStyle = 'bearer' | 'api-key';
export interface CustomModelHost {
id: string;
label: string;
/** Root URL, local or cloud — e.g. "http://192.168.1.50:8080" or an Azure AI Foundry URL. */
baseUrl: string;
apiKey?: string;
/**
* Defaults to 'bearer' (the common `Authorization: Bearer` convention — matches
* llama.cpp, OpenAI-compatible servers, and most gateways). Pick 'api-key' for
* endpoints that specifically want the `api-key` header, e.g. Azure AI Foundry.
*
* ⚠️ There is deliberately NO 'both' option. An earlier design sent BOTH headers
* on every discovery request on the theory that an unused header is harmless —
* live-tested against a real llama-swap server, sending both reliably HUNG the
* request indefinitely (reproduced 3× — Bearer alone: ~500ms, api-key alone:
* ~600ms, both together: no response inside a 15s timeout). Whatever auth
* middleware some servers run apparently does not handle two simultaneous
* credential conventions gracefully, so "send everything and let the server
* ignore what it doesn't need" is not a safe default — it can silently turn a
* working endpoint into one that always times out.
*/
authStyle?: CustomModelAuthStyle;
models?: string[];
lastDiscoveredAt?: string;
}
export function customModelHostsPath(configDir: string): string {
return join(configDir, CUSTOM_MODEL_HOSTS_FILE);
}
export async function readCustomModelHosts(configDir: string): Promise<CustomModelHost[]> {
try {
const raw = await fs.readFile(customModelHostsPath(configDir), 'utf-8');
const parsed = JSON.parse(raw);
return Array.isArray(parsed) ? (parsed as CustomModelHost[]) : [];
} catch {
return [];
}
}
export async function writeCustomModelHosts(configDir: string, hosts: CustomModelHost[]): Promise<void> {
if (!existsSync(configDir)) mkdirSync(configDir, { recursive: true });
const target = customModelHostsPath(configDir);
const tmp = `${target}.${process.pid}.tmp`;
await fs.writeFile(tmp, JSON.stringify(hosts, null, 2), { mode: 0o600 });
await fs.rename(tmp, target);
}
+96
View File
@@ -0,0 +1,96 @@
/**
* @fileoverview The one IO wrapper around `custom-model-injection.ts`'s pure
* `ConfigDirInjection` output — deliberately split out so that file, the
* discovery routes, and `scripts/test-local-llm-harnesses.ts` (via tsx) can
* all share EXACTLY one "write these files, merge this env" implementation.
* Before this existed, the route and the standalone script each carried
* their own copy of this logic, which is exactly the kind of drift the CLI
* registry's "declare once, consume everywhere" design exists to prevent —
* see docs/custom-model-endpoints-plan.md and the "dynamic to support
* cli-registry changes" requirement it was written against.
*/
import { chmodSync, mkdirSync, writeFileSync, rmSync } from 'node:fs';
import { join, dirname } from 'node:path';
import { dataPath } from './config/instance.js';
import type { CliEntry } from './config/cli-registry/types.js';
import {
buildCustomModelInjection,
type ConfigDirInjection,
type CustomModelEndpoint,
} from './custom-model-injection.js';
/** Where a session's isolated `configDir`-kind files live: never the user's real CLI config path. */
export function customModelConfigDir(sessionId: string): string {
return join(dataPath('custom-model-configs'), sessionId);
}
/**
* Writes a `ConfigDirInjection`'s files under `baseDir` and returns the full
* envOverrides object a caller should merge into the session/process env
* (the dir-redirect var plus any `extraEnv` the config file references by
* name). Never touches anything outside `baseDir` — the caller is
* responsible for choosing an isolated directory (never the user's real
* `~/.codex`, `~/.pi`, etc.).
*
* pi and omp embed the API key literally in the file, so the tree is written
* 0700/0600 like every other secret-bearing file under `~/.codeman`; the chmod
* covers a re-apply onto a file that already exists (`mode` only applies at
* creation).
*/
export function applyConfigDirInjection(baseDir: string, injection: ConfigDirInjection): Record<string, string> {
for (const file of injection.files) {
const filePath = join(baseDir, file.relPath);
mkdirSync(dirname(filePath), { recursive: true, mode: 0o700 });
writeFileSync(filePath, file.content, { encoding: 'utf8', mode: 0o600 });
chmodSync(filePath, 0o600);
}
return { [injection.dirEnvVar]: baseDir, ...injection.extraEnv };
}
/** Best-effort recursive removal of a previously-written configDir. Never throws. */
export function removeConfigDir(dir: string | undefined): void {
if (!dir) return;
try {
rmSync(dir, { recursive: true, force: true });
} catch {
// best-effort cleanup only
}
}
/** What applying an endpoint to a session yields, ready for `Session.setCustomModel()`. */
export interface AppliedCustomModel {
envOverrides: Record<string, string>;
envKeys: string[];
configDir?: string;
launchModel?: string;
}
/**
* Compute (and for the `configDir` kind, write) everything a session needs to run
* against `endpoint`/`modelId`. Returns undefined for a CLI with no mechanism.
*
* Idempotent on purpose: the boot-recovery path calls it again for a session that
* was already pointed at an endpoint, so the config files are rewritten in place
* (same content) and the env values, which are never persisted because they carry
* the API key, are re-derived from the endpoint store instead.
*/
export function applyCustomModelInjection(
entry: Pick<CliEntry, 'capabilities'>,
endpoint: CustomModelEndpoint,
modelId: string,
sessionId: string
): AppliedCustomModel | undefined {
const injection = buildCustomModelInjection(entry, endpoint, modelId);
if (injection.kind === 'unsupported') return undefined;
if (injection.kind === 'env') {
return {
envOverrides: injection.envOverrides,
envKeys: Object.keys(injection.envOverrides),
launchModel: injection.launchModel,
};
}
const configDir = customModelConfigDir(sessionId);
const envOverrides = applyConfigDirInjection(configDir, injection);
return { envOverrides, envKeys: Object.keys(envOverrides), configDir, launchModel: injection.launchModel };
}
+252
View File
@@ -0,0 +1,252 @@
/**
* @fileoverview Pure builder for the Custom Model Endpoint Profiles feature
* (docs/custom-model-endpoints-plan.md): turns a CLI registry entry's
* `capabilities.customModelInjection` declaration, a configured endpoint,
* and a chosen model id into the concrete env vars / config-file content
* that would redirect that CLI's session at the endpoint.
*
* No IO here on purpose (mirrors `session-cli-builder.ts`) — a caller
* writes `ConfigDirInjection.files` to disk under an isolated per-session
* directory and points `dirEnvVar` at it; this module only computes what
* those files/env vars should contain.
*
* Confidence: `claude` and `opencode` are verified end-to-end against a real
* llama-swap server (a real "hello world" reply came back). `codex`'s
* config.toml STRUCTURE is now verified (an earlier `[model].default` table
* shape was rejected by a real codex binary with "invalid type: map,
* expected a string" — caught by `scripts/test-local-llm-harnesses.ts`),
* but `wire_api = "responses"` is the only value codex still accepts
* (support for `"chat"` was dropped in Feb 2026), and a plain OpenAI
* Chat-Completions server (llama.cpp, llama-swap, most local setups) does
* NOT implement the Responses API — so codex may still fail at the
* PROTOCOL level even with a correctly-shaped config file. That gap is
* real and current, not a stale warning; see docs/custom-model-endpoints-plan.md. The rest
* (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION flags
* confirmed against real installed binaries' own `--help` output, but
* their custom-endpoint env/config conventions remain web-researched,
* unverified.
*/
import type { CliEntry } from './config/cli-registry/types.js';
export interface CustomModelEndpoint {
id: string;
label: string;
/** Root URL, no trailing slash required — e.g. "http://192.168.1.50:8080" or an Azure AI Foundry URL. */
baseUrl: string;
/** Falls back to a harmless placeholder for endpoints (llama.cpp) that don't check it. */
apiKey?: string;
}
export interface EnvInjection {
kind: 'env';
/** Ready to merge into a session's envOverrides. */
envOverrides: Record<string, string>;
/** See {@link ConfigDirInjection.launchModel}. */
launchModel?: string;
}
export interface ConfigDirInjection {
kind: 'configDir';
/** Env var that must be set to the directory the caller writes `files` under. */
dirEnvVar: string;
files: Array<{ relPath: string; content: string }>;
/**
* Env vars the written config file REFERENCES by name rather than embedding a
* literal value (codex's `env_key = "..."` convention: config.toml never carries
* the API key itself, only the name of an env var codex reads it from). Merge
* these into the session's envOverrides alongside `dirEnvVar` — never skip them,
* or the config points at a credential that was never actually set.
*/
extraEnv?: Record<string, string>;
/**
* The value the CLI's `model` launch param must carry for it to SELECT the injected
* provider (pi/omp: `custom/<modelId>`; grok: the `[model.<name>]` block name). Absent
* when the config alone selects the model. Rendered from the registry entry's
* `customModelInjection.launchModel` template, never hand-built per CLI.
*/
launchModel?: string;
}
export interface UnsupportedInjection {
kind: 'unsupported';
}
export type CustomModelInjectionResult = EnvInjection | ConfigDirInjection | UnsupportedInjection;
const DEFAULT_API_KEY = 'local-dummy-key';
/** Normalizes a base URL to end in exactly one trailing `/v1`, for CLIs whose config expects the OpenAI-style suffix. */
export function withV1Suffix(baseUrl: string): string {
const trimmed = baseUrl.replace(/\/+$/, '');
return /\/v1$/.test(trimmed) ? trimmed : `${trimmed}/v1`;
}
/** JSON-escapes a string for embedding in a TOML/YAML double-quoted scalar — a safe superset of both grammars' basic escapes. */
function quoted(value: string): string {
return JSON.stringify(value);
}
export function buildCustomModelInjection(
entry: Pick<CliEntry, 'capabilities'>,
endpoint: CustomModelEndpoint,
modelId: string
): CustomModelInjectionResult {
const cap = entry.capabilities.customModelInjection;
const apiKey = endpoint.apiKey?.trim() || DEFAULT_API_KEY;
switch (cap.kind) {
case 'env': {
const envOverrides: Record<string, string> = {
[cap.baseUrlVar]: endpoint.baseUrl,
[cap.apiKeyVar]: apiKey,
};
for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId;
return withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId);
}
case 'configContentEnv': {
const content = renderConfigContent(cap.template, endpoint, modelId, apiKey);
return withLaunchModel({ kind: 'env', envOverrides: { [cap.envVar]: content } }, cap.launchModel, modelId);
}
case 'configDir': {
const { content, extraEnv } = renderConfigFile(cap.template, endpoint, modelId, apiKey);
return withLaunchModel(
{ kind: 'configDir', dirEnvVar: cap.dirEnvVar, files: [{ relPath: cap.fileName, content }], extraEnv },
cap.launchModel,
modelId
);
}
case 'unsupported':
return { kind: 'unsupported' };
}
}
/** Render a `launchModel` template (`{modelId}` = the chosen id) onto an injection result. */
function withLaunchModel<T extends EnvInjection | ConfigDirInjection>(
result: T,
template: string | undefined,
modelId: string
): T {
if (!template) return result;
return { ...result, launchModel: template.split('{modelId}').join(modelId) };
}
function renderConfigContent(
template: 'opencode-json',
endpoint: CustomModelEndpoint,
modelId: string,
apiKey: string
): string {
switch (template) {
case 'opencode-json':
return JSON.stringify({
$schema: 'https://opencode.ai/config.json',
provider: {
custom: {
options: { baseURL: withV1Suffix(endpoint.baseUrl), apiKey },
models: { [modelId]: {} },
},
},
model: `custom/${modelId}`,
});
}
}
const CODEX_API_KEY_ENV_VAR = 'CODEMAN_CUSTOM_MODEL_API_KEY';
/** The `[model.<name>]` block name grok's config.toml uses for the injected model — also
* what `-m <name>` in the standalone script's ONE_SHOT argv must reference to select it. */
export const GROK_CUSTOM_MODEL_NAME = 'codeman-custom';
function renderConfigFile(
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' | 'grok-toml',
endpoint: CustomModelEndpoint,
modelId: string,
apiKey: string
): { content: string; extraEnv?: Record<string, string> } {
const baseUrl = withV1Suffix(endpoint.baseUrl);
switch (template) {
case 'codex-toml': {
// Verified against real codex (>= Feb 2026): `model` is a top-level STRING, never
// a `[model].default` table — codex rejects that with "invalid type: map, expected
// a string" (caught by scripts/test-local-llm-harnesses.ts against a real llama-swap
// server). The API key is NEVER a literal TOML field: codex's schema only supports
// `env_key`, the NAME of an env var it reads the credential from at runtime, so the
// actual value must ride along as an extra env var, never embedded in the file.
// ⚠️ `wire_api = "responses"` is the only value codex still accepts (it dropped
// `"chat"` support in Feb 2026) — a plain OpenAI Chat-Completions server (llama.cpp,
// llama-swap, most local setups) does NOT implement the Responses API, so this
// recipe may still fail at the PROTOCOL level even though the file now parses
// correctly. That is a real, currently-unresolved compatibility gap, not a syntax
// bug — track it before calling codex support done.
const content = [
`model = ${quoted(modelId)}`,
`model_provider = "custom"`,
'',
'[model_providers.custom]',
`name = "Custom Endpoint"`,
`base_url = ${quoted(baseUrl)}`,
`env_key = ${quoted(CODEX_API_KEY_ENV_VAR)}`,
`wire_api = "responses"`,
'',
].join('\n');
return { content, extraEnv: { [CODEX_API_KEY_ENV_VAR]: apiKey } };
}
case 'pi-models-json':
// Verified against pi's OWN bundled docs (models.md): `models` is an ARRAY of
// `{id: "..."}` objects, NOT an object keyed by model id — the earlier shape here
// silently loaded zero models ("No models available"), confirmed live. `authHeader:
// true` is required too: pi does not automatically send `Authorization: Bearer
// <apiKey>` just because `apiKey` is set (per the same doc) — without it, a real
// (non-llama.cpp) endpoint that actually checks the key would reject every request.
return {
content: JSON.stringify(
{
providers: {
custom: {
baseUrl,
apiKey,
api: 'openai-completions',
authHeader: true,
models: [{ id: modelId }],
},
},
},
null,
2
),
};
case 'omp-models-yml':
// Mirrors the pi-models-json fix above (omp shares pi's config lineage per
// CLAUDE.md — it reads several of pi's own env vars): a flat list of bare model
// name strings under `models` is UNCONFIRMED against real omp docs (none are
// bundled with the binary) — this now matches pi's `{id: "..."}` object-list
// shape and adds `authHeader: true` on the same reasoning, but has not itself
// been live-tested the way pi's fix was. Verify before raising its confidence.
return {
content: `providers:\n custom:\n baseUrl: ${quoted(baseUrl)}\n apiKey: ${quoted(apiKey)}\n api: openai-completions\n authHeader: true\n models:\n - id: ${quoted(modelId)}\n`,
};
case 'grok-toml': {
// Verified against xAI's own docs (docs.x.ai/build/settings/reference): a
// `[model.<name>]` block, NOT plain env vars — an earlier `env`-kind recipe for
// grok was wrong, not just unverified (see the customModelInjection doc comment
// in cli-registry/types.ts). `api_backend = "chat_completions"` is explicitly
// supported (unlike codex, which dropped it after Feb 2026), so this one CAN
// talk to a plain OpenAI-compatible server directly. `env_key` reuses grok's own
// documented fallback var name (XAI_API_KEY) rather than inventing a new one.
const content = [
`[model.${GROK_CUSTOM_MODEL_NAME}]`,
`model = ${quoted(modelId)}`,
`base_url = ${quoted(baseUrl)}`,
`name = "Custom Endpoint"`,
`env_key = "XAI_API_KEY"`,
`api_backend = "chat_completions"`,
'',
].join('\n');
return { content, extraEnv: { XAI_API_KEY: apiKey } };
}
}
}
+58
View File
@@ -293,6 +293,64 @@ export function toSessionDocker(host: DockerHost, dockerCase: DockerCase): Sessi
return session;
}
/**
* Which existing case, if any, blocks adopting `container` at `containerWorkdir`.
*
* One container may back SEVERAL adopted cases, each pointing at a different
* directory inside it — that is the whole reason to adopt the same container
* twice, and it is safe because the in-container tmux session is named per
* SESSION (`dockerTmuxSessionName`, `codeman-dkr-<id8>`) and not per case, so a
* session teardown kills exactly one session and its siblings on the shared
* in-container tmux server are untouched. Nothing else reaches an adopted
* container's lifecycle either: stop/remove throw at the builder, recreate
* refuses `owned === false`, and the orphan reaper filters on the
* `codeman.managed=1` label that only Codeman-created containers carry.
*
* So the conflicts that remain are NOT about the tmux server:
* - `owned-case` the container backs a case Codeman CREATED, whose lifecycle
* it owns; a recreate or delete there would destroy the
* adopted case's container out from under it.
* - `other-owner` already adopted by a different user. Adoption hands out a
* shell inside someone else's container, so it stays scoped.
* - `duplicate` same container AND same directory: the second case would
* behave identically to the first, so name the first instead
* of silently creating a twin. A DIFFERENT directory is the
* supported case and returns null.
*/
export type AdoptContainerConflict =
| { kind: 'owned-case'; caseName: string }
| { kind: 'other-owner'; caseName: string }
| { kind: 'duplicate'; caseName: string }
| null;
export function classifyAdoptContainerConflict(params: {
container: string;
/** Directory inside the container this adoption targets (already defaulted). */
containerWorkdir: string;
existing: ReadonlyArray<
Pick<DockerCase, 'name' | 'container' | 'containerWorkdir' | 'hostWorkspacePath' | 'owned' | 'owner'>
>;
/** Owner visibility test (canAccessOwned bound to the caller). */
canAccess: (owner?: string) => boolean;
}): AdoptContainerConflict {
const { container, containerWorkdir, existing, canAccess } = params;
const sharing = existing.filter((item) => (item.container ?? dockerContainerName(item.name)) === container);
if (sharing.length === 0) return null;
// `owned` is optional and an ABSENT flag means owned (legacy cases predate the
// field), so this must test `!== false` rather than truthiness.
const owned = sharing.find((item) => item.owned !== false);
if (owned) return { kind: 'owned-case', caseName: owned.name };
const foreign = sharing.find((item) => !canAccess(item.owner));
if (foreign) return { kind: 'other-owner', caseName: foreign.name };
const twin = sharing.find((item) => (item.containerWorkdir ?? item.hostWorkspacePath) === containerWorkdir);
if (twin) return { kind: 'duplicate', caseName: twin.name };
return null;
}
/**
* An ADOPTED container is one the user built and runs themselves. Codeman may
* only exec into it; it must never create, start, stop, restart or remove it.
+21 -7
View File
@@ -13,11 +13,14 @@ import { realpathSync } from 'node:fs';
import { homedir } from 'node:os';
import { join, normalize, sep } from 'node:path';
import { registerExternalAttachment, type AttachmentRegistrationResult } from './attachment-registry.js';
import type { SessionRemote } from './types/session.js';
export interface GeneratedArtifactRegistrationOptions {
sessionId: string;
filePath: string;
sessionWorkingDir: string;
/** Remote (SSH) case: the path lives on the remote host (see attachment-registry). */
remote?: SessionRemote;
}
export async function registerGeneratedArtifactAttachment(
@@ -26,19 +29,30 @@ export async function registerGeneratedArtifactAttachment(
// Decide trust on the symlink-resolved path. If it can't be resolved, fall
// back to the strict force-confined policy (registration will 404 a missing
// file anyway).
let forceWorkspaceConfinement = true;
try {
const resolvedPath = realpathSync(options.filePath);
forceWorkspaceConfinement = !isAllowedGeneratedArtifactPath(resolvedPath, options.sessionWorkingDir);
} catch {
// Keep force confinement.
}
//
// A remote case keeps that strict policy unconditionally: the well-known Codex
// artifact directories are anchored at THIS host's home, which says nothing about
// a remote home, so only a file inside the remote workspace is trusted here.
const resolvedPath = options.remote ? undefined : tryRealpath(options.filePath);
const forceWorkspaceConfinement = !resolvedPath
? true
: !isAllowedGeneratedArtifactPath(resolvedPath, options.sessionWorkingDir);
return registerExternalAttachment(options.sessionId, options.filePath, {
sessionWorkingDir: options.sessionWorkingDir,
forceWorkspaceConfinement,
remote: options.remote,
});
}
/** `realpathSync` without the throw — undefined when the path does not resolve. */
function tryRealpath(path: string): string | undefined {
try {
return realpathSync(path);
} catch {
return undefined;
}
}
/** Well-known Codex generated-artifact directories, anchored at the user's home. */
function codexGeneratedDirs(): string[] {
const home = homedir();
+7
View File
@@ -123,6 +123,13 @@ export interface RespawnPaneOptions {
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (preserved across respawns). */
envOverrides?: Record<string, string>;
/**
* Env vars to REMOVE from the tmux session (`setenv -u`) before `envOverrides` is
* applied. `setenv` persists at the tmux-session level and is inherited by
* `respawn-pane`, so a key that merely disappears from `envOverrides` stays set
* for the relaunched CLI; clearing a custom-model selection has to name it.
*/
unsetEnvKeys?: string[];
/** Claude CLI effort level (preserved across respawns, injected via `--settings`) */
effort?: EffortLevel;
/** Original tmux history-limit retained for config parity; respawn cannot resize the existing pane. */
+278
View File
@@ -0,0 +1,278 @@
/**
* @fileoverview Decide which sessions a host reboot destroyed and may be rebuilt.
*
* A server restart and a host reboot both leave `reconcileSessions()` reporting
* dead sessions, and they need opposite handling. A server restart leaves the
* tmux panes running, so recovery ATTACHES to them. A host reboot takes the tmux
* server down with it, so there is nothing to attach to and the pane has to be
* created again. This module holds the decision half of that second case, kept
* free of tmux and disk access so it can be unit tested without either. Every
* observation it reads is gathered by the caller and passed in.
*
* "Eligible" here means a session the user did not end on purpose. The rule that
* an intentional kill or detach is never auto-revived is enforced at runtime by
* an in-memory guard in `TmuxManager`, and memory does not survive a reboot. The
* durable equivalent is the record `cleanupSession()` leaves behind. An unpinned
* kill deletes the record outright, so it is already absent here. A pinned kill
* goes through `demoteOrRemoveSession()` and lands as `status: 'stopped'`, which
* is the marker this module refuses. Pruning keeps a pinned record WITHOUT
* touching its status, so a pinned session a reboot killed still reads `idle` or
* `busy` and stays eligible.
*
* ⚠️ Ending the AGENT rather than the session is a shape this module CANNOT
* recognise today, and a reboot restores it. `/exit` ends the CLI inside the
* pane, `remain-on-exit` keeps the pane, and the PTY Codeman owns is the
* `tmux attach-session` process, which stays alive throughout — so no exit
* handler runs, no lifecycle `exit` is logged, and the record keeps both its pid
* and `status: 'idle'`. Nothing durable distinguishes it from a session that was
* simply idle when the power went. Ark0N/Codeman#446 covers making Codeman
* notice the dead pane; until a record can say the agent is gone, this pass will
* offer those sessions back, and the user dismisses or closes them.
*
* The `pid` check below is therefore NOT that rule. It refuses a record whose
* attach process was already gone, which is a session that never started or
* whose pane died outright.
*
* @dependencies types (SessionState), config/cli-registry
* @consumedby web/server (plan build at boot), web/routes/reboot-restore-routes
*
* @module reboot-restore
*/
import type { SessionState } from './types.js';
import { getCli } from './config/cli-registry/registry.js';
/** Session statuses a reboot restore may rebuild. `stopped` is the kill marker. */
const RESTORABLE_STATUSES: ReadonlySet<string> = new Set(['idle', 'busy', 'error']);
/** Observations the reboot heuristic reads. Gathered by the caller, never here. */
export interface RebootEvidence {
/** Sessions that still had a live pane during reconciliation. */
livePaneCount: number;
/** Sessions reconciliation just marked dead. */
deadSessionCount: number;
/** `os.uptime()`, in seconds. */
uptimeSeconds: number;
/** Newest `lastActivityAt` across the persisted records, in ms since the epoch. */
newestPersistedActivityAt: number;
/** `Date.now()` when the evidence was gathered, in ms. */
now: number;
}
/**
* Decide whether the machine plausibly rebooted rather than the server restarting.
*
* Two signals have to agree. The socket must hold no panes at all while state
* still lists sessions, which rules out an ordinary server restart. The host
* must also have booted after the newest persisted session activity, which is
* the corroboration `os.uptime()` provides cheaply. A wiped tmux socket on a
* long-uptime host fails the second test, so a user who killed the tmux server
* by hand does not get every session offered back to them.
*
* This heuristic decides whether to ASK, never whether to act. A wrong yes costs
* the user a banner they dismiss, because the restore itself waits for a click.
*
* ⚠️ `os.uptime()` reports the HOST's uptime, which a container shares, and that
* cuts BOTH ways rather than simply switching the feature off in Docker. After a
* genuine host reboot a containerized Codeman sees the host's short uptime, so the
* banner DOES appear and the feature works. What it cannot see is a container-only
* restart: the host uptime is long, the boot test fails, and no banner appears
* although every in-container pane is gone (`docker/server.Dockerfile` installs
* tmux inside the Codeman container, and the self-updater restarts the Compose
* deployment by exiting the container, so that is the case where this would help
* most). Failing quiet is the safe direction, and closing the gap needs a boot
* signal the container owns (PID 1's start time, gated on the existing
* `isRunningInContainer()`) rather than a wider heuristic.
*/
export function looksLikeHostReboot(evidence: RebootEvidence): boolean {
if (evidence.deadSessionCount === 0) return false;
if (evidence.livePaneCount > 0) return false;
if (evidence.newestPersistedActivityAt <= 0) return false;
const bootedAt = evidence.now - evidence.uptimeSeconds * 1000;
return bootedAt > evidence.newestPersistedActivityAt;
}
/**
* Pick the conversation the rebuilt pane should resume.
*
* The chain's tail is the newest conversation the session was holding, which is
* what a compact or a clear leaves behind; `resumeSessionId` covers a session
* that was itself started as a resume, and the session id is the original
* conversation for everything else.
*/
export function resolveResumeConversationId(state: SessionState): string {
const chain = state.claudeSessionChain;
const chainTail = Array.isArray(chain) && chain.length > 0 ? chain[chain.length - 1] : undefined;
return chainTail || state.resumeSessionId || state.id;
}
/**
* Why one session was passed over. Reported for logging and shown to the user.
*
* The first seven are decided before anything is built. `capacity-reached` and
* `rebuild-failed` can only happen once a click is spending the plan, and they
* are the two the banner must not confuse with a missing workspace: one means
* "try again after closing something", the other means the CLI would not start.
*/
export interface RebootRestoreRejection {
sessionId: string;
reason:
| 'no-persisted-record'
| 'intentionally-ended'
| 'not-running'
| 'respawn-blocked'
| 'remote-or-docker'
| 'unsupported-mode'
| 'no-working-dir'
| 'workspace-missing'
| 'workspace-forbidden'
| 'already-live'
| 'capacity-reached'
| 'rebuild-failed';
}
/** One restorable session, as the banner shows it and the rebuild replays it. */
export interface RebootRestoreEntry {
sessionId: string;
name?: string;
workingDir: string;
owner?: string;
mode: string;
/** The conversation the rebuilt pane resumes. */
resumeConversationId: string;
/**
* The persisted record, kept whole so the rebuild can replay what it held.
* Read at boot, before pruning deletes it, and held in memory until the click.
*/
state: SessionState;
}
export interface RebootRestorePlan {
restore: RebootRestoreEntry[];
skipped: RebootRestoreRejection[];
}
/**
* Split the sessions reconciliation just killed into the ones a reboot restore
* may offer and the ones it must leave alone.
*
* @param deadSessionIds Session ids `reconcileSessions()` reported as dead.
* @param persisted The `state.json` session records, which `cleanupStaleSessions()`
* has not pruned yet at the point this runs.
* @param workspaceExists Whether a working directory is still on disk. A tmux
* session can outlive its deleted repo, and rebuilding one there would scaffold
* an empty tree. The caller owns the disk access; the click re-checks, because
* a repo can be deleted between the boot and the click.
*/
export function planRebootRestore(
deadSessionIds: readonly string[],
persisted: Readonly<Record<string, SessionState>>,
workspaceExists: (workingDir: string) => boolean
): RebootRestorePlan {
const restore: RebootRestoreEntry[] = [];
const skipped: RebootRestoreRejection[] = [];
for (const sessionId of deadSessionIds) {
const state = persisted[sessionId];
if (!state) {
// An unpinned kill already deleted the record, so absence IS the guard.
skipped.push({ sessionId, reason: 'no-persisted-record' });
continue;
}
if (!RESTORABLE_STATUSES.has(state.status)) {
// A pinned kill was demoted to `stopped`. Reviving it would undo the kill.
skipped.push({ sessionId, reason: 'intentionally-ended' });
continue;
}
if (state.pid === null || state.pid === undefined) {
// No attach process when the record was last written: the session never
// started, or its pane died outright rather than its agent exiting inside a
// surviving pane. Either way there was nothing running to bring back.
//
// ⚠️ This does NOT catch a session the user ended with `/exit`. See the
// module header: that leaves the pid in place, because the pid is the tmux
// attach process and `remain-on-exit` keeps it alive.
//
// Conservative on purpose. A session that somehow persisted no pid while
// genuinely running is not offered, and its conversation stays reachable
// from the Resume list, which is where every session would be without this
// feature.
skipped.push({ sessionId, reason: 'not-running' });
continue;
}
if (state.respawnBlocked === true) {
// The crash-loop breaker tripped on this pane. Re-creating it restarts the loop.
skipped.push({ sessionId, reason: 'respawn-blocked' });
continue;
}
if (state.remote || state.docker) {
// Both need another host or a container to be up, which a just-booted machine
// cannot promise. The remote reconnect watcher owns the remote case already.
skipped.push({ sessionId, reason: 'remote-or-docker' });
continue;
}
// Capability, not a CLI id: this pass resumes by handing the CLI a conversation
// id through the top-level `resumeSessionId`, which only a CLI whose history the
// claude-jsonl reader understands can consume that way. Others carry their thread
// id in their own `<Mode>Config`, which this pass does not thread through.
if (getCli(state.mode ?? 'claude')?.capabilities.transcript !== 'claude-jsonl') {
skipped.push({ sessionId, reason: 'unsupported-mode' });
continue;
}
if (!state.workingDir) {
skipped.push({ sessionId, reason: 'no-working-dir' });
continue;
}
if (!workspaceExists(state.workingDir)) {
skipped.push({ sessionId, reason: 'workspace-missing' });
continue;
}
restore.push({
sessionId,
name: state.name,
workingDir: state.workingDir,
owner: state.owner,
mode: state.mode ?? 'claude',
resumeConversationId: resolveResumeConversationId(state),
state,
});
}
return { restore, skipped };
}
/**
* Drop the entries whose conversation is already on screen.
*
* Hours can pass between the boot that built the plan and the click that spends
* it, and the Resume list can reach the same conversation in the meantime. Two
* panes running `claude --resume` on one conversation is the failure this
* prevents, so a match on either the session id or the conversation id is enough
* to skip the entry.
*/
export function rejectAlreadyLive(
entries: readonly RebootRestoreEntry[],
liveSessionIds: ReadonlySet<string>,
liveConversationIds: ReadonlySet<string>
): RebootRestorePlan {
const restore: RebootRestoreEntry[] = [];
const skipped: RebootRestoreRejection[] = [];
for (const entry of entries) {
if (liveSessionIds.has(entry.sessionId) || liveConversationIds.has(entry.resumeConversationId)) {
skipped.push({ sessionId: entry.sessionId, reason: 'already-live' });
continue;
}
restore.push(entry);
}
return { restore, skipped };
}
/** Newest `lastActivityAt` across persisted records, or 0 when there are none. */
export function newestPersistedActivity(persisted: Readonly<Record<string, SessionState>>): number {
let newest = 0;
for (const state of Object.values(persisted)) {
const stamp = state.lastActivityAt ?? state.createdAt ?? 0;
if (stamp > newest) newest = stamp;
}
return newest;
}
+398
View File
@@ -0,0 +1,398 @@
/**
* @fileoverview Remote (SSH) file access for remote-SSH cases.
*
* A remote case's `workingDir` is an absolute path on ANOTHER host
* (`Session.workingDir = RemoteCase.remotePath`, see docs/remote-sessions.md). Every
* file route used to read it with local `fs`, which cannot work: the local
* `realpathSync` in `validateSessionFilePath` fails first, so the request died as a
* 404 "File not found" before a byte was read (#415). This module is the ONE place
* that reads remote bytes, mirroring how `remote-hosts.ts` is the one place that
* builds an ssh command line.
*
* Connection options come from `buildSshConnectionArgs()` — never a hand-built ssh
* line (the COD-107 discipline in docs/remote-sessions.md) — so a proxied,
* custom-port or jump-hosted case reaches its files with exactly the credentials the
* launch used, and `BatchMode=yes` means a host that needs a passphrase fails fast
* instead of hanging on a prompt nothing can answer.
*
* ⚠️ The path is the injection surface: it arrives from the browser (`?path=`). It is
* always interpolated as a single `shellescape`d token, and the whole remote command
* is itself shellescaped into the ssh line, so the local shell and the remote shell
* each see one opaque argument. Never build a command here by concatenating a raw
* path into the string.
*
* Read-only by design: previews, text reads and streaming. Writing to a remote file
* is deliberately NOT implemented (docs/file-viewer-edit-plan.md §6), nor are the
* office-conversion/thumbnail paths that would need the bytes on the server's disk.
*/
import { exec, spawn } from 'node:child_process';
import { promisify } from 'node:util';
import { PassThrough, type Readable } from 'node:stream';
import type { SessionRemote } from './types/session.js';
import { buildSshConnectionArgs, remoteSshTarget, shellescape } from './remote-hosts.js';
import { runWithRemoteSshLimit } from './remote-ssh-limiter.js';
const execAsync = promisify(exec);
/**
* Bound on the probe (realpath + stat) round trip. The connect itself is already
* bounded by `buildSshConnectionArgs`'s default `-o ConnectTimeout=10`; this covers
* a host that accepts the TCP connection and then never answers.
*/
const REMOTE_PROBE_TIMEOUT_MS = 20_000;
/** Bound on a buffered remote read (`cat`), on top of the caller's own size cap. */
const REMOTE_READ_TIMEOUT_MS = 30_000;
/** Slack over the caller's byte cap so a file exactly at the limit still fits. */
const READ_BUFFER_SLACK_BYTES = 64 * 1024;
/** Marker a probe prints when the path does not exist on the remote host. */
const NOT_FOUND_MARKER = 'n';
/**
* Marker a probe prints when the path exists but could NOT be canonicalized (no
* `readlink -f`, and the portable fallback hit its hop cap or a `readlink` failure).
* Parsed as `null`, i.e. 404: a path whose real target is unknown must never be
* served, because every containment and blocklist check runs on the resolved path.
*/
const UNRESOLVABLE_MARKER = 'x';
/**
* Paths per ssh round trip. The whole remote script is ONE shellescaped argument,
* and Linux caps a single argv string at 128 KiB, so a 100-entry attachment history
* of long paths is split rather than risking `E2BIG` on the local `sh`.
*/
const REMOTE_PROBE_CHUNK_SIZE = 40;
/** Symlink hops the portable resolver follows before giving up (Linux uses 40). */
const REMOTE_SYMLINK_MAX_HOPS = 40;
/**
* Under vitest no real ssh connection may ever be opened (mirrors
* `checkRemoteTmuxAvailable` and friends in remote-hosts.ts). The route tests mock
* this module, so nothing reaches here today; this is what keeps the NEXT
* remote-session test that touches a file route from opening a connection from CI.
* A clear 502-shaped error, never a fake success: there are no fake bytes to return.
*/
function assertNotUnderTest(): void {
if (process.env.VITEST) {
throw new RemoteFileAccessError('remote file access is disabled under test');
}
}
/** What a remote path turned out to be. `other` = symlink/socket/fifo/device. */
export type RemotePathKind = 'file' | 'directory' | 'other';
export interface RemoteProbe {
/** The path with symlinks resolved on the REMOTE host. */
realPath: string;
kind: RemotePathKind;
/** Size in bytes (0 for anything that is not a regular file). */
size: number;
/** mtime in ms since epoch (0 when the remote `stat` reported none). */
mtimeMs: number;
}
/**
* A remote file access failed for a reason that is NOT "the file is missing" —
* unreachable host, timeout, ssh error, unexpected probe output. Callers map this to
* a 5xx with the remote reason in the message; a missing file is reported separately
* as `null`/404 so the two cannot be confused.
*/
export class RemoteFileAccessError extends Error {
constructor(message: string) {
super(message);
this.name = 'RemoteFileAccessError';
}
}
/**
* Wrap a remote shell command in the shared, shellescaped ssh line.
*
* The single entry point for "run this on the remote host": connection args (port,
* identity, jump host, SOCKS ProxyCommand, extra `-o`) all come from
* `buildSshConnectionArgs`, and the command is ONE shellescaped token, so a path with
* spaces, quotes or `$(…)` cannot escape into the ssh command line.
*/
export function buildRemoteFileCommand(remote: SessionRemote, shellCommand: string): string {
return [...buildSshConnectionArgs(remote), remoteSshTarget(remote), shellescape(shellCommand)].join(' ');
}
/**
* `realpath + stat + existence` for one or more paths, in a SINGLE ssh round trip.
*
* One call instead of three matters: without a shared connection (no ControlMaster)
* every extra `ssh` is a fresh handshake, and the file routes need the path AND the
* workspace root canonicalized to compare them.
*
* Output format: the script first prints a lone NUL, then one NUL-terminated record
* per path, `<index>|n` (missing), `<index>|x` (exists but cannot be canonicalized) or
* `<index>|kind|size|mtime|realPath`. Records are keyed by INDEX and separated by NUL
* rather than newline so that a remote filename containing a newline cannot shift the
* alignment, and the leading NUL is what separates a login banner or an eager rc-file
* `echo` (which land before the script runs) from the records without any "last N
* lines" guesswork. `realPath` is the last field, so a `|` in a path still parses.
*
* Symlink resolution is portable AND fails closed. `readlink -f` where available
* (Linux, macOS >= 12.3); otherwise the fallback canonicalizes the directory chain
* with `cd -P`/`pwd -P` and then follows the LAST component with plain `readlink`
* (which the systems lacking `-f` do have) for a bounded number of hops. A path the
* fallback cannot resolve prints `x`, never the unresolved string: every containment
* and blocklist check downstream runs on `realPath`, and an earlier version of this
* fallback returned the directory-resolved path with the final symlink still in it,
* so `ws/notes.txt -> ~/.ssh/id_rsa` passed containment while `cat` served the key.
*/
export function buildRemoteProbeCommand(paths: readonly string[]): string {
const probes = paths.map((path, index) => `probe ${index} ${shellescape(path)}`).join('\n');
return [
'resolve_last() {',
' q=$1',
' hops=0',
' while :; do',
' d=$(cd -P "$(dirname "$q")" 2>/dev/null && pwd -P) || return 1',
' q=$d/$(basename "$q")',
' [ -L "$q" ] || break',
' hops=$((hops + 1))',
` [ "$hops" -le ${REMOTE_SYMLINK_MAX_HOPS} ] || return 1`,
' l=$(readlink "$q" 2>/dev/null) || return 1',
' [ -n "$l" ] || return 1',
' case $l in /*) q=$l ;; *) q=$d/$l ;; esac',
' done',
' if [ -d "$q" ]; then q=$(cd -P "$q" 2>/dev/null && pwd -P) || return 1; fi',
' printf %s "$q"',
'}',
'probe() {',
' i=$1',
' p=$2',
` if [ ! -e "$p" ]; then printf '%s|${NOT_FOUND_MARKER}\\0' "$i"; return; fi`,
` r=$(readlink -f "$p" 2>/dev/null) || r=$(resolve_last "$p") || { printf '%s|${UNRESOLVABLE_MARKER}\\0' "$i"; return; }`,
` [ -n "$r" ] || { printf '%s|${UNRESOLVABLE_MARKER}\\0' "$i"; return; }`,
' if [ -d "$r" ]; then t=d; elif [ -f "$r" ]; then t=f; else t=o; fi',
' s=0',
' if [ "$t" = f ]; then s=$(stat -c %s "$r" 2>/dev/null || stat -f %z "$r" 2>/dev/null); [ -n "$s" ] || s=0; fi',
' m=$(stat -c %Y "$r" 2>/dev/null || stat -f %m "$r" 2>/dev/null || printf 0)',
` printf '%s|%s|%s|%s|%s\\0' "$i" "$t" "$s" "$m" "$r"`,
'}',
"printf '\\0'",
probes,
].join('\n');
}
/**
* Parse one probe record (index prefix already stripped). `null` for the not-found
* and unresolvable markers or anything malformed.
*/
export function parseRemoteProbeRecord(record: string): RemoteProbe | null {
if (!record || record === NOT_FOUND_MARKER || record === UNRESOLVABLE_MARKER) return null;
const parts = record.split('|');
if (parts.length < 4) return null;
const [kindRaw, sizeRaw, mtimeRaw] = parts;
const kind: RemotePathKind | null =
kindRaw === 'f' ? 'file' : kindRaw === 'd' ? 'directory' : kindRaw === 'o' ? 'other' : null;
if (!kind) return null;
const realPath = parts.slice(3).join('|');
if (!realPath) return null;
const size = Number.parseInt(sizeRaw, 10);
const mtimeSeconds = Number.parseInt(mtimeRaw, 10);
return {
realPath,
kind,
size: Number.isFinite(size) && size > 0 ? size : 0,
mtimeMs: Number.isFinite(mtimeSeconds) && mtimeSeconds > 0 ? mtimeSeconds * 1000 : 0,
};
}
/**
* Parse the output of {@link buildRemoteProbeCommand} into one entry per requested
* path, in order. Throws when a path's record is missing: that means the transport
* or the remote shell did something unexpected, and silently treating it as "not
* found" would turn an infrastructure failure into a wrong 404.
*
* Everything before the first NUL is the remote shell's own chatter (banner, rc-file
* output) and is discarded; records are matched by their index prefix, so neither
* extra output nor a newline inside a filename can shift the mapping.
*/
export function parseRemoteProbeOutput(stdout: string, paths: readonly string[]): Array<RemoteProbe | null> {
const records = stdout.split('\0').slice(1);
const byIndex = new Map<number, string>();
for (const record of records) {
const match = /^(\d+)\|([\s\S]*)$/.exec(record);
if (!match) continue;
const index = Number.parseInt(match[1], 10);
if (!byIndex.has(index)) byIndex.set(index, match[2]);
}
return paths.map((_, index) => {
const record = byIndex.get(index);
if (record === undefined) {
throw new RemoteFileAccessError('remote host returned no usable file information');
}
return parseRemoteProbeRecord(record);
});
}
/**
* Probe one or more remote paths. Entry is `null` for a path that does not exist (or
* could not be canonicalized, which is refused the same way).
*
* Large batches are split into round trips of {@link REMOTE_PROBE_CHUNK_SIZE}, each
* counted against the global ssh limiter, so an attachment history of 100 entries
* costs three connections in sequence rather than 100 at once.
*/
export async function remoteProbePaths(
remote: SessionRemote,
paths: readonly string[]
): Promise<Array<RemoteProbe | null>> {
assertNotUnderTest();
const results: Array<RemoteProbe | null> = [];
for (let offset = 0; offset < paths.length; offset += REMOTE_PROBE_CHUNK_SIZE) {
const chunk = paths.slice(offset, offset + REMOTE_PROBE_CHUNK_SIZE);
const command = buildRemoteFileCommand(remote, buildRemoteProbeCommand(chunk));
let stdout: string;
try {
const result = await runWithRemoteSshLimit(() =>
execAsync(command, { timeout: REMOTE_PROBE_TIMEOUT_MS, maxBuffer: 256 * 1024 })
);
stdout = result.stdout;
} catch (err) {
throw new RemoteFileAccessError(
`remote host ${remote.label || remote.host} unreachable: ${describeExecError(err)}`
);
}
results.push(...parseRemoteProbeOutput(stdout, chunk));
}
return results;
}
/** Read a whole remote file into memory, capped by `maxBytes`. */
export async function remoteReadFile(remote: SessionRemote, remotePath: string, maxBytes: number): Promise<Buffer> {
assertNotUnderTest();
const command = buildRemoteFileCommand(remote, `cat ${shellescape(remotePath)}`);
try {
const result = await runWithRemoteSshLimit(() =>
execAsync(command, {
timeout: REMOTE_READ_TIMEOUT_MS,
maxBuffer: maxBytes + READ_BUFFER_SLACK_BYTES,
encoding: 'buffer',
})
);
return Buffer.isBuffer(result.stdout) ? result.stdout : Buffer.from(result.stdout);
} catch (err) {
throw new RemoteFileAccessError(`failed to read remote file: ${describeExecError(err)}`);
}
}
/**
* Command that writes a remote file's bytes to stdout.
*
* ⚠️ Range reads use `tail -c +N | head -c L` (both POSIX, constant memory) because
* the alternative — `dd bs=1` — issues one read syscall per byte and would make video
* seeking unusable. The trade-off is that a `tail` failure (the file vanished
* mid-request) reports `head`'s exit status, i.e. a short body on an already-sent
* 206; the client retries. The uncompressed path (`cat`) reports its own failure
* correctly, so the streaming error path is still covered by the normal case.
*/
export function buildRemoteReadCommand(remotePath: string, range?: { start: number; end: number }): string {
const quoted = shellescape(remotePath);
if (!range) return `cat ${quoted}`;
const length = range.end - range.start + 1;
return `tail -c +${range.start + 1} ${quoted} | head -c ${length}`;
}
export interface RemoteFileStream {
/** The remote file's bytes, streamed from the ssh child's stdout. */
stream: Readable;
/**
* Abort the transfer and reap the ssh child. The caller MUST call this when the
* HTTP request ends — especially on a client disconnect — or the `ssh` process
* keeps running (and holding a connection open) after nobody is reading it.
*/
close(): void;
}
/**
* Stream a remote file (optionally a byte range) as a Node Readable.
*
* Nothing is buffered in server memory: the bytes go from `ssh`'s stdout straight to
* the HTTP response, which is what makes a multi-GB remote video cost one pipe.
*/
export function remoteCreateReadStream(
remote: SessionRemote,
remotePath: string,
range?: { start: number; end: number }
): RemoteFileStream {
if (process.env.VITEST) {
// Same rule as the buffered calls, in stream form: the consumer sees the error
// through the stream's normal failure path instead of a connection attempt.
const stream = new PassThrough();
process.nextTick(() => stream.destroy(new RemoteFileAccessError('remote file access is disabled under test')));
return { stream, close: () => stream.destroy() };
}
const command = buildRemoteFileCommand(remote, buildRemoteReadCommand(remotePath, range));
const child = spawn(command, { shell: true, stdio: ['ignore', 'pipe', 'pipe'] });
let stderr = '';
child.stderr?.on('data', (chunk: Buffer) => {
if (stderr.length < 2000) stderr += chunk.toString();
});
const stream = child.stdout;
let ended = false;
stream.on('end', () => {
ended = true;
});
stream.on('error', () => {
ended = true;
});
child.on('error', (err: Error) => {
stream.destroy(err);
});
child.on('close', (code: number | null) => {
// Only a truncated transfer is an error. A non-zero exit AFTER the body finished
// (e.g. a signal delivered as the last byte was flushed) must not destroy an
// already-complete response, or the browser reports a broken body for a file it
// received in full.
if (ended || code === 0 || code === null) return;
const detail = stderr.trim().split('\n')[0];
stream.destroy(new RemoteFileAccessError(`remote read failed (ssh exit ${code})${detail ? `: ${detail}` : ''}`));
});
return {
stream,
close(): void {
if (!stream.destroyed) stream.destroy();
child.kill('SIGTERM');
},
};
}
/**
* First useful line of an exec/stderr error, for a user-facing message.
*
* ⚠️ Never Node's `err.message`: for a failed `exec` it is `Command failed: <the whole
* ssh line>`, which carries the identity-file path and the probe script, and this
* string goes out in a 502 body. stderr, the timeout flag and the exit/spawn code are
* everything a user can act on.
*/
function describeExecError(err: unknown): string {
if (typeof err === 'object' && err !== null) {
const record = err as { stderr?: unknown; code?: unknown; killed?: unknown };
const stderr =
typeof record.stderr === 'string' ? record.stderr : Buffer.isBuffer(record.stderr) ? String(record.stderr) : '';
const line = stderr
.split('\n')
.map((entry) => entry.trim())
.find((entry) => entry.length > 0);
if (line) return line.slice(0, 300);
if (record.killed) return 'timed out';
if (typeof record.code === 'number') return `ssh exit ${record.code}`;
if (typeof record.code === 'string') return `ssh could not be started (${record.code})`;
}
return 'unknown error';
}
+7 -1
View File
@@ -147,8 +147,14 @@ export function remoteSshTarget(host: Pick<RemoteHost, 'username' | 'host'>): st
* POSIX single-quote shell-escaping (end-quote, escaped-quote, restart-quote).
* Mirrors the helper in tmux-manager.ts so a value with spaces/metachars stays a
* single shell token. Used here for identity paths and `-o KEY=VALUE` options.
*
* EXPORTED for `remote-files.ts` (#415, remote file access): that module wraps a
* remote shell command in the ssh line built by `buildSshConnectionArgs()`, so it
* needs the same escaping discipline for the remote command itself and for every
* path interpolated into it. A third private copy of this function is exactly how
* two escaping implementations drift apart.
*/
function shellescape(str: string): string {
export function shellescape(str: string): string {
return "'" + str.replace(/'/g, "'\\''") + "'";
}
+89
View File
@@ -0,0 +1,89 @@
/**
* @fileoverview Global concurrency limiter for the short-lived `ssh` children that
* remote-case file access spawns (`src/remote-files.ts`: the realpath+stat probe and
* the buffered text read).
*
* Two paths can fan those out without a human behind each one:
*
* - `GET /api/sessions/:id/attachments` resolves every history entry (up to
* `ATTACHMENT_HISTORY_LIMIT`, 100), and the attachments drawer re-runs it on every
* `attachment:detected` event while it is open, which is exactly when an agent is
* writing files. The route now batches the probes, but a burst of drawers is still
* a burst.
* - A `codeman://attach?path=` magic link in terminal output registers the path
* fire-and-forget, once per distinct link per PTY chunk. In a remote session that
* output is written by a process on the remote host, so a prompt-injected agent can
* print hundreds of links and have the server fork one `ssh` per link, each holding
* a 20s probe timeout.
*
* Without a cap that is the fork-bomb shape `document-conversion-limiter.ts` exists to
* prevent, and it also trips OpenSSH's default `MaxStartups 10:30:100`, which starts
* dropping connections at ten unauthenticated handshakes. This is that limiter for
* ssh: a small fixed pool, FIFO queueing, and a slot handed straight to the next
* waiter on release so the active count can never exceed the cap under interleaved
* async resumption.
*
* Streams (`remoteCreateReadStream`) are deliberately NOT counted: one is opened per
* browser request and held for the life of a media playback, so four open videos
* would otherwise block every preview and the history list. They are already gated
* behind a counted probe (the guard re-probe runs first), so their spawn RATE is
* bounded here even though their concurrency is bounded by the browser.
*
* NOT re-entrant: never acquire from inside a task already holding a slot.
*/
/**
* Max remote probes/reads allowed to run concurrently across the whole process.
* Override with CODEMAN_MAX_REMOTE_FILE_SSH (clamped to >= 1). Four keeps a burst
* well under OpenSSH's ten-handshake default.
*/
const MAX_CONCURRENT_REMOTE_SSH = (() => {
const raw = Number(process.env.CODEMAN_MAX_REMOTE_FILE_SSH);
return Number.isFinite(raw) && raw >= 1 ? Math.floor(raw) : 4;
})();
let active = 0;
const waiters: Array<() => void> = [];
/** Test/diagnostic hook: remote calls currently holding a slot. */
export function getActiveRemoteSshCount(): number {
return active;
}
/** Test/diagnostic hook: remote calls queued behind the cap. */
export function getQueuedRemoteSshCount(): number {
return waiters.length;
}
/** The configured cap, so a test can assert against the real number. */
export function getRemoteSshLimit(): number {
return MAX_CONCURRENT_REMOTE_SSH;
}
function acquire(): Promise<void> {
if (active < MAX_CONCURRENT_REMOTE_SSH) {
active++;
return Promise.resolve();
}
return new Promise<void>((resolve) => waiters.push(resolve));
}
function release(): void {
const next = waiters.shift();
if (next) {
// Hand the slot straight to the next waiter; `active` stays at the cap.
next();
} else {
active--;
}
}
/** Run `task` once an ssh slot is free, releasing the slot afterward. */
export async function runWithRemoteSshLimit<T>(task: () => Promise<T>): Promise<T> {
await acquire();
try {
return await task();
} finally {
release();
}
}
+347
View File
@@ -0,0 +1,347 @@
/**
* @fileoverview Automatic session names from the first prompt.
*
* A new tab is born as `w3-myapp`, which says where it runs and nothing about
* what it is doing. Once the user submits a real prompt the tab can carry a
* title derived from it (`w3-myapp: fix the login redirect`), and this module
* holds the three pure pieces of that: a tracker that reconstructs the composer
* text from the keystrokes Codeman forwards, the title heuristic, and the
* prefix-preserving composition.
*
* Deliberately no LLM: the prompt already passes through the input boundary,
* so a local title is private, deterministic and identical for every CLI.
*
* ⚠️ The tracker sits on the raw keystroke stream, which carries far more than
* the prompt: cursor keys, mouse reports Codeman forwards to the CLI, bracketed
* pastes, Alt chords, the bare Esc that interrupts a turn. Every one of those
* once named a tab something wrong (a lone Esc ate the next prompt's first
* character; a wheel tick mid-word dropped the first half of the prompt), so
* the rules below are explicit per key. The model is a best-effort transcript:
* keys whose effect on the composer is knowable are mirrored, keys that leave
* the text alone are ignored, and keys that replace it with something the
* tracker cannot see (history recall) TAINT the draft so that Enter submits
* nothing rather than a fragment. A prompt that yields no title leaves the
* session eligible for the next one.
*
* Only user-originated input is fed here; the Session decides that. Ralph
* kick-starts, respawn `/clear`s, cron launches and approval answers all go
* through the same write paths and must never become a tab title.
*
* @module session-auto-name
*/
import { MAX_SESSION_NAME_LENGTH } from './config/terminal-limits.js';
/**
* Longest composer draft kept, in code points. The title is cut from the HEAD
* of the prompt, so once the cap is reached further text is counted rather
* than kept (backspaces consume that count first). Keeping the tail instead
* would turn a long paste into a title made of its last line.
*/
const MAX_PROMPT_BUFFER_CODE_POINTS = 8_192;
/** Longest escape sequence collected before the tracker gives up on it. */
const MAX_ESCAPE_SEQUENCE_LENGTH = 64;
/** Longest title, in code points, before it is cut with an ellipsis. */
const MAX_AUTO_NAME_CODE_POINTS = 72;
/**
* A sentence boundary is only honoured this far into the prompt, or "e.g. fix
* this now" becomes "e.g." and "Ok. Fix the bug" becomes "Ok". Short enough
* that a CJK sentence (a dozen code points is a full request) still cuts.
*/
const MIN_SENTENCE_CODE_POINTS = 8;
/** A CSI sequence ends at its first byte in this range. */
const CSI_FINAL_BYTE = /[\x40-\x7e]/;
/** CSI parameter and intermediate bytes; anything else mid-sequence is malformed. */
const CSI_BODY_BYTE = /[\x20-\x3f]/;
/**
* `/clear`, `/model opus`, `/ralph-loop:ralph-loop`: a slash followed by a
* command word and then whitespace or the end. A path (`/home/me/notes.txt
* what is this`) has a second slash where the whitespace should be and so is a
* prompt.
*/
const SLASH_COMMAND_PATTERN = /^\/[a-z][a-z0-9_:-]*(?:\s|$)/i;
// eslint-disable-next-line no-control-regex
const CSI_SEQUENCE_PATTERN = /\x1b\[[\x30-\x3f]*[\x20-\x2f]*[\x40-\x7e]/g;
// eslint-disable-next-line no-control-regex
const CONTROL_CHAR_PATTERN = /[\x00-\x1f\x7f]/g;
const SENTENCE_TERMINATORS = new Set(['.', '!', '?', '。', '!', '?']);
/**
* Reconstructs the composer draft from forwarded keystrokes and reports each
* submitted prompt. Input arrives in arbitrary chunks (one keystroke, a paste,
* an agent's whole prompt plus Enter), so all state lives across calls.
*/
export class SubmittedPromptTracker {
private buffer = '';
private bufferCodePoints = 0;
/** Code points typed past the cap; backspaces eat these before real text. */
private overflow = 0;
/** Escape sequence in progress; a lone ESC means "just saw ESC". */
private sequence = '';
private inPaste = false;
/** The composer holds text the tracker never saw (history recall); Enter submits nothing. */
private tainted = false;
feed(data: string): string[] {
const submitted: string[] = [];
for (const ch of data) {
if (this.sequence) {
this.continueSequence(ch);
continue;
}
if (ch === '\x1b') {
this.sequence = ch;
continue;
}
this.handleKey(ch, submitted);
}
// A chunk that ENDS in a lone ESC is the Esc key, not the start of a
// sequence: xterm hands each key's whole sequence to one write, and the
// programmatic senders (an approval deny sends exactly `\x1b`) send it
// alone. Leaving it pending would make the next prompt's first character
// look like an Alt chord and swallow it.
if (this.sequence === '\x1b') this.sequence = '';
return submitted;
}
private continueSequence(ch: string): void {
if (this.sequence === '\x1b') {
if (ch === '[' || ch === 'O' || ch === ']' || ch === 'P') {
this.sequence += ch;
return;
}
this.sequence = ch === '\x1b' ? ch : '';
// Alt+Enter inserts a newline in the composer; every other Alt chord
// (word movement, Alt+B/F) leaves the text alone.
if (ch === '\r' || ch === '\n') this.appendSeparator();
return;
}
this.sequence += ch;
if (this.sequence.length > MAX_ESCAPE_SEQUENCE_LENGTH) {
// Not a sequence any terminal sends; what follows is unknowable, so the
// draft is tainted rather than titled after the tail of the garbage.
this.sequence = '';
this.tainted = true;
return;
}
const kind = this.sequence[1];
if (kind === '[') {
if (CSI_FINAL_BYTE.test(ch)) {
const sequence = this.sequence;
this.sequence = '';
this.handleCsi(sequence);
} else if (!CSI_BODY_BYTE.test(ch)) {
// Malformed (an ESC [ followed by text): drop the sequence and let the
// character count as typed rather than swallowing up to 64 of them.
this.sequence = '';
this.handleKeyOrEscape(ch);
}
return;
}
if (kind === 'O') {
// SS3 carries exactly one byte (application-mode cursor keys).
this.sequence = '';
if (ch === 'A' || ch === 'B') this.tainted = true;
return;
}
// OSC / DCS run to BEL or ST (ESC \).
if (ch === '\x07' || this.sequence.endsWith('\x1b\\')) this.sequence = '';
}
private handleKeyOrEscape(ch: string): void {
if (ch === '\x1b') {
this.sequence = ch;
return;
}
// Only reached mid-chunk from a malformed sequence, where no submission can
// be reported; a stray Enter there resets the draft like any other Enter.
this.handleKey(ch, []);
}
private handleCsi(sequence: string): void {
if (sequence === '\x1b[200~') {
this.inPaste = true;
return;
}
if (sequence === '\x1b[201~') {
this.inPaste = false;
return;
}
const final = sequence[sequence.length - 1];
// Up/Down (with or without modifiers) recall history: the composer now
// holds a line this tracker never saw. Everything else leaves the text as
// it is: Left/Right/Home/End, Delete (`3~`), Shift+Tab (`Z`), SGR mouse
// reports (`<…M`/`m`, forwarded on every wheel tick), focus reports.
if (final === 'A' || final === 'B') this.tainted = true;
}
private handleKey(ch: string, submitted: string[]): void {
const codePoint = ch.codePointAt(0) ?? 0;
if (this.inPaste) {
// Pasted newlines are newlines IN the composer, never Enter; they and
// the other controls (tabs) become a single separator.
if (codePoint < 0x20 || codePoint === 0x7f) this.appendSeparator();
else this.append(ch);
return;
}
switch (ch) {
case '\r': {
const prompt = this.tainted ? '' : this.buffer.trim();
if (prompt) submitted.push(prompt);
this.reset();
return;
}
case '\n':
// Ctrl+J, and the line feed the send-key route injects for Shift+Enter:
// a newline inside the composer, so the lines join with a separator.
this.appendSeparator();
return;
case '\x7f':
case '\x08':
this.backspace();
return;
case '\x17': // Ctrl+W: word rubout
this.killWord();
return;
case '\x15': // Ctrl+U: line discard
case '\x03': // Ctrl+C: clears the composer (or, empty, arms an exit)
this.reset();
return;
case '\x10': // Ctrl+P
case '\x0e': // Ctrl+N
case '\x12': // Ctrl+R: history search
case '\x1f': // Ctrl+_: undo
this.tainted = true;
return;
default:
// Tab (the @-mention completer, which only ever extends the token),
// cursor chords (Ctrl+A/E/B/F) and the rest of C0 leave the text alone.
if (codePoint < 0x20 || codePoint === 0x7f) return;
this.append(ch);
}
}
/** One space between lines, never a run of them, and none at the start. */
private appendSeparator(): void {
if (this.overflow > 0) return;
if (!this.buffer || /\s$/.test(this.buffer)) return;
this.append(' ');
}
private append(ch: string): void {
if (this.bufferCodePoints >= MAX_PROMPT_BUFFER_CODE_POINTS) {
this.overflow += 1;
return;
}
this.buffer += ch;
this.bufferCodePoints += 1;
}
private backspace(): void {
if (this.overflow > 0) {
this.overflow -= 1;
return;
}
if (!this.buffer) return;
const last = this.buffer.charCodeAt(this.buffer.length - 1);
const units = last >= 0xdc00 && last <= 0xdfff && this.buffer.length >= 2 ? 2 : 1;
this.buffer = this.buffer.slice(0, -units);
this.bufferCodePoints -= 1;
}
private killWord(): void {
this.overflow = 0;
this.buffer = this.buffer.replace(/\S+\s*$/u, '');
this.bufferCodePoints = Array.from(this.buffer).length;
}
private reset(): void {
this.buffer = '';
this.bufferCodePoints = 0;
this.overflow = 0;
this.tainted = false;
}
}
/**
* Turns a submitted prompt into a title, or null when the prompt is not a task:
* empty, a slash command (`/clear`, `/model`), or a `!` shell escape.
*/
export function deriveAutoSessionName(prompt: string): string | null {
const text = prompt.replace(CSI_SEQUENCE_PATTERN, '').replace(CONTROL_CHAR_PATTERN, ' ').replace(/\s+/g, ' ').trim();
if (!text || text.startsWith('!') || SLASH_COMMAND_PATTERN.test(text)) return null;
return truncateCodePoints(firstSentence(text), MAX_AUTO_NAME_CODE_POINTS);
}
/**
* The first sentence, provided it is long enough to be one; a trailing full
* stop is dropped because a tab title is not a sentence.
*/
function firstSentence(text: string): string {
const codePoints = Array.from(text);
for (let i = MIN_SENTENCE_CODE_POINTS - 1; i < codePoints.length; i++) {
if (!SENTENCE_TERMINATORS.has(codePoints[i])) continue;
const next = codePoints[i + 1];
if (next !== undefined && !/\s/.test(next)) continue;
return codePoints
.slice(0, i + 1)
.join('')
.replace(/[.。]+$/, '');
}
return text.replace(/[.。]+$/, '');
}
/** Cuts to `max` code points with an ellipsis, on a word boundary when one is near the end. */
function truncateCodePoints(text: string, max: number): string {
const codePoints = Array.from(text);
if (codePoints.length <= max) return text;
let cut = codePoints.slice(0, max - 1).join('');
const lastSpace = cut.lastIndexOf(' ');
if (lastSpace >= Math.floor(cut.length / 2)) cut = cut.slice(0, lastSpace);
return `${cut.trimEnd()}…`;
}
/**
* The name a placeholder becomes: `<prefix>: <title>`, so the tab keeps its
* case identity and its `w<n>` counter (the tab strip already renders that
* form as the title alone, prefix in the tooltip, and the next-session counter
* still matches it). A session with no name at all just takes the title. The
* result honours `maxLength` in UTF-16 units, the unit the rename route caps.
*/
export function composeAutoSessionName(
currentName: string,
title: string,
maxLength = MAX_SESSION_NAME_LENGTH
): string {
const prefix = currentName.trim();
if (!prefix) return fitTitle(title, maxLength);
const room = maxLength - prefix.length - 2;
if (room <= 0) return prefix;
return `${prefix}: ${fitTitle(title, room)}`;
}
/** Fits a title into `maxUnits` UTF-16 units, ellipsis included. */
function fitTitle(title: string, maxUnits: number): string {
if (title.length <= maxUnits) return title;
let units = 0;
let keep = 0;
for (const codePoint of Array.from(title)) {
if (units + codePoint.length > maxUnits - 1) break;
units += codePoint.length;
keep += 1;
}
return truncateCodePoints(title, keep + 1);
}
/** Codeman's own `w<n>-<case>` / `s<n>-<case>` placeholders, the only names auto-naming replaces. */
export function isGeneratedSessionName(name: string): boolean {
return /^[ws]\d+-[a-zA-Z0-9_-]+$/.test(name);
}
+97
View File
@@ -0,0 +1,97 @@
/**
* @fileoverview The env-var half of the multi-user privilege clamp.
*
* A session's `envOverrides` can hand back privilege that the per-CLI config
* clamp removed, so a non-granted owner's overrides get the privileged keys
* stripped before the session is built. The create and resume routes are what
* this bites on: they clamp what a request asked for.
*
* The reboot-restore route calls it as defence in depth, and today it can strip
* nothing. `Session.getEnvOverridesForPersist()` keeps only `CLAUDE_CODE_*` and
* `CLAUDE_CONFIG_DIR` out of a session's overrides, claude's `privilegedEnvKeys`
* are the five `ANTHROPIC_*` names, and that pass admits claude alone — so a
* persisted record cannot carry a clamped key. The call is there for the day the
* persisted set widens. The grant re-resolution that does bite on that path is
* `resolveClaudeModeForUsername`, which recomputes the permission mode.
*
* This lives outside `web/routes` on purpose. The question it answers is about
* session privilege rather than about HTTP, and `cron/cron-service.ts` sets the
* precedent by importing `canUsernameRunPrivilegedCommands` from `user-store.ts`
* directly and re-resolving the owner's grant when a job fires. Every caller here
* re-resolves the grant at the moment it builds a session, for the same reason.
*
* @dependencies user-store (canUsernameRunPrivilegedCommands), config/cli-registry
* @consumedby web/routes/session-routes, web/routes/reboot-restore-routes
*
* @module session-env-clamp
*/
import { canUsernameRunPrivilegedCommands } from './user-store.js';
import { enabledClis } from './config/cli-registry/registry.js';
/**
* Env-var keys a non-granted owner must not be able to set, because each one
* hands back privilege `clampExternalCliBypassForOwner()` just removed, or redirects a
* credential-resolution endpoint.
*
* The DeepSeek three are reachable because `DSH_*` and `DEEPSEEK_*` are
* allowlisted `envOverrides` prefixes (schemas.ts) — which they have to be, since
* that is also how a user configures the harness's non-privileged knobs.
*
* - `DSH_PERMISSION_MODE` IS the harness's permission switch. Every other CLI's
* bypass is a command-line FLAG, reachable only through the per-CLI config the
* clamp already owns; this one is an env var, so the config clamp alone is
* half a gate.
* - `DSH_HOME` points the launcher at a profile tree, and a profile's plugin code
* executes at BOOT, before any approval row can apply. A user who can write a
* workspace can put a profile in it, so this is the wider of the two.
* - `DEEPSEEK_BASE_URL` aims the provider endpoint, and `_configureCliEnv()`
* forwards the SERVER's own `DEEPSEEK_API_KEY` into every dsh pane before
* `applyEnvOverrides()` runs — so a non-granted owner who could set the base
* URL would have the operator's API key sent as a bearer credential to a host
* of their choosing. (`DEEPSEEK_API_KEY` itself stays overridable: supplying
* your OWN key removes privilege rather than granting it.)
* - `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN` are where omp resolves
* credentials from — the same shape as `DEEPSEEK_BASE_URL` above, reachable
* because `OMP_*` is an allowlisted prefix. Unlike DeepSeek, Codeman does not
* forward any operator-held key into an omp pane today (omp's provider
* credentials live in `~/.omp` config files, not env vars), so there is no
* known concrete exfiltration path yet — clamped defensively anyway, since a
* non-granted owner redirecting where a shared multi-tenant deployment
* resolves auth from is not something to allow silently (found in
* Ark0N/Codeman#353 review; omp's own knobs are otherwise mostly `PI_*`,
* already allowlisted for pi and not addressed here — see resolveOmpHome()).
*/
export function ownerClampedEnvKeys(): string[] {
return enabledClis().flatMap((entry) => entry.capabilities.privilegedEnvKeys);
}
/**
* Env-var half of the multi-user bypass clamp.
*
* `clampExternalCliBypassForOwner()` in `web/routes/session-routes.ts` clamps the
* per-CLI CONFIG, and for every CLI
* but DeepSeek that is the whole story. Here it is not: `applyEnvOverrides()` runs
* AFTER `_configureCliEnv()` in tmux-manager, so an override sent on the SAME
* request lands last and wins, and a non-granted owner could restore
* `danger-full-access` on the very request the config clamp downgraded.
*
* Keys are DROPPED rather than rewritten: dropping falls through to what
* `_configureCliEnv()` exports, which is the clamped config and the server's own
* `DSH_HOME`, i.e. exactly the intended state. No-op in single-user mode and for a
* granted owner, like every other clamp here
* (`canUsernameRunPrivilegedCommands()` returns true when `!isMultiUserMode()`),
* and it returns the caller's own object untouched when there is nothing to strip.
*/
export async function clampEnvOverridesForOwner(
owner: string | undefined,
envOverrides: Record<string, string> | undefined
): Promise<Record<string, string> | undefined> {
if (!envOverrides) return envOverrides;
const keys = ownerClampedEnvKeys();
if (!keys.some((key) => key in envOverrides)) return envOverrides;
if (await canUsernameRunPrivilegedCommands(owner)) return envOverrides;
const clamped = { ...envOverrides };
for (const key of keys) delete clamped[key];
return clamped;
}
+8 -1
View File
@@ -254,7 +254,14 @@ export class SessionManager extends EventEmitter {
// future reader of state.json.
const state = session.toState();
const envOverrides = session.getEnvOverridesForPersist();
const toStore = envOverrides ? { ...state, __envOverrides: envOverrides } : state;
// __customModel: same convention, the disk-only bookkeeping of a custom-model
// selection (env KEYS, config dir, launch model; never the injected values).
const customModel = session.getCustomModelForPersist();
const toStore = {
...state,
...(envOverrides ? { __envOverrides: envOverrides } : {}),
...(customModel ? { __customModel: customModel } : {}),
};
this.store.setSession(session.id, toStore as SessionState);
}
+133
View File
@@ -0,0 +1,133 @@
/**
* @fileoverview Verify that a programmatically sent prompt actually LEFT the composer,
* and press Enter again while it has not.
*
* Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment its
* composer paints but ignores Enter for the first 30 to 50 seconds after it, so the
* `send-keys -l <text>` + `send-keys Enter` pair `TmuxManager.sendInput()` sends 50 ms
* apart leaves the prompt sitting on the composer with `0 tokens`, and every caller
* that then waits for the turn (send-and-wait, the agent skill, the maintainer bot,
* cron, Ralph) burns its whole timeout on a turn that never started. Measured through
* the input route on 2026-09-19: an Enter at 28 s stranded, one at 51 s submitted.
*
* The rule: after a write that carried a carriage return, read the pane on a short
* schedule; while the LAST composer line (the CLI's own prompt glyph) still holds the
* head of what was sent, send Enter again. An empty composer ends it, and so does a
* composer holding anything else, because that text is the user's or the CLI's, never
* ours. A pane with no composer line at all (a shell, a CLI whose glyph is not
* declared, a direct-PTY session with no pane to read) does nothing: this runs for
* EVERY programmatic sender, so a blind Enter here could confirm a dialog nobody asked
* about. The composer is the last glyph line on purpose: Claude Code echoes a submitted
* prompt with the same glyph higher up in the transcript, so only the last one says
* whether the text was taken.
*
* Pure apart from the injected capture, send and log, so the schedule, the cap and
* every stop condition are unit-tested with fake timers (test/session-submit-verifier.test.ts).
*/
import { stripAnsi } from './utils/index.js';
/**
* When to look, counted from the write: 2 s catches the common case (taken) with one
* capture, and the tail reaches 60 s, past twice the longest window measured. Enter is
* re-sent at every check that still finds the prompt, so a 50 s window costs about
* seven Enters and one capture each; a taken prompt costs one capture.
*/
export const SUBMIT_VERIFY_DELAYS_MS: readonly number[] = [
2_000, 3_000, 5_000, 5_000, 5_000, 10_000, 10_000, 10_000, 10_000,
];
/** How many leading characters of the prompt have to match, whitespace removed. */
const PROMPT_HEAD_CHARS = 24;
const compact = (s: string): string => s.replace(/\s+/g, '');
/**
* Whether `prompt` is still sitting unsubmitted in the composer of `screen`.
*
* - `true`: the last `glyph` line holds the prompt's head.
* - `false`: the composer is empty (the prompt was taken) or holds other text.
* - `undefined`: no composer line at all; nothing can be said, so nothing is sent.
*
* Whitespace is removed on both sides before comparing, because the composer wraps a
* long prompt onto indented continuation lines and Claude Code draws a no-break space
* after the glyph; `\s` covers that one in JavaScript.
*/
export function promptStillInComposer(screen: string, prompt: string, glyph: string): boolean | undefined {
if (!glyph) return undefined;
const composerLines = stripAnsi(screen)
.split('\n')
.map((l) => l.trim())
.filter((l) => l.startsWith(glyph));
if (composerLines.length === 0) return undefined;
const composer = compact(composerLines[composerLines.length - 1].slice(glyph.length));
if (!composer) return false;
const head = compact(prompt).slice(0, PROMPT_HEAD_CHARS);
return head.length > 0 && composer.startsWith(head);
}
export interface SubmitVerifierDeps {
/** The rendered pane, or null when there is none to read. */
capture: () => string | null | undefined;
/** Press Enter once. Failures are swallowed; the next check decides again. */
sendEnter: () => Promise<unknown> | unknown;
/** The CLI's composer glyph, resolved at check time (the registry can change). */
glyph: () => string;
log?: (message: string) => void;
/** Test seam; production uses SUBMIT_VERIFY_DELAYS_MS. */
delaysMs?: readonly number[];
}
/**
* One per session. `arm(text)` starts the schedule for the prompt just sent and
* cancels any earlier one: a newer write owns the composer now, and re-sending Enter
* for an older prompt could submit the newer one early. `cancel()` is for teardown.
*/
export class SubmitVerifier {
private timer: NodeJS.Timeout | null = null;
private generation = 0;
constructor(private readonly deps: SubmitVerifierDeps) {}
arm(text: string): void {
this.cancel();
const gen = this.generation;
const delays = this.deps.delaysMs ?? SUBMIT_VERIFY_DELAYS_MS;
let step = 0;
let elapsed = 0;
let resent = 0;
const schedule = (): void => {
if (step >= delays.length) return;
const delay = delays[step++];
elapsed += delay;
this.timer = setTimeout(() => void check(), delay);
this.timer.unref?.();
};
const check = async (): Promise<void> => {
this.timer = null;
if (gen !== this.generation) return;
const screen = this.deps.capture();
if (promptStillInComposer(screen ?? '', text, this.deps.glyph()) !== true) return;
resent++;
this.deps.log?.(
`prompt still in the composer after ${Math.round(elapsed / 1000)}s, re-sending Enter (${resent}/${delays.length})`
);
try {
await this.deps.sendEnter();
} catch {
// The next check re-reads the screen and decides again.
}
if (gen !== this.generation) return;
schedule();
};
schedule();
}
cancel(): void {
this.generation++;
if (this.timer) {
clearTimeout(this.timer);
this.timer = null;
}
}
}
+298 -9
View File
@@ -23,7 +23,7 @@
* ralph-tracker (todo/completion parsing), bash-tool-parser (tool invocation tracking),
* task-tracker (background tasks), mux-interface (tmux abstraction)
* @consumedby session-manager, web/server, respawn-controller
* @emits session:terminal, session:idle, session:working, session:completion, session:exit
* @emits session:terminal, session:idle, session:working, session:completion, session:promptSubmitted, session:exit
*
* @module session
*/
@@ -48,6 +48,8 @@ import {
type OpenCodeConfig,
type CodexConfig,
type EffortLevel,
type CustomModelBookkeeping,
type CustomModelSelection,
type GeminiConfig,
type AntigravityConfig,
type PiConfig,
@@ -56,6 +58,8 @@ import {
type OmpConfig,
type SessionRemote,
type SessionDocker,
type SessionNameSource,
type SessionWriteOptions,
} from './types.js';
import { resolveAndClaimOmpSessionId } from './utils/omp-session-resolver.js';
import { probeDockerCliVersion } from './docker-hosts.js';
@@ -106,6 +110,7 @@ import {
import { DEFAULT_TMUX_HISTORY_LIMIT } from './config/terminal-history.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import { getCli } from './config/cli-registry/registry.js';
import { SubmitVerifier } from './session-submit-verifier.js';
import { compileVersionRegex } from './config/cli-registry/patterns.js';
import { resolveSessionCliVersion } from './utils/cli-resolver.js';
import {
@@ -119,6 +124,7 @@ import { SessionAutoOps } from './session-auto-ops.js';
import { detectUsageLimitPause } from './usage-limit-patterns.js';
import { SessionTaskCache } from './session-task-cache.js';
import { InteractivePtyExitBreaker } from './session-pty-exit-breaker.js';
import { isGeneratedSessionName, SubmittedPromptTracker } from './session-auto-name.js';
import { parseTerminalAttachmentRequests } from './attachment-magic.js';
import {
sanitizeAttachmentHistory,
@@ -421,6 +427,15 @@ export class Session extends EventEmitter {
private _taskCache = new SessionTaskCache();
private _name: string;
private _nameSource: SessionNameSource;
/**
* Reconstructs the composer draft from USER keystrokes so the first real
* prompt can name the tab. Fed only when a write says `fromUser`, and never
* for a CLI whose Enter runs a command rather than submitting a prompt
* (`startMode: 'shell'`), so a shell tab is not renamed after every `ls`.
*/
private readonly _submittedPromptTracker = new SubmittedPromptTracker();
private readonly _acceptsPrompts: boolean;
private ptyProcess: pty.IPty | null = null;
private _pid: number | null = null;
private _status: SessionStatus = 'idle';
@@ -488,6 +503,8 @@ export class Session extends EventEmitter {
private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog
private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read
private _trustDialogTimer: NodeJS.Timeout | null = null; // Re-read after a keystroke (see below)
/** Re-sends Enter while a programmatic prompt still sits in the composer (session-submit-verifier.ts). */
private _submitVerifier: SubmitVerifier | null = null;
private _interactiveStartedAt = 0; // When the interactive pane launched (bounds that scan)
private _taskTracker: TaskTracker;
@@ -577,6 +594,20 @@ export class Session extends EventEmitter {
// the CLAUDE_CODE_EFFORT_LEVEL env var, which would hard-lock the session.
private _effort: EffortLevel | undefined;
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md). `envKeys`,
// `configDir` and `launchModel` are internal bookkeeping ONLY (never surfaced via
// toState()/the customModel getter): they are what setCustomModel() needs to undo a
// previous injection (remove exactly the env keys it added, delete a previous isolated
// config dir) without guessing what it once wrote. Persisted disk-only (`__customModel`).
private _customModel: CustomModelBookkeeping | undefined;
// Env keys a retired custom-model selection injected that the NEXT respawn must
// `tmux setenv -u`. Deleting a key from `_envOverrides` alone does nothing to the
// tmux session, which keeps every `setenv` and hands it to `respawn-pane`, so the
// relaunched CLI would come back still pointed at the old endpoint (measured, see
// TmuxManager.applyEnvOverrides). Drained after a successful respawn.
private _pendingEnvUnsets = new Set<string>();
// tmux history-limit (scrollback lines) allocated when this session's pane is created.
private readonly _tmuxHistoryLimit: number;
@@ -638,6 +669,12 @@ export class Session extends EventEmitter {
workingDir: string;
mode?: SessionMode;
name?: string;
/**
* Who owns the name (see `SessionNameSource`). Omitted, it is inferred
* from the name: Codeman's own `w<n>-<case>` placeholders (or no name)
* stay eligible for auto-naming, anything else counts as the user's.
*/
nameSource?: SessionNameSource;
/** Terminal multiplexer instance (tmux) */
mux?: TerminalMultiplexer;
/** Whether to use multiplexer wrapping */
@@ -707,6 +744,9 @@ export class Session extends EventEmitter {
this.createdAt = config.createdAt || Date.now();
this.mode = config.mode || 'claude';
this._name = config.name || '';
this._nameSource =
config.nameSource ?? (!this._name || isGeneratedSessionName(this._name) ? 'placeholder' : 'manual');
this._acceptsPrompts = getCli(this.mode)?.capabilities.startMode !== 'shell';
this._resumeSessionId = config.resumeSessionId;
// NOW, not `createdAt`: recovery passes the ORIGINAL creation time of a
// days-old tmux session, and seeding last-activity from it would report a
@@ -1238,6 +1278,65 @@ export class Session extends EventEmitter {
}
}
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — public-safe
// subset only (never envKeys/configDir/launchModel, the bookkeeping for setCustomModel).
get customModel(): CustomModelSelection | undefined {
if (!this._customModel) return undefined;
const { endpointId, modelId, label } = this._customModel;
return { endpointId, modelId, label };
}
/**
* The full selection incl. bookkeeping, for state.json ONLY (`__customModel`, the
* same disk-only convention as `getEnvOverridesForPersist()`). Carries no env values,
* so nothing secret lands on disk; recovery re-derives them from the endpoint store.
* Without this a Codeman restart left the pane on the custom endpoint (tmux keeps
* its `setenv`s) while `customModel` came back undefined, so the state was wrong and
* clearing had nothing to unset. Must NOT be included in any API-bound serializer.
*/
getCustomModelForPersist(): CustomModelBookkeeping | undefined {
return this._customModel ? { ...this._customModel, envKeys: [...this._customModel.envKeys] } : undefined;
}
/**
* Update this session's custom-model selection and merge the endpoint's injected env
* vars into `_envOverrides` — first UNDOING whatever the previous selection injected
* (removing exactly those env keys), so switching endpoints, or clearing back to the
* harness's native cloud default, never leaves a stale key behind. Synchronous and
* side-effect-free beyond mutating state, matching `setNice`/`setColor` above — this
* class does no file IO, so it reports the PREVIOUS `configDir` (if any) for the
* caller to clean up on disk (custom-model-injection.ts's configDir kind).
*
* Keys the previous selection injected that the new one does not re-set are queued
* for `tmux setenv -u` on the next respawn (`_pendingEnvUnsets`, threaded through
* `_buildRespawnPaneOptions().unsetEnvKeys`): the tmux session inherits every
* `setenv` into `respawn-pane`, so dropping them from the map alone would relaunch
* the CLI still pointed at the old endpoint — and for the `configDir` kinds, at a
* `HOME`/`CODEX_HOME`/`GROK_HOME` the caller has just deleted.
*/
setCustomModel(
next: CustomModelBookkeeping | undefined,
envOverrides?: Record<string, string>
): { removedEnvKeys: string[]; previousConfigDir: string | undefined } {
const previousConfigDir = this._customModel?.configDir;
const removedEnvKeys: string[] = [];
if (this._customModel) {
for (const key of this._customModel.envKeys) {
if (this._envOverrides) delete this._envOverrides[key];
removedEnvKeys.push(key);
this._pendingEnvUnsets.add(key);
}
}
this._customModel = next ? { ...next, envKeys: [...next.envKeys] } : undefined;
if (envOverrides && Object.keys(envOverrides).length > 0) {
this._envOverrides = { ...(this._envOverrides ?? {}), ...envOverrides };
// A key the new selection sets again does not need an unset (applyEnvOverrides
// would set it right back anyway); keep the list to what actually goes away.
for (const key of Object.keys(envOverrides)) this._pendingEnvUnsets.delete(key);
}
return { removedEnvKeys, previousConfigDir };
}
// Token tracking getters and setters
get totalTokens(): number {
return this._totalInputTokens + this._totalOutputTokens;
@@ -1295,8 +1394,31 @@ export class Session extends EventEmitter {
return this._name;
}
/** An explicit rename: the name is the user's from here on and auto-naming never touches it. */
set name(value: string) {
this._name = value;
this._nameSource = 'manual';
}
/**
* Names the tab after its first prompt. Only a placeholder is eligible, and
* the session stops being one whether or not the string changed: "first
* prompt" means the first, not "every prompt until a rename". Returns
* whether the name changed, so the caller knows whether to persist and
* broadcast.
*/
applyAutoName(value: string): boolean {
if (this._nameSource !== 'placeholder') return false;
const name = value.trim();
if (!name) return false;
this._nameSource = 'auto';
if (this._name === name) return false;
this._name = name;
return true;
}
get nameSource(): SessionNameSource {
return this._nameSource;
}
setAutoClear(enabled: boolean, threshold?: number): void {
@@ -1380,6 +1502,19 @@ export class Session extends EventEmitter {
this._pinnedAt = pinned ? Date.now() : null;
}
/**
* Restore a pin from a persisted record, keeping the moment it was pinned.
*
* `setPinned()` stamps `pinnedAt` with now, which is right for a user pinning a
* session and wrong for a restore: the session-manager orders its pinned group
* by that stamp, so a restored session would jump to the front of a list it had
* been sitting further down.
*/
restorePin(pinned: boolean, pinnedAt?: number): void {
this._pinned = pinned;
this._pinnedAt = pinned ? (pinnedAt ?? Date.now()) : null;
}
get flickerFilterEnabled(): boolean {
return this._flickerFilterEnabled;
}
@@ -1440,6 +1575,7 @@ export class Session extends EventEmitter {
// attach repaint, so the home screens' quiet ordering survives a restart.
lastActivityAt: this._wireActivityAt,
name: this._name,
nameSource: this._nameSource,
mode: this.mode,
autoClearEnabled: this._autoOps.autoClearEnabled,
autoClearThreshold: this._autoOps.autoClearThreshold,
@@ -1478,6 +1614,7 @@ export class Session extends EventEmitter {
ompConfig: this._ompConfig,
resumeSessionId: this._resumeSessionId,
effort: this._effort,
customModel: this.customModel,
// COD-118: runtime-only — surfaced so the frontend can require explicit user
// intent before restarting a crash-looped session. Deliberately NOT restored
// by the constructor: a Codeman restart starts with a fresh breaker so boot
@@ -1617,6 +1754,7 @@ export class Session extends EventEmitter {
console.error('[Session] Failed to respawn pane, will create new session');
needsNewSession = true;
} else {
this._pendingEnvUnsets.clear();
// Wait a moment for the respawned process to fully start
await new Promise((resolve) => setTimeout(resolve, MUX_STARTUP_DELAY_MS));
}
@@ -1710,14 +1848,68 @@ export class Session extends EventEmitter {
return true;
}
/**
* Kill and relaunch this session's CLI process IN PLACE — same pane, same tmux
* session, fresh env/args from current state. Custom Model Endpoint Profiles
* (docs/custom-model-endpoints-plan.md) is the first caller: after `setCustomModel()` merges new
* env vars into `_envOverrides`, the running CLI process still has the OLD env
* (inherited at its own process start, not live-reloaded), so switching a
* session's model/endpoint requires this restart to actually take effect.
*
* A GENERALIZED {@link reattachRemote} with the `!this._remote` guard dropped —
* `_buildRespawnPaneOptions()` already passes `remote: this._remote` through
* unconditionally, so `mux.respawnPane()` builds the right command either way
* (a local session gets `respawn-pane -k` + the real launch line, which is the
* kill-and-relaunch this method exists for; a remote session gets the existing
* reattach-to-durable-tmux behavior). Deliberately does NOT check `isBusy()` —
* that's the caller's job (mirrors `/interactive`'s guard), since a raw restart
* primitive shouldn't itself decide when it's safe to use.
*
* @returns true if the pane was respawned, false otherwise (no mux session, or
* the mux session is gone — see {@link reattachRemote} for that reasoning).
*/
async restartCli(): Promise<boolean> {
if (!this._useMux || !this._mux || !this._muxSession) return false;
const mux = this._mux;
if (!mux.muxSessionExists(this._muxSession.muxName)) {
console.log('[Session] restartCli: mux session gone, skipping:', this._muxSession.muxName);
return false;
}
this._pinOmpRespawnId();
const options = this._buildRespawnPaneOptions();
// Unlike the dead-pane respawn, this one kills a WORKING pane whose conversation
// already has a transcript, and a CLI that launches with `--session-id <id>` refuses
// an id that is already in use (claude: `Error: Session ID ... is already in use.`),
// which turned an endpoint switch into a dead pane and a lost session. A launch that
// declares a `fallback` chain renders `resume || new` once a resume id is set, the
// same `--resume <id> || --session-id <id>` shape the docker and remote pane commands
// already use, so pin the live conversation id for THIS respawn only. The registry
// shape is the gate, not the CLI's name: an entry whose resume id is minted by the
// CLI itself (codex/pi/omp/grok) never declares that chain, and its resume field is
// read from its own `<Mode>Config` rather than this top-level one anyway.
if (!options.resumeSessionId && getCli(this.mode)?.launch.chain === 'fallback') {
options.resumeSessionId = this._claudeSessionId ?? this.id;
}
const newPid = await mux.respawnPane(options);
if (!newPid) {
console.error('[Session] restartCli: respawnPane failed for', this._muxSession.muxName);
return false;
}
this._pendingEnvUnsets.clear();
console.log('[Session] restartCli: restarted CLI for', this._muxSession.muxName, 'pid', newPid);
return true;
}
/**
* Assemble the {@link RespawnPaneOptions} for this session. Single source of
* truth shared by interactive start, shell start (via their inline copies),
* and {@link reattachRemote} so the remote reattach path can never drift from
* {@link reattachRemote}, and {@link restartCli} so no respawn path can drift from
* the spawn path.
*/
private _buildRespawnPaneOptions(): import('./mux-interface.js').RespawnPaneOptions {
return {
const options: import('./mux-interface.js').RespawnPaneOptions = {
sessionId: this.id,
workingDir: this.workingDir,
mode: this.mode,
@@ -1745,12 +1937,37 @@ export class Session extends EventEmitter {
ompConfig: this._ompConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
unsetEnvKeys: this._pendingEnvUnsets.size > 0 ? [...this._pendingEnvUnsets] : undefined,
effort: this._effort,
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
owner: this._owner,
};
return this._withCustomModelLaunchModel(options);
}
/**
* Force the custom-model selection's `launchModel` (pi/omp `custom/<id>`, grok's
* `[model.<name>]` block name) onto the CLI's `model` launch param. Where that param
* lives is registry DATA — the entry's `legacyConfigField` (`piConfig`, `grokConfig`,
* ...) or the top-level `model` for an entry that declares none — so this stays a
* generic reader rather than a branch per CLI. Applied on the OPTIONS only: the stored
* `<Mode>Config` keeps whatever model the user chose at create, which is exactly what a
* later clear must fall back to.
*/
private _withCustomModelLaunchModel(
options: import('./mux-interface.js').RespawnPaneOptions
): import('./mux-interface.js').RespawnPaneOptions {
const launchModel = this._customModel?.launchModel;
if (!launchModel) return options;
const entry = getCli(this.mode);
if (!entry) return options;
const field = entry.launch.legacyConfigField;
if (!field) return { ...options, model: launchModel };
const bag = options as unknown as Record<string, unknown>;
const existing = (bag[field] ?? {}) as Record<string, unknown>;
return { ...options, [field]: { ...existing, model: launchModel } };
}
/**
@@ -2976,6 +3193,9 @@ export class Session extends EventEmitter {
}
private _clearAllTimers(): void {
// Stop re-sending Enter for a prompt this session will never take now
this._submitVerifier?.cancel();
this._submitVerifier = null;
// Clear the workspace-trust follow-up read
if (this._trustDialogTimer) {
clearTimeout(this._trustDialogTimer);
@@ -3357,10 +3577,11 @@ export class Session extends EventEmitter {
* discards the data, but it used to do so with no signal at all — which is how
* input could disappear while the caller believed it had been delivered.
*/
write(data: string): boolean {
this._trackSubmit(data);
write(data: string, options: SessionWriteOptions = {}): boolean {
const submittedPrompt = this._trackSubmit(data, options);
if (!this.ptyProcess) return false;
this.ptyProcess.write(data);
this._emitSubmittedPrompt(submittedPrompt);
return true;
}
@@ -3377,10 +3598,35 @@ export class Session extends EventEmitter {
return this._lastSubmitAt;
}
private _trackSubmit(data: string): void {
/**
* Stamps the pane's last Enter for EVERY write, and feeds the auto-name
* tracker only for user-originated input on a prompt-taking CLI. Ralph
* kick-starts, respawn `/clear`s, cron launches, approval answers and the
* trust-dialog keys all arrive without `fromUser` and so can never name a tab.
*/
private _trackSubmit(data: string, options: SessionWriteOptions): string[] {
const submitted = options.fromUser && this._acceptsPrompts ? this._submittedPromptTracker.feed(data) : [];
if (data.includes('\r') || data.includes('\n')) {
this._lastSubmitAt = Date.now();
}
return submitted;
}
/**
* Feeds user input that reaches the pane AROUND the write paths: the
* send-key route injects Shift+Enter's line feed through `tmux send-keys -H`
* directly, and without this the two lines of a prompt joined with no
* separator. Reports submissions like a write would (a line feed never is one).
*/
trackUserInput(data: string): void {
if (!this._acceptsPrompts) return;
this._emitSubmittedPrompt(this._submittedPromptTracker.feed(data));
}
private _emitSubmittedPrompt(prompts: string[]): void {
for (const prompt of prompts) {
this.emit('promptSubmitted', prompt);
}
}
/**
@@ -3416,6 +3662,21 @@ export class Session extends EventEmitter {
* half-open socket silently drops frames with no error) would type a prompt
* twice whenever an ACK is lost after the write landed.
*/
/**
* The highest input seq recorded for `clientId`, or 0 when this session has
* never seen it.
*
* Reported back on a REJECTED (duplicate) frame so the client can lift its own
* counter above this watermark. Without that number a client whose persisted
* counter fell behind ours has no way to find its way out: every fresh
* keystroke it sends lands at or below the watermark, is dropped as a
* duplicate, and is ACKed anyway — so the UI looks healthy while nothing is
* delivered, and a reload restores the same stale counter from localStorage.
*/
lastInputSeq(clientId: string): number {
return this._appliedInputSeq.get(clientId) ?? 0;
}
shouldApplyInput(clientId: string, seq: number): boolean {
const last = this._appliedInputSeq.get(clientId);
if (last !== undefined && seq <= last) return false;
@@ -3462,19 +3723,47 @@ export class Session extends EventEmitter {
* session.writeViaMux('/init\r'); // Send /init command
* ```
*/
async writeViaMux(data: string): Promise<boolean> {
this._trackSubmit(data);
async writeViaMux(data: string, options: SessionWriteOptions = {}): Promise<boolean> {
const submittedPrompt = this._trackSubmit(data, options);
if (this._mux && this._muxSession) {
return this._mux.sendInput(this.id, data);
const sent = await this._mux.sendInput(this.id, data);
if (sent) {
this._emitSubmittedPrompt(submittedPrompt);
this._verifySubmitted(data);
}
return sent;
}
// Fallback to PTY write
if (this.ptyProcess) {
this.ptyProcess.write(data);
this._emitSubmittedPrompt(submittedPrompt);
return true;
}
return false;
}
/**
* Arm the composer check for a write that carried Enter (session-submit-verifier.ts):
* Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints,
* so the pair `sendInput` just sent can leave the text stranded. Only a mux session
* can read its pane, only text can be stranded, and the glyph is the CLI's own.
*/
private _verifySubmitted(data: string): void {
if (!data.includes('\r') || !this._mux?.capturePaneText || !this._muxSession) return;
const text = data.replace(/[\r\n]/g, '').trimEnd();
if (!text) return;
this._submitVerifier ??= new SubmitVerifier({
capture: () =>
this._isStopped || !this._mux || !this._muxSession
? null
: this._mux.capturePaneText?.(this._muxSession.muxName),
sendEnter: () => this._mux?.sendInput(this.id, '\r'),
glyph: () => getCli(this.mode)?.capabilities.workDetect?.promptGlyph ?? '❯',
log: (m) => console.log(`[Session ${this.id.slice(0, 8)}] ${m}`),
});
this._submitVerifier.arm(text);
}
/** Current PTY dimensions — used to skip no-op resizes that trigger Ink redraws */
private _ptyCols = 120;
private _ptyRows = 40;
+26 -11
View File
@@ -1738,21 +1738,34 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
* Key validation is strict (`/^[A-Z_][A-Z0-9_]*$/`) as defense-in-depth against
* shell-metachar injection even if upstream schema check is bypassed.
*/
private applyEnvOverrides(muxName: string, envOverrides?: Record<string, string>): void {
private applyEnvOverrides(muxName: string, envOverrides?: Record<string, string>, unsetKeys?: string[]): void {
const VALID_KEY = /^[A-Z_][A-Z0-9_]*$/;
// Legacy cleanup: pre-0.7.2 set CLAUDE_CODE_EFFORT_LEVEL via setenv, which persists
// on the tmux session and hard-locks /effort switching in every respawned pane.
// Effort now flows as a `--settings` soft default (see buildEffortSettingsFlag),
// so unconditionally unset the stale var before applying current overrides.
try {
execSync(`${this.tmux()} setenv -t ${shellescape(muxName)} -u CLAUDE_CODE_EFFORT_LEVEL`, {
timeout: EXEC_TIMEOUT_MS,
stdio: ['pipe', 'pipe', 'pipe'],
});
} catch {
/* Non-critical — var may not exist */
//
// The caller's own unsets ride the same path, and run BEFORE the overrides are
// (re)applied: a key that is both unset and present in `envOverrides` ends up set,
// so a stale unset can never clobber a live value. Removing a key from the map is
// not enough on its own — `setenv` persists at the tmux-session level and is
// inherited by `respawn-pane`, measured: `setenv FOO bar` survived two successive
// `respawn-pane -k`. Clearing a custom-model selection is what needs this.
for (const key of ['CLAUDE_CODE_EFFORT_LEVEL', ...(unsetKeys ?? [])]) {
if (!VALID_KEY.test(key)) {
console.warn(`[TmuxManager] Skipping invalid env unset key: ${JSON.stringify(key)}`);
continue;
}
try {
execSync(`${this.tmux()} setenv -t ${shellescape(muxName)} -u ${key}`, {
timeout: EXEC_TIMEOUT_MS,
stdio: ['pipe', 'pipe', 'pipe'],
});
} catch {
/* Non-critical — var may not exist */
}
}
if (!envOverrides) return;
const VALID_KEY = /^[A-Z_][A-Z0-9_]*$/;
for (const [key, value] of Object.entries(envOverrides)) {
if (!value) continue; // Skip empty — nothing to set
if (!VALID_KEY.test(key)) {
@@ -2208,6 +2221,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
ompConfig,
resumeSessionId,
envOverrides,
unsetEnvKeys,
effort,
remote,
docker,
@@ -2269,8 +2283,9 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
);
this._configureStatusLineUserCommand(muxName, userStatusLineCommand);
// Re-apply user env overrides before respawn so the new shell inherits them.
this.applyEnvOverrides(muxName, envOverrides);
// Re-apply user env overrides before respawn so the new shell inherits them,
// dropping the ones the caller retired first (see applyEnvOverrides).
this.applyEnvOverrides(muxName, envOverrides, unsetEnvKeys);
// -c /tmp + cd bounce — see createSession() for rationale (stale FUSE state).
const launchCmd = remote || docker ? fullCmd : `cd ${JSON.stringify(workingDir)} && ${fullCmd}`;
+6
View File
@@ -184,6 +184,8 @@ export interface CaseInfo {
container: string;
image?: string;
path: string;
/** Directory INSIDE the container (defaults to `path` when unset). */
containerWorkdir?: string;
network?: string;
/**
* CLIs available INSIDE the container. A container case runs its agents in
@@ -201,6 +203,10 @@ export interface CaseInfo {
* first session: the container is created on demand by the launch chain, so treating
* "not found" as a fault there hid every agent mode behind an error telling the user
* to start a container Codeman was about to create itself.
*
* It also gates the Add Case panel's "copy an existing case" picker: only an ADOPTED
* container may back several cases at once (`classifyAdoptContainerConflict`), since an
* owned container's lifecycle belongs to its one case.
*/
owned?: boolean;
};
+52
View File
@@ -58,6 +58,24 @@ export type SessionMode =
| 'deepseek'
| 'omp';
/**
* Who owns a session's name. `placeholder`: Codeman's own `w<n>-<case>` (or no
* name at all), still eligible for auto-naming. `auto`: titled after its first
* prompt (`w<n>-<case>: <title>`), which happens once. `manual`: set by a
* person; auto-naming never touches it.
*/
export type SessionNameSource = 'placeholder' | 'auto' | 'manual';
/** Options for `Session.write()` / `Session.writeViaMux()`. */
export interface SessionWriteOptions {
/**
* The bytes were typed by a person, or sent by an agent on their behalf
* (browser keystrokes, `POST /api/sessions/:id/input`). Only such input can
* name a tab; Ralph, respawn, cron and approval writes leave this unset.
*/
fromUser?: boolean;
}
export type RemoteCommandMode = Extract<
SessionMode,
'shell' | 'claude' | 'opencode' | 'codex' | 'gemini' | 'antigravity' | 'pi' | 'grok' | 'deepseek' | 'omp'
@@ -558,6 +576,29 @@ export interface SessionAttachmentHistoryItem {
/**
* Current state of a session
*/
/** The public half of a session's custom-model selection (on the wire, in `SessionState`). */
export interface CustomModelSelection {
endpointId: string;
modelId: string;
label?: string;
}
/**
* The full custom-model selection a session keeps: the public selection plus the
* bookkeeping `Session.setCustomModel()` needs to UNDO it later without guessing what
* it once wrote. Persisted to state.json only as the disk-only `__customModel` field
* (never broadcast); the injected env VALUES are not in here at all, since they carry
* the endpoint's API key, and are re-derived from the endpoint store on recovery.
*/
export interface CustomModelBookkeeping extends CustomModelSelection {
/** Env keys the selection injected into the session's envOverrides / tmux session. */
envKeys: string[];
/** Isolated per-session config directory written for a `configDir`-kind CLI. */
configDir?: string;
/** Value forced onto the CLI's `model` launch param (pi/omp `custom/<id>`, grok's block name). */
launchModel?: string;
}
export interface SessionState {
/** Unique session identifier */
id: string;
@@ -591,6 +632,8 @@ export interface SessionState {
lastActivityAt: number;
/** Session display name */
name?: string;
/** Who owns the name (see `SessionNameSource`); absent on states persisted before auto-naming existed. */
nameSource?: SessionNameSource;
/** Session mode */
mode?: SessionMode;
/** Auto-clear enabled */
@@ -677,6 +720,15 @@ export interface SessionState {
resumeSessionId?: string;
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
effort?: EffortLevel;
/**
* Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): the custom
* OpenAI-compatible endpoint (local or cloud) this session's CLI is currently pointed
* at, if any. Undefined = the harness's native cloud default. No secrets here — the
* endpoint's base URL/api key live only in Session._envOverrides, never in this public
* state. The internal half (which env keys were injected, which config dir was
* written) is {@link CustomModelBookkeeping}, persisted disk-only like `__envOverrides`.
*/
customModel?: CustomModelSelection;
/** Sanitized per-session attachment history. */
attachmentHistory?: SessionAttachmentHistoryItem[];
/**
+71 -1
View File
@@ -23,7 +23,14 @@ import { getHookSecret, HOOK_SECRET_HEADER } from '../../config/hook-secret.js';
import { isMultiUserMode } from '../../config/multiuser.js';
import { findUser, setPassword, touchLastLogin, verifyPassword } from '../../user-store.js';
import { webviewCapabilities } from '../../webview-capabilities.js';
import { capabilityFromProxyPath, capabilityFromReferer } from '../webview-proxy.js';
import {
capabilityFromProxyPath,
capabilityFromReferer,
carriesAuthCredentials,
isLostWebviewFrameNavigation,
lostWebviewFramePage,
LOST_FRAME_PAGE_CSP,
} from '../webview-proxy.js';
import { ApiErrorCode, createErrorResponse, type AuthUser } from '../../types.js';
// Request-scoped identity (multi-user). Single-user leaves it undefined and the
@@ -176,6 +183,65 @@ function hasValidWebviewCapability(req: FastifyRequest, basePath = ''): boolean
return !!fromReferer && webviewCapabilities.resolve(fromReferer) !== undefined;
}
/**
* A web-tab frame that navigated itself off the proxy prefix (see
* isLostWebviewFrameNavigation). It cannot authenticate: opaque origin, no cookie,
* no capability left in the URL. Answer with the static recovery page here, BEFORE
* the credential checks, so the reload of a proxied dashboard neither shows a
* login challenge inside the tab nor counts as a failed attempt against the
* caller's IP — a dev server that full-reloads on every save would otherwise
* rate-limit its own user out of Codeman. Fenced like the Referer exemption: a
* path that resolves to a real route (/api, /q, a registered handler) is never
* answered this way, so a genuine unauthenticated navigation still gets the 401.
*
* `/` is the one registered route that IS answered here, and only when the
* request carries neither the session cookie nor an Authorization header. The
* shim maps `/webview/<cap>/` to exactly `/`, so a dashboard that reloads on its
* landing page (a Vite dev server on a config change) asks for Codeman's root
* as an iframe navigation; answering that with the app shell put Codeman inside
* its own web tab, and with a password it was a 401 in the frame. Nothing in
* Codeman frames its own root and the sandboxed frame has no credentials, so the
* credential-free form can only be that frame; a framed `/` WITH credentials is
* still the shell. Property worth knowing: a non-browser client can set these
* headers too, so an unauthenticated caller can tell a registered route (401)
* from a non-route (200) and enumerate the route table. Accepted, because the
* routes are public in docs/api-reference.md.
*
* @returns true when the reply was sent.
*/
function serveLostWebviewFrame(req: FastifyRequest, reply: FastifyReply): boolean {
if (!isLostWebviewFrameNavigation(req)) return false;
const url = (req.url ?? '').split('?')[0];
if (url.startsWith('/api/') || url.startsWith('/ws/') || url.startsWith('/q/')) return false;
if (url === '/') {
if (carriesAuthCredentials(req.headers, AUTH_COOKIE_NAME)) return false;
} else if (matchesRegisteredRoute(req, url)) {
return false;
}
sendLostWebviewFramePage(reply);
return true;
}
/**
* The landing-page case of serveLostWebviewFrame, for the index route. Without
* CODEMAN_PASSWORD no auth hook runs at all, so a lost frame's reload of `/`
* reaches `GET /` directly and the route asks this before rendering the shell.
* Under a password the hook has already answered a credential-free lost frame,
* so here it only ever sees the credentialed form, which stays the shell.
*/
export function isLostWebviewRootFrame(req: FastifyRequest): boolean {
if (!isLostWebviewFrameNavigation(req)) return false;
if ((req.url ?? '').split('?')[0] !== '/') return false;
return !carriesAuthCredentials(req.headers, AUTH_COOKIE_NAME);
}
/** Send the static recovery page (lostWebviewFramePage) with its own CSP, uncached. */
export function sendLostWebviewFramePage(reply: FastifyReply): FastifyReply {
reply.header('content-security-policy', LOST_FRAME_PAGE_CSP);
reply.header('cache-control', 'no-store');
return reply.type('text/html; charset=utf-8').send(lostWebviewFramePage());
}
/**
* Whether `url` resolves to a route Codeman actually registered.
*
@@ -302,6 +368,8 @@ export function registerAuthMiddleware(app: FastifyInstance, https: boolean, bas
done();
return;
}
// A web-tab frame that lost its prefix: hand it back to its tab, no credentials involved.
if (serveLostWebviewFrame(req, reply)) return;
const clientIp = req.ip;
@@ -439,6 +507,8 @@ function registerMultiUserAuthHook(
// ownership against the identity BOUND TO THE CAPABILITY, which is stricter
// than re-deriving it from a request that carries no credentials.
if (hasValidWebviewCapability(req, basePath)) return;
// A web-tab frame that lost its prefix: hand it back to its tab, no credentials involved.
if (serveLostWebviewFrame(req, reply)) return;
const clientIp = req.ip;
+36
View File
@@ -4,6 +4,7 @@
*/
import type { Session } from '../../session.js';
import type { SessionState } from '../../types.js';
export interface SessionPort {
readonly sessions: ReadonlyMap<string, Session>;
@@ -12,5 +13,40 @@ export interface SessionPort {
setupSessionListeners(session: Session): Promise<void>;
persistSessionState(session: Session): void;
persistSessionStateNow(session: Session): void;
/**
* Re-apply the persisted state a freshly CONSTRUCTED session does not carry.
*
* A `Session` built from a record holds only what its constructor takes, so
* persisting it would otherwise REPLACE the fuller record with the reduced one.
* Two phases: `before-spawn` shapes the pane (the custom-model environment and
* the nice priority) and must precede `startInteractive()`; `after-spawn` is
* the session's own history (the pin, token and cost totals, auto-compact,
* auto-clear, auto-resume, colour, image watcher, flicker filter) and must NOT
* land on a session whose pane failed to start.
*/
reapplyPersistedSessionState(
session: Session,
saved: SessionState,
phase: 'before-spawn' | 'after-spawn',
options?: {
/**
* Re-arm a PENDING auto-resume schedule from the record's `autoResumeAt`.
* Default true, which is what a Codeman restart wants: the limit footer
* will not reprint on its own, so dropping the stamp there strands the
* pause. A reboot restore passes false: the stamp predates the reboot,
* the pane is new, and re-arming means every restored session types
* `continue` into itself about a minute after one click. Auto-resume
* stays ENABLED either way, so it re-arms on fresh evidence.
*/
rearmAutoResumeSchedule?: boolean;
}
): Promise<void>;
/**
* Undo a session that was registered but never got a working pane: the map
* entry, its tab-layout slot, and any pane the launch created before throwing.
* Unlike {@link cleanupSession} it leaves the persisted record, the lifetime
* token totals, the Ralph state and the workspace's own files untouched.
*/
discardPartiallyBuiltSession(sessionId: string): Promise<void>;
getSessionStateWithRespawn(session: Session): unknown;
}
+196 -24
View File
@@ -957,6 +957,10 @@ class CodemanApp {
this.registerServiceWorker();
// Fetch tunnel status for header indicator (desktop only)
this.loadTunnelStatus();
// Ask whether a host reboot left sessions worth rebuilding (banner, never
// automatic). handleInit() re-reads it on every SSE init; this covers the
// path where that event never arrives.
this.initRebootRestoreBanner?.();
// Share a single settings fetch between both consumers
const settingsPromise = fetch('/api/settings').then(r => r.ok ? r.json() : null).then(env => env?.data ?? null).catch(() => null);
this.loadQuickStartCases(null, settingsPromise);
@@ -1906,6 +1910,53 @@ class CodemanApp {
this._onSessionClearTerminal(data);
}
/**
* How a buffer load that just fetched `payload` must end.
*
* A tmux pane capture is a point-in-time frame, so nothing that reached the
* browser after the response headers can already be in it. Such a load
* replays exactly that tail; discarding it drops the CLI's output for the
* rest of the load window, and its next partial redraw then lands on a frame
* the terminal never received.
*
* A `history` payload is the server's byte buffer alone: the direct-PTY
* fallback, or a mux pane whose capture came back empty. The route reads
* that buffer in the same synchronous tick it takes the capture, so it is
* current up to the route's own read and no further, which is the same
* exposure. It deliberately keeps the pre-existing discard all the same:
* both cases are rare, neither has been measured, and a duplicated Ink
* redraw is more visible than a few milliseconds of missing output.
* `capturedFromMux` below is the one line to widen if either turns out to
* matter.
*
* `headersReceivedAt` is the caller's own `performance.now()` reading from
* the moment the response arrived, compared only against other client-side
* readings, so there is no clock skew to worry about.
*
* What this cutoff does NOT cover, and there are two contributors. The
* server appends output to the byte buffer and emits it in the same tick,
* but BROADCASTS on a batch timer (8ms over WebSocket, 16 to 50ms over SSE),
* and the terminal route runs synchronously from `capture-pane` to its
* return, so a batch already pending when the capture ran leaves the server
* after the reply, arrives after `headersReceivedAt`, and is replayed
* although the capture holds it. Separately, `captureActivePaneBuffer` is
* `execSync`, which blocks the event loop for the whole capture: anything
* tmux had already painted into the pane that the server had not yet read
* from the attach PTY is in the capture too, is broadcast only after the
* reply, and replays the same way. The duplicate is one batch interval plus
* one capture wide, against a recovery window that spans the whole chunked
* write. Closing it belongs on the server: flush that session's pending
* batch before taking the capture.
*
* @param {{source?: string}} payload - The parsed `data` of a terminal response.
* @param {number} headersReceivedAt - When that response reached this client.
* @returns {{flushQueued: boolean, since: number}} Options for `_finishBufferLoad`.
*/
_bufferLoadFinishOpts(payload, headersReceivedAt) {
const capturedFromMux = payload?.source === 'mux-visible' || payload?.source === 'mux-full-history';
return { flushQueued: capturedFromMux, since: headersReceivedAt };
}
_onSessionTerminal(data) {
if (data.id === this.activeSessionId) {
if (data.data.length > 32768) _crashDiag.log(`TERMINAL: ${(data.data.length/1024).toFixed(0)}KB`);
@@ -1915,7 +1966,7 @@ class CodemanApp {
// jump over the cap. Dropped data is recovered from the canonical buffer.
const queued = (this.pendingWrites?.reduce((s, w) => s + w.length, 0) || 0)
+ (this.flickerFilterBuffer?.length || 0)
+ (this._loadBufferQueue?.reduce((s, w) => s + w.length, 0) || 0)
+ (this._loadBufferQueue?.reduce((s, w) => s + w.data.length, 0) || 0)
+ (this._terminalWriteInFlightBytes || 0);
if (queued + data.data.length > 131072) { // 128KB — drop to prevent accumulation
// Schedule a self-recovery once the
@@ -2498,9 +2549,11 @@ class CodemanApp {
? `/api/sessions/${sessionId}/terminal?full=1`
: `/api/sessions/${sessionId}/terminal?tail=${TERMINAL_TAIL_SIZE}`
);
let headersReceivedAt = performance.now();
let data = (await res.json())?.data ?? {};
if (useFullHistory && data.terminalBuffer && this._replayWouldShrinkBuffer(data.terminalBuffer)) {
res = await fetch(`/api/sessions/${sessionId}/terminal?tail=${TERMINAL_TAIL_SIZE}`);
headersReceivedAt = performance.now();
data = (await res.json())?.data ?? {};
}
// Bail on a tab switch mid-fetch: writing here would paint this session's
@@ -2516,7 +2569,12 @@ class CodemanApp {
const linesFromBottom = before ? Math.max(0, (before.baseY || 0) - (before.viewportY || 0)) : 0;
this.terminal.clear();
this.terminal.reset();
await this.chunkedTerminalWrite(data.terminalBuffer);
await this.chunkedTerminalWrite(
data.terminalBuffer,
TERMINAL_CHUNK_SIZE,
undefined,
this._bufferLoadFinishOpts(data, headersReceivedAt)
);
// A tail fetch can be partial, and the banner would otherwise keep
// describing the pre-refresh buffer (#258).
this._setHistoryTruncation(sessionId, data);
@@ -2526,6 +2584,10 @@ class CodemanApp {
});
if (target === null || typeof this.terminal.scrollToLine !== 'function') this.terminal.scrollToBottom();
else this.terminal.scrollToLine(target);
// The load's own replay sampled the sticky-scroll baseline while the
// terminal sat at the bottom of a just-rewritten buffer, so the next
// flush would scroll back down and undo the restore above.
this._syncStickyScrollBaseline();
// Re-position local echo overlay at new prompt location
this._localEchoOverlay?.rerender();
// Resize PTY to match actual browser dimensions (critical for OpenCode
@@ -2552,6 +2614,7 @@ class CodemanApp {
// Fetch buffer, clear terminal, write buffer, resize (no Ctrl+L needed)
try {
const res = await fetch(`/api/sessions/${data.id}/terminal`);
const headersReceivedAt = performance.now();
const termData = (await res.json())?.data ?? {};
this.terminal.clear();
@@ -2561,7 +2624,12 @@ class CodemanApp {
// (markers don't help here - this is a static buffer reload, not live Ink redraws)
const cleanBuffer = termData.terminalBuffer.replace(DEC_SYNC_STRIP_RE, '');
// Use chunked write to avoid UI freeze with large buffers (can be 1-2MB)
await this.chunkedTerminalWrite(cleanBuffer);
await this.chunkedTerminalWrite(
cleanBuffer,
TERMINAL_CHUNK_SIZE,
undefined,
this._bufferLoadFinishOpts(termData, headersReceivedAt)
);
}
// Fire-and-forget resize — don't block on it
@@ -2895,7 +2963,7 @@ class CodemanApp {
} else if (msg.t === 'ia') {
// Input ACK — the server applied (or deduped) this seq; drop it from
// the durable queue so it can never be re-delivered/lost.
this._onWsInputAck(msg.seq);
this._onWsInputAck(msg.seq, msg);
}
} catch {
// Ignore malformed messages
@@ -3073,7 +3141,11 @@ class CodemanApp {
this._pendingDeliveries.set(sessionId, list);
}
list.push(rec);
this._persistReliableState();
// ⚠️ SYNCHRONOUS, not the debounced writer: the seq counter is precisely the
// thing that must survive a crash, and a debounce puts it on the path most
// likely to be lost. A counter that comes back BELOW the server's watermark
// makes every later keystroke a silently-dropped duplicate (see _onWsInputAck).
this._persistReliableNow();
this._updateConnectionIndicator();
this._drainSession(sessionId);
}
@@ -3179,9 +3251,40 @@ class CodemanApp {
this.markIdleAlertSeen?.(sessionId);
}
/** Server input-ACK frame ({t:'ia',seq}) over the WebSocket. */
_onWsInputAck(seq) {
if (this._wsSessionId && Number.isInteger(seq)) this._ackDelivery(this._wsSessionId, seq);
/**
* Server input-ACK frame ({t:'ia',seq}) over the WebSocket.
*
* `dup:true` means the server REJECTED the frame as already-seen rather than
* applying it, and `last` is its watermark for this clientId. That combination
* is the escape hatch from a rolled-back counter: our seqs persist on a
* debounced write, so a tab killed between a send and that write comes back
* counting from BELOW the server's watermark, and from then on every keystroke
* is dropped-but-ACKed — a silently dead terminal that a reload cannot fix,
* because the stale counter is restored from localStorage too.
*
* ⚠️ Only a FIRST-attempt record is re-queued. A retry (`tries > 1`) being
* called a duplicate is the mechanism working as designed — the original did
* land — and re-sending it would type the same thing twice.
*/
_onWsInputAck(seq, msg) {
const sessionId = this._wsSessionId;
if (!sessionId || !Number.isInteger(seq)) return;
if (msg && msg.dup) {
const list = this._pendingDeliveries.get(sessionId);
const rec = list && list.find((r) => r.seq === seq);
const watermark = Number.isInteger(msg.last) ? msg.last : seq;
// Lift the counter clear of the server's watermark before anything else, so
// the re-queue below (and every later keystroke) gets an acceptable seq.
if ((this._seqCounters.get(sessionId) || 0) <= watermark) {
this._seqCounters.set(sessionId, watermark);
this._persistReliableNow();
}
const lost = rec && rec.tries <= 1 ? rec.data : null;
this._ackDelivery(sessionId, seq);
if (lost !== null) this._reliableSend(sessionId, lost, rec.useMux);
return;
}
this._ackDelivery(sessionId, seq);
}
/** Called from ws.onopen — flush everything pending over the fresh socket. */
@@ -3601,13 +3704,20 @@ class CodemanApp {
* Reset all app state maps, timers, and handlers to a clean baseline.
* Called by handleInit() on SSE reconnect / page reload to prevent
* memory leaks and stale data.
*
* @param {boolean} [preserveTerminal] Keep the terminal caches. Set when an SSE
* RECONNECT lands back on the session already on screen: the buffers still
* describe that session, and dropping them forces a full refetch + xterm
* reset that throws away the user's scroll position (see handleInit).
*/
_resetAllAppState() {
_resetAllAppState(preserveTerminal = false) {
this.sessions.clear();
this.ralphStates.clear();
this.terminalBuffers.clear();
this.terminalBufferCache.clear();
this._xtermSnapshots?.clear();
if (!preserveTerminal) {
this.terminalBuffers.clear();
this.terminalBufferCache.clear();
this._xtermSnapshots?.clear();
}
this.projectInsights.clear();
this.teams.clear();
this.teamTasks.clear();
@@ -3717,6 +3827,12 @@ class CodemanApp {
// a fresh load / reconnect (authoritative; wins over the localStorage restore).
if (data.planUsage) this.updatePlanUsageChip(data.planUsage);
// A board left open across a host reboot reconnects HERE, to a server that came
// back with an empty session list. The reboot-restore offer is built at boot,
// before any client could be listening, so re-read it on every init rather than
// only on the page-load path.
this.refreshRebootRestoreBanner?.();
// Update version displays (header and toolbar)
if (data.version) {
const versionEl = this.$('versionDisplay');
@@ -3734,7 +3850,23 @@ class CodemanApp {
// Stop any active voice recording on reconnect
VoiceInput.cleanup();
this._resetAllAppState();
// A RECONNECT that lands back on the same session must not become a full
// reload. This used to clear the terminal caches and re-run selectSession()
// unconditionally, so every SSE reconnect refetched the buffer (up to 1 MiB)
// and reset+rewrote xterm. On a link that drops a connection about once a
// minute that reads as the page refreshing itself and losing your place.
// Keep the caches and the active id here; the restore block below resyncs
// through _onSessionNeedsRefresh(), which still reloads the buffer (so
// output produced during the outage is not lost) but preserves the reading
// position.
const activeBefore = this.activeSessionId;
const keepTerminal =
gen > 1 &&
!!activeBefore &&
Array.isArray(data.sessions) &&
data.sessions.some((s) => s.id === activeBefore);
this._resetAllAppState(keepTerminal);
data.sessions.forEach(s => {
this.sessions.set(s.id, s);
@@ -3864,20 +3996,32 @@ class CodemanApp {
}
const previousActiveId = this.activeSessionId;
this.activeSessionId = null;
if (this.sessionOrder.length > 0) {
if (this.sessionOrder.length === 0) {
this.activeSessionId = null;
} else {
// Priority: current active > localStorage > first session
let restoreId = previousActiveId;
if (!restoreId || !this.sessions.has(restoreId)) {
try { restoreId = localStorage.getItem('codeman-active-session'); } catch {}
}
// `auto`: the app is restoring a session on load, not a human opening
// one, so a pending idle alert on that tab stays armed until it is
// actually tapped (see the userInitiated note in selectSession).
if (restoreId && this.sessions.has(restoreId)) {
this.selectSession(restoreId, { auto: true });
if (keepTerminal && restoreId === previousActiveId && this.sessions.has(restoreId)) {
// Reconnect onto the session already on screen. renderSessionTabs() ran
// above and activeSessionId never changed, so the tab strip is already
// correct; only the buffer needs to catch up. The WS has its own
// backoff reconnect, but if it is not on this session (dead socket, or
// a give-up) nothing else would re-establish it from here.
if (this._wsSessionId !== restoreId) this._connectWs(restoreId);
void this._onSessionNeedsRefresh({ id: restoreId });
} else {
this.selectSession(this.sessionOrder[0], { auto: true });
this.activeSessionId = null;
// `auto`: the app is restoring a session on load, not a human opening
// one, so a pending idle alert on that tab stays armed until it is
// actually tapped (see the userInitiated note in selectSession).
if (restoreId && this.sessions.has(restoreId)) {
this.selectSession(restoreId, { auto: true });
} else {
this.selectSession(this.sessionOrder[0], { auto: true });
}
}
}
}
@@ -5782,7 +5926,12 @@ class CodemanApp {
parsedAt,
bufferLength: parsedBufferLength,
completed,
} = await this.chunkedTerminalWrite(buffer, TERMINAL_CHUNK_SIZE, sessionId);
} = await this.chunkedTerminalWrite(
buffer,
TERMINAL_CHUNK_SIZE,
sessionId,
this._bufferLoadFinishOpts(payload, headersReceivedAt)
);
timing.resetAndParseMs = parsedAt - replayStartedAt;
if (!completed || this.activeSessionId !== sessionId) return;
// Keep shell tab restores bounded too. A user-triggered full-history pull
@@ -5800,6 +5949,12 @@ class CodemanApp {
const delta = parsedBufferLength - rowsBefore;
if (delta > 0) this.terminal.scrollToLine(delta);
else this.terminal.scrollToTop();
// The load's own replay sampled the sticky-scroll baseline while the
// terminal sat at the bottom of a just-rewritten buffer, so the next
// flush would scroll back down and undo the restore above. This path is
// reached only from a scroll-up gesture, so being dragged down is the
// exact opposite of what the user asked for.
this._syncStickyScrollBaseline();
timing.totalMs = performance.now() - requestStartedAt;
this._recordTerminalLoadTiming(timing);
} catch {
@@ -6241,6 +6396,15 @@ class CodemanApp {
}
const data = (await res.json())?.data ?? {};
const bodyParsedAt = performance.now();
// How this load must end, decided here because `chunkedTerminalWrite` is
// what actually ends it for a non-empty buffer. A tmux pane capture is a
// point-in-time frame, so nothing that reached the browser after the
// response headers can already be in it. Replay exactly that tail;
// discarding it drops the CLI's output for the rest of the load window,
// and its next partial redraw then lands on a frame the terminal never
// received. `since` keeps the pre-capture events dropped, because the
// capture does hold those and replaying them would duplicate output.
const finishOpts = this._bufferLoadFinishOpts(data, headersReceivedAt);
_crashDiag.log(`FETCH_DONE: ${data.terminalBuffer ? (data.terminalBuffer.length/1024).toFixed(0) + 'KB' : 'empty'} truncated=${data.truncated}`);
let freshResetAndParseMs = 0;
@@ -6267,7 +6431,8 @@ class CodemanApp {
const { parsedAt: freshParsedAt } = await this.chunkedTerminalWrite(
data.terminalBuffer,
TERMINAL_CHUNK_SIZE,
bufferLoadOwner
bufferLoadOwner,
finishOpts
);
freshResetAndParseMs = freshParsedAt - replayStartedAt;
if (this._isStaleSelect(selectGen)) {
@@ -6317,7 +6482,14 @@ class CodemanApp {
// COD-144: when the load painted nothing, FLUSH the queued events instead of
// discarding — a new session's prompt arrives only as a queued SSE event.
if (this._isLoadingBuffer) {
this._finishBufferLoad(bufferLoadOwner, { flushQueued: bufferWasEmpty });
// Only reached when the write was skipped. COD-144 lives here: a new
// session's first prompt exists only as a queued event that predates the
// response, so an empty paint replays its queue WHOLE rather than from
// the header timestamp.
this._finishBufferLoad(
bufferLoadOwner,
bufferWasEmpty ? { flushQueued: true, since: 0 } : finishOpts
);
}
// Drop the guard so user input clears state normally
this._restoringFlushedState = false;
+1
View File
@@ -252,6 +252,7 @@
'Ultracode Agents': 'Ultracode 智能体',
'Ultracode Floating Windows': 'Ultracode 浮动窗口',
'Approvals Inbox': '审批收件箱',
'Auto-name Sessions': '自动命名会话',
Approvals: '审批',
'Prompts waiting on you, across all sessions': '所有会话中等待您处理的提示',
'No pending approvals': '没有待处理的审批',
+33
View File
@@ -213,6 +213,24 @@
<button class="offline-banner-retry" id="offlineBannerRetry" onclick="app.retryConnection()">Retry now</button>
</div>
<!-- Reboot-restore offer: shown when the server found sessions a host reboot
killed and is asking whether to rebuild them. Populated by
reboot-restore-ui.js; nothing is created until the user clicks. -->
<div class="reboot-restore-banner" id="rebootRestoreBanner" role="status" hidden>
<span class="reboot-restore-banner-icon" aria-hidden="true">↺</span>
<span class="reboot-restore-banner-text" id="rebootRestoreBannerText"></span>
<span class="reboot-restore-banner-detail" id="rebootRestoreBannerDetail"></span>
<span class="reboot-restore-banner-note">Conversations return; terminal history does not.</span>
<button
class="reboot-restore-banner-accept"
id="rebootRestoreBannerAccept"
onclick="app.restoreRebootSessions()"
>
Restore
</button>
<button class="reboot-restore-banner-dismiss" onclick="app.dismissRebootRestore()">Dismiss</button>
</div>
<!-- Timer Banner (shown when timed run is active) -->
<div class="timer-banner" id="timerBanner" style="display: none;">
<div class="timer-content">
@@ -2047,6 +2065,13 @@
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsLineageLines" checked><span class="slider"></span></label>
</div>
<div class="set-row" id="appSettingsAutoNameSessionsItem" data-search="auto name session title first prompt tab rename">
<div class="set-row-text">
<span class="set-row-label">Auto-name Sessions <span class="set-tag">synced</span></span>
<span class="set-row-desc">Title a new tab after its first prompt, keeping the case prefix. Renamed tabs are never touched.</span>
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsAutoNameSessions"><span class="slider"></span></label>
</div>
<div class="set-row" id="appSettingsMobileOverviewItem" data-search="overview home screen phone logo">
<div class="set-row-text">
<span class="set-row-label">Overview Home Screen <span class="set-tag">phone</span></span>
@@ -2906,6 +2931,13 @@
<label class="checkbox-row"><input type="checkbox" id="dockerAdoptExisting"> Attach to an existing container</label>
<span class="form-hint">On: Codeman only runs docker exec into a container you already built and run — it never creates, starts, stops or removes it. The CLIs must already be installed and logged in inside it.</span>
</div>
<div class="form-row docker-adopt-only" id="dockerAdoptCloneRow" hidden>
<label>Duplicate an Existing Case</label>
<select id="dockerAdoptCloneFrom" onchange="app.applyDockerCloneSource()">
<option value="">Start from scratch</option>
</select>
<span class="form-hint">Same container, another directory inside it. Picks up the container, host and workspace below &mdash; you only set a new name and container workdir. Adopted containers only: an owned container&#39;s lifecycle belongs to its one case.</span>
</div>
<div class="form-row docker-adopt-only">
<label>Container Name</label>
<input type="text" id="dockerContainerName" list="dockerContainerList" placeholder="my-dev-box" pattern="[a-zA-Z0-9][a-zA-Z0-9_.-]+" autocomplete="off" autocapitalize="off" spellcheck="false">
@@ -3521,6 +3553,7 @@
<script defer src="readmymind-ui.js"></script>
<script defer src="ultracode-panel.js"></script>
<script defer src="approvals-ui.js"></script>
<script defer src="reboot-restore-ui.js"></script>
<script defer src="admin-ui.js"></script>
<script defer src="session-ui.js"></script>
<script defer src="webview-tabs.js"></script>
+46
View File
@@ -120,6 +120,41 @@ const CjkInput = (() => {
c: '\x03', d: '\x04', l: '\x0c', z: '\x1a', a: '\x01', e: '\x05',
};
/** CSI final byte per navigation key, for the modifier-carrying forms below. */
const CSI_NAV_FINAL = {
ArrowUp: 'A',
ArrowDown: 'B',
ArrowRight: 'C',
ArrowLeft: 'D',
End: 'F',
Home: 'H',
};
/**
* The `CSI 1 ; <mod> <final>` form for a Ctrl/Alt-modified navigation key, or
* null when this key is not one.
*
* A modified navigation key is a terminal COMMAND, not text editing — claude's
* own "Jump to bottom (ctrl+End)" is one. PASSTHROUGH_KEYS carries only the
* plain forms, so Ctrl+End used to fail in BOTH directions: with an empty
* field it was sent as a bare `\x1b[F` (the modifier silently dropped, so the
* CLI saw a plain End), and with any text in the field it was not forwarded at
* all and the browser's default moved the caret to the end of the composer,
* which is what the user sees as "the shortcut does something to the input box
* instead".
*
* ⚠️ Shift ALONE is deliberately excluded: Shift+arrow selects text inside the
* composer, which is a real editing gesture worth keeping local. Shift is still
* encoded when it accompanies Ctrl or Alt.
*/
function _modifiedNavSequence(e) {
const final = CSI_NAV_FINAL[e.key];
if (!final) return null;
if (!e.ctrlKey && !e.altKey) return null;
const mod = 1 + (e.shiftKey ? 1 : 0) + (e.altKey ? 2 : 0) + (e.ctrlKey ? 4 : 0);
return `\x1b[1;${mod}${final}`;
}
function _strip(str) {
return str.replace(/​/g, '');
}
@@ -321,6 +356,17 @@ const CjkInput = (() => {
return;
}
// Ctrl/Alt-modified navigation keys go to the PTY REGARDLESS of whether
// the field has text: they are commands for the CLI, and the composer has
// no editing behaviour for them worth preserving (plain Home/End still
// edit locally through the table below).
const modNav = _modifiedNavSequence(e);
if (modNav) {
e.preventDefault();
_send(modNav);
return;
}
// Arrow/function keys: forward to PTY when no real text
if (PASSTHROUGH_KEYS[e.key] && _isEffectivelyEmpty()) {
e.preventDefault();
+38 -3
View File
@@ -5,7 +5,9 @@
*
* - KeyboardAccessoryBar (singleton object) — Quick action buttons shown above the virtual
* keyboard on mobile: arrow up/down, /init, Tab, paste, Esc, and dismiss (the extended
* bar adds /clear, /compact, Shift+Tab and more). Tab flushes any locally-buffered
* bar adds /clear, /compact, Shift+Tab and more). Shift+Left/Right ship in both agent
* layouts but are revealed only on Codex sessions (`codex-enabled` marker class on the
* bar, synced on every session switch), since they are Codex bindings. Tab flushes any locally-buffered
* prompt text to the PTY before sending \t, so completion applies to what was typed.
* The paste button opens a dialog that handles both text paste and image attach
* (native picker + best-effort image paste, routed through app._uploadAndInsertImages).
@@ -661,6 +663,8 @@ const KeyboardAccessoryBar = {
</button>
<button class="accessory-btn" data-action="init" title="/init">/init</button>
<button class="accessory-btn" data-action="tab" title="Tab">Tab</button>
<button class="accessory-btn accessory-btn-codex" data-action="shift-left" title="Shift+Left (Codex: edit queued message)" aria-label="Shift+Left (Codex: edit queued message)">⇧←</button>
<button class="accessory-btn accessory-btn-codex" data-action="shift-right" title="Shift+Right (Codex: prompt stack back)" aria-label="Shift+Right (Codex: prompt stack back)">⇧→</button>
<button class="accessory-btn" data-action="paste" title="Paste from clipboard">
<svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
<path d="M16 4h2a2 2 0 0 1 2 2v14a2 2 0 0 1-2 2H6a2 2 0 0 1-2-2V6a2 2 0 0 1 2-2h2"/>
@@ -746,6 +750,8 @@ const KeyboardAccessoryBar = {
<button class="accessory-btn" data-action="clear-input" title="Clear the current unsent input">&#x232B; All</button>
<button class="accessory-btn accessory-btn-rmm" data-action="readmymind" title="Read My Mind: predict your next prompt">🧠</button>
<button class="accessory-btn" data-action="tab" title="Tab">Tab</button>
<button class="accessory-btn accessory-btn-codex" data-action="shift-left" title="Shift+Left (Codex: edit queued message)" aria-label="Shift+Left (Codex: edit queued message)">⇧←</button>
<button class="accessory-btn accessory-btn-codex" data-action="shift-right" title="Shift+Right (Codex: prompt stack back)" aria-label="Shift+Right (Codex: prompt stack back)">⇧→</button>
<button class="accessory-btn" data-action="shift-tab" title="Shift+Tab">⇧Tab</button>
<button class="accessory-btn" data-action="effort-max" title="/effort max">Max</button>
<button class="accessory-btn" data-action="ctrl-o" title="Ctrl+O">⌃O</button>
@@ -772,6 +778,9 @@ const KeyboardAccessoryBar = {
// The 🧠 key is opt-in (`readMyMindEnabled`, synced): it ships in both
// templates but stays display:none until the bar carries the marker class.
this.syncReadMyMind();
// The ⇧←/⇧→ keys are Codex bindings: same shape, gated on the active
// session's mode instead of a setting.
this.syncCodexKeys();
// Add click handlers — preventDefault stops event from reaching terminal
this.element.addEventListener('click', (e) => {
@@ -784,7 +793,7 @@ const KeyboardAccessoryBar = {
this.handleAction(action, btn);
// Refocus terminal so keyboard stays open (tap blurs terminal → keyboard dismisses → toolbar shifts)
const refocusActions = new Set(['scroll-up', 'scroll-down', 'arrow-left', 'arrow-right', 'tab', 'shift-tab', 'ctrl', 'ctrl-o', 'opt-enter', 'esc', 'effort-max', 'clear-input']);
const refocusActions = new Set(['scroll-up', 'scroll-down', 'arrow-left', 'arrow-right', 'tab', 'shift-tab', 'shift-left', 'shift-right', 'ctrl', 'ctrl-o', 'opt-enter', 'esc', 'effort-max', 'clear-input']);
if (refocusActions.has(action) ||
((action === 'clear' || action === 'compact') && this._confirmAction)) {
if (typeof app !== 'undefined' && app.terminal) {
@@ -815,6 +824,7 @@ const KeyboardAccessoryBar = {
refreshForActiveSession() {
this.clearCtrl();
this._applyLayout(this._resolveMode());
this.syncCodexKeys();
},
/** Which layout the current state calls for. */
@@ -827,6 +837,11 @@ const KeyboardAccessoryBar = {
return app.sessions?.get(app.activeSessionId)?.mode === 'shell';
},
_isCodexSession() {
if (typeof app === 'undefined' || !app.activeSessionId) return false;
return app.sessions?.get(app.activeSessionId)?.mode === 'codex';
},
/** Swap the button set in the DOM. */
_applyLayout(mode) {
if (!this.element || mode === this._mode) return;
@@ -913,6 +928,12 @@ const KeyboardAccessoryBar = {
case 'arrow-right':
this.sendNavKey('\x1b[C');
break;
case 'shift-left':
this.sendNavKey('\x1b[1;2D');
break;
case 'shift-right':
this.sendNavKey('\x1b[1;2C');
break;
case 'esc':
this.sendKey('\x1b');
break;
@@ -1009,6 +1030,20 @@ const KeyboardAccessoryBar = {
this.element.classList.toggle('rmm-enabled', enabled === true);
},
/** Reveal the ⇧←/⇧→ keys only while the active session runs Codex. They are
* Codex bindings (edit the last queued message / prompt stack back) and do
* nothing in any other CLI, yet a tap still goes through sendNavKey(), which
* hands the session to plain PTY echo for the rest of the prompt, so on a
* phone a dead key would also switch off local echo. Same marker-class
* shape as syncReadMyMind(): the class lives on the BAR because setMode()
* rebuilds the buttons' innerHTML. Synced at init and on every session
* switch (refreshForActiveSession); a session's mode is fixed at create, so
* no other event can change the answer. */
syncCodexKeys() {
if (!this.element) return;
this.element.classList.toggle('codex-enabled', this._isCodexSession());
},
/** Send a slash command to the active session.
* Sends text and Enter separately so Ink processes them as distinct events. */
sendCommand(command) {
@@ -1044,7 +1079,7 @@ const KeyboardAccessoryBar = {
},
/**
* A composer nav key (the four arrows) from the bar, under the SAME contract
* A composer nav key (arrows, including Shift+Left/Right) from the bar, under the SAME contract
* as pressing one on a hardware keyboard (the `isComposerNavKey` branch of
* terminal-ui.js's onData): flush the unsent draft so the key edits the real
* composer, then hand the session to plain PTY echo until Enter or Ctrl+C,
+37
View File
@@ -3241,6 +3241,43 @@ html:is([data-skin="paper-gray"], [data-skin="solarized-light"], [data-skin="cat
other banners. The overlay is fixed and handles its own insets.
============================================================================ */
@media (max-width: 599px) {
/* Reboot-restore banner: the same treatment as the offline banner below. Its
text and note are nowrap and the two buttons cannot shrink, so without this
the actions are pushed off a phone-width viewport and become unreachable. */
.reboot-restore-banner {
padding: 0.4rem 0.5rem;
padding-left: calc(0.5rem + var(--safe-area-left));
padding-right: calc(0.5rem + var(--safe-area-right));
font-size: 0.7rem;
gap: 0.4rem;
}
/* The session names and the scrollback note are the first things to go. The
count plus the two buttons carry the message on their own, and the note
survives as the accept button's title. */
.reboot-restore-banner-detail,
.reboot-restore-banner-note {
display: none;
}
/* A flex item will not shrink below its content width at the default
`min-width: auto`, so without this the nowrap text pushes the buttons off a
360px viewport and the ellipsis never engages. */
.reboot-restore-banner-text {
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
}
.reboot-restore-banner-accept,
.reboot-restore-banner-dismiss {
padding: 0.25rem 0.5rem;
}
.reboot-restore-banner-accept {
margin-left: auto;
}
.offline-banner {
padding: 0.4rem 0.5rem;
padding-left: calc(0.5rem + var(--safe-area-left));
+142
View File
@@ -0,0 +1,142 @@
/**
* @fileoverview Reboot-restore banner: offer back the sessions a host reboot destroyed.
*
* A host reboot takes the tmux server down with it, so every session's pane dies
* and the board comes up empty. The server works out what was running from the
* records it still holds at boot, and this banner asks the user whether to
* rebuild them. Nothing is created until they click, because the server's
* reboot guess is a heuristic and a wrong automatic restore would spawn CLI
* processes nobody asked for.
*
* Seeded from `GET /api/reboot-restore` on init and again on every SSE reconnect,
* because the tab most likely to want this is one that was open across the reboot
* and reconnects to a server that came back up with an empty board. Restore posts to
* `POST /api/reboot-restore/restore` and Dismiss posts to
* `POST /api/reboot-restore/dismiss`. Dismiss always clears the banner; Restore
* re-reads the plan afterwards, because the server puts back anything it could
* not build for a reason that may pass, such as a session limit or an agent that
* would not start. The restored sessions arrive as ordinary `session:created`
* events, so no extra rendering is needed here.
*
* The banner says that terminal history did not survive, because a restored
* session is a new pane: the conversation continues and the scrollback does not.
* Saying so is what keeps an empty pane from reading as a broken restore.
* Backend: src/web/reboot-restore-registry.ts, src/web/routes/reboot-restore-routes.ts.
*
* @mixin Extends CodemanApp.prototype via Object.assign
* @dependency app.js (CodemanApp class, showToast)
* @dependency api-client.js at runtime (this._api / this._apiJson)
* @loadorder 11.65, after approvals-ui.js and before admin-ui.js (11.7)
*/
/** Plain-language wording for one skip reason, for the toast after a restore. */
function rebootSkipReason(reason) {
switch (reason) {
case 'workspace-missing':
return 'workspace is gone';
case 'workspace-forbidden':
return 'workspace is outside your space';
case 'already-live':
return 'already open';
case 'capacity-reached':
return 'session limit reached';
case 'rebuild-failed':
return 'the agent would not start';
default:
return reason;
}
}
Object.assign(CodemanApp.prototype, {
/** Ask the server whether a reboot left anything on offer, and show the banner if so. */
async initRebootRestoreBanner() {
const data = await this._apiJson('/api/reboot-restore');
const sessions = data?.sessions ?? [];
if (sessions.length === 0) return;
this._rebootRestoreSessions = sessions;
this.renderRebootRestoreBanner();
},
renderRebootRestoreBanner() {
const banner = this.$('rebootRestoreBanner');
if (!banner) return;
const sessions = this._rebootRestoreSessions ?? [];
if (sessions.length === 0) {
banner.hidden = true;
return;
}
const count = sessions.length;
const text = this.$('rebootRestoreBannerText');
if (text) {
const noun = count === 1 ? 'session' : 'sessions';
text.textContent = `Restore ${count} ${noun} from before the reboot`;
}
const detail = this.$('rebootRestoreBannerDetail');
if (detail) {
// Names, so the user can tell what they are about to relaunch.
const names = sessions
.map((s) => s.name || s.workingDir?.split('/').pop() || s.id.slice(0, 8))
.slice(0, 4)
.join(', ');
detail.textContent = count > 4 ? `${names}, …` : names;
detail.title = sessions.map((s) => `${s.name || s.id}\n${s.workingDir}`).join('\n\n');
}
const accept = this.$('rebootRestoreBannerAccept');
// The note is hidden at phone width, so the warning travels on the button too.
if (accept) accept.title = 'Conversations return; terminal history does not.';
banner.hidden = false;
},
/** Rebuild everything on offer. The panes are new, so scrollback does not come back. */
async restoreRebootSessions() {
const button = this.$('rebootRestoreBannerAccept');
if (button) button.disabled = true;
const res = await this._api('/api/reboot-restore/restore', { method: 'POST', body: {} });
if (res && res.status === 409) {
if (button) button.disabled = false;
this.showToast?.('A restore is already running', 'info');
return;
}
// The uniform envelope wraps every /api payload; reading the outer object
// would report every count as zero.
const body = res && res.ok ? (await res.json().catch(() => null))?.data : null;
if (!body) {
if (button) button.disabled = false;
this.showToast?.('Could not restore the sessions', 'error');
return;
}
const restored = body.restored?.length ?? 0;
const skipped = body.skipped?.length ?? 0;
// Re-read rather than clearing: the server puts back anything it could not
// build for a reason that may pass, such as a session limit or an agent that
// would not start, and blanking the banner here would put those entries out
// of reach until a reload.
await this.refreshRebootRestoreBanner();
if (button) button.disabled = false;
if (restored > 0) {
const noun = restored === 1 ? 'conversation' : 'conversations';
this.showToast?.(`Restored ${restored} ${noun}. Terminal history did not survive the reboot.`, 'success');
}
if (skipped > 0) {
// Each reason means a different next step for the user, so they are not
// collapsed into one message: capacity clears by closing something, a
// failed start usually means the CLI is not on the server's PATH.
const reasons = new Set((body.skipped ?? []).map((s) => s.reason));
this.showToast?.(`${skipped} not restored: ${[...reasons].map(rebootSkipReason).join('; ')}`, 'warning');
}
},
/** Re-read the offer after a reconnect, for a tab that was open across the reboot. */
async refreshRebootRestoreBanner() {
const data = await this._apiJson('/api/reboot-restore');
this._rebootRestoreSessions = data?.sessions ?? [];
this.renderRebootRestoreBanner();
},
/** Drop the offer. The Resume list still reaches every one of these conversations. */
async dismissRebootRestore() {
this._rebootRestoreSessions = [];
this.renderRebootRestoreBanner();
await this._apiPost('/api/reboot-restore/dismiss', {});
},
});
+129 -1
View File
@@ -3153,7 +3153,10 @@ Object.assign(CodemanApp.prototype, {
const adopting = document.getElementById('dockerAdoptExisting')?.checked;
if (adopting) modal.setAttribute('data-docker-adopt', '1');
else modal.removeAttribute('data-docker-adopt');
if (adopting) void this._loadDockerContainerOptions();
if (adopting) {
void this._loadDockerContainerOptions();
void this._loadDockerCloneOptions();
}
},
/**
@@ -3165,6 +3168,120 @@ Object.assign(CodemanApp.prototype, {
* Best-effort by design — the endpoint returns [] for an unreachable daemon,
* and an empty list simply leaves the field as plain text input.
*/
/**
* Fill the "Duplicate an Existing Case" picker with the ADOPTED docker cases.
*
* One adopted container can back several cases, each pointing at a different
* directory inside it (classifyAdoptContainerConflict) — but re-typing the
* container, host and workspace by hand for every directory is exactly the
* friction that makes the capability go unused. Picking a case here fills those
* three and leaves only the two fields that MUST differ: the case name and the
* container workdir.
*
* ⚠️ Adopted cases only (`docker.owned === false`). An owned container's
* lifecycle belongs to its one case — a second case on it would be destroyed
* out from under itself by that case's recreate or delete — and the server
* refuses it, so offering it here would only produce a confusing error.
*/
async _loadDockerCloneOptions() {
const select = document.getElementById('dockerAdoptCloneFrom');
const row = document.getElementById('dockerAdoptCloneRow');
if (!select || !row) return;
let cases = [];
try {
const res = await fetch('/api/cases');
const data = await res.json();
cases = (Array.isArray(data) ? data : data?.data || []).filter(
(c) => c?.docker && c.docker.owned === false
);
} catch {
cases = [];
}
select.textContent = '';
const blank = document.createElement('option');
blank.value = '';
blank.textContent = 'Start from scratch';
select.appendChild(blank);
for (const c of cases) {
const option = document.createElement('option');
option.value = c.name;
// Server-supplied strings: textContent, never markup.
option.textContent = `${c.name} — ${c.docker.container}:${c.docker.containerWorkdir || c.docker.path}`;
option.dataset.container = c.docker.container;
option.dataset.hostId = c.docker.hostId;
option.dataset.path = c.docker.path;
option.dataset.workdir = c.docker.containerWorkdir || c.docker.path;
select.appendChild(option);
}
// Nothing to duplicate yet: an empty picker is noise on the first adoption.
row.hidden = cases.length === 0;
},
/**
* Apply the picked case: carry over what STAYS the same, clear what must not.
*
* The two cleared fields are the point of the feature — a duplicate that kept
* the original's name would be rejected as an existing case, and one that kept
* its container workdir would be rejected as an exact twin (both by the server,
* with a clear message, but a form that pre-fills a value it knows will be
* refused is just a trap).
*/
applyDockerCloneSource() {
const select = document.getElementById('dockerAdoptCloneFrom');
const option = select?.selectedOptions?.[0];
if (!option || !option.value) return;
const set = (id, value) => {
const el = document.getElementById(id);
if (el) el.value = value || '';
};
set('dockerContainerName', option.dataset.container);
set('dockerHostId', option.dataset.hostId);
set('dockerWorkspacePath', option.dataset.path);
// Pre-filled, NOT cleared: these two must differ from the source, but editing
// `/srv/app/api` into `/srv/app/web` beats retyping a long path, and the same
// goes for the name. What keeps a duplicate from being submitted unchanged is
// the guard below (dockerCloneGuard), which is a better trade than an empty
// field: the form stays a starting point instead of a blank form with three
// fields mysteriously filled in.
set('dockerCaseName', option.value);
set('dockerAdoptWorkdir', option.dataset.workdir);
// Remembered so the guard can tell "unchanged" from "happens to look similar".
select.dataset.appliedName = option.value;
select.dataset.appliedWorkdir = option.dataset.workdir || '';
const workdir = document.getElementById('dockerAdoptWorkdir');
workdir?.focus();
// Caret at the end: the tail is the part that changes.
if (workdir) workdir.setSelectionRange(workdir.value.length, workdir.value.length);
},
/**
* Refuse a duplicate that still carries the source case's name or directory.
*
* Both are pre-filled so they can be EDITED, which means both can also be left
* alone by accident. The server refuses either (an existing case name, or an
* exact same-container-same-directory twin) with a clear message, but a
* round-trip to be told "you forgot to change the field you were looking at" is
* worse than saying so here, next to the field, before anything is sent.
*
* Returns the offending element, or null when the form is fine.
*/
dockerCloneGuard() {
const select = document.getElementById('dockerAdoptCloneFrom');
if (!select || !select.value) return null;
const name = document.getElementById('dockerCaseName');
const workdir = document.getElementById('dockerAdoptWorkdir');
if (name && name.value.trim() === (select.dataset.appliedName || '')) {
return { el: name, message: `"${name.value.trim()}" is the case you copied from — give this one a new name.` };
}
if (workdir && workdir.value.trim() === (select.dataset.appliedWorkdir || '')) {
return {
el: workdir,
message: 'Same container and same directory as the case you copied from — point this one at another directory.',
};
}
return null;
},
async _loadDockerContainerOptions() {
const list = document.getElementById('dockerContainerList');
if (!list) return;
@@ -3265,6 +3382,17 @@ Object.assign(CodemanApp.prototype, {
this.showToast('Enter the name of the running container to attach to', 'error');
return;
}
// A duplicate that still carries the source's name or directory: say so here,
// beside the field, rather than sending a request certain to come back refused.
const cloneIssue = adopting ? this.dockerCloneGuard() : null;
if (cloneIssue) {
this.showToast(cloneIssue.message, 'error');
const statusEl = document.getElementById('dockerLinkStatus');
if (statusEl) statusEl.textContent = cloneIssue.message;
cloneIssue.el.focus();
cloneIssue.el.select?.();
return;
}
try {
if (statusEl) {
+3
View File
@@ -408,6 +408,8 @@ Object.assign(CodemanApp.prototype, {
// header), so the row is hidden elsewhere rather than offering a toggle that
// changes nothing. Default ON — only an explicit false turns it off.
document.getElementById('appSettingsLineageLines').checked = settings.sessionLineageLines ?? defaults.sessionLineageLines ?? true;
// Auto-name sessions: synced, default OFF (opt-in; only an explicit true enables).
document.getElementById('appSettingsAutoNameSessions').checked = settings.autoNameSessions === true;
const lineageItem = document.getElementById('appSettingsLineageLinesItem');
if (lineageItem) lineageItem.style.display = MobileDetection.getDeviceType() === 'desktop' ? '' : 'none';
document.getElementById('appSettingsMobileOverview').checked = settings.mobileOverviewEnabled ?? defaults.mobileOverviewEnabled ?? false;
@@ -2111,6 +2113,7 @@ Object.assign(CodemanApp.prototype, {
showRedrawButton: document.getElementById('appSettingsShowRedrawButton').checked,
mobileOverviewEnabled: document.getElementById('appSettingsMobileOverview').checked,
sessionLineageLines: document.getElementById('appSettingsLineageLines').checked,
autoNameSessions: document.getElementById('appSettingsAutoNameSessions').checked,
showSessionButton: document.getElementById('appSettingsShowSessionButton').checked,
showAwayDigestButton: document.getElementById('appSettingsShowAwayDigestButton').checked,
showCronButton: document.getElementById('appSettingsShowCronButton').checked,

Some files were not shown because too many files have changed in this diff Show More