Compare commits

..
Author SHA1 Message Date
github-actions[bot]Claude Opus 5github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
20fc7b3c3d chore: version packages (#447)
* chore: version packages

* chore: sync the CLAUDE.md version line to 1.30.0

The changesets bot does not touch this line, and pushing it to master
after merging the version PR starts a second Release run that has raced
the first before. Riding the bot's own branch keeps it to one push.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
2026-09-18 14:09:55 +02:00
Codeman maintainer 0e1191b774 chore(changeset): trim the contributor entries and add the 1.30.0 thanks
Changeset text becomes user-facing CHANGELOG, so the #429 entry is cut
from five bullets of internal bash-array detail down to what the change
does for someone running the installer, as promised on the PR. The #441
entry loses its em-dashes, which are not house style. Adds an entry for
the maintainer fixes applied while landing #442, and the Thanks section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:55:22 +02:00
Codeman maintainer bb8ada7e5f fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does.

1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred
   it from the name: a session the user renamed by hand to something shaped
   like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the
   next prompt overwrote their name. The route persists right after, so the
   loss went to disk. `restoreMuxSessions()` already passes it.

2. The already-live sets were snapshotted once before a loop that awaits a
   real `startInteractive()` per entry, so by the tenth entry the snapshot
   was tens of seconds old and a conversation resumed by hand from the
   Resume list in that window was invisible to it: two panes on one
   transcript, the exact thing the check exists to prevent. Both sets are
   now read per iteration, and the late case is spent rather than re-offered
   for the same reason the batch case is.

3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path.
   The stamp predates the reboot and the pane is new, so honouring it meant
   one click had every restored session type `continue` into itself about a
   minute later, unattended, against the route header's own promise that a
   restored session comes back idle and disarmed. The setting stays ENABLED,
   so it re-arms on the next real limit message. A Codeman restart still
   re-arms from the stamp, because the limit footer will not reprint on its
   own; the new option exists only to tell the two paths apart.

4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()`
   and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession`
   performs that it was missing. Cosmetic, but a run left open reads as
   still going in the away digest.

5. A restored claude session gets `seedAgentSessionPreamble()` like both
   create paths, so the agent skill's bootstrap stays a two-line loader.

6. The heuristic's container comment was wrong in one direction and quiet
   about the real gap: after a genuine host reboot a containerized Codeman
   sees the host's short uptime and the banner does appear. What it cannot
   see is a container-only restart, which is where this would help most.

7. The banner is hidden in a solo window, which shows one session and has
   no tab strip to put restored ones in.

Also reverts 17 of the 18 hunks in docs/api-reference.md, which were
Prettier reformatting of prose the PR does not otherwise touch (docs/ is
outside the format glob), keeping only the Reboot restore section and
repairing the two continuation lines that reformat de-indented; renumbers
reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which
loads after it; and gives the feature its CLAUDE.md entry plus a route
test for the multi-user workspace-forbidden branch, the only new rule that
had nothing behind it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:46:04 +02:00
Codeman maintainer ea5323d990 test(input): pin the batched commit-plus-Enter ordering #441 fixes
The unit harness proves WHICH candidate gets forwarded; the ordering is
the half that shipped the bug, and only a real xterm shows it. The new
browser case dispatches the character's keydown, its composed insertText
and Enter's keydown in ONE page task, the shape an Android soft keyboard
delivers through a single InputConnection transaction, and asserts what
reaches the send path.

Verified in both directions on this machine: with the drain in place the
wire is `o\r`; with the drain removed (master's behaviour) it is `\r` and
the character is gone entirely, because by the time the zero-delay timer
runs xterm has emitted the `\r` and bumped the canonical counter past the
candidate's snapshot, so the candidate stands down. The other four cases
pass in both states.

CLAUDE.md now names the decision point, what it costs (a keydown decides
with less evidence than the timer did) and why that is safe for Enter,
and says that the pin lives in a suite the CI gate does not run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:42:44 +02:00
Codeman maintainer dee674d3e2 fix(install): point the launcher-only caveat at the thing that resolves it
The caveat #429 added ends with "see the docs above", and "the docs
above" is CLI_DOCS[$i], which for DeepSeek is the upstream harness repo.
Per docs/deepseek-integration.md the harness ships only the web,
headless and base profiles, so following that link and running
`npm install -g @deepseek-ai/dsh` leaves the reader exactly where the
caveat is warning them about: a dsh that cannot drive a pane. What
actually resolves it is Codeman's own Run dropdown, which offers
"DeepSeek: add a terminal profile..." and installs one in a click.

The new wording stays generic for any future launcherProfile entry,
since Codeman is the thing being installed at all three call sites.

Also flips one word in the generator: the comment said "see
installCommandFor below" and that function is defined above it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:41:45 +02:00
Ark0N f32c4f60d5 Merge pull request #442 from irisitymichaelgrundberg/feat/restore-sessions-after-reboot
feat(sessions): offer to rebuild the sessions a host reboot destroyed
2026-09-18 13:41:24 +02:00
Ark0N 9a503872d9 Merge pull request #441 from shenlvkang-collab/fix/android-last-char
fix(input): deliver a recovered keystroke before the Enter that submits it
2026-09-18 13:41:21 +02:00
Ark0N 9d7b29d899 Merge pull request #429 from opticon454/chore/cli-catalog-followups
chore(cli-registry): clean up dead code and stale claims left after #380
2026-09-18 13:41:14 +02:00
Ark0N ff8dc92187 Merge pull request #424 from Ark0N/fix/terminal-history-anchor-after-parse
fix(terminal): restore the history anchor after xterm parses, not before
2026-09-18 13:41:08 +02:00
Codeman maintainer 1f61d21298 docs: correct six stale counts and claims in CLAUDE.md
Each of these was measurable and wrong: the CI note listed 5 excluded
Playwright tests where config/test-suites.ts has 9, never mentioned the
packages/xterm-zerolag-input run that follows the gate, and never
mentioned wiki-sync.yml at all; the format glob note omitted that lint
covers only src/**/*.ts; app.js is ~6.9K lines, not ~6.7K, and
voice-pcm-worklet.js is fetched from JS rather than sitting in the load
order; src/config/ holds 23 files plus the cli-registry/ subdir, not 21,
and nothing said that the repo-root config/ is a different directory;
the route count is ~232 with cases at 34, not ~228 with cases at 30.

Also adds the pointer to docs/wiki/ as the user-facing manual, which the
header describes every other doc surface but not that one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:41:02 +02:00
Michael GrundbergandClaude Opus 5 62ceb4e87b fix(sessions): correct what the missing-pid rule actually recognises
A second real reboot disproved the mechanism the previous commit was built
on. Typing `/exit` does not persist `pid: null`, and the session was
restored anyway.

The pid a session record carries is its `tmux attach-session` process, not
the agent. `/exit` ends the CLI inside the pane, `remain-on-exit` keeps the
pane, and the attach process stays alive throughout — so Codeman's PTY never
exits, no exit handler runs, and the record keeps both its pid and
`status: 'idle'`. The lifecycle log for the session that came back shows
created, started, stale_cleaned and recovered, with no exit event at all,
which is the proof: Codeman never learned the agent was gone.

So nothing durable distinguishes an exited agent from a session that was
idle when the power went, and this pass restores both. Ark0N/Codeman#446 is
about making Codeman notice the dead pane; contrary to what the previous
commit's message claimed, this genuinely does wait on that. Until a record
can say the agent is gone, the user dismisses or closes those sessions.

The rule itself is kept, because a record with no attach process does
describe a session that never started or whose pane died outright, and
refusing it is right. Only its documentation was wrong. The module header,
the branch comment and the test names now say what it recognises instead of
claiming the case it cannot see.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 19:49:06 +02:00
Michael GrundbergandClaude Opus 5 5108a24bf0 fix(sessions): never restore a session whose agent was already exited
Found by a real reboot, which is the first thing to catch it. Typing `/exit`
ends the CLI process and leaves the session record behind, and the
process-exit handler persists `pid: null` with `status: 'idle'` before
anything else runs. By status alone that is indistinguishable from a session
sitting idle when the power went, so the boot pass offered those sessions
back and a click spawned the agents the user had deliberately closed — the
exact case the eligibility rule exists to exclude.

The absent pid is what tells the two apart, and the plan step now refuses a
record without one, under its own `not-running` reason so the boot log says
why. On a healthy board every running session carries a pid; a record with
none describes an agent that is already gone.

Deliberately the conservative direction. A session that somehow persisted no
pid while genuinely running is not offered, and its conversation stays
reachable from the Resume list, which is where every session would be
without this feature. The opposite error spawns processes nobody asked for.

Ark0N/Codeman#446 covers the dead panes those exits leave behind, but this
does not wait on it: the rule belongs here whether or not the record's shape
changes later.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 18:46:15 +02:00
Michael GrundbergandClaude Opus 5 5f55f9cb65 fix(sessions): never let the reboot-restore plan fail recovery
The plan build runs inside the try that decides whether restoreMuxSessions()
succeeded, so a throw would be caught there, report restoration as failed,
and block the stale cleanup and layout reconciliation that follow. An
optional convenience would then break the recovery it exists to help. It is
guarded on its own now: the correct way for this to fail is an offer nobody
gets.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 14:11:58 +02:00
Michael GrundbergandClaude Opus 5 18ab2ab595 docs(sessions): correct what a failed rebuild is actually likely to be
Ran the feature against a real server for the first time, on an isolated
instance, and two claims in the code turned out to be wrong.

A rebuild that fails after the session is registered was documented as
commonly caused by a CLI binary missing from a freshly booted machine's
PATH. It is not: the resolver finds its binary by absolute path, so PATH
never enters into it, and a server started without claude on PATH restored
every session normally. Nor does an un-enterable workspace fail — tmux falls
back to another directory and the pane comes up there. Neither obvious cause
throws, so the discard path is defended rather than expected, and the
comments now say that instead of naming a cause that cannot happen.

The four review rounds that shaped this path all reasoned about a trigger
none of them could test. The path itself is still worth having, since a mux
failure would reach it, but its comments should not claim a likelihood the
machine disagrees with.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 08:26:29 +02:00
Michael GrundbergandClaude Opus 5 39976041e0 fix(sessions): let a dismiss reach the entries a restore is holding
Fourth review of the reboot-restore branch, and the third to find a defect
in the previous round's fix. This one is the same shape as its predecessor:
a counter keyed on one thing, compared against a set keyed on another.

The generation counter was indexed by the entry's owner, while the in-flight
set holds the caller doing the restoring. Those are the same person exactly
when a user restores their own sessions, which is every case the tests
covered. The route deliberately supports the other case: an admin may spend
another user's entries. So when an admin restored Bob's sessions and Bob
dismissed the banner, nothing matched, the entries came back, and a plan Bob
had explicitly dismissed was re-armed for another twenty-four hours.

Rather than reconcile the two key spaces, the counter is gone. `take()` now
parks the entries it hands out, remembering which caller is spending them,
and they stay parked until that restore ends. A dismiss filters the parked
entries by `canAccess(entry.owner)` — the same predicate it already applies
to the plan — so it reaches them wherever they are. `releaseFlight()` puts
back only what is still parked. Expiry and a fresh boot plan unpark
everything, for the same reason. There is one key space now, the entry's
owner, and the spender is only ever used to tell two concurrent flights
apart. That removes `generations`, `snapshotGenerations()`, `bump()`,
`bumpAll()` and the argument threaded through the route.

The discard grew the teardown it still lacked. A rebuild can fail after
startInteractive() resolved, and a restored workspace still carries
Codeman's hooks, so the CLI can post a hook event within milliseconds; the
transcript watcher that starts from it, the attachment registry, the wait
registry and the approvals inbox all outlive the listeners and would meet
the retry, which reuses the session id by design. Its steps also run in
reverse order now, so no live listener can reach a tracker that has already
stopped, and the mux kill has its own guard, because stop() kills the pane
in its last block after destroying four trackers.

Tests. The run-summary test named an interval and asserted a map entry, so
dropping stop() left it green; it now spies on stop(). Nothing pinned that
before-spawn must precede setupSessionListeners, which reads the flag that
phase restores, so swapping the two lines was silent; the ordering test now
includes the listener setup. The retry assertion was a tautology and now
asserts a different refs object. Both strengthened tests were verified by
reverting their fix. Two new tests cover the admin-restores-another-owner
cases this round was about. The server in the discard test is built once and
stopped, since its constructor registers handlers on module-level watchers,
and the workspace is removed through safeRmHomeTree.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 16:51:00 +02:00
Michael GrundbergandClaude Opus 5 71ed7b127c fix(sessions): make the discard a real inverse of the construction
Third review of the reboot-restore branch. The narrow discard the previous
commit introduced avoided everything cleanupSession() did wrongly, and in
dropping so much of it also dropped four things it had to keep.

The worst broke the retry the whole design rests on. setupSessionListeners()
returns early while sessionListenerRefs still holds the session id, and the
discard never cleared that entry. So the advertised flow — a rebuild fails
because the agent binary is missing, the user fixes their PATH and clicks
again — reused the same id, wired no listeners at all, and produced a tab
that never showed output, never updated its status and never persisted. That
is worse than the leak the discard was added to prevent. Three more
registrations leaked with it: a RunSummaryTracker and its interval, an image
watcher on the workspace, and the Ralph fix-plan watcher. The discard now
undoes each registration setupSessionListeners() makes, in its order, and
the per-session custom-model config directory, which holds the endpoint's
API key literally and which nothing else would ever remove.

The image-watcher flag was restored after the code that reads it, so a
session came back reporting the feature as on with nothing watching. It
moves to the before-spawn phase, and that phase now runs before the
listeners rather than after them.

The generation counter that lets a mid-restore dismiss win was global while
clear() is ownership-scoped, so one user's dismiss discarded another user's
unspent entries, permanently, because nothing rebuilds an in-memory plan. It
is now per owner. Bumping only the owners of entries the dismiss removed was
not enough either: take() has already emptied the plan by then, so a dismiss
landing mid-restore saw nothing of that owner's to remove and invalidated
nothing. The owners that matter are those with a restore in flight, filtered
by what the dismissing user may access, and that is what clear() now bumps.
Plan expiry bumps too, so a restore straddling the 24-hour boundary cannot
hand entries back and give an expired plan another full day.

Tests. discardPartiallyBuiltSession had no test at all: the only
implementation any test ran was the mock's one-line stub, which is why every
defect above was invisible. test/discard-partially-built-session.ts drives
the real WebServer, and the retry assertion fails if the listener refs are
left behind — verified by reverting the fix. The dismiss-race test drove the
registry by hand, so deleting the route's generation argument left it green;
it now goes through the route, and two further tests cover the multi-user
cases.

The mock context has now gone stale twice, because route tests pass it as
`ctx as never` and tsconfig.json includes only src, so nothing ever compares
it to the ports. A type-level guard is therefore inert — I wrote one and
confirmed it never fires. test/mocks/mock-route-context-completeness.ts
compares the mock's keys against WebServer.createRouteContext() at runtime
instead, and names what is missing.

Also: the API reference now says workspace-forbidden is judged against the
owner's grant, the banner's module header no longer claims Restore always
dismisses it, and the detail span gets the same min-width: 0 the phone rule
already needed.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:34:47 +02:00
Michael GrundbergandClaude Opus 5 fa52753e8b fix(sessions): undo a failed rebuild without deleting the user's data
A second review of the previous commit found that its own repair for the
session leak introduced three defects, all from reaching for
cleanupSession() to undo a half-built session. That function is the
user-initiated delete, not an undo.

It banked the session's historical token and cost totals into the lifetime
figures, and a reboot never runs cleanup, so those totals had never been
counted before; every failed rebuild added them again. It saw the pin that
had just been restored and demoted the record to `stopped`, which this pass
reads as the durable marker of a deliberate kill, so a pinned session whose
rebuild failed became permanently unrestorable. And it recursively removed
`.claude-images` from the working directory, which belongs to the workspace
rather than to the session, so a failed rebuild destroyed the pasted images
of any other live session in that repo.

discardPartiallyBuiltSession() now undoes only what the construction did:
the map entry, the tab-layout slot, the listeners and any pane the launch
created before throwing. The persisted record, the lifetime totals, the
Ralph state and the workspace's files are left alone.

Re-applying the persisted state also splits in two, which removes the first
two defects at the root rather than only at the call site. The half that
shapes the pane, the custom-model environment and the nice priority, still
runs before the spawn. The half that is the session's own history now runs
after it, so a session whose pane never started carries no totals and no pin
for anything downstream to misread.

The rest of that review. The multi-user workspace confinement re-check read
the requesting user's grant, and returns true for an admin, so the case its
own comment described was the one it missed; it now resolves the entry
owner's grant through isWorkingDirAllowedForUsername, the way cron does. A
forbidden workspace goes back on offer, matching both the registry's stated
contract and the API reference. The client re-reads the plan after a restore
instead of blanking the banner, so entries the server put back stay
reachable, and a 409 now says a restore is already running rather than
reporting a failure. A dismiss arriving mid-restore wins, through a
generation counter the route carries across its take. The re-application
also restores the tab colour, the image-watcher flag and the original
pinnedAt, via a new Session.restorePin that does not re-stamp the pin time.
The phone breakpoint gains min-width: 0, without which a nowrap flex item
never shrinks and the buttons still overflow, and it folds into the existing
phone block.

Ralph's loop configuration still does not survive a restore, because
toState() reads it off a live tracker and there is no way to keep it without
arming the loop. The method now says so rather than leaving it implied.

Tests. The capacity test could not fail on the property it existed for: it
filled the board past the cap before the loop, so a single pre-loop check
would have passed it. It now leaves one seat, so only a per-iteration check
restores exactly one entry. New tests cover the ordering around the spawn,
a throw before the loop returning the whole plan and releasing the flight,
the dismiss-during-restore race, and that the failure path calls the narrow
discard rather than the delete. The shared mock context gains the port
method it was missing, which is what made the first run of these tests fail
for the wrong reason.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:04:35 +02:00
Michael GrundbergandClaude Opus 5 fbede5cd2a fix(sessions): act on the dual review of the reboot-restore route
Fifteen findings from two independent reviews of #442, three of them
blocking. Every one is addressed here.

The three blockers all sat in the restore route. A rebuild that threw after
addSession left a registered session with no pane behind it, visible on the
board, holding a layout slot and written to state.json, with its plan entry
already spent; the catch now cleans the session up and puts the entry back.
The loop checked neither the global nor the per-user session cap, so one
click could take a board past a documented limit; capacity is now re-checked
per iteration, because the loop is itself creating the sessions it counts.
Worst of the three, a rebuilt session carried none of the state its
constructor has no parameter for and then persisted itself over the record
that held it, zeroing token and cost totals and dropping the pin. The pin
matters most: pruning keeps a record only while it is pinned, so discarding
it handed the record to the next stale sweep. A new
reapplyPersistedSessionState() on the session port restores the pin, the
token totals, auto-compact, auto-clear, auto-resume, nice priority, the
flicker filter and the custom-model selection, and it runs before both
startInteractive and the first persist.

The rest, in the order they bite a user. Every rebuild failure was reported
as workspace-missing, so the banner told users their repo was gone when the
agent had simply failed to start; there are now distinct reasons, and the
toast names each one. The client read restored and skipped off the outer
response object rather than through the uniform envelope, so every count
came back zero and neither toast ever fired. A board left open across the
reboot never learned an offer existed, because the banner was seeded only on
the page-load path; it now re-reads on every SSE init. The workspace check
was existence-only, skipping the multi-user confinement that the create
route applies, so a withdrawn grant would not be noticed. The banner had no
phone breakpoint while its text was nowrap and its buttons could not shrink.

Smaller: a missing workspace is now re-offered rather than dropped, while an
already-open conversation is dropped rather than re-offered forever; a throw
anywhere in the route returns the unspent entries instead of discarding the
plan; the single flight is keyed by owner, since take() already stops two
callers receiving one entry; the env clamp's header no longer claims a
protection it cannot provide on this path today, and names the check that
does bite; the three endpoints are documented in docs/api-reference.md; and
the module header now says that os.uptime() reads the host's clock, so the
feature is effectively off inside a container.

The review also explained why the tests missed all of this: they proved the
construction claim through their own copy of the construction rather than
through the route, and the route tests used workspaces that did not exist,
so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts
mocks the Session module to drive the route's real path, and covers the
cleanup, the reason reported, the re-application ordering, the broadcast and
the caps. The mock route context gains the port method and the mux call the
route needs.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 11:44:37 +02:00
Michael GrundbergandClaude Opus 5 da933d70be feat(sessions): offer to rebuild the sessions a host reboot destroyed
A host reboot takes the tmux server down with it, so every pane dies,
reconciliation finds nothing to attach to, and the board comes up empty.
Picking yesterday's work back up meant finding each conversation in history
and resuming it by hand, one at a time.

The boot pass now works out what the reboot killed and leaves it on offer.
It runs inside restoreMuxSessions(), in the window where reconciliation has
reported the dead sessions and cleanupStaleSessions() has not pruned their
records yet, which is the only place the records can still be read. The
board shows a banner, and nothing is created until the user clicks it.

A click rather than an automatic restore is what makes the reboot heuristic
acceptable. The heuristic cannot tell a reboot from a crash that took tmux
down inside the same window, so it decides whether to ASK, never whether to
act: a wrong yes costs a line of text the user dismisses instead of N CLI
processes nobody asked for.

Four things are re-checked when the click arrives rather than trusted from
boot, because hours can pass and the board moves on. The owner's privilege
grant re-resolves through the env clamp. The workspace must still be on
disk. A conversation the user already resumed by hand from the Resume list
is skipped, since two panes running --resume on one conversation would
fight over the same transcript. Entries leave the plan synchronously before
the first await, and the route is single-flighted, so a double-click or two
devices cannot both reach the same entry.

A restored session comes back attached, idle and disarmed. Respawn
controllers and Ralph loops are deliberately not re-armed: a machine that
just came up is the worst moment to turn an autonomous run loose. Its
workspace hooks are installed by the restore route itself, because the
boot-time sweep sits behind a gate that is false after a reboot and has
finished long before the click; without them a session goes silently blind,
with no stop or idle events for respawn, no Approvals Inbox item and no red
tab on a blocking dialog. Stats collection starts the same way.

The pane is new, so the conversation continues and the terminal scrollback
does not. The banner says so rather than letting an empty pane read as a
broken restore.

The plan lives in memory only. A server restart drops it, which costs the
convenience this adds and never the conversation: the conversation is the
transcript under ~/.claude/projects, which the Welcome screen's Resume list
and the Session Manager already read, so a dropped plan returns the user to
resuming by hand.

clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the
question it answers is about session privilege rather than about HTTP and
it now has a caller outside the route layer. Its test hook stays re-exported
from session-routes.ts.

Claude sessions only for this pass. The other CLIs name their thread in
their own config object, which this does not thread through yet. Remote and
docker sessions are skipped on purpose, because both need another host or a
container to be up and a freshly booted machine cannot promise either.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 08:05:55 +02:00
codeman-localandClaude Opus 5 a1c35da0d8 fix(input): deliver a recovered keystroke before the Enter that submits it
Every message typed on an Android phone lost its last character.

An Android soft keyboard commits the last typed character and sends the
Enter key in ONE InputConnection transaction, so the committed-text
`input` event and the Enter keydown are both processed before any
zero-delay timer runs. The orphaned-input recovery from #388 resolved
its candidate only on such a timer, and that lost the character twice
over:

  * ORDER — xterm emits `\r` synchronously from the Enter keydown, and
    the local-echo composer submits `pendingText` right there. The
    recovered character arrived one macrotask too late to be part of the
    prompt.
  * LOSS — that same `\r` bumps the canonical counter, so by the time
    the candidate resolved, `canonicalCount > snapshot` read as "xterm
    spoke for this keystroke" and stood the recovery down. The character
    was not merely late, it was dropped.

Drain pending candidates synchronously at the next keydown instead, from
xterm's custom key handler, which runs before xterm processes that key.
The counter then still holds the value it had while the candidate's own
keystroke was current, so the stand-down decision is made against the
right keystroke, and the recovered byte reaches the composer ahead of
whatever the new key emits. The timer stays as the fallback for a
keystroke with no key after it.

Physical keyboards are unaffected: there the timer has already resolved
the candidate long before the next key arrives.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 09:26:30 +08:00
Codeman maintainer bd286bf502 docs(wiki): catch the manual up to 1.29.0 and add the three run modes it never had
The wiki was written for seven run modes and never received Grok Build, DeepSeek
Harness or OMP. They now appear everywhere the others do: the modes table and
per-CLI notes, install commands, environment prefixes, the Quick Start table, the
requirements rows, the vocabulary, and every "seven modes" count.

The 1.27 to 1.29.0 changes land on the pages that own them: attaching a case to an
existing container, multi-case adoption and the copy-a-case picker (Docker Cases);
file reads over ssh in remote cases and what stays unavailable (Remote SSH Sessions,
Working With Files, Security); single-page app routing, frame recovery, localhost
links as tabs and the egress guard (Web Tabs); DeepSeek as the one non-Claude mode
with real stop/blocked signals and Approvals items, Codex's own work detection,
last-response, the model-endpoint routes and refreshed counts (HTTP API, Driving
From An Agent, Hooks, Notifications, Keeping Agents Running, Core Concepts);
Shift+drag, right-click copy, Auto Copy, the Ctrl+Z guard, font weight, the vertical
rail and its activity sort (Keyboard Shortcuts, Input And Voice, The Dashboard,
Settings Reference); the 600px phone cutoff, Codex shift arrows and iPhone Duo
(Mobile Guide); the Docker Compose route and its update rule (Installation, Running
As A Service); four new symptom entries and a "which CLIs" question (Troubleshooting,
FAQ).

Custom model endpoints are deliberately left to #430, which adds that page and edits
Agent CLIs, Settings Reference and the sidebar; these edits stay out of the regions
#430, #428 and #376 touch, and all three still merge cleanly on top.

Both READMEs: the web-tab menu entry is labelled "Add URL" in the UI, not
"Add dashboard".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 19:05:59 +02:00
DevvynandClaude Sonnet 5 3f2928ae73 chore(cli-registry): clean up dead code and stale claims left after #380
Addresses the "left as they are"/"worth knowing" items Ark0N named when
merging #380 (the CLI-catalogue-driven install.sh + Docker agent image
PR), none of which were correctness-blocking but all of which were real:

- Removed install.sh's dead _cli_index/check_cli/get_cli_path helpers:
  the catalogue-driven menu and hints stopped calling them and nothing
  else ever did.
- The generator no longer emits CLI_KIND/CLI_NPM, two bash arrays
  install.sh never read (the .mjs/docker-hosts.ts producers already
  read the JSON catalogue's kind/npmPackage fields directly, so only
  the bash copies were dead).
- detect_all_clis now skips a disabled entry's probe entirely instead
  of running it and filtering the result downstream. No stock entry
  ships disabled today, so this closes a latent inefficiency before it
  is a latent bug rather than fixing an observed one.
- The install hint for a launcherProfile entry (DeepSeek today) now
  explains in one line why it's a docs link and not a command: its own
  docs page documents `npm install -g @deepseek-ai/dsh`, which installs
  the launcher only and can't drive a pane, the exact trap the menu
  already avoids by withholding the command. Driven by a new generated
  CLI_LAUNCHER_ONLY array (from discovery.launcherProfile), not an id
  check, so any future launcherProfile entry gets the same caveat free.
- Corrected the non-interactive-default comment: on a wget-only host,
  Claude's curl one-liner is filtered out of the offered list first, so
  the default becomes whichever npm-based entry sorts earliest instead
  (Codex today), not always Claude. Behaviour is unchanged — it was
  already printed, never silent — only the comment overclaimed.

Tests: extended test/install-sh-invariants.test.ts with a positive
guard for the new array and the trimmed array list, a negative guard
that CLI_KIND/CLI_NPM/the three dead helpers cannot come back, and two
real-bash tests (driven the same way the existing skip-menu tests are)
proving a disabled entry is genuinely never probed rather than merely
filtered after the fact.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011WzDjJnbK7zug8iQWnCc9z
2026-09-15 09:10:37 +08:00
Codeman maintainer de864e7d63 fix(terminal): restore the history anchor after xterm parses, not before
flushPendingWrites() captured the viewport of a user who was reading
scrollback, called terminal.write(), and restored the anchor on the next
line. xterm parses on its own schedule, so at that point the buffer has not
moved: the guard `viewportY !== preserveViewportY` was false, scrollToLine
was never called at all, and the Codex redraw landed a tick later and took
the viewport to the live bottom with nothing left to pull it back. Scrolling
up during a stream still got dragged down, which is what #358 reports, and a
refresh was the only way back to a coherent view.

The restore moves inside xterm's write callback, the first moment the
redraw's effect exists, and runs before _scheduleTerminalWriteFlush() so a
deferred remainder re-captures the restored anchor rather than the bottom.

Two things follow from it running later:

- A live anchor now wins over the sticky scroll-to-bottom. The two are
  captured at different moments (_wasAtBottomBeforeWrite at the frame's
  first batchTerminalWrite, the anchor at flush time), so a scroll-up in
  between leaves both set, and running both would jump to the bottom and
  come back a frame later instead of staying put.
- The anchor is dropped if the active session changed or a buffer load
  started while the write was in flight. It indexes the buffer it was
  captured from, and selectSession() resets the terminal and chunk-loads a
  different scrollback.

The existing regression passed throughout, because its write mock moved the
viewport synchronously, which real xterm never does. The harness now models
an asynchronous parse (redraw lands, then the callback fires), and all five
of the anchor tests fail against the old code.

Fixes #358

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-15 00:10:41 +02:00
62 changed files with 3555 additions and 224 deletions
+1 -1
View File
@@ -10,7 +10,7 @@
"name": "codeman",
"source": "./plugins/codeman",
"description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.29.1",
"version": "1.30.0",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
+18
View File
@@ -1,5 +1,23 @@
# aicodeman
## 1.30.0
### Minor Changes
- da933d7: Offer to rebuild the sessions a host reboot destroyed. A reboot takes the tmux server down with it, so every pane dies and the board comes up empty. Codeman now works out what was running, and the board offers to restore it behind a click. The conversations come back; the terminal scrollback does not, and the banner says so.
### Patch Changes
- a1c35da: Stop a phone keyboard losing the last character of every message it sends. Android soft keyboards commit the last typed character and send the Enter key in one InputConnection transaction, so the `input` event and the Enter keydown are both processed before any zero-delay timer runs. The orphaned-input recovery from #388 only resolved its candidate on such a timer, and lost it both ways: xterm emits `\r` synchronously from the Enter keydown, so the local-echo composer submitted the prompt before the recovered character existed, and that `\r` bumped the "did xterm speak for this keystroke" counter, so the candidate then stood itself down and dropped the character outright. Pending candidates are now drained synchronously at the next keydown, from xterm's custom key handler, which runs before xterm processes that key, so the counter still holds the value it had while the candidate's own keystroke was current, and the recovered byte reaches the composer ahead of the Enter. Typing on a physical keyboard is unaffected: there, the timer has already resolved the candidate before the next key arrives.
- 3f2928a: The installer's hint for a launcher-only CLI (DeepSeek today) now says why it is a docs link rather than a command you can run, and points at the thing that resolves it: the package installs a launcher that still needs a terminal profile, and Codeman's Run menu can add one in a click. Driven by a generated `CLI_LAUNCHER_ONLY` flag rather than an id check, so it covers any future entry of that shape. Also removes three dead lookup helpers and two never-read generated arrays from `install.sh`, skips a disabled entry's probe instead of filtering it afterwards, and corrects a comment that claimed the non-interactive default is always Claude Code (on a wget-only host its curl one-liner is filtered out first).
- 0e1191b: Maintainer fixes applied while landing the above. A session restored after a reboot keeps the name you gave it (the rebuild dropped the field that records who named a session, so a hand-renamed session came back looking auto-named and the next prompt overwrote it), and no longer types `continue` into itself on its own: a pending auto-resume stamp from before the reboot is dropped rather than re-armed, since the pane is new and one click could otherwise arm several unattended prompts at once. Auto-resume itself stays on and re-arms on the next real usage-limit message. The restore offer is also hidden in a detached single-session window, which has no tab strip to put restored sessions in, and a conversation that goes live while an earlier session in the same batch is starting is no longer restored a second time.
- 0e1191b: ### Thanks
- @irisitymichaelgrundberg for the reboot-restore banner (#442), and for the three real reboots behind it rather than a mocked one.
- @shenlvkang-collab for tracking down why Android keyboards lost the last character of every message (#441), including the half where the character was not late but gone.
- @opticon454 for going back and closing out the loose ends left as "worth knowing rather than fixing" after #380 (#429).
- de864e7: Keep the terminal anchored where you are reading while an agent streams (#358). Scrolling up during a Codex response could still be dragged back to the live bottom by the next redraw: the flush captured the viewport before writing and restored it immediately after, but xterm parses asynchronously, so at that moment the buffer had not moved yet, the restore compared the anchor against itself and did nothing, and the redraw landed a tick later with nothing left to pull the view back. The restore now runs inside xterm's own write callback, which is the first point at which the redraw's effect exists, and it holds across consecutive and chunked redraws. It is dropped if you switch sessions or a history replay starts before the write parses, since the anchor indexes the buffer it was captured from.
## 1.29.1
### Patch Changes
+11 -9
View File
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -443,7 +443,7 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
- **Clone a GitHub repo as a case** — paste a repository URL into **Add Case → Clone Repo** and Codeman clones it into `~/codeman-cases/<name>` and registers it as a normal case, ready to run an agent in. It preflights the URL while you type (tells you whether it can be cloned anonymously and offers the repo's real branches and tags for the optional branch/tag field), fills the case name in from the URL, and lets you pick which CLI the Run button should use. Public repositories over `https://`; Codeman never collects or stores credentials
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, **Pi**, **Grok**, **DeepSeek Harness**, or **OMP** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*` vs `GROK_*`/`XAI_*` vs `DSH_*`/`DEEPSEEK_*` vs `OMP_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md), [`docs/pi-integration.md`](docs/pi-integration.md), [`docs/grok-integration.md`](docs/grok-integration.md), [`docs/deepseek-integration.md`](docs/deepseek-integration.md) and [`docs/omp-integration.md`](docs/omp-integration.md)
- **Custom model endpoints** _(new in 1.29.0, HTTP API for now)_ — point a session's CLI at any OpenAI-compatible endpoint instead of its native backend: a local llama.cpp, llama-swap, Ollama or vLLM box, or a cloud gateway such as Azure AI Foundry or OpenRouter. Save an endpoint once (`POST /api/model-endpoints`; its models are discovered from `/v1/models`), apply it to a session (`POST /api/sessions/:id/custom-model`), and the CLI restarts in place on that endpoint. Verified live for Claude, OpenCode, Pi, Grok and OMP; Codex, Gemini and DeepSeek have documented gaps, Antigravity has no mechanism. A toolbar picker is the follow-up. See [`docs/custom-model-endpoints.md`](docs/custom-model-endpoints.md)
- **Web tabs** — open Grafana, Uptime Kuma, a Vite dev server or any dashboard URL as a tab beside your sessions (Run dropdown → **Web / URL** → **Add dashboard**). Dashboards are proxied through Codeman's own origin, so an `http://` target works from a phone over HTTPS and through the tunnel, single-page apps route on their own paths, and a frame that reloads recovers itself. A `localhost` link an agent prints opens as a web tab automatically. See [`docs/web-tabs.md`](docs/web-tabs.md)
- **Web tabs** — open Grafana, Uptime Kuma, a Vite dev server or any dashboard URL as a tab beside your sessions (Run dropdown → **Web / URL** → **Add URL**). Dashboards are proxied through Codeman's own origin, so an `http://` target works from a phone over HTTPS and through the tunnel, single-page apps route on their own paths, and a frame that reloads recovers itself. A `localhost` link an agent prints opens as a web tab automatically. See [`docs/web-tabs.md`](docs/web-tabs.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container, or attach a case to a container you already run; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host; file previews and downloads come over the same ssh connection. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
+1 -1
View File
@@ -445,7 +445,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
- **把 GitHub 仓库克隆成 case** —— 在 **Add Case → Clone Repo** 里粘贴一个仓库 URL,Codeman 会把它克隆到 `~/codeman-cases/<name>` 并注册为普通 case,随时可以跑智能体。输入时它会预检 URL(告诉你能否匿名克隆,并为可选的分支/标签字段提供仓库真实的分支与标签),从 URL 里填好 case 名,还让你选 Run 按钮该用哪个 CLI。支持 `https://` 的公开仓库;Codeman 绝不收集或保存凭据
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity**、**Gemini**、**Pi**、**Grok**、**DeepSeek Harness** 或 **OMP**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*`、`GEMINI_*`/`GOOGLE_*`、`PI_*`、`GROK_*`/`XAI_*`、`DSH_*`/`DEEPSEEK_*` 与 `OMP_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)、[`docs/pi-integration.md`](docs/pi-integration.md)、[`docs/grok-integration.md`](docs/grok-integration.md)、[`docs/deepseek-integration.md`](docs/deepseek-integration.md) 与 [`docs/omp-integration.md`](docs/omp-integration.md)
- **自定义模型端点**(1.29.0 新增,目前仅 HTTP API)—— 让某个会话的 CLI 指向任意 OpenAI 兼容端点,而不是它自己的官方后端:本地的 llama.cpp、llama-swap、Ollama 或 vLLM 机器,也可以是 Azure AI Foundry、OpenRouter 这类云端网关。端点只需保存一次(`POST /api/model-endpoints`,模型列表从它的 `/v1/models` 自动发现),再应用到会话(`POST /api/sessions/:id/custom-model`),CLI 就会在原地重启并接上该端点。Claude、OpenCode、Pi、Grok 与 OMP 已实测通过;Codex、Gemini 与 DeepSeek 存在已记录的缺口,Antigravity 没有可用机制。工具栏选择器是下一步。详见 [`docs/custom-model-endpoints.md`](docs/custom-model-endpoints.md)
- **Web 标签页** —— 把 Grafana、Uptime Kuma、一个 Vite 开发服务器或任何仪表盘 URL 作为标签页打开在会话旁边(Run 下拉菜单 → **Web / URL** → **Add dashboard**)。仪表盘通过 Codeman 自己的源代理,因此 `http://` 目标在手机上走 HTTPS 也能用、走隧道也能用;单页应用能在自己的路径上正常路由,页面自己重载后也能自行恢复。智能体打印出的 `localhost` 链接会自动以 Web 标签页打开。详见 [`docs/web-tabs.md`](docs/web-tabs.md)
- **Web 标签页** —— 把 Grafana、Uptime Kuma、一个 Vite 开发服务器或任何仪表盘 URL 作为标签页打开在会话旁边(Run 下拉菜单 → **Web / URL** → **Add URL**)。仪表盘通过 Codeman 自己的源代理,因此 `http://` 目标在手机上走 HTTPS 也能用、走隧道也能用;单页应用能在自己的路径上正常路由,页面自己重载后也能自行恢复。智能体打印出的 `localhost` 链接会自动以 Web 标签页打开。详见 [`docs/web-tabs.md`](docs/web-tabs.md)
- **Docker 会话** —— 在隔离且加固的容器中运行 case。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一 case 的多个会话共享一个容器,也可以把 case 挂到你已经在跑的容器上;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
- **远程 SSH 会话** —— 把 case 指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话;文件预览与下载走同一条 ssh 连接。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
+42
View File
@@ -479,6 +479,48 @@ re-captured, or the item acknowledged), `approval:resolved` (`{ id, sessionId, k
`resolution` one of `answered | resolved_in_terminal | superseded |
session_ended | dismissed | expired`).
## Reboot restore
A host reboot takes the tmux server down with it, so every pane dies and the
board comes up empty. At boot Codeman works out which sessions the reboot
destroyed and holds that plan in memory, and these endpoints let a client offer
it to the user. Nothing creates a pane until the user asks: the boot-time reboot
heuristic decides whether to ASK, never whether to act.
Claude-mode sessions only (others carry their conversation id in their own
config object); remote and docker sessions are never offered, because both need
another host or container to be up. The plan is in-memory, so a server restart
drops it and the offer is gone; the conversations themselves are unaffected,
since they live in the CLI's own transcript store and stay reachable from the
Resume list. A plan nobody spends expires after 24 hours.
- `GET /api/v1/reboot-restore` → `{ sessions: RestorableSession[],
scrollbackRestored: false }`, ownership-scoped in multi-user mode.
`RestorableSession`: `{ id, name?, workingDir, mode, owner? }`. The persisted
record itself is never sent. `scrollbackRestored` is always `false` and exists
so a client states it: a restored session is a NEW pane, so the conversation
continues and the terminal history does not.
- `POST /api/v1/reboot-restore/restore` with `{ sessionIds?: string[] }` (omit
to restore everything the caller can see) → `{ restored: RestorableSession[],
skipped: { sessionId, reason }[] }`. `reason` is one of `workspace-missing`
(the directory is gone), `workspace-forbidden` (in multi-user mode it is
outside the workspace of the user the session belongs to, re-checked against
that owner's current grant rather than the caller's), `already-live` (the conversation is already
open, typically resumed by hand from the Resume list), `capacity-reached`
(the global or per-user session cap), or `rebuild-failed` (the agent would not
start, most often a CLI binary missing from the server's PATH).
`409 CONFLICT` when that caller already has a restore running. Entries are
removed from the plan before any pane is built, so a double-click cannot put
two panes on one conversation; anything that never became a pane goes back on
offer, except `already-live`, which cannot stop being true. A restored session
comes back attached, idle and disarmed: respawn controllers and Ralph loops
are never re-armed automatically.
- `POST /api/v1/reboot-restore/dismiss` → `{ dismissed: n }`. Drops the offer
for everything the caller can see.
Each rebuilt session also emits the ordinary `session:created` SSE event, so
clients other than the one that clicked pick it up without refetching.
## Read My Mind intent profiles
Per-case profiles of what the user is trying to accomplish: user/agent-stated
+85 -11
View File
@@ -1,9 +1,9 @@
# Agent CLIs
Codeman drives seven run modes: six agent CLIs plus a plain shell. This page covers picking
Codeman drives ten run modes: nine agent CLIs plus a plain shell. This page covers picking
one, setting it up, and the differences that actually change how you work.
## The seven modes
## The ten modes
| Mode | CLI | Get it |
| -------------------- | ---------------------------- | ---------------------------------------------------------------------- |
@@ -13,6 +13,9 @@ one, setting it up, and the differences that actually change how you work.
| **Gemini** | `gemini` | [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) |
| **Antigravity** | `agy` | [antigravity.google](https://antigravity.google) |
| **Pi** | `pi` | [pi.dev](https://pi.dev) |
| **Grok Build** | `grok` | [github.com/xai-org/grok-build](https://github.com/xai-org/grok-build) |
| **DeepSeek Harness** | `dsh` | [github.com/deepseek-ai/deepseek-harness](https://github.com/deepseek-ai/deepseek-harness) |
| **OMP** | `omp` | [github.com/can1357/oh-my-pi](https://github.com/can1357/oh-my-pi) |
| **Terminal / Shell** | your `$SHELL` | Already installed. |
Any combination works, including all of them. The run mode is chosen per session from the
@@ -47,8 +50,12 @@ If a CLI is installed but a Run button for it never appears:
precisely to avoid this; a hand-written plist or unit will not.
3. Restart the server after installing a new CLI.
`pi` is additionally version-probed rather than trusted by name, because `pi` is a generic
enough command that something else on your PATH may answer to it.
`pi`, `grok`, `omp` and `dsh` are additionally identity-probed rather than trusted by name:
`pi` and `omp` are generic enough that something else on your PATH may answer to them,
`grok` has npm squatters, and Debian ships an unrelated `dsh` (dancer's shell). Each has a
status endpoint (`/api/grok/status`, `/api/deepseek/status`, `/api/omp/status`) that reports
the path and version that actually resolved, so a misresolution is visible rather than
presenting as "the mode just does not work".
## Claude is the reference mode
@@ -62,15 +69,15 @@ output. The other CLIs expose no equivalent.
| Respawn cycling and unattended runs | Yes | Yes |
| Cron jobs | Yes | Yes |
| Docker cases, remote SSH cases | Yes | Yes |
| Precise idle detection (hook-driven) | Yes | Output-stabilization fallback, coarser |
| Precise idle detection | Yes | Codex: same screen check, via its own prompt and working line. DeepSeek: reports its state itself. Others: output stabilization, coarser |
| Auto-resume when a usage limit resets | Yes | No |
| Plan usage chip | Yes | No |
| Approvals Inbox | Yes | No |
| Approvals Inbox | Yes | DeepSeek yes; others no |
| Read My Mind | Yes | No |
| Ralph loop and its task tracker | Yes | No |
| Subagent and team windows | Yes | No |
| Model, effort, and ultracode controls | Yes | No |
| `stop` and `blocked` wait signals | Yes | 400 if you ask for them explicitly |
| `stop` and `blocked` wait signals | Yes | DeepSeek yes; elsewhere 400 if you ask for them explicitly |
| The bundled agent skill | Yes | No |
Everything that makes a session a session works everywhere. What is Claude-only is mostly
@@ -124,6 +131,11 @@ Two behaviours that are deliberate and worth knowing:
- **The wheel is not forwarded** into its transcript. Codex ignores the mouse reports
Codeman would send, so forwarding produced a dead wheel. Scrolling in a Codex session is
local scrollback.
- **Work detection is Codex's own.** Codex declares its `›` composer glyph and its
`esc to interrupt` working line, so it gets the same screen-checked idle detection Claude
does; before 1.26.1 every Codex session reported idle for its whole life. Codex
conversations also appear in Past Sessions and can be resumed, and on phones the keyboard
bar grows `⇧←` / `⇧→` for Codex's queued-message editing and prompt stack.
### Gemini
@@ -157,6 +169,60 @@ Pi needs the opposite instincts from every other CLI here.
Guide: [`docs/pi-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/pi-integration.md).
### Grok Build
xAI's `grok`, installed with `curl -fsSL https://x.ai/cli/install.sh | bash` into
`~/.grok/bin`. Codex-shaped on permissions and OpenCode-shaped on rendering:
- **Its bypass switch is `--always-approve`**, Grok's own `bypassPermissions` mode, and the
Run button sends it the way it sends Codex's. In multi-user mode a user without a grant
has it stripped.
- **Authentication is Grok's own**: browser OAuth on first run (a device-code screen inside
a Codeman pane), `grok login --device-auth` for headless hosts, or `XAI_API_KEY` as a
per-session environment override.
- It renders a full-screen TUI, so scrolling is local scrollback.
Guide: [`docs/grok-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/grok-integration.md).
### DeepSeek Harness
The mode wired least like the others, for two reasons worth knowing before you use it.
**`dsh` is a launcher, not an agent.** It boots a *profile*, and the three DeepSeek ships
(`web`, `headless`, `base`) cannot drive a terminal pane. So "installed" and "runnable" are
different questions: the Run menu offers **DeepSeek** only once a pane-capable profile
exists, and until then shows **DeepSeek — add a terminal profile…**, which installs the
community `dsh-tui` with one click (`pnpm` must be on PATH, because the launcher spawns it
directly).
**Permissions are an environment variable, not a flag.** The harness has no
skip-permissions switch. `DSH_PERMISSION_MODE` (`read-only`, `workspace-write`,
`danger-full-access`) is the whole control, and it is the one setting Codeman deliberately
carries as an environment variable, because the harness reads it as a soft boot-time
default. In multi-user mode a user without a grant is clamped to `workspace-write`.
The reward for the odd wiring: **DeepSeek is the one non-Claude mode with real signals.**
Its terminal front door reports idle, working and blocked to Codeman, so a DeepSeek
session gets precise idle detection, the `stop` and `blocked` wait signals, and Approvals
Inbox items. Answers are read from the harness's own transcript on disk rather than
scraped off the pane. The model is not a session setting; it is part of the profile.
Guide: [`docs/deepseek-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/deepseek-integration.md).
### OMP
Oh My Pi, installed with `curl -fsSL https://omp.sh/install | sh` into `~/.local/bin`.
OMP owns its auth, provider routing and approval mode entirely in `~/.omp`: there is no
Codeman-side login, key field, or bypass switch. Run `omp` once outside Codeman to finish
its own onboarding, and every session started through Codeman inherits that config. Its
documented default approval mode is `yolo`, so an OMP pane auto-approves tool use with no
flag from Codeman; change that in OMP's own config, not here.
OMP conversations appear in Past Sessions and can be resumed, and a respawn continues the
same conversation with `--continue`.
Guide: [`docs/omp-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/omp-integration.md).
### Terminal / Shell
A plain shell in a tmux session. No agent, no hooks, no idle detection.
@@ -180,9 +246,15 @@ respawns. Which variables are accepted depends on the mode:
| Gemini | `GEMINI_*`, `GOOGLE_*` |
| Antigravity | `ANTIGRAVITY_*` |
| Pi | `PI_*` |
| Grok | `GROK_*`, `XAI_*` |
| DeepSeek | `DSH_*`, `DEEPSEEK_*` |
| OMP | `OMP_*` |
Anything outside the allowlist is rejected at the schema. This is intentional: the allowlist
is one global list, so widening it for one CLI widens it for all of them.
is one global list, so widening it for one CLI widens it for all of them. In multi-user mode
the keys that could redirect a CLI's traffic or move its config home (`DSH_PERMISSION_MODE`,
`DSH_HOME`, `DEEPSEEK_BASE_URL`, `OMP_AUTH_BROKER_URL`, and the base URLs and config
directories of the others) are dropped for a user without the bypass grant.
Two things that deliberately do **not** travel as environment variables: **effort**, because
an environment variable hard-locks it and blocks `/effort`, and **model**, which is written
@@ -192,9 +264,11 @@ into the case's `.claude/settings.local.json` so that `/model` keeps working.
- **Claude Code** if you want every Codeman feature. Unattended overnight runs, usage-limit
auto-resume, the Approvals Inbox, and subagent visualization all assume it.
- **Codex, OpenCode, Gemini, Antigravity** when you prefer that agent or that model. You get
the session layer, respawn, cron, Docker, and remote SSH; you do not get the hook-driven
features.
- **Codex, OpenCode, Gemini, Antigravity, Grok, OMP** when you prefer that agent or that
model. You get the session layer, respawn, cron, Docker, and remote SSH; you do not get the
hook-driven features.
- **DeepSeek Harness** if you want DeepSeek's models with real status signals. It is the one
non-Claude mode that reports idle, working and blocked to Codeman itself.
- **Pi** if you want a fast, unsandboxed agent and you understand what project trust does.
- **Shell** for the times you want a terminal on your phone with no agent at all. It is a
genuinely useful mode, not a fallback.
+1 -1
View File
@@ -108,7 +108,7 @@ Conventions for wiki pages:
- Images are referenced from the main repository over raw URLs rather than being copied into
the wiki.
- Say what the default is, especially when it is off. Most of Codeman is opt-in.
- Label Claude-only behaviour every time it appears. Six of the seven run modes are not
- Label Claude-only behaviour every time it appears. Nine of the ten run modes are not
Claude.
## Conduct
+11 -9
View File
@@ -50,10 +50,10 @@ A session carries state the case does not:
## Run mode
The **run mode** is which CLI the session runs: `claude`, `opencode`, `codex`, `gemini`,
`antigravity`, `pi`, or `shell`. It is chosen at start and does not change afterwards; to
`antigravity`, `pi`, `grok`, `deepseek`, `omp`, or `shell`. It is chosen at start and does not change afterwards; to
switch, start another session.
Claude is the reference mode. Six of the seven are not Claude, and a number of Codeman
Claude is the reference mode. Nine of the ten are not Claude, and a number of Codeman
features are Claude-only for structural reasons rather than missing effort: they depend on
Claude Code's hook system or on parsing its terminal output. Every such feature is labelled
Claude-only where it appears, and [Agent CLIs](Agent-CLIs) lists them in one place.
@@ -68,8 +68,8 @@ Where a case runs is **separate from** which CLI it runs. There are three locati
| **Docker** | One long-lived container per case; sessions `docker exec` into it. See [Docker Cases](Docker-Cases). |
| **Remote SSH** | A durable tmux server on the remote host, fronted by a local pane running `ssh`. See [Remote SSH Sessions](Remote-SSH-Sessions). |
This matters because it is a common source of confusion: Docker is **not** an eighth run
mode. All seven run modes work in all three locations. A case is docker-backed or
This matters because it is a common source of confusion: Docker is **not** an eleventh run
mode. All ten run modes work in all three locations. A case is docker-backed or
ssh-backed; a session is claude or codex or shell.
**Web tabs** are the other thing that is not a session. A saved dashboard URL renders as a
@@ -155,9 +155,11 @@ report events back: a permission prompt appeared, the turn finished, the agent w
task completed. Those events drive tab alerts, the Approvals Inbox, notifications, and the
wait primitives.
This is why some features are Claude-only. The other CLIs have no equivalent hook system,
so for them Codeman falls back to watching terminal output, which is coarser: it can see
that something happened, not what it was.
This is why some features are Claude-only. The one partial exception is DeepSeek Harness,
whose terminal front door reports idle, working and blocked to Codeman over the harness's
own supervisor contract, so it gets the hook-driven signals without a hook file. The other
CLIs have no equivalent, so for them Codeman falls back to watching terminal output, which
is coarser: it can see that something happened, not what it was.
See [Hooks And Integrations](Hooks-And-Integrations).
@@ -167,7 +169,7 @@ See [Hooks And Integrations](Hooks-And-Integrations).
| --------------- | ---------------------------------------------------------------------------- |
| **Case** | Named working directory. |
| **Session** | One CLI in one tmux session. |
| **Run mode** | Which CLI: claude, opencode, codex, gemini, antigravity, pi, shell. |
| **Run mode** | Which CLI: claude, opencode, codex, gemini, antigravity, pi, grok, deepseek, omp, shell. |
| **Respawn** | Restarting the CLI on idle to keep an unattended run going. |
| **Ralph loop** | An autonomous single-session task loop. |
| **Orchestrator**| A phased plan driven across multiple agents. |
@@ -178,6 +180,6 @@ See [Hooks And Integrations](Hooks-And-Integrations).
## Read next
- [The Dashboard](The-Dashboard) - what the UI is showing you.
- [Agent CLIs](Agent-CLIs) - the seven run modes in detail.
- [Agent CLIs](Agent-CLIs) - the ten run modes in detail.
- [Keeping Agents Running](Keeping-Agents-Running) - respawn, idle detection, usage limits.
- [`docs/architecture-invariants.md`](https://github.com/Ark0N/Codeman/blob/master/docs/architecture-invariants.md) - the mechanisms behind all of this, for contributors.
+27 -6
View File
@@ -4,7 +4,7 @@ Run a case inside its own container instead of directly on your host: for isolat
reproducible toolchain, and for the ability to pick the whole environment up and move it to
another machine.
A docker case is a **location overlay**, not a run mode. All seven run modes work inside a
A docker case is a **location overlay**, not a run mode. All ten run modes work inside a
container. See [Core Concepts](Core-Concepts).
## One-time setup: the base image
@@ -26,7 +26,7 @@ A zero exit code proves the layers ran, not that the toolchain works. Verify:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy pi; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
'for c in claude codex gemini opencode agy pi grok dsh omp; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
The image is secret-free. Credentials are delivered at runtime, never baked in, so exports
@@ -79,6 +79,25 @@ Exactly one long-lived container per case, shared by every session in it.
conversation** from the bind-mounted transcript.
- Deleting the case removes the container. The workspace on the host survives.
## Attaching to a container you already run
Tick **Attach to an existing container** on **Add Case → Docker** to link a case to a
container that already exists instead of creating one. Codeman only `exec`s into it and
never creates, starts, stops, restarts or removes it, so a container that is missing or
stopped fails with a message rather than being fixed for you. Drift detection does not
apply (the container carries no Codeman configuration label). The full-image export is
refused, since it would `docker commit` someone else's container, and the workspace export
skips the pause that keeps an owned container consistent during the capture.
One adopted container can back several cases at different in-container directories, and
**copy an existing case** pre-fills the form from a sibling on the same container. An exact
twin (the same container and the same directory) is refused, as is a container another
user adopted.
Adoption is **admin-only in multi-user mode**. Linking creates Codeman's own container
with one bind mount that has already been checked; an adopted container's mounts belong to
whoever started it, and one that mounts `/` hands the adopter the host.
## Credentials
Your existing host logins work inside the container without logging in again. Credentials
@@ -92,10 +111,12 @@ the container instead.
Bind mounts are excluded from image capture, so exports stay secret-free.
One consequence worth knowing: Pi's credentials are seeded per file rather than as a whole
directory, because that directory also holds sessions, extensions, and installed packages,
which can be gigabytes. So in-container Pi sessions are invisible from the host, and `pi -c`
inside a docker case sees only that container's history.
One consequence worth knowing: Pi, Grok and OMP credentials are seeded per file rather than
as whole directories, because those directories also hold sessions, extensions, downloads and
installed packages, which can be gigabytes. So in-container Pi and Grok sessions are
invisible from the host (`pi -c` and `grok -c` inside a docker case see only that
container's history). OMP's `sessions/` is the exception and is shared read-write, because
Codeman reads it host-side for history and resume.
## Isolation
+18 -5
View File
@@ -54,7 +54,10 @@ create-time sweep would yank the skill out from under other live sessions sharin
directory. Remove them per case with `codeman skill uninstall --case <name>`.
The skill ships with the verb index always loaded, plus on-demand references for the verbs,
worked multi-worker recipes, endpoint tables, and cross-session messaging.
worked multi-worker recipes, endpoint tables, and cross-session messaging. It drives
DeepSeek Harness workers the same way it drives Claude ones (`spawn_workers alpha
beta:deepseek` is a mixed fleet in one call), since those are the two modes with real
completion signals.
## The manual path
@@ -92,8 +95,9 @@ Read these before writing any code. Each one has cost somebody an afternoon.
5. **Wait instead of polling, and a timeout is not an error.** The wait endpoints answer
`200` with `wait.timedOut: true`. Loop over short waits rather than one long call, because
tunnels cut idle connections.
6. **Only `claude` sessions emit `stop` and `blocked`.** They come from Claude Code hooks.
Shell and the external CLIs accept only `idle`, `working`, and `exit`; asking for `stop`
6. **Only `claude` and `deepseek` sessions emit `stop` and `blocked`.** Claude's come from
Claude Code hooks, DeepSeek's from the harness reporting its state to Codeman. Shell and
the other external CLIs accept only `idle`, `working`, and `exit`; asking for `stop`
explicitly there is a `400`, while omitting `until` is always safe. On a shell session
`idle` fires **once at startup and never again**, so synchronize hook-less sessions with an
output marker instead.
@@ -130,7 +134,10 @@ curl -s -X POST "$API/api/sessions/$ID/input" \
# Or wait for a marker in the output, which works on shell sessions too
curl -s "$API/api/sessions/$ID/wait-output?contains=DONE_17909&from=buffer" | jq
# Read the terminal back
# Read the last answer as clean text (claude, codex, deepseek sessions)
curl -s "$API/api/sessions/$ID/last-response" | jq -r '.data.text'
# Or read the terminal back
curl -s "$API/api/sessions/$ID/terminal?tail=4000" | jq -r '.data.output'
# Clean up, by exact id
@@ -157,7 +164,13 @@ Make it unique per call, because tmux repaints replay old screen text.
### Reading output
Use `terminal?tail=`, not `/output`. The latter's text field is empty for every tmux-backed
For `claude`, `codex` and `deepseek` sessions, read the answer from the transcript rather
than the screen: `GET /api/sessions/:id/last-response` returns the last reply as clean text
with no TUI frames or repaint noise. Poll it briefly rather than reading once, because the
transcript lands slightly after the `stop` signal, so a read immediately after send-and-wait
returns often comes back empty.
For everything else, use `terminal?tail=`, not `/output`. The latter's text field is empty for every tmux-backed
session, which is every interactive session. `tail` counts **bytes**, and what comes back is
terminal data with ANSI sequences included.
+6
View File
@@ -21,6 +21,12 @@ No. Codeman drives agent CLIs you have already installed and logged in yourself.
subscription or key that CLI uses is what pays for the tokens. Codeman never collects,
stores, or refreshes your credentials.
### Which agent CLIs does it support?
Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi, Grok Build, DeepSeek Harness and
OMP, plus a plain shell, chosen per session. Claude is the reference mode and a few features
are Claude-only; [Agent CLIs](Agent-CLIs) has the table.
### Does Codeman send my code or prompts anywhere?
No. There is no telemetry, no analytics, and no phone-home. The only network traffic
+18 -9
View File
@@ -66,14 +66,14 @@ self-signed certificate, add `-k`.
## Endpoint map
Roughly 200 handlers across 24 route modules. By domain:
Roughly 235 handlers across 26 route modules. By domain:
| Domain | Handlers | Covers |
| ------------------- | -------- | --------------------------------------------------- |
| System | 45 | Status, settings, search, digest, updates. |
| Sessions | 34 | Create, input, terminal, wait, kill. |
| Cases | 29 | Create, link, clone, remote and docker cases. |
| Files | 16 | Preview, edit, raw, attachments, path picker. |
| System | 56 | Status, settings, digest, updates, tunnel. |
| Sessions | 34 | Create, input, terminal, wait, last response, kill. |
| Cases | 34 | Create, link, clone, remote and docker cases. |
| Files | 17 | Preview, edit, raw, attachments, path picker. |
| Orchestrator | 10 | Plans and phases. |
| Ralph | 9 | Loop control and configuration. |
| Cron | 9 | Jobs and run history. |
@@ -82,10 +82,12 @@ Roughly 200 handlers across 24 route modules. By domain:
| Respawn | 7 | Respawn configuration and presets. |
| Webviews | 6 | Saved dashboards, plus the proxy. |
| Mux | 5 | tmux operations. |
| Custom model endpoints | 5 | Saved OpenAI-compatible endpoints, and applying one to a session. |
| Push | 4 | Web push subscriptions. |
| Read My Mind | 4 | Intent profiles and prediction. |
| Scheduled | 4 | The legacy scheduled-run concept. |
| Approvals | 3 | The inbox and answering. |
| Approvals | 4 | The inbox, answering, acknowledging. |
| Tab layout | 2 | Named tab groups per owner. |
| Teams, me, search, hooks, clipboard, telemetry, voice, ws | 1-2 each | |
Each route module documents its own endpoints in its file header.
@@ -114,12 +116,13 @@ Three semantics that break callers who assume otherwise:
`wait-output` matches a **literal substring, never a regex.** That is deliberate: no regex
means no catastrophic backtracking on attacker-influenced output.
Only `claude` sessions emit `stop` and `blocked`, because those come from Claude Code hooks.
Shell and external CLI sessions accept `idle`, `working`, and `exit`.
Only `claude` and `deepseek` sessions emit `stop` and `blocked`: Claude's come from Claude
Code hooks, DeepSeek's from the harness reporting its state to Codeman. Shell and the other
external CLI sessions accept `idle`, `working`, and `exit`.
## SSE
`GET /api/events` is the live event stream. 156 event names, kept in sync between server and
`GET /api/events` is the live event stream. 158 event names, kept in sync between server and
client with a test that fails on drift.
The heartbeat is a **named** `sse:heartbeat` event rather than an SSE comment, because
@@ -141,6 +144,12 @@ curl -s "$API/api/sessions" | jq '.data[].name' # live sessions
curl -s "$API/api/sessions/unified" | jq # live + historical, deduped
curl -s "$API/api/subagents" | jq # background agents
curl -s "$API/api/search?q=deploy" | jq # cross-session search
# with ID set to a session id:
curl -s "$API/api/sessions/$ID/last-response" | jq -r '.data.text' # last answer, from the transcript (claude, codex, deepseek)
curl -s "$API/api/model-endpoints" | jq # saved custom OpenAI-compatible endpoints
curl -s -X POST "$API/api/sessions/$ID/custom-model" -H 'Content-Type: application/json' \
-d '{"endpointId":"local-llama","modelId":"qwen3-27b"}' | jq # restart the CLI on that endpoint; {"clear":true} undoes it
```
## Limits
+4 -4
View File
@@ -5,8 +5,8 @@
<h3 align="center">Mission control for AI coding agents</h3>
Codeman runs your coding agents on your own machine and puts them behind one dashboard you
can open from any device. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, or
Pi inside persistent tmux sessions, streams the real terminal to the browser, and keeps
can open from any device. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi,
Grok, DeepSeek Harness, or OMP inside persistent tmux sessions, streams the real terminal to the browser, and keeps
working while you are away from the keyboard: it re-prompts idle agents, resumes when a
subscription limit resets, runs jobs on a schedule, and shows every background subagent
live.
@@ -33,7 +33,7 @@ codeman web # then open http://localhost:3000
**Already running it**
- [Agent CLIs](Agent-CLIs) - the seven run modes, their setup, and which features are Claude-only.
- [Agent CLIs](Agent-CLIs) - the ten run modes, their setup, and which features are Claude-only.
- [Mobile Guide](Mobile-Guide) - phone and tablet use, QR login, the touch keyboard bar.
- [Remote Access](Remote-Access) - Tailscale, Cloudflare tunnel, LAN plus password, QR login.
- [Keeping Agents Running](Keeping-Agents-Running) - idle detection, respawn cycling, auto-resume on usage limits.
@@ -122,7 +122,7 @@ codeman web # then open http://localhost:3000
| OS | macOS or Linux. Windows works through WSL2. |
| Node.js | 22 or newer. |
| tmux | Required. Sessions live in tmux, which is what makes them survive restarts. |
| An agent CLI | At least one of Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi. Plain shell sessions need none. |
| An agent CLI | At least one of Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi, Grok Build, DeepSeek Harness, OMP. Plain shell sessions need none. |
| Network | Binds to `127.0.0.1` by default. Reaching it from another device is a deliberate step: see [Remote Access](Remote-Access). |
Codeman is MIT licensed, self-hosted, and sends no telemetry. Everything runs on your
+6 -4
View File
@@ -19,9 +19,11 @@ terminal into something that can notify you.
| `teammate_idle` | An agent-team member goes idle. | Team surfaces. |
| `task_completed` | A task finishes. | Task tracking, run summary. |
This is why several Codeman features are Claude-only. The other CLIs have no hook system, so
for them Codeman watches terminal output, which reveals that something happened but not what
it was.
This is why several Codeman features are Claude-only. The one partial exception is DeepSeek
Harness, whose terminal front door reports idle, working and blocked to Codeman over the
harness's own supervisor contract, so it gets the hook-driven surfaces without any hook
file. The other CLIs have no equivalent, so for them Codeman watches terminal output, which
reveals that something happened but not what it was.
### How hooks get installed
@@ -67,7 +69,7 @@ sit beside the agents with no code at all. See [Web Tabs](Web-Tabs).
### 2. SSE events
`GET /api/events` streams everything Codeman knows: session lifecycle, output, agent
activity, approvals, cron runs. 155 named events, stable under semantic versioning.
activity, approvals, cron runs. 158 named events, stable under semantic versioning.
This is the seam for anything that reacts. A bot that pings your chat channel when an agent
needs a human is a short script over this stream.
+11
View File
@@ -27,6 +27,15 @@ The result is the property you want on a phone: a connection that drops mid-prom
loses the prompt and never delivers it twice. Two browser tabs on the same session coexist,
and only a reconnect from the *same* tab supersedes the old connection.
## Selecting and copying
Agent CLIs hold the mouse: clicks and drags are reported into the transcript rather than
selecting text. `Shift+drag` starts a selection anyway, right-click copies it (with nothing
selected the native context menu is left alone), and `Ctrl+Shift+C` copies without ever
interrupting. **Auto Copy Selection** in **App Settings → Terminal & Input**, off by
default, copies the moment you release the mouse. On phones, long-press selects; see
[Mobile Guide](Mobile-Guide).
## Zero-lag local echo
On touch devices, keystrokes are painted in the terminal immediately and sent when you press
@@ -55,6 +64,8 @@ reconcile against the real buffer and only apply while the cursor is on the comp
Chinese, Japanese, and Korean input needs an IME, and an IME needs a real text field.
Turning on CJK input in **App Settings → Terminal & Input** puts an always-visible textarea
below the terminal that owns composition, then delivers the composed text to the session.
Ctrl- and Alt-modified navigation keys typed through it reach the CLI as the modified
sequences, so word jumps and history keys keep working.
## Voice dictation
+26 -4
View File
@@ -9,7 +9,7 @@ Getting Codeman onto a machine, verifying it works, updating it, and removing it
| **macOS or Linux** | Windows works through WSL2. See [Windows](#windows-wsl) below. |
| **Node.js 22+** | The installer offers to install it if missing. |
| **tmux** | Not optional. Sessions live inside tmux, which is what makes them survive a server restart, a dropped connection, or a closed laptop. |
| **An agent CLI** | At least one of [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev). Plain shell sessions need none. See [Agent CLIs](Agent-CLIs). |
| **An agent CLI** | At least one of [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), [OMP](https://github.com/can1357/oh-my-pi). Plain shell sessions need none. See [Agent CLIs](Agent-CLIs). |
Codeman itself sends no telemetry and phones no home. The only network traffic is your
browser to your server, and whatever the agent CLI you chose does on its own.
@@ -20,13 +20,16 @@ browser to your server, and whatever the agent CLI you chose does on its own.
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if they are missing, clones Codeman into `~/.codeman/app`,
and builds it.
This installs Node.js, tmux and a build toolchain if they are missing (node-pty ships no
Linux prebuild, so it compiles from source), clones Codeman into `~/.codeman/app`, and
builds it.
What it asks you:
1. **Permission for every system change.** Package installs and agent CLI downloads are
prompted individually. Nothing is installed silently.
prompted individually. Nothing is installed silently. If no agent CLI is found, a menu
offers to install any of them (DeepSeek excepted: its npm package installs only a
launcher with no runnable profile), or you skip and install one yourself later.
2. **How the dashboard should be reachable.** Three choices:
- **Tailscale** (recommended for phone access): keeps the loopback bind and walks you
through `tailscale serve`, including the tailnet HTTPS toggle, then verifies the result
@@ -101,6 +104,21 @@ at server start, so markup changes need a restart.
See [Contributing](Contributing) for the rest of the development loop.
## Route D: Docker Compose
Codeman itself can run in a container and spawn Docker cases as sibling containers through
the host's Docker socket. Copy `docker/.env.example` to `docker/.env`, set
`CODEMAN_PASSWORD`, then:
```bash
bash docker/Start-Codeman.sh
```
Run the script again after updating rather than a plain `docker compose up`, so the rebuilt
image, the refreshed volumes and the entrypoint arrive together. The full guide, including
storage and networking options, is
[`docker/README.md`](https://github.com/Ark0N/Codeman/blob/master/docker/README.md).
## Installing an agent CLI
Codeman drives CLIs, it does not bundle them. Install at least one:
@@ -113,6 +131,9 @@ Codeman drives CLIs, it does not bundle them. Install at least one:
| **Antigravity** | See [antigravity.google](https://antigravity.google) | Google's successor to the consumer Gemini CLI. |
| **Gemini CLI** | See [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) | Enterprise only since Google's June 2026 consumer cutover. |
| **Pi** | See [pi.dev](https://pi.dev) | No permission prompts and no sandbox by design. Read [Agent CLIs](Agent-CLIs) before using it on a repo you care about. |
| **Grok Build** | `curl -fsSL https://x.ai/cli/install.sh \| bash` | xAI. Lands in `~/.grok/bin`; `grok login --device-auth` for headless hosts. |
| **DeepSeek Harness** | `npm i -g @deepseek-ai/dsh pnpm`, then a terminal profile | The npm package is only a launcher. Codeman's Run menu installs the community terminal profile for you. See [Agent CLIs](Agent-CLIs). |
| **OMP** | `curl -fsSL https://omp.sh/install \| sh` | Oh My Pi. Run it once by hand to finish its own onboarding. |
Log each CLI in once, by hand, before pointing Codeman at it. Codeman never collects or
stores your CLI credentials.
@@ -161,6 +182,7 @@ Full detail, including logs and the self-updater, is in
| Installer | Re-run the one-liner, or **App Settings → System → Updates** in the UI. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart. |
| Docker Compose | Re-run `Start-Codeman.sh`. The in-app updater works too, and refuses a release that changes the container definition until you re-run the script. |
The in-app updater covers git-clone installs supervised by systemd or launchd. It restarts
the process that is running it, so the actual work happens in a detached script and the
+16 -8
View File
@@ -27,9 +27,12 @@ keystroke echo. Idle now lands a few seconds after a turn genuinely ends.
There are several layers stacked on that: a completion message from the CLI, an AI check,
output silence, and token stability.
**For every other CLI**, there are no hooks to lean on, so detection is output
stabilization: the session is idle when output stops changing. Coarser, and it is why the
features further down this page are Claude-only.
**For the other CLIs** it depends on what the CLI tells Codeman. Codex declares its own
prompt glyph and working line, so it gets the same screen check Claude does (before 1.26.1
every Codex session reported idle for its whole life). DeepSeek Harness reports idle,
working and blocked to Codeman itself, which is as precise as hooks. Everything else is
output stabilization: the session is idle when output stops changing. Coarser, and it is
why the features further down this page are Claude-only.
## The Respawn Controller
@@ -101,13 +104,18 @@ subscription plan.
**Claude only.** A header chip showing live subscription usage, on by default on desktop and
off on phones.
It works by installing a status line exporter into Claude Code, which posts Claude's own
rate limit data back to Codeman. The exporter is marker-identified, so it only ever touches
a status line Codeman installed, never one you wrote yourself, and it prints your footer
through so the in-terminal status line still works.
It works through a status line exporter that Codeman hands to `claude` as an ephemeral
setting when it spawns the session, never written to disk, which posts Claude's own rate
limit data back to Codeman. Your own status line (project-local, project, then
`~/.claude/settings.json`) is wrapped and printed through, and a `claude` you run by hand
outside Codeman sees nothing of it. Workspaces an older Codeman wrote the exporter into are
cleaned up the first time a session starts there. Codex limits come from a read-only poll of
its own app-server. Known limit: sessions inside a Docker case do not feed the chip yet.
The chip and the exporter are the same setting. Turning the chip on without the exporter
would leave it showing a dash forever, so resolve it in one place: **App Settings**.
would leave it showing a dash forever, so resolve it in one place: **App Settings**. A
device writes the switch only when it flips the chip, so a phone (chip off by default)
saving its font size cannot switch collection off for your desktop.
## Circuit breakers
+3
View File
@@ -30,6 +30,9 @@ Press `Ctrl+?` in the app for the same list in a floating overlay.
| `Ctrl+Shift+R` | Restore terminal size. |
| `Ctrl` `+` / `Ctrl` `-` | Font size. |
| `Shift+Wheel` | Scroll the local buffer, even where the wheel is forwarded to the CLI. |
| `Shift+drag` | Start a selection in a pane whose mouse events go to the CLI. |
| Right-click | Copy the selection. With nothing selected the native menu is left alone. |
| `Ctrl+Z` | Swallowed in agent sessions so a running CLI cannot be suspended. Normal job control in a shell. |
## Everything else
+9 -3
View File
@@ -30,8 +30,12 @@ require a secure context.
| Toolbar | Bottom: Run, Stop, **Enter**, case picker, voice, settings. |
| Keyboard bar | Above the on-screen keyboard when it is open. |
Layout respects notch and home-indicator safe areas, touch targets are 44px, and the case
picker is a bottom sheet rather than a dropdown.
The phone layout applies up to 599px of viewport width, so the Plus and Pro Max iPhones,
the Pixel Pro and a folded Z Fold get it too; wider devices get the tablet layout. Layout
respects notch and home-indicator safe areas, touch targets are 44px, and the case picker is
a bottom sheet rather than a dropdown. On a folding phone (iPhone Duo) dialogs stay clear of
the hinge, and opening or closing the device is treated as the device changing shape, never
as the keyboard appearing.
**Swipe left and right** on the terminal to switch sessions.
@@ -58,7 +62,9 @@ A row of keys above the virtual keyboard, and what it contains depends on the se
**Agent sessions** get quick actions: `/init`, `/clear`, `/compact`, a clipboard key, `Esc`,
a path picker, an image key, and 🧠 when Read My Mind is on. Destructive commands need a
double press, so you cannot fire `/clear` with a stray thumb.
double press, so you cannot fire `/clear` with a stray thumb. On Codex sessions the bar also
shows `⇧←` and `⇧→`, the Shift-modified arrows Codex binds to editing the last queued
message and walking the prompt stack.
**Shell sessions** automatically swap it for terminal controls: `Ctrl`, `Esc`, `Tab`, four
arrows, paste, and dismiss. Your normal preference is remembered and restored when you
+6 -4
View File
@@ -33,8 +33,8 @@ reloading the dashboard while a permission dialog is blocking a session does not
with a normal-looking tab.
For Claude sessions, these come from Claude Code's hooks and are precise about *why* the
session stopped. For other CLIs there are no hooks, so you get the coarser output-based
signal.
session stopped; DeepSeek Harness sessions report the same states themselves. For the other
CLIs there are no hooks, so you get the coarser output-based signal.
## Window title and OS notifications
@@ -62,7 +62,8 @@ Once subscribed, a blocking prompt reaches your phone even from a locked screen.
## The Approvals Inbox
**Opt-in, off by default. Claude sessions only.**
**Opt-in, off by default. Claude sessions, plus DeepSeek Harness sessions, whose terminal
front door reports its prompts to Codeman.**
One queue of every prompt currently waiting on a human, across all your sessions, answerable
in place. When you have eight workers running, this is the difference between checking eight
@@ -136,7 +137,8 @@ from the lock screen.
- **No push over plain HTTP.** It is a browser requirement, not a Codeman one.
- **iOS needs the home screen install.** A Safari tab will never receive push.
- **The bell is invisible at zero.** That is deliberate, not a broken setting.
- **Approvals are Claude-only.** They are built on hook events the other CLIs do not emit.
- **Approvals need real signals.** They are built on hook events, which Claude emits and
DeepSeek Harness reports itself; the other CLIs do neither.
- **A stale menu answer is refused, not sent.** If you answer a card for a dialog that has
since gone away, Codeman declines rather than typing a digit into the composer.
+3
View File
@@ -67,6 +67,9 @@ one:
| **Gemini** | Enterprise only since Google's consumer cutover. |
| **Antigravity** | Google's successor to the consumer Gemini CLI. |
| **Pi** | No permission prompts and no sandbox by design. |
| **Grok Build** | xAI's CLI. |
| **DeepSeek Harness** | Needs a terminal profile; the menu offers to install one. |
| **OMP** | Oh My Pi, configured entirely through its own `~/.omp`. |
| **Terminal / Shell** | A plain shell, no agent. Also the **Run Shell** button. |
The dropdown also lists any saved dashboard URLs ([Web Tabs](Web-Tabs)) and your recent
+13 -2
View File
@@ -4,7 +4,7 @@ Point a case at another machine and the agent runs **there**, with the same dash
mobile UI, and autonomy features. Your laptop becomes a window onto a session living on the
remote host.
Like Docker, this is a **location overlay** on a case, not a run mode. All seven run modes
Like Docker, this is a **location overlay** on a case, not a run mode. All ten run modes
work remotely. See [Core Concepts](Core-Concepts).
## Why bother
@@ -52,7 +52,10 @@ A watcher with bounded backoff notices a dead SSH pane and quietly reattaches to
running remote session. On by default; the kill switch is in
**App Settings → Agents & CLIs → Remote auto-reconnect**.
Intentional kills are never revived. Closing a session means closing it.
Intentional kills are never revived. Closing a session means closing it. Neither is a clean
exit inside the pane (Ctrl-D, `exit`, Ctrl-C at the CLI's prompt): that tears the remote
tmux session down, and the watcher revives a session only when that durable session is
verifiably still alive. Only a transport drop is reconnected.
## Discover and attach
@@ -70,6 +73,14 @@ Attaching to someone else's session and closing your tab must not end their run,
not. Several clients can attach the same remote session at different window sizes without
clamping each other, and discovery shows a shared badge with the client count.
## Files
Previews, downloads and text reads in a remote case go over the same ssh connection the
session uses, so a clicked path opens the file on the machine the agent is on, `Range`
seeking included. Nothing is copied to the Codeman host. Editing, Office previews,
thumbnails, the file tree and the tail viewer are not available remotely and answer a clear
400 rather than a misleading 404. Details in [Working With Files](Working-With-Files).
## Security
Every SSH command line in Codeman flows through one builder that shell-escapes every
+13
View File
@@ -139,6 +139,7 @@ log stream --predicate 'process == "node"' # macOS, noisy
| Installer | Re-run the one-liner, or **App Settings → System → Updates**. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart. |
| Docker Compose | Re-run `Start-Codeman.sh`, or the in-app updater, which restarts the container in place. |
### The in-app updater
@@ -170,6 +171,18 @@ service without colliding with the main one. `CODEMAN_DATA_DIR` and `CODEMAN_TMU
exist for the rare case where they need to differ, but setting only one of them recreates
exactly the problem you were avoiding.
## Running Codeman itself in Docker
The Compose deployment in `docker/` runs the server in a container and spawns Docker cases
as sibling containers through the mounted host socket. Start it with
`bash docker/Start-Codeman.sh` rather than a bare `docker compose up`: the script pre-creates
the bind-mounted directories with the right owner, honours a `docker-compose.override.yml`,
and refreshes the build volumes when the checkout moved under them. The in-app updater
applies code only and restarts by letting the container exit, so it refuses a release that
changes the Dockerfile, the compose file, or adds a new `.env` key, until you re-run the
script. Guide:
[`docker/README.md`](https://github.com/Ark0N/Codeman/blob/master/docker/README.md).
## The tunnel as a service
```bash
+1
View File
@@ -67,6 +67,7 @@ be wrong for at least one of them:
| **File Viewer** | Real path resolution before boundary checks, so symlinks cannot escape. Sensitive trees blocked. Edit mode adds an extension allowlist, a size cap, `.git` denial, and optimistic concurrency. It never creates files. |
| **Attachments** | An id-based registry, so browser requests never carry absolute paths. The magic-link scanner is prompt-injectable by nature and is therefore force-confined to the session's workspace. Extension allowlist, not a blocklist. |
| **Path picker** | Its own root allowlist rather than the workspace confinement. In multi-user mode a non-admin gets only their own user space, because per-user spaces live inside the home directory. |
| **Remote cases** | Reads go over the session's own ssh connection and are resolved and contained on the remote host, with a bounded number of ssh children. Nothing is copied to the Codeman host; writes, Office previews and thumbnails are refused. |
Downloads block sensitive paths outright (`.env`, credentials files, `~/.ssh`, AWS
credentials), and SVG and HTML are served as downloads with `nosniff` so they cannot execute
+7 -1
View File
@@ -46,6 +46,7 @@ supervised by systemd or launchd; npm installs report as non-updatable. See
| Extended Keyboard Bar | Per device | Which accessory bar phones get. Shell sessions override it while they are active. |
| Wheel Scrolls Local History | Off | Keeps the wheel on the local buffer instead of forwarding it to the CLI. |
| Auto Copy Selection | Off | Copies highlighted terminal text to the clipboard the moment you finish selecting it. Ctrl+C still copies on demand. |
| Normal / Bold font weight | xterm defaults | Per device, each slot from 100 to 900. The bundled JetBrains Mono renders every step, so a lighter normal weight makes Claude's bold headings stand out. Applies live to the terminal, both echo overlays and open team panes. |
| WebGL Renderer | On | With a GPU-stall watchdog that falls back to DOM rendering. |
| Gesture Control | Off | Camera hand tracking. Also needs `CODEMAN_GESTURE=1` on the server. |
@@ -72,7 +73,9 @@ every session or only the active tab.
| Entrance Animations | Per-surface animation styles for tabs, terminals, windows, and lineage lines. All default to the legacy no-animation behaviour. |
| Display Name | Your name in the UI. Cosmetic only; it never renames the package, CLI, API, or storage. |
| Interface Language | English or Simplified Chinese. Per device. |
| Session List Layout | Header tab strip (default) or a collapsible left sidebar. See [The Dashboard](The-Dashboard#session-list-layout). |
| Session List Layout | Header tab strip (default), a collapsible left sidebar, or the sidebar with detailed rows. See [The Dashboard](The-Dashboard#session-list-layout). |
| Tab Orientation | Keeps the header list but turns the strip vertical beside the terminal, resizable, with detailed rows by default. Desktop and tablet only. |
| Vertical Rail Order | *By activity* (default) sorts the rail the way the home screens are sorted; *Manual* keeps your tab order and drag-reordering. |
| Tall Tabs | Taller tab strip. |
| Pop-out Button on Tabs | Adds the detach control to tabs, with a per-tab override. |
| Spawn Lineage Lines | Arcs from a parent tab to sessions it spawned. Desktop only, on by default. |
@@ -156,6 +159,9 @@ Some things are configured before the server starts, not in the UI:
| `CODEMAN_DOCKER_BRIDGE_HOOKS` | Lets in-container hooks reach the host on a loopback bind. |
| `CODEMAN_FILE_PICKER_ROOTS` | Extra roots for the path picker. |
| `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK` | Acknowledges exposing the server with no password. |
| `CODEMAN_BASE_URL` | Mounts Codeman under a sub-path behind a reverse proxy that forwards the prefix unchanged. See [Remote Access](Remote-Access). |
| `CODEMAN_MAX_DOWNLOAD_BYTES` | Cap on raw file bodies and downloads. 2 GB by default, `0` for none. |
| `CODEMAN_MAX_REMOTE_FILE_SSH` | Concurrent ssh reads for files in remote cases. 4 by default. |
## Gotchas
+14 -5
View File
@@ -22,12 +22,14 @@ page says so and names the setting.
The session list lives in the header as a horizontal strip by default. With a lot of
sessions open that strip stops being scannable, so **App Settings → Appearance → Tabs →
Session List Layout** can move it into a vertical sidebar on the left instead.
Session List Layout** can move it into a vertical sidebar on the left instead, and
**Tab Orientation** can turn the strip itself into a vertical rail.
| Layout | Behaviour |
| -------------------- | --------------------------------------------------------------------------------- |
| **Header tab strip** | The default. Wraps to a second row on desktop, scrolls sideways on a phone. |
| **Left sidebar** | A vertical list with a filter box and a live session count. `Alt+B` collapses it to a narrow rail that keeps the status dots and task badges visible. On a phone it is an off-canvas drawer rather than a docked rail. |
| **Left sidebar** | A vertical list with a filter box and a live session count. `Alt+B` collapses it to a narrow rail that keeps the status dots and task badges visible. On a phone it is an off-canvas drawer rather than a docked rail. A detailed variant adds the home screen's per-session line (`created 3d ago · working 12m`) and a status pill. |
| **Vertical rail** | The strip turned vertical beside the terminal, resizable, with detailed rows by default. **Vertical Rail Order** sorts it by activity (blocked on you first, then longest running, then most recently quiet), the same order as the home screens; pick *Manual* to get your own order and drag-reordering back. Desktop and tablet only. |
It is the same list either way, just re-hosted: tab order, drag-to-reorder, the `Alt+1`
to `Alt+9` numbers and every status colour below behave identically in both. The setting is
@@ -154,6 +156,10 @@ Worth knowing:
always local scrollback. Other CLIs scroll locally.
- **Selection copy.** `Ctrl+C` copies when text is selected and interrupts when it is not.
`Ctrl+Shift+C` always copies.
- **Selecting where the CLI owns the mouse.** `Shift+drag` starts a selection even in a pane
whose mouse events are forwarded to the CLI, and right-click copies the selection (with
nothing selected the native menu is left alone). **Auto Copy Selection** in App Settings
copies the moment you release.
- **Zero-lag input.** On touch devices, keystrokes paint locally before the round trip. See
[Input And Voice](Input-And-Voice).
- **Renderer.** WebGL by default, with a watchdog that falls back to DOM rendering if the
@@ -167,8 +173,9 @@ which lists past sessions including Claude conversations started outside Codeman
Two extras depending on the device:
- **Desktop, wide windows**: your open tabs appear as a rail docked to the left edge, in tab
order, with created and last-active stamps. It needs at least 1180px of width; below that
- **Desktop, wide windows**: your open tabs appear as a rail docked to the left edge, in
overview order (blocked on you first, then longest running, then most recently quiet),
with created and state-duration stamps. It needs at least 1180px of width; below that
it is hidden so it cannot overlap the search panel.
- **Phones**: tapping the "C" logo gives a session overview instead: NEEDS YOU first, then
current sessions, then past ones. On by default.
@@ -204,7 +211,9 @@ so it is fast and cannot be turned into a traversal.
## Appearance
**App Settings → Appearance** carries the theme skins, including light ones. The choice is
applied before the first paint, so there is no flash of the wrong theme on load.
applied before the first paint, so there is no flash of the wrong theme on load. Terminal
font family and weight are per device too: a normal and a bold weight, each from 100 to
900, and the bundled JetBrains Mono renders every step.
The same section has the entrance animations for tabs, terminals, agent windows, and
lineage lines. All of them default to the legacy no-animation behaviour, so an untouched
+32 -2
View File
@@ -138,6 +138,12 @@ That is the PTY-exit circuit breaker. Repeated rapid PTY exits trip it, and it b
automatic restarts so a broken configuration does not spin forever. Reset it explicitly from
the session's controls. Reattaching does not clear it, deliberately.
### Typed prompts are silently ignored after restoring a tab
Update. A browser whose input sequence counter fell behind the server's (a restored tab,
cleared site data) used to have every prompt deduplicated away. Since 1.29.0 the duplicate
acknowledgement carries the watermark and the client re-sends.
### Sessions I did not create appeared, or my session resized itself
Two Codeman servers are running against the same data directory and tmux socket. The second
@@ -167,6 +173,16 @@ Things to try:
Codex ignores the mouse reports that forwarding would send, so Codeman does not forward
there. Scrolling is local, and `Shift+Wheel` behaves the same way.
### Selected text is invisible on a light skin
Update. Every skin named its selection colour under a key xterm renamed in v5, so the four
light skins painted white at 30% over near-white. Fixed in 1.29.0.
### `Ctrl+Z` suspended my agent
Update. Since 1.28.0 `Ctrl+Z` is swallowed in agent sessions, so a running CLI cannot be
stopped by job control. Shell sessions keep it.
### `Ctrl+C` copies when I wanted to interrupt
With a selection, `Ctrl+C` copies. With no selection, it interrupts. Clear the selection
@@ -252,11 +268,25 @@ node scripts/build-agent-image.mjs --no-cache
A plain rebuild reuses the cached `npm install -g` layer and keeps the CLIs frozen at their
original versions while reporting success.
### Every file in a remote case says "File not found"
Update. Before 1.29.0 the file routes resolved every path on the Codeman host, so in a
remote case every click failed while the file plainly existed on the other machine. Reads
now go over ssh; see [Working With Files](Working-With-Files). Editing and Office previews
stay unavailable remotely and say so with a 400.
### Compose: the server crash-loops with `EACCES` on first start
Start the stack with `bash docker/Start-Codeman.sh` rather than a plain `docker compose up`,
and update: since 1.29.0 the entrypoint corrects a root-owned bind mount before dropping
privileges. See [Running As A Service](Running-As-A-Service).
### A remote SSH session dropped and did not come back
A bounded-backoff watcher reattaches dropped sessions, and it is on by default. Intentional
kills are never revived. Check the host is reachable and that the remote tmux server is
still running.
kills are never revived, and neither is a clean exit inside the pane (Ctrl-D, `exit`): only
a transport drop is reconnected. Check the host is reachable and that the remote tmux server
is still running.
## Gathering diagnostics
+23
View File
@@ -24,6 +24,23 @@ Switching tabs does not reload a dashboard. Frames stay alive in the background,
took a while to authenticate is still there when you come back. Past six live frames, the
least recently viewed is dropped to bound memory.
## Single-page apps, reloads and links
A history-routed dashboard (React Router, Vue Router, a Vite dev server) sees the path it
would see on its own origin, not the proxy prefix, so it renders its real route instead of
its own "page not found". A navigation the page starts itself afterwards, a dev server's
full reload or a root-absolute `location.href`, would land outside the proxy with no
capability; Codeman recognises it, answers with a small recovery page, and remounts the
frame at the path that was lost, bounded to five recoveries a minute per frame. A reload on
the dashboard's landing page is recovered the same way.
A `localhost` or `127.0.0.1` link in agent output opens as a web tab automatically, reusing
a saved dashboard for the same server or saving one under its `host:port`. On a phone that
address only exists on the Codeman box, so the link would otherwise be a guaranteed
connection error. LAN and tailnet addresses still open directly. `*.localhost` names are
deliberately not auto-routed: they are DNS names rather than address literals, and the link
came from agent output. Add such a dashboard by hand instead.
## Why dashboards are proxied
A plain cross-origin iframe fails three ways at once in the setup Codeman actually ships in:
@@ -89,6 +106,12 @@ The proxy authenticates on an in-memory capability embedded in the path, which i
exempt from the cookie and Origin checks that every API route enforces. That exemption is
fenced to safe methods and non-API paths, and there is a test pinning it in place.
Saved URLs are refused when they point at a link-local or cloud-metadata address, at save
time and again against the address the name resolves to at connect time; loopback and
private ranges stay allowed, because a `localhost` Grafana is the feature. Capabilities are
revoked on logout, and proxied responses carry a same-origin referrer policy so a dashboard
cannot hand the capability-bearing URL to a third party.
Two failure modes that only appear inside a sandboxed frame, and that curl can never
reproduce, are handled: runtime-built root-absolute URLs escaping the injected base, and
same-host requests being CORS-checked with a null origin. Both present as the dashboard's own
+16
View File
@@ -110,6 +110,22 @@ it is written. Outside the workspace they open in the preview instead: the tail
Nothing is registered until you click. Opening a file this way does not add an attachment card.
## Remote (SSH) cases
In a remote case the workspace lives on the other machine, and so do the files. Previews,
downloads, text reads and the clicked-path route all go over the same ssh connection the
session uses: one `realpath` plus `stat` probe for the file and the workspace root, then a
streamed `cat` (or a slice of it, so video seeking works). Symlinks are resolved on the host
that can resolve them, the size cap applies to the remote size before a byte is requested,
and an unreachable host answers 502 rather than pretending the file is missing. Nothing is
ever copied onto the Codeman host, and a same-named local file is never served under a
remote name.
Not available over ssh, and said so with a 400 instead of a misleading 404: editing in
place, Office previews and generated thumbnails (both need the bytes on the server's disk),
the file tree and path picker, and the tail viewer. Docker cases are unaffected, because
their workspace is bind-mounted at the same path.
## The path picker
For choosing a path rather than typing one. It appears in two places:
+28 -37
View File
@@ -93,8 +93,7 @@ export PUPPETEER_SKIP_DOWNLOAD="${PUPPETEER_SKIP_DOWNLOAD:-1}"
CLI_IDS=('claude' 'shell' 'opencode' 'codex' 'gemini' 'antigravity' 'pi' 'grok' 'deepseek' 'omp')
CLI_LABELS=('Claude' 'Shell' 'OpenCode' 'Codex' 'Gemini' 'Antigravity' 'Pi' 'Grok' 'DeepSeek' 'OMP')
CLI_ENABLED=(1 1 1 1 1 1 1 1 1 1)
CLI_KIND=('agent' 'shell' 'agent' 'agent' 'agent' 'agent' 'agent' 'agent' 'agent' 'agent')
CLI_NPM=('@anthropic-ai/claude-code' '' 'opencode-ai' '@openai/codex' '@google/gemini-cli' '' '@earendil-works/pi-coding-agent' '' '@deepseek-ai/dsh' '')
CLI_LAUNCHER_ONLY=(0 0 0 0 0 0 0 0 1 0)
CLI_DOCS=('https://docs.claude.com/claude-code' '' 'https://opencode.ai/docs' 'https://developers.openai.com/codex/cli' 'https://github.com/google-gemini/gemini-cli' 'https://antigravity.google/cli' 'https://pi.dev' 'https://github.com/xai-org/grok-build' 'https://github.com/deepseek-ai/deepseek-harness' 'https://omp.sh')
CLI_CMD_LINUX=('curl -fsSL https://claude.ai/install.sh | bash' '' 'curl -fsSL https://opencode.ai/install | bash' 'npm install -g @openai/codex' 'npm install -g @google/gemini-cli' 'curl -fsSL https://antigravity.google/cli/install.sh | bash' 'npm install -g --ignore-scripts @earendil-works/pi-coding-agent' 'curl -fsSL https://x.ai/cli/install.sh | bash' '' 'curl -fsSL https://omp.sh/install | sh')
CLI_CMD_DARWIN=('curl -fsSL https://claude.ai/install.sh | bash' '' 'curl -fsSL https://opencode.ai/install | bash' 'npm install -g @openai/codex' 'npm install -g @google/gemini-cli' 'curl -fsSL https://antigravity.google/cli/install.sh | bash' 'npm install -g --ignore-scripts @earendil-works/pi-coding-agent' 'curl -fsSL https://x.ai/cli/install.sh | bash' '' 'brew install can1357/tap/omp')
@@ -410,22 +409,6 @@ check_build_tools() {
# test/install-sh-detection-parity.test.ts: the process PATH first (each declared
# binary name in turn), then each known install path, dir-major.
# Index of "$1" in CLI_IDS -> CLI_IDX, returning 1 with CLI_IDX=-1 when unknown.
# A global rather than an echo because this runs inside loops, and a subshell per
# lookup is a fork per CLI per call site.
CLI_IDX=-1
_cli_index() {
local want="$1" i
CLI_IDX=-1
for ((i = 0; i < ${#CLI_IDS[@]}; i++)); do
if [[ "${CLI_IDS[$i]}" == "$want" ]]; then
CLI_IDX=$i
return 0
fi
done
return 1
}
# `dsh` is the hardest name of the lot: Debian ships an unrelated `dsh`
# (dancer's shell). The server-side resolver settles it by demanding the
# harness's own help banner; detection here only feeds the "you have no AI CLI"
@@ -465,7 +448,8 @@ _cli_candidate_ok() {
# Resolve every CLI in ONE pass, memoized.
#
# CLI_FOUND_PATH is parallel to CLI_IDS ('' when not found). CLI_FOUND_COUNT
# CLI_FOUND_PATH is parallel to CLI_IDS ('' when not found, and also '' for a
# DISABLED entry — it is never probed at all, see below). CLI_FOUND_COUNT
# counts only ENABLED entries that have a binary to look for, which is what the
# "no AI CLI found" gate asks about — `shell` has no binary and must never make
# that gate think an agent is installed.
@@ -485,6 +469,16 @@ detect_all_clis() {
for ((i = 0; i < ${#CLI_IDS[@]}; i++)); do
found=""
# A disabled entry is never even probed: every consumer already filters
# on CLI_ENABLED before showing anything, so the command-v/stat calls
# below would be pure waste — and, unlike filtering downstream, skipping
# the probe here is what makes CLI_ENABLED mean "look for it" rather
# than just "offer it once found".
if [[ "${CLI_ENABLED[$i]}" != "1" ]]; then
CLI_FOUND_PATH[$i]=""
continue
fi
# 1. The process PATH, each declared binary name in turn.
bin_end=$((${CLI_BIN_OFF[$i]} + ${CLI_BIN_LEN[$i]}))
for ((j = ${CLI_BIN_OFF[$i]}; j < bin_end; j++)); do
@@ -520,20 +514,6 @@ detect_all_clis() {
return 0
}
# Is this CLI installed? Unknown id is "no", never an error.
check_cli() {
detect_all_clis
_cli_index "$1" || return 1
[[ -n "${CLI_FOUND_PATH[$CLI_IDX]}" ]]
}
# Where it was found, or nothing.
get_cli_path() {
detect_all_clis
_cli_index "$1" || return 1
printf '%s\n' "${CLI_FOUND_PATH[$CLI_IDX]}"
}
# ----------------------------------------------------------------------------
# Catalogue helpers
# ----------------------------------------------------------------------------
@@ -589,7 +569,12 @@ cli_catalog_names() {
# the registry but an empty one here: installing the launcher alone leaves
# nothing that can drive a pane, so the generator withholds the command for
# any launcherProfile entry (see installCommandFor in generate-cli-catalog.mts)
# and this hint falls through to the docs URL instead.
# and this hint falls through to the docs URL instead — CLI_LAUNCHER_ONLY adds
# one line explaining WHY it is a docs link and not a command, so a user who
# follows that link straight to `npm install -g @deepseek-ai/dsh` (which the
# docs page itself documents) does not land back in the same "installed but
# cannot drive a pane" trap the menu exists to avoid. Data-driven, not an id
# check: any future launcherProfile entry gets the same caveat for free.
cli_catalog_print_install_hints() {
detect_all_clis
local i
@@ -601,6 +586,9 @@ cli_catalog_print_install_hints() {
echo -e " ${CYAN}${CLI_INSTALL_CMD_TRUSTED[$i]}${NC} # ${CLI_LABELS[$i]}"
elif [[ -n "${CLI_DOCS[$i]}" ]]; then
echo -e " ${CLI_LABELS[$i]}: see ${CYAN}${CLI_DOCS[$i]}${NC}"
if [[ "${CLI_LAUNCHER_ONLY[$i]}" == "1" ]]; then
echo -e " (installs a launcher only: it still needs a terminal profile, and Codeman's Run menu can add one)"
fi
fi
done
}
@@ -679,9 +667,12 @@ offer_ai_cli_install() {
local cli_choice=""
if [[ "$NONINTERACTIVE" == "1" ]] || ! has_tty; then
# Explicit automation opt-in: default to the first offered entry,
# which is registry order, which is Claude Code (order 0) — the
# same default this prompt has always taken non-interactively.
# Explicit automation opt-in: default to the first OFFERED entry.
# That is registry order, which is Claude Code (order 0), UNLESS
# this is a wget-only host and Claude's curl one-liner was just
# filtered out of offer_idx above — there, the first survivor is
# whichever npm-based entry sorts earliest (Codex today), not
# Claude. Printed either way so the choice is never silent.
cli_choice="1"
info "CODEMAN_NONINTERACTIVE=1: defaulting to ${CLI_LABELS[${offer_idx[0]}]}"
else
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.29.1",
"version": "1.30.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.29.1",
"version": "1.30.0",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.29.1",
"version": "1.30.0",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
+1 -1
View File
@@ -1,7 +1,7 @@
{
"name": "codeman",
"description": "Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.29.1",
"version": "1.30.0",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
+7 -6
View File
@@ -135,8 +135,7 @@ export function renderInstallShBlock(entries: CliEntry[] = STOCK_CLIS): string {
const ids: string[] = [];
const labels: string[] = [];
const enabled: string[] = [];
const kinds: string[] = [];
const npm: string[] = [];
const launcherOnly: string[] = [];
const docs: string[] = [];
const cmdLinux: string[] = [];
const cmdDarwin: string[] = [];
@@ -151,8 +150,11 @@ export function renderInstallShBlock(entries: CliEntry[] = STOCK_CLIS): string {
ids.push(shQuote(entry.id as string));
labels.push(shQuote(entry.label));
enabled.push(entry.enabled ? '1' : '0');
kinds.push(shQuote(entry.kind));
npm.push(shQuote(entry.discovery.install.npmPackage ?? ''));
// Parallel to CLI_IDS: 1 when this entry's install command installs a launcher rather
// than something that can drive a pane on its own (see installCommandFor above). Purely
// derived from discovery.launcherProfile — install.sh's hint printer reads this to add a
// caveat instead of hardcoding which id it means.
launcherOnly.push(entry.discovery.launcherProfile ? '1' : '0');
docs.push(shQuote(entry.discovery.install.docsUrl ?? ''));
cmdLinux.push(shQuote(installCommandFor(entry, 'linux')));
cmdDarwin.push(shQuote(installCommandFor(entry, 'darwin')));
@@ -195,8 +197,7 @@ export function renderInstallShBlock(entries: CliEntry[] = STOCK_CLIS): string {
arr('CLI_IDS', ids),
arr('CLI_LABELS', labels),
arr('CLI_ENABLED', enabled),
arr('CLI_KIND', kinds),
arr('CLI_NPM', npm),
arr('CLI_LAUNCHER_ONLY', launcherOnly),
arr('CLI_DOCS', docs),
arr('CLI_CMD_LINUX', cmdLinux),
arr('CLI_CMD_DARWIN', cmdDarwin),
+278
View File
@@ -0,0 +1,278 @@
/**
* @fileoverview Decide which sessions a host reboot destroyed and may be rebuilt.
*
* A server restart and a host reboot both leave `reconcileSessions()` reporting
* dead sessions, and they need opposite handling. A server restart leaves the
* tmux panes running, so recovery ATTACHES to them. A host reboot takes the tmux
* server down with it, so there is nothing to attach to and the pane has to be
* created again. This module holds the decision half of that second case, kept
* free of tmux and disk access so it can be unit tested without either. Every
* observation it reads is gathered by the caller and passed in.
*
* "Eligible" here means a session the user did not end on purpose. The rule that
* an intentional kill or detach is never auto-revived is enforced at runtime by
* an in-memory guard in `TmuxManager`, and memory does not survive a reboot. The
* durable equivalent is the record `cleanupSession()` leaves behind. An unpinned
* kill deletes the record outright, so it is already absent here. A pinned kill
* goes through `demoteOrRemoveSession()` and lands as `status: 'stopped'`, which
* is the marker this module refuses. Pruning keeps a pinned record WITHOUT
* touching its status, so a pinned session a reboot killed still reads `idle` or
* `busy` and stays eligible.
*
* ⚠️ Ending the AGENT rather than the session is a shape this module CANNOT
* recognise today, and a reboot restores it. `/exit` ends the CLI inside the
* pane, `remain-on-exit` keeps the pane, and the PTY Codeman owns is the
* `tmux attach-session` process, which stays alive throughout — so no exit
* handler runs, no lifecycle `exit` is logged, and the record keeps both its pid
* and `status: 'idle'`. Nothing durable distinguishes it from a session that was
* simply idle when the power went. Ark0N/Codeman#446 covers making Codeman
* notice the dead pane; until a record can say the agent is gone, this pass will
* offer those sessions back, and the user dismisses or closes them.
*
* The `pid` check below is therefore NOT that rule. It refuses a record whose
* attach process was already gone, which is a session that never started or
* whose pane died outright.
*
* @dependencies types (SessionState), config/cli-registry
* @consumedby web/server (plan build at boot), web/routes/reboot-restore-routes
*
* @module reboot-restore
*/
import type { SessionState } from './types.js';
import { getCli } from './config/cli-registry/registry.js';
/** Session statuses a reboot restore may rebuild. `stopped` is the kill marker. */
const RESTORABLE_STATUSES: ReadonlySet<string> = new Set(['idle', 'busy', 'error']);
/** Observations the reboot heuristic reads. Gathered by the caller, never here. */
export interface RebootEvidence {
/** Sessions that still had a live pane during reconciliation. */
livePaneCount: number;
/** Sessions reconciliation just marked dead. */
deadSessionCount: number;
/** `os.uptime()`, in seconds. */
uptimeSeconds: number;
/** Newest `lastActivityAt` across the persisted records, in ms since the epoch. */
newestPersistedActivityAt: number;
/** `Date.now()` when the evidence was gathered, in ms. */
now: number;
}
/**
* Decide whether the machine plausibly rebooted rather than the server restarting.
*
* Two signals have to agree. The socket must hold no panes at all while state
* still lists sessions, which rules out an ordinary server restart. The host
* must also have booted after the newest persisted session activity, which is
* the corroboration `os.uptime()` provides cheaply. A wiped tmux socket on a
* long-uptime host fails the second test, so a user who killed the tmux server
* by hand does not get every session offered back to them.
*
* This heuristic decides whether to ASK, never whether to act. A wrong yes costs
* the user a banner they dismiss, because the restore itself waits for a click.
*
* ⚠️ `os.uptime()` reports the HOST's uptime, which a container shares, and that
* cuts BOTH ways rather than simply switching the feature off in Docker. After a
* genuine host reboot a containerized Codeman sees the host's short uptime, so the
* banner DOES appear and the feature works. What it cannot see is a container-only
* restart: the host uptime is long, the boot test fails, and no banner appears
* although every in-container pane is gone (`docker/server.Dockerfile` installs
* tmux inside the Codeman container, and the self-updater restarts the Compose
* deployment by exiting the container, so that is the case where this would help
* most). Failing quiet is the safe direction, and closing the gap needs a boot
* signal the container owns (PID 1's start time, gated on the existing
* `isRunningInContainer()`) rather than a wider heuristic.
*/
export function looksLikeHostReboot(evidence: RebootEvidence): boolean {
if (evidence.deadSessionCount === 0) return false;
if (evidence.livePaneCount > 0) return false;
if (evidence.newestPersistedActivityAt <= 0) return false;
const bootedAt = evidence.now - evidence.uptimeSeconds * 1000;
return bootedAt > evidence.newestPersistedActivityAt;
}
/**
* Pick the conversation the rebuilt pane should resume.
*
* The chain's tail is the newest conversation the session was holding, which is
* what a compact or a clear leaves behind; `resumeSessionId` covers a session
* that was itself started as a resume, and the session id is the original
* conversation for everything else.
*/
export function resolveResumeConversationId(state: SessionState): string {
const chain = state.claudeSessionChain;
const chainTail = Array.isArray(chain) && chain.length > 0 ? chain[chain.length - 1] : undefined;
return chainTail || state.resumeSessionId || state.id;
}
/**
* Why one session was passed over. Reported for logging and shown to the user.
*
* The first seven are decided before anything is built. `capacity-reached` and
* `rebuild-failed` can only happen once a click is spending the plan, and they
* are the two the banner must not confuse with a missing workspace: one means
* "try again after closing something", the other means the CLI would not start.
*/
export interface RebootRestoreRejection {
sessionId: string;
reason:
| 'no-persisted-record'
| 'intentionally-ended'
| 'not-running'
| 'respawn-blocked'
| 'remote-or-docker'
| 'unsupported-mode'
| 'no-working-dir'
| 'workspace-missing'
| 'workspace-forbidden'
| 'already-live'
| 'capacity-reached'
| 'rebuild-failed';
}
/** One restorable session, as the banner shows it and the rebuild replays it. */
export interface RebootRestoreEntry {
sessionId: string;
name?: string;
workingDir: string;
owner?: string;
mode: string;
/** The conversation the rebuilt pane resumes. */
resumeConversationId: string;
/**
* The persisted record, kept whole so the rebuild can replay what it held.
* Read at boot, before pruning deletes it, and held in memory until the click.
*/
state: SessionState;
}
export interface RebootRestorePlan {
restore: RebootRestoreEntry[];
skipped: RebootRestoreRejection[];
}
/**
* Split the sessions reconciliation just killed into the ones a reboot restore
* may offer and the ones it must leave alone.
*
* @param deadSessionIds Session ids `reconcileSessions()` reported as dead.
* @param persisted The `state.json` session records, which `cleanupStaleSessions()`
* has not pruned yet at the point this runs.
* @param workspaceExists Whether a working directory is still on disk. A tmux
* session can outlive its deleted repo, and rebuilding one there would scaffold
* an empty tree. The caller owns the disk access; the click re-checks, because
* a repo can be deleted between the boot and the click.
*/
export function planRebootRestore(
deadSessionIds: readonly string[],
persisted: Readonly<Record<string, SessionState>>,
workspaceExists: (workingDir: string) => boolean
): RebootRestorePlan {
const restore: RebootRestoreEntry[] = [];
const skipped: RebootRestoreRejection[] = [];
for (const sessionId of deadSessionIds) {
const state = persisted[sessionId];
if (!state) {
// An unpinned kill already deleted the record, so absence IS the guard.
skipped.push({ sessionId, reason: 'no-persisted-record' });
continue;
}
if (!RESTORABLE_STATUSES.has(state.status)) {
// A pinned kill was demoted to `stopped`. Reviving it would undo the kill.
skipped.push({ sessionId, reason: 'intentionally-ended' });
continue;
}
if (state.pid === null || state.pid === undefined) {
// No attach process when the record was last written: the session never
// started, or its pane died outright rather than its agent exiting inside a
// surviving pane. Either way there was nothing running to bring back.
//
// ⚠️ This does NOT catch a session the user ended with `/exit`. See the
// module header: that leaves the pid in place, because the pid is the tmux
// attach process and `remain-on-exit` keeps it alive.
//
// Conservative on purpose. A session that somehow persisted no pid while
// genuinely running is not offered, and its conversation stays reachable
// from the Resume list, which is where every session would be without this
// feature.
skipped.push({ sessionId, reason: 'not-running' });
continue;
}
if (state.respawnBlocked === true) {
// The crash-loop breaker tripped on this pane. Re-creating it restarts the loop.
skipped.push({ sessionId, reason: 'respawn-blocked' });
continue;
}
if (state.remote || state.docker) {
// Both need another host or a container to be up, which a just-booted machine
// cannot promise. The remote reconnect watcher owns the remote case already.
skipped.push({ sessionId, reason: 'remote-or-docker' });
continue;
}
// Capability, not a CLI id: this pass resumes by handing the CLI a conversation
// id through the top-level `resumeSessionId`, which only a CLI whose history the
// claude-jsonl reader understands can consume that way. Others carry their thread
// id in their own `<Mode>Config`, which this pass does not thread through.
if (getCli(state.mode ?? 'claude')?.capabilities.transcript !== 'claude-jsonl') {
skipped.push({ sessionId, reason: 'unsupported-mode' });
continue;
}
if (!state.workingDir) {
skipped.push({ sessionId, reason: 'no-working-dir' });
continue;
}
if (!workspaceExists(state.workingDir)) {
skipped.push({ sessionId, reason: 'workspace-missing' });
continue;
}
restore.push({
sessionId,
name: state.name,
workingDir: state.workingDir,
owner: state.owner,
mode: state.mode ?? 'claude',
resumeConversationId: resolveResumeConversationId(state),
state,
});
}
return { restore, skipped };
}
/**
* Drop the entries whose conversation is already on screen.
*
* Hours can pass between the boot that built the plan and the click that spends
* it, and the Resume list can reach the same conversation in the meantime. Two
* panes running `claude --resume` on one conversation is the failure this
* prevents, so a match on either the session id or the conversation id is enough
* to skip the entry.
*/
export function rejectAlreadyLive(
entries: readonly RebootRestoreEntry[],
liveSessionIds: ReadonlySet<string>,
liveConversationIds: ReadonlySet<string>
): RebootRestorePlan {
const restore: RebootRestoreEntry[] = [];
const skipped: RebootRestoreRejection[] = [];
for (const entry of entries) {
if (liveSessionIds.has(entry.sessionId) || liveConversationIds.has(entry.resumeConversationId)) {
skipped.push({ sessionId: entry.sessionId, reason: 'already-live' });
continue;
}
restore.push(entry);
}
return { restore, skipped };
}
/** Newest `lastActivityAt` across persisted records, or 0 when there are none. */
export function newestPersistedActivity(persisted: Readonly<Record<string, SessionState>>): number {
let newest = 0;
for (const state of Object.values(persisted)) {
const stamp = state.lastActivityAt ?? state.createdAt ?? 0;
if (stamp > newest) newest = stamp;
}
return newest;
}
+97
View File
@@ -0,0 +1,97 @@
/**
* @fileoverview The env-var half of the multi-user privilege clamp.
*
* A session's `envOverrides` can hand back privilege that the per-CLI config
* clamp removed, so a non-granted owner's overrides get the privileged keys
* stripped before the session is built. The create and resume routes are what
* this bites on: they clamp what a request asked for.
*
* The reboot-restore route calls it as defence in depth, and today it can strip
* nothing. `Session.getEnvOverridesForPersist()` keeps only `CLAUDE_CODE_*` and
* `CLAUDE_CONFIG_DIR` out of a session's overrides, claude's `privilegedEnvKeys`
* are the five `ANTHROPIC_*` names, and that pass admits claude alone — so a
* persisted record cannot carry a clamped key. The call is there for the day the
* persisted set widens. The grant re-resolution that does bite on that path is
* `resolveClaudeModeForUsername`, which recomputes the permission mode.
*
* This lives outside `web/routes` on purpose. The question it answers is about
* session privilege rather than about HTTP, and `cron/cron-service.ts` sets the
* precedent by importing `canUsernameRunPrivilegedCommands` from `user-store.ts`
* directly and re-resolving the owner's grant when a job fires. Every caller here
* re-resolves the grant at the moment it builds a session, for the same reason.
*
* @dependencies user-store (canUsernameRunPrivilegedCommands), config/cli-registry
* @consumedby web/routes/session-routes, web/routes/reboot-restore-routes
*
* @module session-env-clamp
*/
import { canUsernameRunPrivilegedCommands } from './user-store.js';
import { enabledClis } from './config/cli-registry/registry.js';
/**
* Env-var keys a non-granted owner must not be able to set, because each one
* hands back privilege `clampExternalCliBypassForOwner()` just removed, or redirects a
* credential-resolution endpoint.
*
* The DeepSeek three are reachable because `DSH_*` and `DEEPSEEK_*` are
* allowlisted `envOverrides` prefixes (schemas.ts) — which they have to be, since
* that is also how a user configures the harness's non-privileged knobs.
*
* - `DSH_PERMISSION_MODE` IS the harness's permission switch. Every other CLI's
* bypass is a command-line FLAG, reachable only through the per-CLI config the
* clamp already owns; this one is an env var, so the config clamp alone is
* half a gate.
* - `DSH_HOME` points the launcher at a profile tree, and a profile's plugin code
* executes at BOOT, before any approval row can apply. A user who can write a
* workspace can put a profile in it, so this is the wider of the two.
* - `DEEPSEEK_BASE_URL` aims the provider endpoint, and `_configureCliEnv()`
* forwards the SERVER's own `DEEPSEEK_API_KEY` into every dsh pane before
* `applyEnvOverrides()` runs — so a non-granted owner who could set the base
* URL would have the operator's API key sent as a bearer credential to a host
* of their choosing. (`DEEPSEEK_API_KEY` itself stays overridable: supplying
* your OWN key removes privilege rather than granting it.)
* - `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN` are where omp resolves
* credentials from — the same shape as `DEEPSEEK_BASE_URL` above, reachable
* because `OMP_*` is an allowlisted prefix. Unlike DeepSeek, Codeman does not
* forward any operator-held key into an omp pane today (omp's provider
* credentials live in `~/.omp` config files, not env vars), so there is no
* known concrete exfiltration path yet — clamped defensively anyway, since a
* non-granted owner redirecting where a shared multi-tenant deployment
* resolves auth from is not something to allow silently (found in
* Ark0N/Codeman#353 review; omp's own knobs are otherwise mostly `PI_*`,
* already allowlisted for pi and not addressed here — see resolveOmpHome()).
*/
export function ownerClampedEnvKeys(): string[] {
return enabledClis().flatMap((entry) => entry.capabilities.privilegedEnvKeys);
}
/**
* Env-var half of the multi-user bypass clamp.
*
* `clampExternalCliBypassForOwner()` in `web/routes/session-routes.ts` clamps the
* per-CLI CONFIG, and for every CLI
* but DeepSeek that is the whole story. Here it is not: `applyEnvOverrides()` runs
* AFTER `_configureCliEnv()` in tmux-manager, so an override sent on the SAME
* request lands last and wins, and a non-granted owner could restore
* `danger-full-access` on the very request the config clamp downgraded.
*
* Keys are DROPPED rather than rewritten: dropping falls through to what
* `_configureCliEnv()` exports, which is the clamped config and the server's own
* `DSH_HOME`, i.e. exactly the intended state. No-op in single-user mode and for a
* granted owner, like every other clamp here
* (`canUsernameRunPrivilegedCommands()` returns true when `!isMultiUserMode()`),
* and it returns the caller's own object untouched when there is nothing to strip.
*/
export async function clampEnvOverridesForOwner(
owner: string | undefined,
envOverrides: Record<string, string> | undefined
): Promise<Record<string, string> | undefined> {
if (!envOverrides) return envOverrides;
const keys = ownerClampedEnvKeys();
if (!keys.some((key) => key in envOverrides)) return envOverrides;
if (await canUsernameRunPrivilegedCommands(owner)) return envOverrides;
const clamped = { ...envOverrides };
for (const key of keys) delete clamped[key];
return clamped;
}
+13
View File
@@ -1499,6 +1499,19 @@ export class Session extends EventEmitter {
this._pinnedAt = pinned ? Date.now() : null;
}
/**
* Restore a pin from a persisted record, keeping the moment it was pinned.
*
* `setPinned()` stamps `pinnedAt` with now, which is right for a user pinning a
* session and wrong for a restore: the session-manager orders its pinned group
* by that stamp, so a restored session would jump to the front of a list it had
* been sitting further down.
*/
restorePin(pinned: boolean, pinnedAt?: number): void {
this._pinned = pinned;
this._pinnedAt = pinned ? (pinnedAt ?? Date.now()) : null;
}
get flickerFilterEnabled(): boolean {
return this._flickerFilterEnabled;
}
+36
View File
@@ -4,6 +4,7 @@
*/
import type { Session } from '../../session.js';
import type { SessionState } from '../../types.js';
export interface SessionPort {
readonly sessions: ReadonlyMap<string, Session>;
@@ -12,5 +13,40 @@ export interface SessionPort {
setupSessionListeners(session: Session): Promise<void>;
persistSessionState(session: Session): void;
persistSessionStateNow(session: Session): void;
/**
* Re-apply the persisted state a freshly CONSTRUCTED session does not carry.
*
* A `Session` built from a record holds only what its constructor takes, so
* persisting it would otherwise REPLACE the fuller record with the reduced one.
* Two phases: `before-spawn` shapes the pane (the custom-model environment and
* the nice priority) and must precede `startInteractive()`; `after-spawn` is
* the session's own history (the pin, token and cost totals, auto-compact,
* auto-clear, auto-resume, colour, image watcher, flicker filter) and must NOT
* land on a session whose pane failed to start.
*/
reapplyPersistedSessionState(
session: Session,
saved: SessionState,
phase: 'before-spawn' | 'after-spawn',
options?: {
/**
* Re-arm a PENDING auto-resume schedule from the record's `autoResumeAt`.
* Default true, which is what a Codeman restart wants: the limit footer
* will not reprint on its own, so dropping the stamp there strands the
* pause. A reboot restore passes false: the stamp predates the reboot,
* the pane is new, and re-arming means every restored session types
* `continue` into itself about a minute after one click. Auto-resume
* stays ENABLED either way, so it re-arms on fresh evidence.
*/
rearmAutoResumeSchedule?: boolean;
}
): Promise<void>;
/**
* Undo a session that was registered but never got a working pane: the map
* entry, its tab-layout slot, and any pane the launch created before throwing.
* Unlike {@link cleanupSession} it leaves the persisted record, the lifetime
* token totals, the Ralph state and the workspace's own files untouched.
*/
discardPartiallyBuiltSession(sessionId: string): Promise<void>;
getSessionStateWithRespawn(session: Session): unknown;
}
+10
View File
@@ -957,6 +957,10 @@ class CodemanApp {
this.registerServiceWorker();
// Fetch tunnel status for header indicator (desktop only)
this.loadTunnelStatus();
// Ask whether a host reboot left sessions worth rebuilding (banner, never
// automatic). handleInit() re-reads it on every SSE init; this covers the
// path where that event never arrives.
this.initRebootRestoreBanner?.();
// Share a single settings fetch between both consumers
const settingsPromise = fetch('/api/settings').then(r => r.ok ? r.json() : null).then(env => env?.data ?? null).catch(() => null);
this.loadQuickStartCases(null, settingsPromise);
@@ -3759,6 +3763,12 @@ class CodemanApp {
// a fresh load / reconnect (authoritative; wins over the localStorage restore).
if (data.planUsage) this.updatePlanUsageChip(data.planUsage);
// A board left open across a host reboot reconnects HERE, to a server that came
// back with an empty session list. The reboot-restore offer is built at boot,
// before any client could be listening, so re-read it on every init rather than
// only on the page-load path.
this.refreshRebootRestoreBanner?.();
// Update version displays (header and toolbar)
if (data.version) {
const versionEl = this.$('versionDisplay');
+19
View File
@@ -213,6 +213,24 @@
<button class="offline-banner-retry" id="offlineBannerRetry" onclick="app.retryConnection()">Retry now</button>
</div>
<!-- Reboot-restore offer: shown when the server found sessions a host reboot
killed and is asking whether to rebuild them. Populated by
reboot-restore-ui.js; nothing is created until the user clicks. -->
<div class="reboot-restore-banner" id="rebootRestoreBanner" role="status" hidden>
<span class="reboot-restore-banner-icon" aria-hidden="true">↺</span>
<span class="reboot-restore-banner-text" id="rebootRestoreBannerText"></span>
<span class="reboot-restore-banner-detail" id="rebootRestoreBannerDetail"></span>
<span class="reboot-restore-banner-note">Conversations return; terminal history does not.</span>
<button
class="reboot-restore-banner-accept"
id="rebootRestoreBannerAccept"
onclick="app.restoreRebootSessions()"
>
Restore
</button>
<button class="reboot-restore-banner-dismiss" onclick="app.dismissRebootRestore()">Dismiss</button>
</div>
<!-- Timer Banner (shown when timed run is active) -->
<div class="timer-banner" id="timerBanner" style="display: none;">
<div class="timer-content">
@@ -3535,6 +3553,7 @@
<script defer src="readmymind-ui.js"></script>
<script defer src="ultracode-panel.js"></script>
<script defer src="approvals-ui.js"></script>
<script defer src="reboot-restore-ui.js"></script>
<script defer src="admin-ui.js"></script>
<script defer src="session-ui.js"></script>
<script defer src="webview-tabs.js"></script>
+37
View File
@@ -3241,6 +3241,43 @@ html:is([data-skin="paper-gray"], [data-skin="solarized-light"], [data-skin="cat
other banners. The overlay is fixed and handles its own insets.
============================================================================ */
@media (max-width: 599px) {
/* Reboot-restore banner: the same treatment as the offline banner below. Its
text and note are nowrap and the two buttons cannot shrink, so without this
the actions are pushed off a phone-width viewport and become unreachable. */
.reboot-restore-banner {
padding: 0.4rem 0.5rem;
padding-left: calc(0.5rem + var(--safe-area-left));
padding-right: calc(0.5rem + var(--safe-area-right));
font-size: 0.7rem;
gap: 0.4rem;
}
/* The session names and the scrollback note are the first things to go. The
count plus the two buttons carry the message on their own, and the note
survives as the accept button's title. */
.reboot-restore-banner-detail,
.reboot-restore-banner-note {
display: none;
}
/* A flex item will not shrink below its content width at the default
`min-width: auto`, so without this the nowrap text pushes the buttons off a
360px viewport and the ellipsis never engages. */
.reboot-restore-banner-text {
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
}
.reboot-restore-banner-accept,
.reboot-restore-banner-dismiss {
padding: 0.25rem 0.5rem;
}
.reboot-restore-banner-accept {
margin-left: auto;
}
.offline-banner {
padding: 0.4rem 0.5rem;
padding-left: calc(0.5rem + var(--safe-area-left));
+142
View File
@@ -0,0 +1,142 @@
/**
* @fileoverview Reboot-restore banner: offer back the sessions a host reboot destroyed.
*
* A host reboot takes the tmux server down with it, so every session's pane dies
* and the board comes up empty. The server works out what was running from the
* records it still holds at boot, and this banner asks the user whether to
* rebuild them. Nothing is created until they click, because the server's
* reboot guess is a heuristic and a wrong automatic restore would spawn CLI
* processes nobody asked for.
*
* Seeded from `GET /api/reboot-restore` on init and again on every SSE reconnect,
* because the tab most likely to want this is one that was open across the reboot
* and reconnects to a server that came back up with an empty board. Restore posts to
* `POST /api/reboot-restore/restore` and Dismiss posts to
* `POST /api/reboot-restore/dismiss`. Dismiss always clears the banner; Restore
* re-reads the plan afterwards, because the server puts back anything it could
* not build for a reason that may pass, such as a session limit or an agent that
* would not start. The restored sessions arrive as ordinary `session:created`
* events, so no extra rendering is needed here.
*
* The banner says that terminal history did not survive, because a restored
* session is a new pane: the conversation continues and the scrollback does not.
* Saying so is what keeps an empty pane from reading as a broken restore.
* Backend: src/web/reboot-restore-registry.ts, src/web/routes/reboot-restore-routes.ts.
*
* @mixin Extends CodemanApp.prototype via Object.assign
* @dependency app.js (CodemanApp class, showToast)
* @dependency api-client.js at runtime (this._api / this._apiJson)
* @loadorder 11.65, after approvals-ui.js and before admin-ui.js (11.7)
*/
/** Plain-language wording for one skip reason, for the toast after a restore. */
function rebootSkipReason(reason) {
switch (reason) {
case 'workspace-missing':
return 'workspace is gone';
case 'workspace-forbidden':
return 'workspace is outside your space';
case 'already-live':
return 'already open';
case 'capacity-reached':
return 'session limit reached';
case 'rebuild-failed':
return 'the agent would not start';
default:
return reason;
}
}
Object.assign(CodemanApp.prototype, {
/** Ask the server whether a reboot left anything on offer, and show the banner if so. */
async initRebootRestoreBanner() {
const data = await this._apiJson('/api/reboot-restore');
const sessions = data?.sessions ?? [];
if (sessions.length === 0) return;
this._rebootRestoreSessions = sessions;
this.renderRebootRestoreBanner();
},
renderRebootRestoreBanner() {
const banner = this.$('rebootRestoreBanner');
if (!banner) return;
const sessions = this._rebootRestoreSessions ?? [];
if (sessions.length === 0) {
banner.hidden = true;
return;
}
const count = sessions.length;
const text = this.$('rebootRestoreBannerText');
if (text) {
const noun = count === 1 ? 'session' : 'sessions';
text.textContent = `Restore ${count} ${noun} from before the reboot`;
}
const detail = this.$('rebootRestoreBannerDetail');
if (detail) {
// Names, so the user can tell what they are about to relaunch.
const names = sessions
.map((s) => s.name || s.workingDir?.split('/').pop() || s.id.slice(0, 8))
.slice(0, 4)
.join(', ');
detail.textContent = count > 4 ? `${names}, …` : names;
detail.title = sessions.map((s) => `${s.name || s.id}\n${s.workingDir}`).join('\n\n');
}
const accept = this.$('rebootRestoreBannerAccept');
// The note is hidden at phone width, so the warning travels on the button too.
if (accept) accept.title = 'Conversations return; terminal history does not.';
banner.hidden = false;
},
/** Rebuild everything on offer. The panes are new, so scrollback does not come back. */
async restoreRebootSessions() {
const button = this.$('rebootRestoreBannerAccept');
if (button) button.disabled = true;
const res = await this._api('/api/reboot-restore/restore', { method: 'POST', body: {} });
if (res && res.status === 409) {
if (button) button.disabled = false;
this.showToast?.('A restore is already running', 'info');
return;
}
// The uniform envelope wraps every /api payload; reading the outer object
// would report every count as zero.
const body = res && res.ok ? (await res.json().catch(() => null))?.data : null;
if (!body) {
if (button) button.disabled = false;
this.showToast?.('Could not restore the sessions', 'error');
return;
}
const restored = body.restored?.length ?? 0;
const skipped = body.skipped?.length ?? 0;
// Re-read rather than clearing: the server puts back anything it could not
// build for a reason that may pass, such as a session limit or an agent that
// would not start, and blanking the banner here would put those entries out
// of reach until a reload.
await this.refreshRebootRestoreBanner();
if (button) button.disabled = false;
if (restored > 0) {
const noun = restored === 1 ? 'conversation' : 'conversations';
this.showToast?.(`Restored ${restored} ${noun}. Terminal history did not survive the reboot.`, 'success');
}
if (skipped > 0) {
// Each reason means a different next step for the user, so they are not
// collapsed into one message: capacity clears by closing something, a
// failed start usually means the CLI is not on the server's PATH.
const reasons = new Set((body.skipped ?? []).map((s) => s.reason));
this.showToast?.(`${skipped} not restored: ${[...reasons].map(rebootSkipReason).join('; ')}`, 'warning');
}
},
/** Re-read the offer after a reconnect, for a tab that was open across the reboot. */
async refreshRebootRestoreBanner() {
const data = await this._apiJson('/api/reboot-restore');
this._rebootRestoreSessions = data?.sessions ?? [];
this.renderRebootRestoreBanner();
},
/** Drop the offer. The Resume list still reaches every one of these conversations. */
async dismissRebootRestore() {
this._rebootRestoreSessions = [];
this.renderRebootRestoreBanner();
await this._apiPost('/api/reboot-restore/dismiss', {});
},
});
+83
View File
@@ -2548,6 +2548,10 @@ body.solo-mode .header-tokens,
body.solo-mode .btn-notifications,
body.solo-mode .btn-multimonitor,
body.solo-mode .header-plan-usage,
/* A solo window shows ONE session and has no tab strip to put restored ones in,
so offering to rebuild a list of them there is an offer it cannot show the
result of. The dashboard that spawned this window carries the banner. */
body.solo-mode .reboot-restore-banner,
body.solo-mode .btn-lifecycle-log {
display: none !important;
}
@@ -15243,6 +15247,85 @@ html[data-skin="daylight-blue"] .welcome-btn-tunnel.active:hover {
skin, including the light ones. Visibility is driven by the `hidden`
attribute, so the display rules need !important to lose to it. */
/* Reboot-restore offer. Amber rather than red: nothing is wrong, the board is
asking a question, and the user can ignore it. See reboot-restore-ui.js. */
.reboot-restore-banner {
display: flex;
align-items: center;
gap: 0.6rem;
padding: 0.45rem 1rem;
background: linear-gradient(90deg, #b45309, #92400e);
border-bottom: 1px solid rgba(0, 0, 0, 0.35);
color: #fff;
font-size: 0.78rem;
font-weight: 600;
letter-spacing: 0.01em;
flex-shrink: 0;
z-index: 1250;
}
.reboot-restore-banner[hidden] {
display: none !important;
}
.reboot-restore-banner-icon {
flex-shrink: 0;
font-size: 0.95rem;
line-height: 1;
}
.reboot-restore-banner-text {
white-space: nowrap;
}
.reboot-restore-banner-detail {
color: rgba(255, 255, 255, 0.8);
font-weight: 500;
/* A flex item will not shrink below its content width at the default
`min-width: auto`, so without this the session names push the buttons out of
the line between the phone breakpoint and full width. */
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.reboot-restore-banner-note {
color: rgba(255, 255, 255, 0.75);
font-weight: 500;
white-space: nowrap;
margin-left: auto;
}
.reboot-restore-banner-accept,
.reboot-restore-banner-dismiss {
flex-shrink: 0;
padding: 0.2rem 0.6rem;
border-radius: 5px;
border: 1px solid rgba(255, 255, 255, 0.55);
background: rgba(255, 255, 255, 0.12);
color: #fff;
font-size: 0.72rem;
font-weight: 600;
cursor: pointer;
}
.reboot-restore-banner-accept:hover,
.reboot-restore-banner-dismiss:hover {
background: rgba(255, 255, 255, 0.24);
}
.reboot-restore-banner-accept:disabled {
opacity: 0.6;
cursor: default;
}
.reboot-restore-banner-dismiss {
border-color: rgba(255, 255, 255, 0.3);
background: transparent;
font-weight: 500;
}
.offline-banner {
display: flex;
align-items: center;
@@ -66,6 +66,38 @@
let composing = false;
const pending = [];
/**
* Resolve every candidate still pending, right now, instead of waiting for
* its zero-delay timer.
*
* Android soft keyboards commit the last character and send the Enter key
* in ONE InputConnection transaction: the `input` event and the Enter
* keydown are both processed before any timer runs. Left on its timer the
* candidate lost BOTH ways — xterm emits '\r' synchronously from the Enter
* keydown (so the local-echo composer submitted the prompt without the
* character), and that '\r' bumps `canonicalCount`, so the candidate then
* read "xterm spoke for this keystroke" and stood down, dropping the
* character outright. That is the "every message loses its last character"
* report from phones.
*
* Draining at the next keydown is correct on both counts: the counter still
* holds the value it had while this candidate's keystroke was current, and
* the byte reaches the composer ahead of whatever the new key emits.
*/
function flushPending() {
for (const candidate of pending.splice(0)) {
if (candidate.timer !== null) {
try {
clearTimer(candidate.timer);
} catch {
// A broken timer host must not break input handling.
}
candidate.timer = null;
}
resolveCandidate(candidate);
}
}
function cancelPending() {
for (const candidate of pending.splice(0)) {
candidate.active = false;
@@ -111,6 +143,11 @@
*/
function handleKeyEvent(event) {
if (destroyed || event?.type !== 'keydown') return;
// Settle the PREVIOUS keystroke before this one can move the counter or
// reach the PTY — see flushPending(). This runs from xterm's custom key
// handler, i.e. before xterm processes the key, so a recovered character
// is always ordered ahead of the bytes this keydown produces.
flushPending();
keydownSnapshot = canonicalCount;
}
+42 -8
View File
@@ -3543,6 +3543,29 @@ Object.assign(CodemanApp.prototype, {
this._sendInputAsync(this.activeSessionId, text);
},
/**
* Re-assert a history anchor captured before a terminal write (#358).
*
* Called from xterm's write callback, never synchronously after write():
* xterm parses on its own schedule, so the buffer only carries the redraw's
* effect once that callback fires. A null anchor means the user was following
* live output and nothing needs restoring.
*/
_restoreTerminalViewport(preserveViewportY, sessionId) {
if (preserveViewportY === null || preserveViewportY === undefined) return;
// The anchor is a row index into the buffer it was captured from. Now that
// this runs a parse later instead of synchronously, a session switch can land
// in between: selectSession() resets the terminal and chunk-loads the new
// session's scrollback, and scrolling THAT buffer to a row that meant
// something in the previous one is not a restore, it is a jump to an
// arbitrary place. Both checks cover one half of that window.
if (sessionId !== undefined && sessionId !== this.activeSessionId) return;
if (this._isLoadingBuffer) return;
if (typeof this.terminal?.scrollToLine !== 'function') return;
if (this.terminal.buffer?.active?.viewportY === preserveViewportY) return;
this.terminal.scrollToLine(preserveViewportY);
},
/**
* Flush pending writes to terminal, processing DEC 2026 sync markers.
* Strips markers and writes content atomically within a single frame.
@@ -3578,6 +3601,8 @@ Object.assign(CodemanApp.prototype, {
// scroll-to-bottom below, where it protects against a mid-flush race.
const preserveViewportY =
this.terminal.buffer?.active && !this.isTerminalAtBottom() ? this.terminal.buffer.active.viewportY : null;
// Which buffer the anchor belongs to, checked again when the write parses.
const flushSessionId = this.activeSessionId;
const writeChunk = joined.slice(0, MAX_FRAME_BYTES);
if (_joinedLen > MAX_FRAME_BYTES) {
@@ -3592,6 +3617,16 @@ Object.assign(CodemanApp.prototype, {
this.terminal.write(writeChunk, () => {
this._terminalWriteInFlight = false;
this._terminalWriteInFlightBytes = 0;
// Restore INSIDE the callback (#358). xterm parses asynchronously, so
// the moment write() returns the buffer has not moved yet: the old
// restore ran here, found viewportY still equal to the anchor, and did
// nothing at all — then the parse landed and a cursor-addressed Codex
// redraw dragged the viewport to the live bottom with nothing left to
// pull it back. The callback is xterm's own "this chunk is parsed"
// signal, which is the earliest point the anchor can actually be
// reasserted. (The synchronous version passed its regression test only
// because the test's write mock moved the viewport synchronously.)
this._restoreTerminalViewport(preserveViewportY, flushSessionId);
this._scheduleTerminalWriteFlush();
});
} catch (err) {
@@ -3599,13 +3634,6 @@ Object.assign(CodemanApp.prototype, {
this._terminalWriteInFlightBytes = 0;
throw err;
}
if (
preserveViewportY !== null &&
this.terminal.buffer?.active?.viewportY !== preserveViewportY &&
typeof this.terminal.scrollToLine === 'function'
) {
this.terminal.scrollToLine(preserveViewportY);
}
const bytesThisFrame = deferred ? MAX_FRAME_BYTES : _joinedLen;
const _dt = performance.now() - _t0;
if (_dt > 100 || deferred)
@@ -3617,7 +3645,13 @@ Object.assign(CodemanApp.prototype, {
// Give manual scroll-up gestures a short grace window so high-frequency
// Codex status ticks do not snap the viewport back while the user is
// trying to inspect earlier output.
if (this._wasAtBottomBeforeWrite && !this._hasRecentUserScrollUp()) {
//
// A live anchor wins outright. The two flags are captured at different
// moments (_wasAtBottomBeforeWrite at the frame's first batchTerminalWrite,
// the anchor at flush time), so a scroll-up in between leaves both set; now
// that the anchor is reasserted after the parse, running both would jump to
// the bottom and then back one frame later instead of simply staying put.
if (preserveViewportY === null && this._wasAtBottomBeforeWrite && !this._hasRecentUserScrollUp()) {
this.terminal.scrollToBottom();
}
+198
View File
@@ -0,0 +1,198 @@
/**
* @fileoverview The pending restore plan: what a host reboot destroyed, waiting on a click.
*
* The boot pass builds this plan inside `restoreMuxSessions()`, in the window
* where reconciliation has reported the dead sessions and `cleanupStaleSessions()`
* has not pruned their records yet. The board then offers "restore N sessions
* from before the reboot", and `web/routes/reboot-restore-routes` spends the plan
* when the user clicks.
*
* Invariants:
* - Entries are in-memory only. A server restart drops the plan, and nothing
* re-builds it, because the records it was built from are pruned by then.
* That costs the convenience this feature adds and never the conversation:
* the conversation IS the transcript under `~/.claude/projects`, which
* `services/unified-session-service.ts` reads for the Welcome screen's Resume
* list and the Session Manager, and `resumeHistorySession()` in
* `web/public/terminal-ui.js` resumes from a row there with no persisted
* session record involved. A dropped plan therefore returns the user to
* resuming by hand, one at a time, which is where they are without this
* feature. What the plan held that a transcript does not is the owner, the
* name, the env overrides, the effort and the lineage.
* - Module-level singleton in the style of `web/approval-inbox.ts`: no `Session`
* import and no IO, which keeps it unit-testable and cycle-free.
* - Spending is take-then-build: `take()` removes entries synchronously, before
* the route's first `await`, so a double-click or two devices cannot both
* reach the same entry and put two panes on one conversation.
* - One restore runs at a time per owner. `beginSpending()` single-flights the
* route, so two concurrent clicks cannot interleave pane creation for the same
* user, while two different users never block each other.
*
* @dependencies reboot-restore (RebootRestoreEntry)
* @consumedby web/server (plan build at boot), web/routes/reboot-restore-routes
*
* @module web/reboot-restore-registry
*/
import type { RebootRestoreEntry } from '../reboot-restore.js';
/**
* A plan older than this is dropped on read. A machine that rebooted yesterday
* has moved on, and an offer nobody took by then is noise rather than a rescue.
*/
const PLAN_TTL_MS = 24 * 60 * 60 * 1000;
export class RebootRestoreRegistry {
/** Keyed by session id, in the order the boot pass found them. */
private entries = new Map<string, RebootRestoreEntry>();
/** When the boot pass built the plan, in ms since the epoch. */
private builtAt = 0;
/**
* Entries handed to a restore that has not finished, by session id, each
* remembering which caller is spending it.
*
* A taken entry is still part of the offer until its restore resolves it, so
* it has to stay reachable by everything that can invalidate an offer. Holding
* the entries themselves — rather than a counter to compare against later —
* means `clear()` filters them by the SAME `canAccess(entry.owner)` predicate
* it already applies to the plan. A counter cannot do that, because the caller
* spending an entry need not be its owner: an admin may restore another user's
* sessions, and then the spender and the owner are different keys.
*/
private parked = new Map<string, { entry: RebootRestoreEntry; spender: string | undefined }>();
/**
* Owners with a restore in flight, between its take and its last pane.
* Keyed by owner so one user's restore does not turn another user's click into
* a conflict; `take()` already guarantees no two callers get the same entry.
* Single-user mode has one key, `undefined`, so it behaves as one global flight.
*/
private spending = new Set<string | undefined>();
/** Replace the plan with what the boot pass found. An empty list clears it. */
set(entries: readonly RebootRestoreEntry[]): void {
this.entries = new Map(entries.map((entry) => [entry.sessionId, entry]));
this.builtAt = entries.length > 0 ? Date.now() : 0;
// A fresh boot plan supersedes anything an in-flight restore still holds.
this.parked.clear();
}
/**
* The entries a viewer may see, newest plan first-come order preserved.
*
* @param canAccess Ownership predicate, so a user sees their own entries and
* an admin sees all. Applied here rather than in the route so the count the
* banner shows and the entries a click spends come from one filter.
*/
list(canAccess: (owner: string | undefined) => boolean): RebootRestoreEntry[] {
this.dropIfExpired();
return [...this.entries.values()].filter((entry) => canAccess(entry.owner));
}
/**
* Remove and return the entries a click is about to spend.
*
* Synchronous and total: an entry leaves the plan here, before any pane is
* created, so a second click finds nothing to spend. Entries a caller may not
* access are left in place, and unknown ids are ignored.
*
* @param sessionIds The ids to spend, or undefined for every visible entry.
*/
take(
canAccess: (owner: string | undefined) => boolean,
sessionIds: readonly string[] | undefined,
spender: string | undefined
): RebootRestoreEntry[] {
this.dropIfExpired();
const wanted = sessionIds ? new Set(sessionIds) : undefined;
const taken: RebootRestoreEntry[] = [];
for (const entry of [...this.entries.values()]) {
if (wanted && !wanted.has(entry.sessionId)) continue;
if (!canAccess(entry.owner)) continue;
this.entries.delete(entry.sessionId);
// Parked rather than forgotten: until this restore resolves the entry, a
// dismiss still has to be able to reach and cancel it.
this.parked.set(entry.sessionId, { entry, spender });
taken.push(entry);
}
return taken;
}
/**
* Put entries back after a rebuild never got as far as creating a pane.
*
* Used for the click-time rejections that may resolve themselves: a workspace
* that comes back, a capacity limit the user makes room under, a CLI that
* starts once its binary is on the PATH. A conversation the user resumed by
* hand is NOT put back, because that one cannot stop being true, and an entry
* the banner keeps re-offering forever is noise only Dismiss can clear.
*/
releaseFlight(spender: string | undefined, keep: readonly RebootRestoreEntry[]): void {
const wanted = new Set(keep.map((entry) => entry.sessionId));
let added = 0;
for (const [sessionId, held] of [...this.parked]) {
if (held.spender !== spender) continue;
this.parked.delete(sessionId);
// Still parked means nothing cancelled it while the restore ran. A dismiss,
// an expiry or a fresh boot plan removes it from `parked`, and then it does
// not come back however the restore ended.
if (wanted.has(sessionId)) {
this.entries.set(sessionId, held.entry);
added += 1;
}
}
if (added > 0 && this.builtAt === 0) this.builtAt = Date.now();
}
/** Drop the entries a viewer can see. Returns how many went. */
clear(canAccess: (owner: string | undefined) => boolean): number {
const removable = [...this.entries.values()].filter((entry) => canAccess(entry.owner));
for (const entry of removable) this.entries.delete(entry.sessionId);
// Entries a restore is holding are dismissed by the same rule, so a dismiss
// that lands mid-restore wins. Judged on the ENTRY's owner, exactly as above,
// rather than on who happens to be restoring it.
let parkedRemoved = 0;
for (const [sessionId, held] of [...this.parked]) {
if (!canAccess(held.entry.owner)) continue;
this.parked.delete(sessionId);
parkedRemoved += 1;
}
if (this.entries.size === 0) this.builtAt = 0;
return removable.length + parkedRemoved;
}
/**
* Claim the right to run a restore for one owner, or report that owner already
* has one running. Callers that get `true` must call `endSpending()` in a
* `finally` with the same owner.
*/
beginSpending(owner?: string): boolean {
if (this.spending.has(owner)) return false;
this.spending.add(owner);
return true;
}
endSpending(owner?: string): void {
this.spending.delete(owner);
}
/** Test hook: forget everything, including the single-flight claim. */
reset(): void {
this.entries.clear();
this.parked.clear();
this.builtAt = 0;
this.spending.clear();
}
private dropIfExpired(): void {
if (this.builtAt > 0 && Date.now() - this.builtAt > PLAN_TTL_MS) {
// A restore that took entries just before the expiry must not hand them
// back afterwards and give an expired plan another full day of life.
this.parked.clear();
this.entries.clear();
this.builtAt = 0;
}
}
}
/** Process-wide singleton, mirroring `approvalInbox`. */
export const rebootRestoreRegistry = new RebootRestoreRegistry();
+1
View File
@@ -11,6 +11,7 @@ export { registerCronRoutes } from './cron-routes.js';
export { registerSystemRoutes } from './system-routes.js';
export { registerHookEventRoutes } from './hook-event-routes.js';
export { registerApprovalRoutes } from './approval-routes.js';
export { registerRebootRestoreRoutes } from './reboot-restore-routes.js';
export { registerReadMyMindRoutes } from './readmymind-routes.js';
export { registerStatusTelemetryRoutes } from './status-telemetry-routes.js';
export { registerCaseRoutes } from './case-routes.js';
+313
View File
@@ -0,0 +1,313 @@
/**
* @fileoverview Reboot-restore routes: offer back the sessions a host reboot destroyed.
*
* The boot pass leaves a plan in `web/reboot-restore-registry` when the machine
* plausibly rebooted. The board reads it, shows a banner, and the user decides:
* - `GET /api/reboot-restore`: what is on offer, ownership-scoped
* - `POST /api/reboot-restore/restore`: rebuild some or all of it
* - `POST /api/reboot-restore/dismiss`: drop the offer
*
* A click, not the heuristic, is what creates panes. The heuristic only decides
* whether the banner appears, so a wrong yes costs a line of text the user
* dismisses rather than N CLI processes nobody asked for.
*
* Rebuilding is take-then-build: entries leave the plan synchronously at the top
* of the route, before the first `await`, and the whole route is single-flighted,
* so a double-click or two devices cannot put two panes on one conversation.
* Three things are re-checked at click time rather than trusted from boot: the
* owner's privilege grant, the workspace still being on disk, and the
* conversation not already being live because the user resumed it by hand.
*
* A rebuilt session comes back attached, idle and disarmed. Respawn controllers
* and Ralph loops are deliberately not re-armed, and its terminal scrollback is
* gone, because the pane is new. The banner says so.
*/
import { FastifyInstance } from 'fastify';
import { existsSync } from 'node:fs';
import { ApiErrorCode, createErrorResponse, getErrorMessage } from '../../types.js';
import { RebootRestoreRequestSchema } from '../schemas.js';
import {
parseBody,
getAuthUser,
canAccessOwned,
ownerFor,
isWorkingDirAllowedForUsername,
sessionCapacityMessage,
} from '../route-helpers.js';
import { rebootRestoreRegistry } from '../reboot-restore-registry.js';
import { rejectAlreadyLive, type RebootRestoreEntry, type RebootRestoreRejection } from '../../reboot-restore.js';
import { clampEnvOverridesForOwner } from '../../session-env-clamp.js';
import { Session } from '../../session.js';
import { resolveClaudeModeForUsername } from '../../user-store.js';
import { getCli } from '../../config/cli-registry/registry.js';
import { applyWorkspaceHooks, seedAgentSessionPreamble } from '../../hooks-config.js';
import { getLifecycleLog } from '../../session-lifecycle-log.js';
import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js';
import { SseEvent } from '../sse-events.js';
import type { SessionAttachmentHistoryItem } from '../../types.js';
import type { SessionPort, EventPort, ConfigPort, InfraPort } from '../ports/index.js';
type RebootRestoreCtx = SessionPort & EventPort & ConfigPort & InfraPort;
/** The banner's view of one restorable session. The record itself never leaves the server. */
function toBannerItem(entry: RebootRestoreEntry) {
return {
id: entry.sessionId,
name: entry.name,
workingDir: entry.workingDir,
mode: entry.mode,
owner: entry.owner,
};
}
export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRestoreCtx): void {
const accessorFor = (req: Parameters<typeof getAuthUser>[0]) => {
const user = getAuthUser(req);
return (owner: string | undefined) => canAccessOwned(user, owner);
};
// ========== What is on offer ==========
app.get('/api/reboot-restore', async (req) => {
const entries = rebootRestoreRegistry.list(accessorFor(req));
return {
sessions: entries.map(toBannerItem),
// Said plainly here so the banner never implies a full restore: the pane is
// new, so the conversation continues and the terminal history does not.
scrollbackRestored: false,
};
});
// ========== Spend it ==========
app.post('/api/reboot-restore/restore', async (req, reply) => {
const body = parseBody(RebootRestoreRequestSchema, req.body, 'Invalid reboot restore request');
const canAccess = accessorFor(req);
const owner = ownerFor(req);
// Take BEFORE the first await: a second click must find nothing to spend.
// The flight is per owner, because `take()` already guarantees two callers
// never receive the same entry, so one user's restore need not block another's.
if (!rebootRestoreRegistry.beginSpending(owner)) {
return reply.code(409).send(createErrorResponse(ApiErrorCode.CONFLICT, 'A reboot restore is already running'));
}
const taken = rebootRestoreRegistry.take(canAccess, body.sessionIds, owner);
// Entries nothing built a pane for, returned to the plan on every exit path
// including a throw. Without this a failure between here and the loop would
// spend the offer and rebuild nothing, and the plan cannot be rebuilt.
const unspent = new Set(taken);
try {
if (taken.length === 0) return { restored: [], skipped: [] };
// The plan was built at boot and the board has moved on since. A conversation
// the user resumed by hand from the Resume list is already on screen, and a
// second pane on it would fight the first for the same transcript. This one
// is never re-offered: unlike a missing workspace, it cannot stop being true.
// Read fresh each time rather than snapshotted once: the loop below awaits a
// real `startInteractive()` per entry, so by the tenth entry a snapshot taken
// here is tens of seconds old, and a conversation the user resumed by hand in
// that window would be invisible to it.
const liveSessionIds = () => new Set(ctx.sessions.keys());
const liveConversationIds = () =>
new Set(
[...ctx.sessions.values()].map((session) => session.claudeSessionId).filter((id): id is string => !!id)
);
const { restore, skipped } = rejectAlreadyLive(taken, liveSessionIds(), liveConversationIds());
for (const entry of taken) {
if (skipped.some((s) => s.sessionId === entry.sessionId)) unspent.delete(entry);
}
const restored: ReturnType<typeof toBannerItem>[] = [];
const failures: RebootRestoreRejection[] = [...skipped];
const workspaceHooksEnabled = await ctx.getWorkspaceHooksEnabled();
for (const entry of restore) {
// The already-live check, re-run against the board as it is NOW. The pass
// above decided the batch; this catches a conversation that went live while
// an earlier entry in this same batch was starting. Spent rather than
// returned to the plan, for the same reason as the batch pass: unlike a
// missing workspace or a withdrawn grant, an open conversation is not a
// condition that stops being true.
const [lateLive] = rejectAlreadyLive([entry], liveSessionIds(), liveConversationIds()).skipped;
if (lateLive) {
failures.push(lateLive);
unspent.delete(entry);
continue;
}
// Capacity is re-checked per iteration, because this loop is itself
// creating the sessions it counts. The offer can be a day old, so the
// board may be fuller now than the plan assumed.
const capMsg = sessionCapacityMessage(ctx.sessions, entry.owner);
if (capMsg) {
failures.push({ sessionId: entry.sessionId, reason: 'capacity-reached' });
continue;
}
// A repo can be deleted between the boot that planned this and the click.
if (!existsSync(entry.workingDir)) {
failures.push({ sessionId: entry.sessionId, reason: 'workspace-missing' });
continue;
}
// Multi-user workspace separation: the create route confines a non-admin's
// workingDir to their own case space, and a grant can be withdrawn between
// the session's creation and this restore, so the confinement is re-run
// rather than inherited from the record. Keyed on the OWNER, not on the
// caller: an admin spending another user's entry must be held to that
// user's confinement, and `isWorkingDirAllowed` would wave an admin
// through. The same reason the two grant re-checks below read
// `saved.owner`.
if (!(await isWorkingDirAllowedForUsername(entry.owner, entry.workingDir))) {
// Left on offer: a withdrawn grant can be restored, unlike an already-open
// conversation, so this is not the permanent kind of refusal.
failures.push({ sessionId: entry.sessionId, reason: 'workspace-forbidden' });
continue;
}
try {
const saved = entry.state;
const claudeModeConfig = await ctx.getClaudeModeConfig();
const session = new Session({
// The old id is reused on purpose: a pinned record, subagent parents,
// window states and the lifecycle log all key off it, and the unpinned
// record is gone, so there is nothing to collide with.
id: saved.id,
workingDir: saved.workingDir,
mode: saved.mode,
name: saved.name,
// Without this the constructor re-infers ownership from the name, so a
// session the user renamed by hand to something shaped like `w<n>-<case>`
// comes back as `placeholder` and auto-naming overwrites their name on
// the next prompt. The route persists below, so the loss would go to
// disk. `restoreMuxSessions()` passes it for the same reason.
nameSource: saved.nameSource,
createdAt: saved.createdAt,
mux: ctx.mux,
useMux: true,
// No `muxSession`: the reboot took the pane with it, so `startInteractive()`
// takes its create branch and makes a fresh one.
claudeMode: await resolveClaudeModeForUsername(claudeModeConfig.claudeMode, saved.owner),
allowedTools: claudeModeConfig.allowedTools,
resumeSessionId: entry.resumeConversationId,
// Re-resolved against the owner's CURRENT grant, never replayed from the
// record: a grant held when the record was written may be gone now.
envOverrides: await clampEnvOverridesForOwner(
saved.owner,
(saved as { __envOverrides?: Record<string, string> }).__envOverrides
),
effort: saved.effort,
attachmentHistory:
(saved as { __attachmentHistory?: SessionAttachmentHistoryItem[] }).__attachmentHistory ??
saved.attachmentHistory,
lastSubmitAt: saved.lastSubmitAt,
claudeSessionChain: saved.claudeSessionChain,
lastActivityAt: saved.lastActivityAt,
owner: saved.owner,
parentSessionId: saved.parentSessionId,
});
await ctx.addSession(session);
// Before the listeners, because setupSessionListeners() reads the
// image-watcher flag this phase restores; before the spawn, because the
// custom-model environment and the nice priority shape the process.
await ctx.reapplyPersistedSessionState(session, saved, 'before-spawn');
await ctx.setupSessionListeners(session);
await session.startInteractive();
// The session's own history, applied only once the pane exists: on a
// failed start these totals would belong to a session that never ran.
// Both halves precede the route's OWN persist, which matters because a
// constructed session carries none of this and `toState()` is written
// wholesale, so persisting first would replace the fuller record with
// the reduced one and drop the pin that keeps it from being pruned. A
// listener-driven persist can still land inside the debounce window
// while the pane starts; the write below repairs the record.
// `rearmAutoResumeSchedule: false`: the saved stamp predates the reboot and
// the pane is new, so honouring it would have every restored session type
// `continue` into itself about a minute after one click. Auto-resume stays
// enabled and re-arms on the next real limit message. This is also what the
// module header promises ("comes back attached, idle and disarmed").
await ctx.reapplyPersistedSessionState(session, saved, 'after-spawn', {
rearmAutoResumeSchedule: false,
});
ctx.persistSessionState(session);
// A session without its workspace hooks goes silently blind: no stop or
// idle events for respawn, no Approvals Inbox item, no red tab on a
// blocking dialog. The boot-time sweep finished hours ago, so the click
// path installs them itself. `hooks: 'always'` is the capability that says
// this CLI installs Codeman's hooks into the workspace.
if (workspaceHooksEnabled && getCli(session.mode)?.capabilities.hooks === 'always') {
await applyWorkspaceHooks(session.workingDir, true).catch((err: unknown) =>
console.warn(`[reboot-restore] hook install failed for ${session.workingDir}: ${getErrorMessage(err)}`)
);
}
// Both create paths seed this; without it a restored claude session's agent
// skill falls back to writing out the whole ~150-line §0 preamble. Remote and
// docker sessions never reach here (the plan rejects them as
// `remote-or-docker`), so the local-only condition is structural.
if (getCli(session.mode)?.capabilities.agentSkillInjection && (await ctx.getAgentSkillEnabled())) {
await seedAgentSessionPreamble(session.id).catch((err: unknown) =>
console.warn(`[agent-skill] preamble seed failed for ${session.id}: ${getErrorMessage(err)}`)
);
}
getLifecycleLog().log({ event: 'recovered', sessionId: session.id, name: session.name });
// Every other open tab and phone needs this; the clicking tab already has
// the response, and the client's handler is an idempotent upsert.
ctx.broadcast(SseEvent.SessionCreated, ctx.getSessionStateWithRespawn(session));
restored.push(toBannerItem(entry));
} catch (err) {
// One entry that will not start must not stop the rest of the pass, and
// must not leave a registered session with no pane behind it: by this
// point the session is in `ctx.sessions`, holds a tab-layout slot and has
// listeners.
//
// Reaching this is rarer than it looks, measured against a real server:
// the CLI resolver finds its binary by absolute path rather than through
// PATH, and tmux falls back to another directory rather than failing when
// it cannot enter the workspace, so neither of the two obvious "freshly
// booted machine" failures throws. What is left is the mux layer itself
// failing, which is why this path is defended rather than expected.
console.error(`[reboot-restore] failed to rebuild ${entry.sessionId}:`, err);
// Not cleanupSession(): that is the user-initiated delete, and it would
// count this session's historical tokens into the lifetime totals, demote
// a pinned record to `stopped` (which this pass reads as an intentional
// kill, making the session permanently unrestorable) and delete the
// workspace's `.claude-images`. This undoes only the construction.
await ctx
.discardPartiallyBuiltSession(entry.sessionId)
.catch((discardErr: unknown) =>
console.error(`[reboot-restore] discarding a failed rebuild failed: ${getErrorMessage(discardErr)}`)
);
failures.push({ sessionId: entry.sessionId, reason: 'rebuild-failed' });
// Left on offer: the user can put the binary back and click again.
continue;
}
unspent.delete(entry);
}
if (restored.length > 0) {
// A reboot leaves recovery with nothing alive to find, so its own block never
// started the stats collector. This clears and re-arms its interval, so it is
// safe to call whether or not the collector is already running.
ctx.mux.startStatsCollection(STATS_COLLECTION_INTERVAL_MS);
}
return { restored, skipped: failures };
} finally {
// Anything that never became a pane goes back on offer, including after a
// throw, so a transient failure costs a retry rather than the whole plan.
// Ends the flight: entries still parked for it come back if they are in
// `unspent`, and a Dismiss that unparked them meanwhile wins.
rebootRestoreRegistry.releaseFlight(owner, [...unspent]);
rebootRestoreRegistry.endSpending(owner);
}
});
// ========== Drop it ==========
app.post('/api/reboot-restore/dismiss', async (req) => {
const dismissed = rebootRestoreRegistry.clear(accessorFor(req));
return { dismissed };
});
}
+1 -66
View File
@@ -89,6 +89,7 @@ import {
} from '../route-helpers.js';
import { buildAgentCaseMarker, writeAgentCaseMarker } from '../../agent-case-marker.js';
import { canUsernameRunPrivilegedCommands, resolveClaudeModeForUsername } from '../../user-store.js';
import { clampEnvOverridesForOwner } from '../../session-env-clamp.js';
import { enabledClis, getCli } from '../../config/cli-registry/registry.js';
import { resolveCliLaunchError } from '../../utils/cli-launcher.js';
import { legacyConfigForMode } from '../../session-cli-registry-bridge.js';
@@ -442,72 +443,6 @@ export async function _clampExternalCliBypassForOwner(
};
}
/**
* Env-var keys a non-granted owner must not be able to set, because each one
* hands back privilege the config clamp above just removed, or redirects a
* credential-resolution endpoint.
*
* The DeepSeek three are reachable because `DSH_*` and `DEEPSEEK_*` are
* allowlisted `envOverrides` prefixes (schemas.ts) — which they have to be, since
* that is also how a user configures the harness's non-privileged knobs.
*
* - `DSH_PERMISSION_MODE` IS the harness's permission switch. Every other CLI's
* bypass is a command-line FLAG, reachable only through the per-CLI config the
* clamp already owns; this one is an env var, so the config clamp alone is
* half a gate.
* - `DSH_HOME` points the launcher at a profile tree, and a profile's plugin code
* executes at BOOT, before any approval row can apply. A user who can write a
* workspace can put a profile in it, so this is the wider of the two.
* - `DEEPSEEK_BASE_URL` aims the provider endpoint, and `_configureCliEnv()`
* forwards the SERVER's own `DEEPSEEK_API_KEY` into every dsh pane before
* `applyEnvOverrides()` runs — so a non-granted owner who could set the base
* URL would have the operator's API key sent as a bearer credential to a host
* of their choosing. (`DEEPSEEK_API_KEY` itself stays overridable: supplying
* your OWN key removes privilege rather than granting it.)
* - `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN` are where omp resolves
* credentials from — the same shape as `DEEPSEEK_BASE_URL` above, reachable
* because `OMP_*` is an allowlisted prefix. Unlike DeepSeek, Codeman does not
* forward any operator-held key into an omp pane today (omp's provider
* credentials live in `~/.omp` config files, not env vars), so there is no
* known concrete exfiltration path yet — clamped defensively anyway, since a
* non-granted owner redirecting where a shared multi-tenant deployment
* resolves auth from is not something to allow silently (found in
* Ark0N/Codeman#353 review; omp's own knobs are otherwise mostly `PI_*`,
* already allowlisted for pi and not addressed here — see resolveOmpHome()).
*/
function ownerClampedEnvKeys(): string[] {
return enabledClis().flatMap((entry) => entry.capabilities.privilegedEnvKeys);
}
/**
* Env-var half of the multi-user bypass clamp.
*
* `clampExternalCliBypassForOwner()` clamps the per-CLI CONFIG, and for every CLI
* but DeepSeek that is the whole story. Here it is not: `applyEnvOverrides()` runs
* AFTER `_configureCliEnv()` in tmux-manager, so an override sent on the SAME
* request lands last and wins, and a non-granted owner could restore
* `danger-full-access` on the very request the config clamp downgraded.
*
* Keys are DROPPED rather than rewritten: dropping falls through to what
* `_configureCliEnv()` exports, which is the clamped config and the server's own
* `DSH_HOME`, i.e. exactly the intended state. No-op in single-user mode and for a
* granted owner, like every other clamp here
* (`canUsernameRunPrivilegedCommands()` returns true when `!isMultiUserMode()`),
* and it returns the caller's own object untouched when there is nothing to strip.
*/
async function clampEnvOverridesForOwner(
owner: string | undefined,
envOverrides: Record<string, string> | undefined
): Promise<Record<string, string> | undefined> {
if (!envOverrides) return envOverrides;
const keys = ownerClampedEnvKeys();
if (!keys.some((key) => key in envOverrides)) return envOverrides;
if (await canUsernameRunPrivilegedCommands(owner)) return envOverrides;
const clamped = { ...envOverrides };
for (const key of keys) delete clamped[key];
return clamped;
}
/** Test hook: the env-var half of the same multi-user safety gate. */
export const _clampEnvOverridesForOwner = clampEnvOverridesForOwner;
+14
View File
@@ -1161,6 +1161,20 @@ const NotificationEventSchema = z
})
.optional();
/**
* Body of `POST /api/reboot-restore/restore`.
*
* `sessionIds` restores a subset, and omitting it restores everything the caller
* can see. The ids are session ids from `GET /api/reboot-restore`, and an id the
* caller does not own is ignored rather than refused, matching how the session
* list scopes rather than 403s.
*/
export const RebootRestoreRequestSchema = z
.object({
sessionIds: z.array(z.string().max(128)).max(200).optional(),
})
.strict();
export const SettingsUpdateSchema = z
.object({
// User-facing product branding. This changes browser/UI copy only; package,
+246 -1
View File
@@ -39,7 +39,9 @@ import { fileURLToPath } from 'node:url';
import { existsSync, mkdirSync, readFileSync, chmodSync, rmSync, statSync } from 'node:fs';
import fs from 'node:fs/promises';
import { execSync } from 'node:child_process';
import { hostname as getHostname } from 'node:os';
import { hostname as getHostname, uptime as osUptime } from 'node:os';
import { looksLikeHostReboot, newestPersistedActivity, planRebootRestore } from '../reboot-restore.js';
import { rebootRestoreRegistry } from './reboot-restore-registry.js';
import { dataPath, getDataDir, CODEMAN_INSTANCE } from '../config/instance.js';
import { normalizeBasePath, stripBasePath, joinBasePath } from '../config/base-path.js';
import { GLYPH, palette } from '../cli-style.js';
@@ -171,6 +173,7 @@ import {
registerScheduledRoutes,
registerHookEventRoutes,
registerApprovalRoutes,
registerRebootRestoreRoutes,
registerReadMyMindRoutes,
registerStatusTelemetryRoutes,
registerSystemRoutes,
@@ -668,6 +671,8 @@ export class WebServer extends EventEmitter {
setupSessionListeners: this.setupSessionListeners.bind(this),
persistSessionState: this.persistSessionState.bind(this),
persistSessionStateNow: this._persistSessionStateNow.bind(this),
reapplyPersistedSessionState: this.reapplyPersistedSessionState.bind(this),
discardPartiallyBuiltSession: this.discardPartiallyBuiltSession.bind(this),
getSessionStateWithRespawn: this.getSessionStateWithRespawn.bind(this),
// EventPort
broadcast: this.broadcast.bind(this),
@@ -1060,6 +1065,7 @@ export class WebServer extends EventEmitter {
registerScheduledRoutes(this.app, ctx);
registerHookEventRoutes(this.app, ctx);
registerApprovalRoutes(this.app, ctx);
registerRebootRestoreRoutes(this.app, ctx);
registerReadMyMindRoutes(this.app, ctx);
registerStatusTelemetryRoutes(this.app, ctx);
registerSystemRoutes(this.app, ctx);
@@ -2857,6 +2863,229 @@ export class WebServer extends EventEmitter {
return false;
}
/**
* Work out what a host reboot destroyed, and leave it on offer for the board.
*
* Runs inside `restoreMuxSessions()`, in the window after `reconcileSessions()`
* has reported the dead sessions and before `finalizeRestoredState()` prunes
* their records, so `state.json` is still the full picture here. That window is
* the only place the plan can be built, which is why the boot pass builds it
* even though nothing is rebuilt until a user clicks.
*
* Nothing is created here. The plan goes to `rebootRestoreRegistry`, the board
* offers it as a banner, and `web/routes/reboot-restore-routes` rebuilds what
* the user asks for. A wrong reboot guess therefore costs a line of text the
* user dismisses, not N CLI processes nobody asked for.
*
* @returns how many sessions are on offer.
*/
private planRebootRestoreOffer(dead: string[], livePaneCount: number): number {
if (dead.length === 0) return 0;
const persisted = this.store.getSessions();
if (
!looksLikeHostReboot({
livePaneCount,
deadSessionCount: dead.length,
uptimeSeconds: osUptime(),
newestPersistedActivityAt: newestPersistedActivity(persisted),
now: Date.now(),
})
) {
return 0;
}
const { restore, skipped } = planRebootRestore(dead, persisted, (workingDir) => existsSync(workingDir));
if (skipped.length > 0) {
console.log(`[Server] Reboot restore is passing over ${skipped.length} dead session(s):`);
for (const rejection of skipped) {
console.log(`[Server] ${rejection.sessionId}: ${rejection.reason}`);
}
}
rebootRestoreRegistry.set(restore);
if (restore.length > 0) {
console.log(`[Server] Host reboot detected; offering ${restore.length} session(s) for restore`);
}
return restore.length;
}
/**
* Re-apply the persisted state that a `Session` constructor does not take.
*
* The reboot-restore route builds a session from a record rather than
* attaching to a surviving pane, so everything the constructor has no
* parameter for starts at its default. Persisting such a session writes
* `toState()` wholesale, which would REPLACE the record with the reduced
* version — and for a pinned session that is worse than losing a setting,
* because `cleanupSessionsByIds()` keeps a record only while it is pinned, so
* dropping the pin hands the record to the next stale sweep.
*
* Split in two phases because the two halves have opposite timing needs:
*
* - `before-spawn` shapes the pane itself, so it has to land before the CLI
* process starts, and before `setupSessionListeners()`, which reads the
* image-watcher flag. The custom-model selection is an environment injection
* and the nice priority is applied to the spawn.
* - `after-spawn` is the session's own accumulated history. It must NOT land
* on a session whose pane failed to start: the totals would then belong to a
* session that never ran, and any later cleanup would add them to the
* lifetime figures a second time.
*
* Respawn and Ralph are deliberately NOT re-armed: a machine that just came up
* is the worst moment to turn an autonomous run loose, and the user re-arms
* what they want. Ralph's loop CONFIGURATION does not survive either, because
* `toState()` reads `ralphEnabled` and the completion phrase off a live
* tracker, and there is no way to hold them without arming the loop.
*/
async reapplyPersistedSessionState(
session: Session,
saved: SessionState,
phase: 'before-spawn' | 'after-spawn',
options?: { rearmAutoResumeSchedule?: boolean }
): Promise<void> {
if (phase === 'before-spawn') {
// The custom-model env has to be rebuilt from the endpoint store: the persist
// deliberately keeps the injected VALUES out of state.json, so only the
// bookkeeping survives a restart and the values are re-derived here.
const savedCustomModel = (saved as { __customModel?: CustomModelBookkeeping }).__customModel;
if (savedCustomModel) {
session.setCustomModel(savedCustomModel, await this._rebuildCustomModelEnv(session, savedCustomModel));
}
if (saved.niceEnabled !== undefined || saved.niceValue !== undefined) {
session.setNice({ enabled: saved.niceEnabled, niceValue: saved.niceValue });
}
// `setupSessionListeners()` READS this flag to decide whether to start the
// watcher, so setting it later would leave the session reporting the feature
// as on with nothing watching.
if (saved.imageWatcherEnabled !== undefined) session.imageWatcherEnabled = saved.imageWatcherEnabled;
return;
}
if (saved.pinned) session.restorePin(true, saved.pinnedAt);
if (saved.autoCompactEnabled !== undefined || saved.autoCompactThreshold !== undefined) {
session.setAutoCompact(saved.autoCompactEnabled ?? false, saved.autoCompactThreshold, saved.autoCompactPrompt);
}
if (saved.autoClearEnabled !== undefined || saved.autoClearThreshold !== undefined) {
session.setAutoClear(saved.autoClearEnabled ?? false, saved.autoClearThreshold);
}
if (saved.autoResumeEnabled) {
// The stamp is re-armed by default, because a Codeman restart leaves the
// limit footer un-reprinted and dropping it there would strand the pause.
// A reboot restore opts out: that stamp predates the reboot, the pane is
// new, and honouring it means every session the user restored types
// `continue` into itself about a minute later, unattended. The setting
// itself stays on either way, so it re-arms on the next limit message.
const rearm = options?.rearmAutoResumeSchedule !== false;
session.restoreAutoResume(true, rearm ? saved.autoResumeAt : undefined);
}
if (saved.inputTokens !== undefined || saved.outputTokens !== undefined || saved.totalCost !== undefined) {
session.restoreTokens(saved.inputTokens ?? 0, saved.outputTokens ?? 0, saved.totalCost ?? 0);
// Seed the daily-usage baseline, or the restored totals are counted again as new usage.
this.lastRecordedTokens.set(session.id, {
input: saved.inputTokens ?? 0,
output: saved.outputTokens ?? 0,
});
}
if (saved.color) session.setColor(saved.color);
if (saved.flickerFilterEnabled !== undefined) session.flickerFilterEnabled = saved.flickerFilterEnabled;
}
/**
* Undo a session that was registered but never got a working pane.
*
* Deliberately NOT `cleanupSession()`, which is the user-initiated delete: that
* path adds the session's token totals to the lifetime figures, demotes a
* pinned record to `stopped` (the durable marker of an intentional kill, which
* would make the session permanently ineligible for a reboot restore), drops
* the persisted Ralph state, and recursively removes `.claude-images` from the
* WORKING DIRECTORY, which belongs to the workspace rather than to this session
* and may hold another live session's pasted images.
*
* Everything else `_doCleanupSession()` does, this has to do as well. It is the
* inverse of `registerSessionWithLayout()` plus `setupSessionListeners()`, and
* every registration those two make has to come back out — above all
* `sessionListenerRefs`, whose presence makes `setupSessionListeners()` return
* early. Leaving that entry behind is worse than the leak this function exists
* to prevent: the retry reuses the same session id, wires no listeners at all,
* and the user gets a tab that never shows output.
*
* The persisted record, the lifetime totals, the stored Ralph state and the
* workspace's own files are left exactly as they were, so the session stays
* restorable on the next attempt.
*/
async discardPartiallyBuiltSession(sessionId: string): Promise<void> {
const session = this.sessions.get(sessionId);
if (!session) return;
this.sessions.delete(sessionId);
// --- the inverse of setupSessionListeners(), in reverse order ---
// Listeners first: while they are attached, one of them can still reach a
// tracker this is about to stop.
const listeners = this.sessionListenerRefs.get(sessionId);
if (listeners) {
detachSessionListeners(session, listeners);
this.sessionListenerRefs.delete(sessionId);
}
// An FSWatcher on the workspace that nothing else closes.
imageWatcher.unwatchSession(sessionId);
// An fs.watch on the workspace (or on @fix_plan.md), likewise.
session.ralphTracker.stopWatchingFixPlan();
const summaryTracker = this.runSummaryTrackers.get(sessionId);
if (summaryTracker) {
// Closes the run's own record before the tracker goes, the way
// `_doCleanupSession()` does. Cosmetic rather than load-bearing, but a
// run left open reads as still going in the away digest.
summaryTracker.recordSessionStopped();
summaryTracker.stop();
this.runSummaryTrackers.delete(sessionId);
}
// Also mirrors `_doCleanupSession()`. The PERSISTED Ralph state is left
// alone on purpose (that is one of the things separating this from
// cleanupSession); this only clears the in-memory tracker the failed
// construction built, which the retry reuses the id of.
session.ralphTracker.fullReset();
// --- what anything else may have attached to this id in the meantime ---
// A rebuild can fail AFTER startInteractive() resolved, and a restored
// workspace still carries Codeman's hooks, so the CLI can post a hook event
// within milliseconds. Each of these outlives the listeners and would
// otherwise meet the retry, which reuses the same session id by design.
this.stopTranscriptWatcher(sessionId);
attachmentRegistry.clearSession(sessionId);
sessionWaits.notifySignal(sessionId, 'exit');
sessionWaits.cancelAll(sessionId);
approvalInbox.resolveForSession(sessionId, 'session_ended');
// --- the inverse of the construction itself ---
this.sse.cleanupSessionBatches(sessionId);
this.persistDeb.cancelKey(sessionId);
fileStreamManager.closeSessionStreams(sessionId);
// `lastRecordedTokens` is deliberately NOT deleted: the `after-spawn` phase
// seeds it as the daily-usage baseline for these restored totals, and the
// retry reuses the id, so dropping it would count them as new usage.
// The per-session custom-model config dir carries the endpoint's API key, and
// `before-spawn` may already have written it. Nothing else would ever remove
// it: the stale sweep only touches state.json. A retry rewrites it.
removeConfigDir(customModelConfigDir(sessionId));
try {
session.removeAllListeners();
await session.stop(true);
} catch (err) {
console.warn(`[Server] stopping a partially built session failed: ${getErrorMessage(err)}`);
// `stop()` kills the mux session in its last block, after destroying its
// trackers, so a throw on the way there leaves the pane running.
await this.mux.killSession(sessionId).catch(() => {});
}
try {
await this.tabLayouts.sessionsRemoved([{ id: sessionId, owner: session.owner }]);
} catch (err) {
console.warn(`[Server] releasing the tab layout slot failed: ${getErrorMessage(err)}`);
}
// Any `session:updated` the half-built session emitted before it failed left a
// tab on every other open board, and the client's handler is an upsert.
this.broadcast(SseEvent.SessionDeleted, { id: sessionId });
}
private async restoreMuxSessions(): Promise<boolean> {
try {
// Reconcile mux sessions to find which ones are still alive (also discovers unknown ones)
@@ -2866,6 +3095,22 @@ export class WebServer extends EventEmitter {
console.log(`[Server] Discovered ${discovered.length} unknown mux session(s)`);
}
// Build the reboot-restore offer HERE: `dead` is only known after
// reconciliation, and the records it reads are pruned by
// `cleanupStaleSessions()` as soon as `finalizeRestoredState()` runs.
//
// Guarded on its own, because this runs inside the try that decides whether
// RECOVERY succeeded. A throw here would otherwise be caught below, report
// restoration as failed, and block the stale cleanup and layout
// reconciliation that follow — turning an optional convenience into a
// failure of the thing it is supposed to help. An offer nobody gets is the
// correct way for this to fail.
try {
this.planRebootRestoreOffer(dead, alive.length);
} catch (err) {
console.error('[Server] Building the reboot-restore offer failed; continuing recovery:', err);
}
if (alive.length > 0 || discovered.length > 0) {
console.log(`[Server] Found ${alive.length + discovered.length} alive mux session(s) from previous run`);
@@ -0,0 +1,131 @@
/**
* `WebServer.discardPartiallyBuiltSession()` against the real server object.
*
* The reboot-restore route calls this when a rebuild registers a session and
* then fails to start its pane. It has to be the exact inverse of
* `registerSessionWithLayout()` plus `setupSessionListeners()`, and it must NOT
* be the user-initiated delete: banking the session's token totals, demoting a
* pinned record or deleting the workspace's files would all be wrong for a
* session that never ran.
*
* These tests drive the real method rather than the route, because the route
* tests run against a mock context whose `discardPartiallyBuiltSession` is a
* one-line stub — an earlier version of this function left four registrations
* behind and every route test still passed.
*
* The retry assertion is the important one. `setupSessionListeners()` returns
* early when `sessionListenerRefs` still holds the session id, so a discard that
* leaves that entry makes the next attempt wire nothing at all, and the user
* gets a tab that never shows output.
*/
import { mkdirSync } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
import { afterAll, afterEach, beforeAll, describe, expect, it, vi } from 'vitest';
import { safeRmHomeTree } from './mocks/test-helpers.js';
import { WebServer } from '../src/web/server.js';
import { Session } from '../src/session.js';
import { TmuxManager } from '../src/tmux-manager.js';
/** Reach the private collections the discard is responsible for emptying. */
interface ServerInternals {
sessions: Map<string, Session>;
sessionListenerRefs: Map<string, unknown>;
runSummaryTrackers: Map<string, unknown>;
registerSessionWithLayout(session: Session): Promise<void>;
setupSessionListeners(session: Session): Promise<void>;
discardPartiallyBuiltSession(sessionId: string): Promise<void>;
}
const WORKSPACE = join(homedir(), '.codeman-test-discard');
const SESSION_ID = 'a1b2c3d4e5f60718';
let server: WebServer;
let internals: ServerInternals;
let mux: TmuxManager;
function buildSession(): Session {
return new Session({
id: SESSION_ID,
workingDir: WORKSPACE,
mode: 'claude',
name: 'rebuilt session',
mux,
useMux: true,
});
}
beforeAll(() => {
mkdirSync(WORKSPACE, { recursive: true });
// Test mode: no port is opened and no CLI is launched. One server for the file,
// stopped at the end: the constructor registers handlers on the module-level
// image, subagent, team and workflow watchers, and only stop() removes them.
server = new WebServer(0, false, true);
internals = server as unknown as ServerInternals;
mux = new TmuxManager();
});
afterEach(async () => {
await internals.discardPartiallyBuiltSession(SESSION_ID).catch(() => {});
});
afterAll(async () => {
await server.stop().catch(() => {});
safeRmHomeTree(WORKSPACE);
});
describe('discarding a session whose pane never started', () => {
it('takes the session back out of the server', async () => {
const session = buildSession();
await internals.registerSessionWithLayout(session);
await internals.setupSessionListeners(session);
expect(internals.sessions.has(SESSION_ID)).toBe(true);
await internals.discardPartiallyBuiltSession(SESSION_ID);
expect(internals.sessions.has(SESSION_ID)).toBe(false);
});
it('releases the listener registration, so a retry can wire itself again', async () => {
const first = buildSession();
await internals.registerSessionWithLayout(first);
await internals.setupSessionListeners(first);
expect(internals.sessionListenerRefs.has(SESSION_ID)).toBe(true);
const firstRefs = internals.sessionListenerRefs.get(SESSION_ID);
await internals.discardPartiallyBuiltSession(SESSION_ID);
expect(internals.sessionListenerRefs.has(SESSION_ID)).toBe(false);
// The retry reuses the id by design. `setupSessionListeners()` returns early
// while the refs are still there, so a session built now would run blind: no
// terminal output, no status updates, no exit broadcast. Asserting a DIFFERENT
// refs object is what distinguishes wiring the retry from finding the corpse
// of the first attempt still in place.
const retry = buildSession();
await internals.registerSessionWithLayout(retry);
await internals.setupSessionListeners(retry);
const retryRefs = internals.sessionListenerRefs.get(SESSION_ID);
expect(retryRefs).toBeDefined();
expect(retryRefs).not.toBe(firstRefs);
});
it('stops the run-summary tracker, whose interval would otherwise keep firing', async () => {
const session = buildSession();
await internals.registerSessionWithLayout(session);
await internals.setupSessionListeners(session);
const tracker = internals.runSummaryTrackers.get(SESSION_ID) as { stop: () => void };
expect(tracker).toBeDefined();
// Dropping the map entry is not enough: the tracker arms a setInterval in its
// constructor, and only stop() clears it, so a discard that merely forgot the
// entry would leave the timer running for the life of the process.
const stopped = vi.spyOn(tracker, 'stop');
await internals.discardPartiallyBuiltSession(SESSION_ID);
expect(stopped).toHaveBeenCalled();
expect(internals.runSummaryTrackers.has(SESSION_ID)).toBe(false);
});
it('does nothing at all for a session it never registered', async () => {
await expect(internals.discardPartiallyBuiltSession('never-existed')).resolves.toBeUndefined();
});
});
+62 -3
View File
@@ -48,13 +48,12 @@ describe('install.sh generated-catalogue block', () => {
);
});
it('declares every array the detection code indexes', () => {
it('declares every array install.sh actually reads', () => {
for (const name of [
'CLI_IDS',
'CLI_LABELS',
'CLI_ENABLED',
'CLI_KIND',
'CLI_NPM',
'CLI_LAUNCHER_ONLY',
'CLI_DOCS',
'CLI_CMD_LINUX',
'CLI_CMD_DARWIN',
@@ -69,6 +68,17 @@ describe('install.sh generated-catalogue block', () => {
}
});
it('declares no array install.sh never reads', () => {
// CLI_KIND and CLI_NPM were generated and read by nothing (the .mjs/docker-hosts.ts
// producers read the JSON's `kind`/`npmPackage` fields directly; only these two bash
// arrays were dead). A generated-but-unread array is a maintenance trap the generator
// itself cannot warn about — it has no reader to check against — so this pins the
// opposite of the test above: naming what must NOT come back rather than what must.
for (const name of ['CLI_KIND', 'CLI_NPM']) {
expect(new RegExp(`^${name}=\\(`, 'm').test(SOURCE), `${name} is declared but nothing reads it`).toBe(false);
}
});
it('keeps no hand-written per-CLI detection behind', () => {
// The nine `*_SEARCH_PATHS` arrays and eighteen `check_<cli>`/`get_<cli>_path` pairs are
// what this change removes. One left behind would be a second source of truth that the
@@ -93,6 +103,16 @@ describe('install.sh generated-catalogue block', () => {
[]
);
});
it('keeps no dead generic-lookup helpers behind', () => {
// _cli_index/check_cli/get_cli_path were the ungenericized precursor to the per-CLI
// helpers above: same shape, one level of indirection, called from nowhere once the
// catalogue-driven menu and hints stopped needing a lookup-by-id. Unlike the per-CLI
// pairs these are exact names, not derived from the catalogue.
for (const fn of ['_cli_index()', 'check_cli()', 'get_cli_path()']) {
expect(CODE.includes(fn), `${fn} should have been removed as dead code`).toBe(false);
}
});
});
describe('install.sh trust boundary', () => {
@@ -245,3 +265,42 @@ describe('install.sh AI CLI install menu', () => {
expect(run.status).toBe(1);
});
});
describe('install.sh detect_all_clis and a disabled entry', () => {
// No stock entry ships disabled today, so this is characterization rather than a regression
// pin on real data: it drives the real function in a real bash with entry 0 fabricated
// disabled, and points its binary at `bash` — guaranteed resolvable via `command -v` — to
// prove the entry is genuinely never PROBED (CLI_FOUND_PATH stays empty) rather than merely
// filtered out downstream by every consumer's own `CLI_ENABLED` check.
function driveDetect(disableEntry0: boolean) {
const driver = `
set -euo pipefail
export CODEMAN_INSTALL_SH_LIB=1
. "$1"
k=0; while [[ $k -lt \${#CLI_ALL_BINS[@]} ]]; do CLI_ALL_BINS[$k]="codeman-test-no-such-bin-$k"; k=$((k + 1)); done
k=0; while [[ $k -lt \${#CLI_ALL_PATHS[@]} ]]; do CLI_ALL_PATHS[$k]="/nonexistent/codeman-test/$k"; k=$((k + 1)); done
# Point entry 0's first declared binary at something that WILL resolve, so a probe that
# runs at all finds it.
CLI_ALL_BINS[\${CLI_BIN_OFF[0]}]="bash"
${disableEntry0 ? 'CLI_ENABLED[0]="0"' : ''}
CLI_DETECT_DONE=""
detect_all_clis
echo "path0=[\${CLI_FOUND_PATH[0]}]"
echo "found=$CLI_FOUND_COUNT"
`;
const result = spawnSync('bash', ['-c', driver, 'bash', INSTALL_SH], { encoding: 'utf-8', timeout: 30_000 });
return { status: result.status, stdout: result.stdout ?? '', stderr: result.stderr ?? '' };
}
it('probes an enabled entry (control case)', () => {
const run = driveDetect(false);
expect(run.stdout, run.stderr).not.toContain('path0=[]');
expect(run.stdout).toContain('found=1');
});
it('never probes a disabled entry', () => {
const run = driveDetect(true);
expect(run.stdout, run.stderr).toContain('path0=[]');
expect(run.stdout).toContain('found=0');
});
});
@@ -0,0 +1,28 @@
/**
* The mock route context must offer everything the real one does.
*
* Route tests pass their context as `ctx as never`, and `tsconfig.json` includes
* only `src/**`, so no type check ever compares the mock against the ports. A
* port that gained a method left this mock missing it twice; both times the
* route under test threw a TypeError inside its own catch, and the suite
* reported a plausible-looking failure for an unrelated reason.
*
* So the comparison is made at runtime, against `WebServer.createRouteContext()`
* rather than against the port types, which is what keeps it from drifting: the
* server's own context object is the thing route modules are really given.
*/
import { describe, expect, it } from 'vitest';
import { WebServer } from '../../src/web/server.js';
import { createMockRouteContext } from './mock-route-context.js';
describe('the mock route context', () => {
it('offers every member the real route context does', () => {
const server = new WebServer(0, false, true);
const real = (server as unknown as { createRouteContext(): Record<string, unknown> }).createRouteContext();
const mock = createMockRouteContext() as unknown as Record<string, unknown>;
const missing = Object.keys(real).filter((key) => !(key in mock));
expect(missing, `mock-route-context.ts is missing: ${missing.join(', ')}`).toEqual([]);
});
});
+5
View File
@@ -61,6 +61,10 @@ export function createMockRouteContext(options?: {
setupSessionListeners: vi.fn(async () => {}),
persistSessionState: vi.fn(),
persistSessionStateNow: vi.fn(),
reapplyPersistedSessionState: vi.fn(async () => {}),
discardPartiallyBuiltSession: vi.fn(async (id: string) => {
sessions.delete(id);
}),
getSessionStateWithRespawn: vi.fn((s: MockSession) => s.toState()),
// -- EventPort --
@@ -149,6 +153,7 @@ export function createMockRouteContext(options?: {
clearRespawnConfig: vi.fn(),
updateRespawnConfig: vi.fn(),
setHistoryLimit: vi.fn(async () => {}),
startStatsCollection: vi.fn(),
},
runSummaryTrackers: new Map(),
activePlanOrchestrators: new Map(),
+395
View File
@@ -0,0 +1,395 @@
/**
* @fileoverview The decision half of reboot restore, and proof that the existing
* recovery construction path can CREATE a resumed pane.
*
* Three things are under test. `src/reboot-restore.ts` decides whether the
* machine rebooted and which dead sessions may be offered back. The plan
* registry in `src/web/reboot-restore-registry.ts` holds that offer between the
* boot that builds it and the click that spends it. The third is the claim the
* whole feature rests on: a `Session` built the way `restoreMuxSessions()`
* already builds one, but given no `muxSession` and a `resumeSessionId`, creates
* a fresh pane that resumes the old conversation. If that holds, the restore
* needs no new session-creation service.
*
* `reconcileSessions()` reports every session ALIVE under vitest, so the
* server's own boot pass cannot be reached from here. The decision logic is
* therefore driven directly, and the construction claim is driven through a real
* `Session` against the in-memory tmux layer vitest substitutes.
*/
import { mkdirSync, rmSync } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
import { afterEach, describe, expect, it } from 'vitest';
import { Session } from '../src/session.js';
import { TmuxManager } from '../src/tmux-manager.js';
import type { SessionState } from '../src/types.js';
import {
looksLikeHostReboot,
newestPersistedActivity,
planRebootRestore,
rejectAlreadyLive,
resolveResumeConversationId,
type RebootRestoreEntry,
} from '../src/reboot-restore.js';
import { RebootRestoreRegistry } from '../src/web/reboot-restore-registry.js';
const HOUR = 60 * 60 * 1000;
const NOW = 1_760_000_000_000;
function persistedSession(overrides: Partial<SessionState> & { id: string }): SessionState {
return {
// A live agent's record carries its process id; `/exit` persists null instead.
pid: 99999,
status: 'idle',
workingDir: '/tmp/spike',
currentTaskId: null,
createdAt: NOW - 4 * HOUR,
lastActivityAt: NOW - 2 * HOUR,
mode: 'claude',
...overrides,
} as SessionState;
}
describe('reboot detection', () => {
const base = {
livePaneCount: 0,
deadSessionCount: 2,
// The host came up 10 minutes ago, well after the sessions were last active.
uptimeSeconds: 600,
newestPersistedActivityAt: NOW - 2 * HOUR,
now: NOW,
};
it('calls it a reboot when the socket is empty and the host booted after the last activity', () => {
expect(looksLikeHostReboot(base)).toBe(true);
});
it('refuses when some panes survived, which is an ordinary server restart', () => {
expect(looksLikeHostReboot({ ...base, livePaneCount: 3 })).toBe(false);
});
it('refuses on a long-uptime host, where someone wiped the tmux socket by hand', () => {
// Up for 30 days: the sessions were active long AFTER this boot, so the panes
// went away for some reason other than the machine restarting.
expect(looksLikeHostReboot({ ...base, uptimeSeconds: 30 * 24 * 60 * 60 })).toBe(false);
});
it('refuses when nothing died', () => {
expect(looksLikeHostReboot({ ...base, deadSessionCount: 0 })).toBe(false);
});
it('reads the newest activity stamp across the persisted records', () => {
const persisted = {
a: persistedSession({ id: 'a', lastActivityAt: NOW - 5 * HOUR }),
b: persistedSession({ id: 'b', lastActivityAt: NOW - 1 * HOUR }),
};
expect(newestPersistedActivity(persisted)).toBe(NOW - 1 * HOUR);
});
});
describe('which dead sessions may be rebuilt', () => {
it('rebuilds a session that was simply running when the power went out', () => {
const persisted = { live: persistedSession({ id: 'live', status: 'busy' }) };
const plan = planRebootRestore(['live'], persisted, () => true);
expect(plan.restore.map((s) => s.sessionId)).toEqual(['live']);
});
it('never revives a session the user killed while pinned (COD-142 demotes it to stopped)', () => {
const persisted = { killed: persistedSession({ id: 'killed', status: 'stopped', pinned: true }) };
const plan = planRebootRestore(['killed'], persisted, () => true);
expect(plan.restore).toEqual([]);
expect(plan.skipped).toEqual([{ sessionId: 'killed', reason: 'intentionally-ended' }]);
});
it('never revives a session whose record an unpinned kill already deleted', () => {
const plan = planRebootRestore(['gone'], {}, () => true);
expect(plan.restore).toEqual([]);
expect(plan.skipped).toEqual([{ sessionId: 'gone', reason: 'no-persisted-record' }]);
});
it('never revives a pane whose PTY-exit breaker had tripped', () => {
const persisted = { crashy: persistedSession({ id: 'crashy', respawnBlocked: true }) };
expect(planRebootRestore(['crashy'], persisted, () => true).skipped[0].reason).toBe('respawn-blocked');
});
it('leaves remote sessions to the COD-108 reconnect watcher', () => {
const persisted = {
r: persistedSession({
id: 'r',
remote: { hostId: 'h', host: 'example.test', username: 'u', sessionName: 'n', owned: true },
} as Partial<SessionState> & { id: string }),
};
expect(planRebootRestore(['r'], persisted, () => true).skipped[0].reason).toBe('remote-or-docker');
});
it('leaves docker sessions alone, since the container may not be up', () => {
const persisted = {
d: persistedSession({ id: 'd', docker: { containerId: 'abc', caseId: 'c' } } as Partial<SessionState> & {
id: string;
}),
};
expect(planRebootRestore(['d'], persisted, () => true).skipped[0].reason).toBe('remote-or-docker');
});
it('skips a CLI whose history the claude transcript reader does not understand', () => {
const persisted = { c: persistedSession({ id: 'c', mode: 'codex' }) };
expect(planRebootRestore(['c'], persisted, () => true).skipped[0].reason).toBe('unsupported-mode');
});
});
describe('a session with no attach process in its record', () => {
it('is refused, because there was nothing running to bring back', () => {
// A session that never started, or whose pane died outright. NOT a session
// the user ended with `/exit`: that keeps its pid, because the pid is the
// tmux attach process and `remain-on-exit` keeps the pane alive.
const persisted = { exited: persistedSession({ id: 'exited', status: 'idle', pid: null }) };
const plan = planRebootRestore(['exited'], persisted, () => true);
expect(plan.restore).toEqual([]);
expect(plan.skipped).toEqual([{ sessionId: 'exited', reason: 'not-running' }]);
});
it('still restores the session beside it that was attached when the power went', () => {
const persisted = {
exited: persistedSession({ id: 'exited', pid: null }),
running: persistedSession({ id: 'running', pid: 4242 }),
};
const plan = planRebootRestore(['exited', 'running'], persisted, () => true);
expect(plan.restore.map((entry) => entry.sessionId)).toEqual(['running']);
expect(plan.skipped.map((s) => s.reason)).toEqual(['not-running']);
});
it('refuses a record with no pid field at all', () => {
const persisted = { odd: persistedSession({ id: 'odd', pid: undefined as unknown as null }) };
expect(planRebootRestore(['odd'], persisted, () => true).skipped[0].reason).toBe('not-running');
});
});
describe('a workspace that is no longer on disk', () => {
it('is kept out of the offer, so a click cannot scaffold a deleted repo', () => {
const persisted = { gone: persistedSession({ id: 'gone', workingDir: '/tmp/deleted-repo' }) };
const plan = planRebootRestore(['gone'], persisted, () => false);
expect(plan.restore).toEqual([]);
expect(plan.skipped).toEqual([{ sessionId: 'gone', reason: 'workspace-missing' }]);
});
it('is judged per session, not for the batch', () => {
const persisted = {
kept: persistedSession({ id: 'kept', workingDir: '/tmp/still-here' }),
gone: persistedSession({ id: 'gone', workingDir: '/tmp/deleted-repo' }),
};
const plan = planRebootRestore(['kept', 'gone'], persisted, (dir) => dir === '/tmp/still-here');
expect(plan.restore.map((entry) => entry.sessionId)).toEqual(['kept']);
expect(plan.skipped.map((s) => s.reason)).toEqual(['workspace-missing']);
});
});
describe('a conversation that came back on its own before the click', () => {
const entry: RebootRestoreEntry = {
sessionId: 'abc',
workingDir: '/tmp/spike',
mode: 'claude',
resumeConversationId: 'conv-1',
state: persistedSession({ id: 'abc' }),
};
it('is skipped when the user resumed it by hand from the Resume list', () => {
// Same conversation, different session id: the Resume list creates a NEW id.
const result = rejectAlreadyLive([entry], new Set(['other']), new Set(['conv-1']));
expect(result.restore).toEqual([]);
expect(result.skipped).toEqual([{ sessionId: 'abc', reason: 'already-live' }]);
});
it('is skipped when a session with that id is already on the board', () => {
const result = rejectAlreadyLive([entry], new Set(['abc']), new Set());
expect(result.skipped).toEqual([{ sessionId: 'abc', reason: 'already-live' }]);
});
it('is rebuilt when neither its id nor its conversation is live', () => {
const result = rejectAlreadyLive([entry], new Set(['other']), new Set(['conv-other']));
expect(result.restore.map((e) => e.sessionId)).toEqual(['abc']);
expect(result.skipped).toEqual([]);
});
});
describe('the plan the banner spends', () => {
const all = () => true;
const entryFor = (sessionId: string, owner?: string): RebootRestoreEntry => ({
sessionId,
owner,
workingDir: '/tmp/spike',
mode: 'claude',
resumeConversationId: `conv-${sessionId}`,
state: persistedSession({ id: sessionId, owner }),
});
it('hands an entry to the first caller and nothing to the second', () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('a'), entryFor('b')]);
expect(registry.take(all, undefined, undefined).map((e) => e.sessionId)).toEqual(['a', 'b']);
// The double-click: two panes on one conversation is what this prevents.
expect(registry.take(all, undefined, undefined)).toEqual([]);
});
it('spends only the ids a caller asked for', () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('a'), entryFor('b')]);
expect(registry.take(all, ['b'], undefined).map((e) => e.sessionId)).toEqual(['b']);
expect(registry.list(all).map((e) => e.sessionId)).toEqual(['a']);
});
it("shows a user their own sessions and leaves another owner's alone", () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('mine', 'alice'), entryFor('theirs', 'bob')]);
const asAlice = (owner: string | undefined) => owner === 'alice';
expect(registry.list(asAlice).map((e) => e.sessionId)).toEqual(['mine']);
expect(registry.take(asAlice, undefined, 'alice').map((e) => e.sessionId)).toEqual(['mine']);
// Bob's entry is still on offer for Bob.
expect(registry.list(() => true).map((e) => e.sessionId)).toEqual(['theirs']);
});
it('puts back an entry that no pane was created for', () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('a')]);
const taken = registry.take(all, undefined, undefined);
registry.releaseFlight(undefined, taken);
expect(registry.list(all).map((e) => e.sessionId)).toEqual(['a']);
});
it('runs one restore at a time', () => {
const registry = new RebootRestoreRegistry();
expect(registry.beginSpending()).toBe(true);
expect(registry.beginSpending()).toBe(false);
registry.endSpending();
expect(registry.beginSpending()).toBe(true);
});
it('drops what a dismiss cleared', () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('a'), entryFor('b')]);
expect(registry.clear(all)).toBe(2);
expect(registry.list(all)).toEqual([]);
});
it('forgets a plan nobody took for a day', () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('a')]);
const dayLater = Date.now() + 25 * HOUR;
const realNow = Date.now;
Date.now = () => dayLater;
try {
expect(registry.list(all)).toEqual([]);
} finally {
Date.now = realNow;
}
});
});
describe('which conversation a rebuilt pane resumes', () => {
it('prefers the chain tail, the conversation the CLI reported last', () => {
const state = persistedSession({
id: 'sess-1',
resumeSessionId: 'launch-id',
claudeSessionChain: ['launch-id', 'after-clear'],
});
expect(resolveResumeConversationId(state)).toBe('after-clear');
});
it('falls back to the id the session originally resumed', () => {
const state = persistedSession({ id: 'sess-1', resumeSessionId: 'resumed-id' });
expect(resolveResumeConversationId(state)).toBe('resumed-id');
});
it('falls back to the session id, which is what Claude was launched with', () => {
expect(resolveResumeConversationId(persistedSession({ id: 'sess-1' }))).toBe('sess-1');
});
});
describe('the recovery construction path can create a resumed pane', () => {
const workingDir = join(homedir(), 'codeman-cases', 'reboot-restore-spike');
const sessions: Session[] = [];
afterEach(() => {
for (const s of sessions.splice(0)) s.stop();
rmSync(workingDir, { recursive: true, force: true });
});
/** Built exactly as the reboot pass builds one: no `muxSession`, plus a resume id. */
function rebuildFromPersistedState(state: SessionState, mux: TmuxManager): Session {
mkdirSync(workingDir, { recursive: true });
const session = new Session({
id: state.id,
workingDir,
mode: state.mode,
name: state.name,
createdAt: state.createdAt,
mux,
useMux: true,
resumeSessionId: resolveResumeConversationId(state),
owner: state.owner,
lastActivityAt: state.lastActivityAt,
claudeSessionChain: state.claudeSessionChain,
});
sessions.push(session);
return session;
}
it('creates a NEW mux session rather than needing one to attach to', async () => {
const mux = new TmuxManager();
const state = persistedSession({ id: 'aaaaaaa1-1111-4111-8111-111111111111', name: 'w1-spike' });
const session = rebuildFromPersistedState(state, mux);
expect(mux.getSessions()).toHaveLength(0);
await session.startInteractive();
const created = mux.getSessions();
expect(created).toHaveLength(1);
expect(created[0].sessionId).toBe('aaaaaaa1-1111-4111-8111-111111111111');
expect(created[0].workingDir).toBe(workingDir);
});
it('comes back pointed at the conversation the pane was holding', async () => {
const mux = new TmuxManager();
const state = persistedSession({
id: 'aaaaaaa2-2222-4222-8222-222222222222',
resumeSessionId: 'launch-id',
claudeSessionChain: ['launch-id', 'after-clear'],
});
const session = rebuildFromPersistedState(state, mux);
await session.startInteractive();
// The chain tail wins: a `/clear` before the reboot moved the CLI off the launch id.
expect(session.claudeSessionId).toBe('after-clear');
});
it('comes back idle, with no prompt sent and no autonomous loop armed', async () => {
const mux = new TmuxManager();
const state = persistedSession({
id: 'aaaaaaa3-3333-4333-8333-333333333333',
ralphEnabled: true,
respawnEnabled: true,
});
const session = rebuildFromPersistedState(state, mux);
await session.startInteractive();
// No prompt was queued: nothing is waiting on a task. The status itself is not
// assertable here, because the test PTY echoes and the activity detector reads
// that echo as work; in production the pane settles once the CLI finishes booting.
expect(session.currentTaskId).toBeNull();
// The pass never touches the tracker, so a persisted Ralph loop stays cold.
expect(session.ralphTracker.enabled).toBe(false);
});
it('keeps the owner it was persisted with, there being no request to read one from', async () => {
const mux = new TmuxManager();
const state = persistedSession({ id: 'aaaaaaa4-4444-4444-8444-444444444444', owner: 'alice' });
const session = rebuildFromPersistedState(state, mux);
await session.startInteractive();
expect(session.owner).toBe('alice');
expect(mux.getSessions()[0].owner).toBe('alice');
});
});
@@ -0,0 +1,377 @@
/**
* Reboot-restore route: what happens when a rebuild gets part-way and then fails.
*
* The other route test file deliberately uses workspaces that do not exist, so it
* never reaches `new Session()`. This one mocks the `Session` module so the route
* runs its whole construction path — `addSession`, `setupSessionListeners`,
* `reapplyPersistedSessionState`, `startInteractive` — and then throws.
*
* The mock is the only way in. Driven against a real server, `startInteractive()`
* does not throw for either obvious cause: the CLI resolver finds its binary by
* absolute path rather than through PATH, and tmux falls back to another
* directory rather than failing when it cannot enter the workspace. A mux-layer
* failure is what is left, and it cannot be provoked from a test. Without the
* mock this path would go unexercised, which is how the original version of this
* route shipped a session leak the tests could not see.
*
* It also covers the session caps, because those too are only reachable once the
* route is actually willing to build something.
*/
import { describe, it, expect, afterEach, vi, beforeEach } from 'vitest';
import Fastify, { type FastifyInstance } from 'fastify';
import fastifyCookie from '@fastify/cookie';
/** Set per test: whether the mocked `startInteractive()` rejects. */
let startShouldThrow = false;
/** Ordering log, so a test can assert what ran before the pane spawned. */
const callOrder: string[] = [];
vi.mock('../../src/session.js', () => ({
Session: class {
id: string;
mode: string;
name?: string;
workingDir: string;
owner?: string;
claudeSessionId: string | null = null;
constructor(config: { id: string; mode?: string; name?: string; workingDir: string; owner?: string }) {
this.id = config.id;
this.mode = config.mode ?? 'claude';
this.name = config.name;
this.workingDir = config.workingDir;
this.owner = config.owner;
}
async startInteractive() {
callOrder.push('startInteractive');
if (startShouldThrow) throw new Error('spawn claude ENOENT');
}
/** The mock route context projects a session through this on broadcast. */
toState() {
return { id: this.id, mode: this.mode, name: this.name, workingDir: this.workingDir, owner: this.owner };
}
},
}));
const { registerRebootRestoreRoutes } = await import('../../src/web/routes/reboot-restore-routes.js');
const { rebootRestoreRegistry } = await import('../../src/web/reboot-restore-registry.js');
const { installRouteErrorHandler } = await import('../../src/web/route-error-handler.js');
const { httpStatusForErrorCode } = await import('../../src/types.js');
const { createMockRouteContext } = await import('../mocks/index.js');
type ApiErrorCode = import('../../src/types.js').ApiErrorCode;
type RebootRestoreEntry = import('../../src/reboot-restore.js').RebootRestoreEntry;
type SessionState = import('../../src/types.js').SessionState;
/** A real directory, so the route's workspace checks pass and it reaches the build. */
const WORKSPACE = process.cwd();
function offerEntry(sessionId: string, owner?: string): RebootRestoreEntry {
return {
sessionId,
name: `session ${sessionId}`,
workingDir: WORKSPACE,
owner,
mode: 'claude',
resumeConversationId: `conv-${sessionId}`,
state: {
id: sessionId,
pid: null,
status: 'idle',
workingDir: WORKSPACE,
currentTaskId: null,
createdAt: 1_760_000_000_000,
mode: 'claude',
owner,
} as SessionState,
};
}
async function createHarness(ctx: ReturnType<typeof createMockRouteContext>): Promise<FastifyInstance> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
registerRebootRestoreRoutes(app, ctx as never);
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
if (!req.url.startsWith('/api')) return done(null, payload);
if (payload === null || typeof payload !== 'object') return done(null, payload);
const p = payload as { success?: unknown; errorCode?: unknown };
if (p.success === false) {
if (reply.statusCode === 200 && typeof p.errorCode === 'string') {
reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
}
return done(null, payload);
}
if (p.success === true) return done(null, payload);
return done(null, { success: true, data: payload });
});
installRouteErrorHandler(app);
await app.ready();
return app;
}
beforeEach(() => {
startShouldThrow = false;
callOrder.length = 0;
});
afterEach(() => {
rebootRestoreRegistry.reset();
vi.clearAllMocks();
});
describe('a rebuild that fails after the session is registered', () => {
it('reports why it failed rather than blaming the workspace', async () => {
startShouldThrow = true;
rebootRestoreRegistry.set([offerEntry('a')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
const app = await createHarness(ctx);
const res = (await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} })).json().data;
expect(res.restored).toEqual([]);
// Not `workspace-missing`: the directory is there, the agent would not start.
expect(res.skipped).toEqual([{ sessionId: 'a', reason: 'rebuild-failed' }]);
await app.close();
});
it('does not leave a registered session with no pane behind it', async () => {
startShouldThrow = true;
rebootRestoreRegistry.set([offerEntry('a')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
const app = await createHarness(ctx);
await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
// The session reached ctx.sessions via addSession; the route has to take it
// back out, or the board shows a tab whose pane never existed.
expect(ctx.discardPartiallyBuiltSession).toHaveBeenCalledWith('a');
expect(ctx.sessions.has('a')).toBe(false);
// NOT the user-initiated delete: that would bank this session's historical
// tokens into the lifetime totals, demote a pinned record to `stopped`, and
// delete the workspace's .claude-images.
expect(ctx.cleanupSession).not.toHaveBeenCalled();
await app.close();
});
it('keeps the entry on offer, so the user can fix the PATH and click again', async () => {
startShouldThrow = true;
rebootRestoreRegistry.set([offerEntry('a')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
const app = await createHarness(ctx);
await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions.map((s: { id: string }) => s.id)).toEqual(['a']);
await app.close();
});
});
describe('a rebuild that succeeds', () => {
it('re-applies the persisted state before the record is written again', async () => {
rebootRestoreRegistry.set([offerEntry('a')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
const app = await createHarness(ctx);
const res = (await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} })).json().data;
expect(res.restored.map((s: { id: string }) => s.id)).toEqual(['a']);
// A session built from a record carries none of the pin, token totals or
// custom-model selection, so persisting it first would replace the fuller
// record with the reduced one.
expect(ctx.reapplyPersistedSessionState).toHaveBeenCalled();
const reapplyOrder = (ctx.reapplyPersistedSessionState as ReturnType<typeof vi.fn>).mock.invocationCallOrder[0];
const persistOrder = (ctx.persistSessionState as ReturnType<typeof vi.fn>).mock.invocationCallOrder[0];
expect(reapplyOrder).toBeLessThan(persistOrder);
await app.close();
});
it('shapes the pane before it spawns, and restores the history after', async () => {
rebootRestoreRegistry.set([offerEntry('a')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
(ctx.reapplyPersistedSessionState as ReturnType<typeof vi.fn>).mockImplementation(
async (_s: unknown, _saved: unknown, phase: string) => {
callOrder.push(`reapply:${phase}`);
}
);
(ctx.setupSessionListeners as ReturnType<typeof vi.fn>).mockImplementation(async () => {
callOrder.push('setupSessionListeners');
});
const app = await createHarness(ctx);
await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
// `setupSessionListeners()` READS the image-watcher flag that `before-spawn`
// restores, so the phase has to precede it or the session comes back
// reporting the watcher as on with nothing watching. The custom-model
// environment has to reach the process, and the token totals must not land
// on a session whose pane never started.
expect(callOrder).toEqual([
'reapply:before-spawn',
'setupSessionListeners',
'startInteractive',
'reapply:after-spawn',
]);
await app.close();
});
it('tells every other board about the rebuilt session', async () => {
rebootRestoreRegistry.set([offerEntry('a')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
const app = await createHarness(ctx);
await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
expect(ctx.broadcast).toHaveBeenCalledWith('session:created', expect.anything());
await app.close();
});
it('spends the entry, so it is no longer on offer', async () => {
rebootRestoreRegistry.set([offerEntry('a')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
const app = await createHarness(ctx);
await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions).toEqual([]);
await app.close();
});
});
describe('the session caps', () => {
it('counts the sessions it is itself creating, not just the ones it started with', async () => {
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
// One seat short of the documented maximum of 50, counting the session the
// mock context seeds. A check that ran once before the loop would restore
// BOTH entries; only a per-iteration check refuses the second.
for (let i = 0; i < 48; i += 1) {
ctx.sessions.set(`filler-${i}`, { id: `filler-${i}`, owner: undefined } as never);
}
const app = await createHarness(ctx);
const res = (await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} })).json().data;
expect(res.restored.map((s: { id: string }) => s.id)).toEqual(['a']);
expect(res.skipped).toEqual([{ sessionId: 'b', reason: 'capacity-reached' }]);
// Refused rather than lost: closing a session and clicking again works.
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions.map((s: { id: string }) => s.id)).toEqual(['b']);
await app.close();
});
it('refuses every entry when the board is already at the cap', async () => {
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
for (let i = 0; i < 50; i += 1) {
ctx.sessions.set(`filler-${i}`, { id: `filler-${i}`, owner: undefined } as never);
}
const app = await createHarness(ctx);
const res = (await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} })).json().data;
expect(res.restored).toEqual([]);
expect(res.skipped.map((s: { reason: string }) => s.reason)).toEqual(['capacity-reached', 'capacity-reached']);
await app.close();
});
});
describe('a failure before any entry is considered', () => {
it('returns the whole plan rather than spending it', async () => {
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
(ctx.getWorkspaceHooksEnabled as ReturnType<typeof vi.fn>).mockRejectedValue(new Error('settings unreadable'));
const app = await createHarness(ctx);
const res = await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
expect(res.statusCode).toBeGreaterThanOrEqual(500);
// The plan cannot be rebuilt once boot has pruned the records, so a throw
// anywhere in the route has to hand the entries back.
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions.map((s: { id: string }) => s.id).sort()).toEqual(['a', 'b']);
await app.close();
});
it('releases the single flight, so the next click is not refused', async () => {
rebootRestoreRegistry.set([offerEntry('a')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
(ctx.getWorkspaceHooksEnabled as ReturnType<typeof vi.fn>).mockRejectedValue(new Error('settings unreadable'));
const app = await createHarness(ctx);
await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
expect(rebootRestoreRegistry.beginSpending(undefined)).toBe(true);
rebootRestoreRegistry.endSpending(undefined);
await app.close();
});
});
describe('a dismiss that lands while a restore is running', () => {
it('wins when an admin is restoring the entries and their owner dismisses', async () => {
const theirs = offerEntry('theirs', 'bob');
rebootRestoreRegistry.set([theirs]);
// An admin may spend another user's entries, so the caller doing the restore
// and the owner of what is being restored are different people.
const taken = rebootRestoreRegistry.take(() => true, undefined, 'admin');
expect(taken.map((e) => e.sessionId)).toEqual(['theirs']);
// Bob dismisses his own banner. Nothing of his is in the plan any more, and
// the restore is running under a different name than his.
rebootRestoreRegistry.clear((owner) => owner === 'bob');
rebootRestoreRegistry.releaseFlight('admin', taken);
expect(rebootRestoreRegistry.list(() => true)).toEqual([]);
});
it('wins when an admin dismisses everything mid-restore', async () => {
rebootRestoreRegistry.set([offerEntry('theirs', 'bob')]);
const taken = rebootRestoreRegistry.take(() => true, undefined, 'admin');
rebootRestoreRegistry.clear(() => true);
rebootRestoreRegistry.releaseFlight('admin', taken);
expect(rebootRestoreRegistry.list(() => true)).toEqual([]);
});
it('wins, rather than being undone when the route hands its entries back', async () => {
rebootRestoreRegistry.set([offerEntry('a')]);
const ctx = createMockRouteContext({ workspaceHooksEnabled: false });
// The user clicks Dismiss while the restore is between its take and its
// return. Driven through the ROUTE, so removing the generation argument from
// the route would make this fail.
(ctx.getWorkspaceHooksEnabled as ReturnType<typeof vi.fn>).mockImplementation(async () => {
rebootRestoreRegistry.clear(() => true);
throw new Error('settings unreadable');
});
const app = await createHarness(ctx);
await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions).toEqual([]);
await app.close();
});
it('reaches an in-flight restore the dismisser can see, even once its entries are taken', async () => {
const mine = offerEntry('mine', 'alice');
rebootRestoreRegistry.set([mine]);
const taken = rebootRestoreRegistry.take((owner) => owner === 'alice', undefined, 'alice');
expect(taken).toHaveLength(1);
// The plan is empty now, so the dismiss has nothing of Alice's left in the
// plan; it has to reach the entry the restore is holding.
rebootRestoreRegistry.clear((owner) => owner === 'alice');
rebootRestoreRegistry.releaseFlight('alice', taken);
expect(rebootRestoreRegistry.list(() => true)).toEqual([]);
});
it('does not reach another owner, whose unspent entries still come back', async () => {
const mine = offerEntry('mine', 'alice');
const theirs = offerEntry('theirs', 'bob');
rebootRestoreRegistry.set([mine, theirs]);
// Bob is mid-restore, holding his own entry.
const bobsTaken = rebootRestoreRegistry.take((owner) => owner === 'bob', undefined, 'bob');
expect(bobsTaken.map((e) => e.sessionId)).toEqual(['theirs']);
// Alice dismisses her own banner meanwhile.
rebootRestoreRegistry.clear((owner) => owner === 'alice');
// Bob's restore finishes and hands his entry back. Alice's dismiss covered
// her entries, not his, so his offer survives.
rebootRestoreRegistry.releaseFlight('bob', bobsTaken);
expect(rebootRestoreRegistry.list(() => true).map((e) => e.sessionId)).toEqual(['theirs']);
});
});
+248
View File
@@ -0,0 +1,248 @@
/**
* Reboot-restore route tests (src/web/routes/reboot-restore-routes.ts) via
* app.inject(), no live port.
*
* Every entry these tests put on offer names a workspace that does not exist, so
* the route's click-time workspace check rejects it before any `Session` is
* constructed. That keeps the tests on the route's own guards — taking, scoping,
* single-flighting and re-checking — and leaves pane creation to
* test/reboot-restore.test.ts, which drives a real `Session` for it.
*
* The routes read the process-wide `rebootRestoreRegistry` singleton, so every
* test resets it; a leaked entry would bleed into the next one.
*/
import { describe, it, expect, afterEach, beforeEach } from 'vitest';
import { mkdtempSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import Fastify, { type FastifyInstance } from 'fastify';
import fastifyCookie from '@fastify/cookie';
import { registerRebootRestoreRoutes } from '../../src/web/routes/reboot-restore-routes.js';
import { rebootRestoreRegistry } from '../../src/web/reboot-restore-registry.js';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
import { httpStatusForErrorCode, type ApiErrorCode } from '../../src/types.js';
import { createMockRouteContext } from '../mocks/index.js';
import type { RebootRestoreEntry } from '../../src/reboot-restore.js';
import type { SessionState } from '../../src/types.js';
async function createHarness(authUser?: { username: string; role: 'admin' | 'user' }): Promise<FastifyInstance> {
return createHarnessWithCtx(createMockRouteContext(), authUser);
}
async function createHarnessWithCtx(
ctx: ReturnType<typeof createMockRouteContext>,
authUser?: { username: string; role: 'admin' | 'user' }
): Promise<FastifyInstance> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
if (authUser) {
app.addHook('onRequest', async (req) => {
(req as unknown as { authUser: typeof authUser }).authUser = authUser;
});
}
registerRebootRestoreRoutes(app, ctx as never);
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
if (!req.url.startsWith('/api')) return done(null, payload);
if (payload === null || typeof payload !== 'object') return done(null, payload);
const p = payload as { success?: unknown; errorCode?: unknown };
if (p.success === false) {
if (reply.statusCode === 200 && typeof p.errorCode === 'string') {
reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
}
return done(null, payload);
}
if (p.success === true) return done(null, payload);
return done(null, { success: true, data: payload });
});
installRouteErrorHandler(app);
await app.ready();
return app;
}
/** An entry whose workspace is deliberately absent, so no pane is ever created. */
function offerEntry(sessionId: string, owner?: string): RebootRestoreEntry {
return {
sessionId,
name: `session ${sessionId}`,
workingDir: `/tmp/codeman-reboot-restore-missing/${sessionId}`,
owner,
mode: 'claude',
resumeConversationId: `conv-${sessionId}`,
state: {
id: sessionId,
pid: null,
status: 'idle',
workingDir: `/tmp/codeman-reboot-restore-missing/${sessionId}`,
currentTaskId: null,
createdAt: 1_760_000_000_000,
mode: 'claude',
owner,
} as SessionState,
};
}
afterEach(() => {
rebootRestoreRegistry.reset();
});
describe('GET /api/reboot-restore', () => {
it('reports nothing when no reboot left anything behind', async () => {
const app = await createHarness();
const res = await app.inject({ method: 'GET', url: '/api/reboot-restore' });
expect(res.statusCode).toBe(200);
expect(res.json().data.sessions).toEqual([]);
await app.close();
});
it('names what is on offer, and says the scrollback is not coming back', async () => {
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);
const app = await createHarness();
const body = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(body.sessions.map((s: { id: string }) => s.id)).toEqual(['a', 'b']);
expect(body.scrollbackRestored).toBe(false);
await app.close();
});
it('never carries the persisted record itself to the browser', async () => {
rebootRestoreRegistry.set([offerEntry('a', 'alice')]);
const app = await createHarness({ username: 'alice', role: 'admin' });
const body = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(Object.keys(body.sessions[0]).sort()).toEqual(['id', 'mode', 'name', 'owner', 'workingDir']);
expect(body.sessions[0].state).toBeUndefined();
await app.close();
});
});
describe('POST /api/reboot-restore/restore', () => {
it('reports a workspace that is gone, and keeps offering it in case it comes back', async () => {
rebootRestoreRegistry.set([offerEntry('a')]);
const app = await createHarness();
const first = (await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} })).json().data;
expect(first.restored).toEqual([]);
expect(first.skipped).toEqual([{ sessionId: 'a', reason: 'workspace-missing' }]);
// Nothing was built, so the entry goes back: a repo can be restored from a
// backup between two clicks, and losing the offer would be unrecoverable.
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions.map((s: { id: string }) => s.id)).toEqual(['a']);
await app.close();
});
it('never re-offers a conversation that is already open', async () => {
const entry = offerEntry('a');
rebootRestoreRegistry.set([entry]);
const app = await createHarness();
const ctx = createMockRouteContext({ sessionId: entry.sessionId });
// A session with that id is live, which is what the Resume list would produce.
const liveApp = await createHarnessWithCtx(ctx);
const res = (await liveApp.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} })).json().data;
expect(res.skipped).toEqual([{ sessionId: 'a', reason: 'already-live' }]);
// Unlike a missing workspace, this one is dropped: it cannot stop being true.
const left = (await liveApp.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions).toEqual([]);
await liveApp.close();
await app.close();
});
it('spends only the sessions the click named', async () => {
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);
const app = await createHarness();
const res = await app.inject({
method: 'POST',
url: '/api/reboot-restore/restore',
payload: { sessionIds: ['b'] },
});
expect(res.json().data.skipped).toEqual([{ sessionId: 'b', reason: 'workspace-missing' }]);
// 'a' was never taken, and 'b' came back because no pane was built for it.
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions.map((s: { id: string }) => s.id).sort()).toEqual(['a', 'b']);
await app.close();
});
it('refuses a body it does not recognise rather than guessing', async () => {
const app = await createHarness();
const res = await app.inject({
method: 'POST',
url: '/api/reboot-restore/restore',
payload: { sessionIds: 'not-an-array' },
});
expect(res.statusCode).toBeGreaterThanOrEqual(400);
await app.close();
});
it('turns a second concurrent restore away rather than interleaving it', async () => {
rebootRestoreRegistry.set([offerEntry('a')]);
// Claimed by a restore already in flight for this same owner (undefined in
// single-user mode, which is what the harness runs as).
expect(rebootRestoreRegistry.beginSpending(undefined)).toBe(true);
const app = await createHarness();
const res = await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
expect(res.statusCode).toBe(409);
rebootRestoreRegistry.endSpending(undefined);
await app.close();
});
});
describe('POST /api/reboot-restore/restore: multi-user workspace confinement', () => {
const saved: Record<string, string | undefined> = {};
let realDir: string;
beforeEach(() => {
saved.CODEMAN_MULTIUSER = process.env.CODEMAN_MULTIUSER;
process.env.CODEMAN_MULTIUSER = '1';
// This branch sits AFTER the existsSync check, so the workspace has to be
// real for the confinement rule to be the thing that rejects the entry.
realDir = mkdtempSync(join(tmpdir(), 'codeman-reboot-restore-real-'));
});
afterEach(() => {
if (saved.CODEMAN_MULTIUSER === undefined) delete process.env.CODEMAN_MULTIUSER;
else process.env.CODEMAN_MULTIUSER = saved.CODEMAN_MULTIUSER;
rmSync(realDir, { recursive: true, force: true });
});
it("refuses a workspace outside the OWNER's case space, and leaves it on offer", async () => {
const entry = offerEntry('a', 'alice');
entry.workingDir = realDir;
(entry.state as { workingDir: string }).workingDir = realDir;
rebootRestoreRegistry.set([entry]);
// An admin does the clicking. The confinement is still resolved against
// alice, the entry's OWNER: `isWorkingDirAllowed` waves an admin through, so
// reading the caller here would hand an admin the power to rebuild another
// user's session anywhere on the box.
const app = await createHarness({ username: 'root-user', role: 'admin' });
const res = await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
expect(res.statusCode).toBe(200);
expect(res.json().data.restored).toEqual([]);
expect(res.json().data.skipped).toEqual([{ sessionId: 'a', reason: 'workspace-forbidden' }]);
// A withdrawn grant can be given back, so unlike `already-live` this is not
// the permanent kind of refusal and the entry stays claimable.
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions.map((s: { id: string }) => s.id)).toEqual(['a']);
await app.close();
});
});
describe('POST /api/reboot-restore/dismiss', () => {
it('drops the offer and leaves the banner with nothing to show', async () => {
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);
const app = await createHarness();
const res = await app.inject({ method: 'POST', url: '/api/reboot-restore/dismiss', payload: {} });
expect(res.json().data.dismissed).toBe(2);
const after = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(after.sessions).toEqual([]);
await app.close();
});
});
+132 -8
View File
@@ -49,6 +49,39 @@ function loadTerminalUiHarness(mode: string) {
return { app, writes };
}
/**
* Swap in a terminal whose write() parses ASYNCHRONOUSLY, the way xterm.js does.
*
* The real renderer queues the chunk and applies it later, firing the write
* callback once it has been parsed; a redraw that addresses a row past the
* viewport (Codex's status line) drags the viewport to the live bottom at that
* point, not when write() returns. `parse()` runs that pending work.
*/
function attachAsyncParsingTerminal(app: any, opts: { viewportY: number; baseY: number }) {
const buffer = { viewportY: opts.viewportY, baseY: opts.baseY };
const pending: Array<() => void> = [];
app.terminal.buffer = { active: buffer };
app.terminal.write = vi.fn((_data: string, callback?: () => void) => {
pending.push(() => {
buffer.viewportY = buffer.baseY; // the redraw lands
callback?.();
});
});
app.terminal.scrollToLine = vi.fn((line: number) => {
buffer.viewportY = line;
});
app.terminal.scrollToBottom = vi.fn(() => {
buffer.viewportY = buffer.baseY;
});
return {
buffer,
parse: () => {
const queued = pending.splice(0, pending.length);
for (const run of queued) run();
},
};
}
function loadAppHarness() {
const dir = resolve(import.meta.dirname, '../src/web/public');
const fetchMock = vi.fn();
@@ -337,20 +370,111 @@ describe('terminal flush budget', () => {
it('restores the user scroll position when Codex Working redraws move the viewport', () => {
const { app } = loadTerminalUiHarness('codex');
const buffer = { viewportY: 40, baseY: 100 };
app.terminal.buffer = { active: buffer };
app.terminal.write = vi.fn(() => {
buffer.viewportY = buffer.baseY;
});
app.terminal.scrollToLine = vi.fn((line: number) => {
buffer.viewportY = line;
});
const { buffer, parse } = attachAsyncParsingTerminal(app, { viewportY: 40, baseY: 100 });
app._wasAtBottomBeforeWrite = true;
app._lastUserScrollUpAt = 0;
app.pendingWrites.push('\x1b[55;1H\x1b[2m• Working (6s)');
app.flushPendingWrites();
parse();
expect(buffer.viewportY).toBe(40);
});
// Issue #358. xterm.js parses on its own schedule, so the buffer still holds
// the pre-write viewport the instant write() returns: restoring there compared
// the anchor against itself, did nothing, and left the redraw free to drag the
// viewport to the live bottom a tick later. The previous regression passed
// because its write mock moved the viewport synchronously, which real xterm
// never does. These drive the callback explicitly instead.
it('restores the history anchor only AFTER xterm has parsed the write (#358)', () => {
const { app } = loadTerminalUiHarness('codex');
const { buffer, parse } = attachAsyncParsingTerminal(app, { viewportY: 40, baseY: 100 });
app.pendingWrites.push('\x1b[55;1H\x1b[2m• Working (6s)');
app.flushPendingWrites();
// Nothing has parsed yet, so nothing may have been restored yet either.
expect(app.terminal.scrollToLine).not.toHaveBeenCalled();
expect(buffer.viewportY).toBe(40);
parse();
expect(app.terminal.scrollToLine).toHaveBeenCalledWith(40);
expect(buffer.viewportY).toBe(40);
});
it('holds the anchor across consecutive Codex redraws', () => {
const { app } = loadTerminalUiHarness('codex');
const { buffer, parse } = attachAsyncParsingTerminal(app, { viewportY: 40, baseY: 100 });
for (const frame of ['\x1b[55;1H\x1b[2m• Working (6s)', '\x1b[55;1H\x1b[2m• Working (7s)']) {
app.pendingWrites.push(frame);
app.flushPendingWrites();
parse();
expect(buffer.viewportY).toBe(40);
}
});
it('holds the anchor across a chunked write whose remainder is deferred', () => {
const { app } = loadTerminalUiHarness('codex');
const { buffer, parse } = attachAsyncParsingTerminal(app, { viewportY: 40, baseY: 100 });
// Over the 32KB codex frame budget, so the flush defers a remainder and the
// second chunk goes out from the write callback's reschedule.
app.pendingWrites.push('x'.repeat(40000));
app.flushPendingWrites();
parse();
expect(buffer.viewportY).toBe(40);
app.flushPendingWrites();
parse();
expect(buffer.viewportY).toBe(40);
expect(app.pendingWrites).toHaveLength(0);
});
it('drops the anchor when the user switched sessions before the write parsed', () => {
// The anchor indexes the buffer it came from. selectSession() resets the
// terminal and chunk-loads a different scrollback, so replaying row 40 into
// that one is a jump to an arbitrary place, not a restore. Only reachable now
// that the restore runs a parse later than the write.
const { app } = loadTerminalUiHarness('codex');
const { buffer, parse } = attachAsyncParsingTerminal(app, { viewportY: 40, baseY: 100 });
app.pendingWrites.push('\x1b[55;1H\x1b[2m• Working (6s)');
app.flushPendingWrites();
app.activeSessionId = 'session-2'; // the user clicked another tab
parse();
expect(app.terminal.scrollToLine).not.toHaveBeenCalled();
expect(buffer.viewportY).toBe(buffer.baseY);
});
it('drops the anchor while a buffer load is replaying history', () => {
const { app } = loadTerminalUiHarness('codex');
const { parse } = attachAsyncParsingTerminal(app, { viewportY: 40, baseY: 100 });
app.pendingWrites.push('\x1b[55;1H\x1b[2m• Working (6s)');
app.flushPendingWrites();
app._isLoadingBuffer = true; // chunkedTerminalWrite owns the viewport now
parse();
expect(app.terminal.scrollToLine).not.toHaveBeenCalled();
});
it('does not bounce off the bottom when the sticky flag and an anchor disagree', () => {
// _wasAtBottomBeforeWrite is captured at the frame's first batchTerminalWrite
// and the anchor at flush time, so a scroll-up in between leaves both live.
// The anchor wins: scrolling to the bottom and back would be a visible jump.
const { app } = loadTerminalUiHarness('codex');
const { buffer, parse } = attachAsyncParsingTerminal(app, { viewportY: 40, baseY: 100 });
app._wasAtBottomBeforeWrite = true;
app._lastUserScrollUpAt = 0;
app.pendingWrites.push('\x1b[55;1H\x1b[2m• Working (6s)');
app.flushPendingWrites();
parse();
expect(app.terminal.scrollToBottom).not.toHaveBeenCalled();
expect(buffer.viewportY).toBe(40);
});
});
@@ -9,7 +9,10 @@
* (stopPropagation, not stopImmediatePropagation) does not silence it;
* - a `composed: true` insertText preceded by a keydown — the shape Chrome on
* Android delivers — is dropped by xterm and recovered by us, exactly once;
* - a keystroke xterm DOES handle is delivered exactly once, not twice.
* - a keystroke xterm DOES handle is delivered exactly once, not twice;
* - a character committed in the SAME page task as Enter reaches the send
* path ahead of the `\r`, which is the ordering the zero-delay timer
* alone cannot produce.
*
* Browser-driven, so it is excluded from `npm run test:ci` like the other
* Playwright suites. Run locally:
@@ -146,6 +149,84 @@ describe('orphaned terminal input recovery wiring', () => {
expect(second.sent.join('')).toBe('z');
});
/**
* The batched shape an Android soft keyboard actually delivers when the user
* taps the last character and then Enter: the character's keydown, its
* `composed: true` insertText, and Enter's keydown all land in ONE page task,
* before any zero-delay timer can run.
*
* This is the ordering half of the fix, and the half the unit harness cannot
* reach: the unit tests prove WHICH candidate is forwarded, this proves WHEN.
* Resolving the pending candidate only on its 0 ms timer loses the character
* outright here, because by the time that timer runs xterm has already
* emitted the `\r` and bumped the canonical counter past the candidate's
* snapshot, so it stands down. Draining at the next keydown, from xterm's
* custom key handler (which runs before xterm processes that key), puts the
* character on the wire ahead of the `\r`.
*/
async function batchedCommitThenEnter(data: string) {
return page.evaluate(async (text) => {
const app = (window as any).app;
const textarea = document.querySelector('.xterm-helper-textarea') as HTMLTextAreaElement;
const originalSessionId = app.activeSessionId;
const originalLocalEcho = app._localEchoEnabled;
const originalSendInput = app._sendInputAsync;
const originalPendingInput = app._pendingInput;
const originalLastKeystrokeTime = app._lastKeystrokeTime;
const sent: string[] = [];
try {
app.activeSessionId = 'cod388-browser-batched';
app._localEchoEnabled = false;
app._pendingInput = '';
app._lastKeystrokeTime = 0;
app._sendInputAsync = (_sessionId: string, chunk: string) => sent.push(chunk);
textarea.focus();
// One task, no awaits between the three dispatches.
const charDown = new KeyboardEvent('keydown', {
key: 'Unidentified',
bubbles: true,
cancelable: true,
composed: true,
});
Object.defineProperties(charDown, { keyCode: { value: 65 }, which: { value: 65 } });
textarea.dispatchEvent(charDown);
textarea.value = text;
textarea.dispatchEvent(
new InputEvent('input', { data: text, inputType: 'insertText', bubbles: true, composed: true })
);
const enterDown = new KeyboardEvent('keydown', {
key: 'Enter',
code: 'Enter',
bubbles: true,
cancelable: true,
composed: true,
});
Object.defineProperties(enterDown, { keyCode: { value: 13 }, which: { value: 13 } });
textarea.dispatchEvent(enterDown);
await new Promise((resolve) => setTimeout(resolve, 80));
return { wire: sent.join('') };
} finally {
app.activeSessionId = originalSessionId;
app._localEchoEnabled = originalLocalEcho;
app._sendInputAsync = originalSendInput;
app._pendingInput = originalPendingInput;
app._lastKeystrokeTime = originalLastKeystrokeTime;
textarea.value = '';
}
}, data);
}
it('delivers a character committed in the same task as Enter BEFORE the carriage return', async () => {
const { wire } = await batchedCommitThenEnter('o');
// Not '\r' (character lost, the defect) and not '\ro' (recovered too late).
expect(wire).toBe('o\r');
});
it('sends nothing for a keydown that produces no input event', async () => {
const { sent } = await keystroke({ data: 'q', dispatchInput: false, keyCode: 65 });
expect(sent).toEqual([]);
+46
View File
@@ -218,6 +218,52 @@ describe('orphaned terminal input recovery', () => {
expect(reads).toEqual([]);
});
it('delivers the last character BEFORE the Enter that submits it (defect 4)', () => {
// Android soft keyboards commit the last character and send the Enter key in
// ONE InputConnection transaction, so the `input` event and the Enter keydown
// are processed before any zero-delay timer runs. Two things then went wrong
// with a candidate that only resolved on its timer:
//
// 1. ORDER — xterm emits '\r' synchronously from the Enter keydown, and the
// local-echo composer submits `pendingText` right there. The recovered
// character arrived one macrotask too late to be part of the prompt.
// 2. LOSS — that '\r' bumps the canonical counter, so by the time the
// candidate resolved, `canonicalCount > snapshot` read as "xterm spoke
// for this keystroke" and stood the recovery down. The character was
// dropped outright: every message sent from the phone lost its last
// character.
//
// Resolving pending candidates synchronously at the NEXT keydown fixes both:
// the counter still holds the value it had when that candidate was created,
// and the byte reaches the composer ahead of the Enter.
const h = harness();
h.keydown();
h.input('o');
expect(h.emitted).toEqual([]);
h.keydown({ key: 'Enter' });
expect(h.emitted).toEqual(['o']);
// xterm now emits '\r' for the Enter. The already-resolved candidate must
// not fire a second time when its timer is flushed.
h.controller.notifyCanonicalData();
h.flushTimers();
expect(h.emitted).toEqual(['o']);
expect(h.pendingTimers()).toBe(0);
});
it('still stands down at the next keydown when xterm spoke for the candidate', () => {
// The synchronous resolve must not become a "forward everything" path: a
// keystroke xterm delivered itself is still a duplicate if recovered.
const h = harness();
h.keydown();
h.input('x');
h.controller.notifyCanonicalData();
h.keydown({ key: 'Enter' });
h.flushTimers();
expect(h.emitted).toEqual([]);
});
it('ignores input events that are not committed text', () => {
const h = harness();
for (const inputType of ['insertCompositionText', 'deleteContentBackward', 'insertLineBreak', 'insertFromPaste']) {