Commit Graph
100 Commits
Author SHA1 Message Date
Codeman maintainer 1e5a53830f Merge pull request #408 from shenlvkang-collab/feat/codex-shift-arrow-keys
feat(mobile): add Shift arrows for Codex queued input and prompt navigation
2026-09-14 23:30:14 +02:00
Codeman maintainer c03714eb74 chore: version packages
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 16:21:09 +02:00
Codeman maintainer 1ca0a33830 chore: version packages
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 16:20:36 +02:00
Codeman maintainer 21dcec5d24 test(mobile): follow the 600px phone cut on the Duo branch
Rebased over #390, which moved the phone tier's cutoff from 430px to
600px. The palette's compound fold rule now lives in the 600-768px band
mobile.css pads, the cascade samples the palette inside that band, and
the closed iPhone Duo (466pt) is a phone rather than a small tablet while
the open one (626pt) stays a tablet. Comments in both stylesheets, the
device registry, CLAUDE.md and architecture-invariants say 600.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 16:02:02 +02:00
Codeman maintainer e46089bc7f fix(statusline): print nothing instead of the bare word codeman
Ported from #416 (discussion #405): a statusline reading just `codeman`
is what a hand-run claude in a managed repo showed, and it reads as a
broken config rather than a footer. Three paths produced it and all
three now yield an empty footer: the exporter's `|| echo codeman`
fallback (now `curl -sfk ... || true`, with -f keeping an HTTP error
body off stdout), the unknown-session answer of POST /api/status-telemetry,
and formatSessionStatusText() with nothing to show. The exporter script
marker moves to V4 so live installs pick the new content up on the next
spawn.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:59:00 +02:00
Codeman maintainer fa1ea8d9fe fix(statusline): unset a stale user statusline var, write the exporter script atomically
Three small follow-ups from the #361 review.

A tmux setenv survives respawn-pane, so _configureStatusLineUserCommand
returning early when the user has no statusline left a previously
exported CODEMAN_USER_STATUSLINE_CMD in place: a user who deleted their
own statusline kept getting the stale one wrapped, and lost Codeman's
footer print-through, until the tmux session was recreated. It now
issues `setenv -u` in that case, the same shape as the effort-level
cleanup in applyEnvOverrides.

ensureStatusLineExporterScript truncated and rewrote a script that live
sessions execute on every statusline render, and chmod'd it after the
write. It now writes a temp file next to the target, chmods that, and
rename()s it into place.

The non-tmux direct-PTY fallback carries no exporter; that is now stated
at the spawn site and in the architecture-invariants paragraph rather
than left as a silent gap.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:59:00 +02:00
Codeman maintainer 707ea345eb fix(statusline): GET /api/settings never writes, and a save sends the collection switch only on a flip
Two follow-ups to #361's sticky telemetry switch.

GET /api/settings reconciled an absent showPlanUsageLimits by persisting
true, but readJsonConfig() answers {} for ANY read failure (a parse
error, EACCES, EMFILE, a read landing inside PUT's non-atomic write), not
only ENOENT, and every page load calls this route, so one unlucky read
replaced the whole settings file with a one-key file. The route is a
plain read again and the default moved into the reader:
readPlanUsageTelemetryEnabled() treats an absent key as ON, the same way
readWorkspaceHooksEnabled() does, which is what the desktop chip already
shows for an install that never touched the setting.

saveAppSettings() sent showPlanUsageLimits on every save. The chip
defaults OFF on handhelds, so a phone saving its font size persisted
false and switched collection off for every desktop, whose chip then
went stale with no error anywhere. The key is now stripped like the
other per-device display keys and re-added only when the save FLIPS the
chip relative to what the device had (planUsageCollectionFlip), so an
explicit toggle on any device still writes it in either direction.

Tests pin both: the GET route with a mocked filesystem (absent, missing,
EACCES, garbage, explicit), the reader default, and the flip helper plus
its wiring in saveAppSettings.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:59:00 +02:00
Codeman maintainer 2f9fc72252 docs(mobile): record the fold cascade traps and the keyboard-free baseline
CLAUDE.md's folding-devices rule gains the two new invariants (a shape change
with the keyboard up baselines to window.innerHeight; a base gutter overridden
by a later @media block needs its own zero-base fold restatement, and a
compound rule written against a mobile.css shorthand is scoped to that band)
plus the architecture-invariants pointer it lacked; the new Folding devices
section there carries the mechanisms and the measurements. The device count
is 138 since the two Duo profiles landed (68 Playwright + 70 custom).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:57:08 +02:00
Codeman maintainer b6dbbbcfe0 fix(mobile): fold padding keeps phone sheets flush, scopes the palette rule, caps the response viewer under 430px
Three cascade problems in the fold reserved-region CSS, each measured by
computed style in headless Chromium (styles.css + mobile.css in index.html
link order):

- The unconditional .path-picker-overlay / .path-preview-overlay fold rules at
  the end of the file beat the `padding: 0` both overlays set under 600px, so
  every phone got a 16px and 18px gutter on dialogs built flush (393 and 500px:
  edges floating off the screen). The fold strip is now restated on a ZERO
  base inside the same media query: 0/0 without a fold, the strip alone with
  one, 16/18 plus the strip from 626px up as before.
- .modal.command-palette-modal was unscoped, so outside the 430-768px band
  (where mobile.css pads the palette with a shorthand) it ADDED 0.75rem with
  no gutter to compose with and pushed the shell 6px off centre at 393, 900
  and 1400px, while inside the band the shorthand beat the generic .modal rule
  on the bottom side and the palette lost its block-end gutter. The compound
  rule now lives inside that band and restates both sides.
- The tabletop cap on .response-viewer lost to mobile.css's `max-height:
  92dvh` under 430px (same specificity, later file). mobile.css now carries an
  identical twin at its end.

test/foldable-layout.test.ts simulates the padding cascade across both files
at every breakpoint, with and without the fold rules, and requires the two to
differ by exactly the fold strip; it also pins the palette rule to the band
mobile.css keys on and the response-viewer twin to the styles.css value. Its
model reproduces the Chromium numbers, and against the pre-fix stylesheets it
fails on all three problems.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:57:08 +02:00
Codeman maintainer ef15768e5f fix(mobile): keep the keyboard layout through a fold or rotation with the keyboard up
A shape change with the keyboard up re-baselined initialViewportHeight to the
SHRUNK visual height, so heightDiff was 0 and the settle event the OS fires at
the new width (or any later address-bar drift) satisfied the hide branch and
ran onKeyboardHide() with the keyboard still on screen: accessory bar hidden,
toolbar lift dropped, main's padding cleared. It could not recover, since no
further 150px drop re-arms the show branch against a baseline already sitting
at the shrunk height.

Baseline to window.innerHeight instead when the keyboard is up: the page sets
no interactive-widget, so the keyboard shrinks only the visual viewport and
the layout viewport stays the display's full height on both engines, the same
fact updateLayoutForKeyboard() relies on.

The vm harness now models the two heights separately (resizeTo takes an
optional layout height) and pins the fold flavour (626x590, 466x378, 466x378),
the rotation flavour (393x359, 852x150, 852x160) and the eventual close. All
three fail against the old line.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:57:08 +02:00
Codeman maintainer 8389423459 feat(mobile): iPhone Duo support (fold-aware dialogs, no phantom keyboard)
Apple's "Designing for iPhone Duo" asks an app to adapt to both displays,
to stay continuous as the device opens and closes, and to treat the band a
partly-open display folds through as a reserved region. Three things here.

1. A visual-viewport resize that changes the WIDTH is the device changing
   shape (a rotation, or a foldable opening or closing) and is never the
   virtual keyboard, which only ever takes height. handleViewportResize()
   read any height drop over 150px as the keyboard appearing, so closing a
   Duo (890 to 678pt tall) latched keyboardVisible with no keyboard on
   screen: the accessory bar appeared, main grew 84px of dead padding, and
   updateAppHeight() stopped refreshing --app-height. The latch was sticky,
   because clearing it needs the height back within 100px of a baseline
   belonging to a display the user is no longer looking at. Rotating any
   phone hit the same latch. The shape branch re-baselines instead, which
   is also what lets a keyboard opened after the fold be detected.

2. The hinge is now a reserved region in CSS. --fold-inline-end and
   --fold-block-end measure the strip to keep clear from the Viewport
   Segments env() variables, and are 0px everywhere else, so the seven
   centred overlays are inert by construction off a foldable. Each shrinks
   its content box with padding rather than the box itself, so the backdrop
   still covers the far side of the fold and still swallows taps there.

3. iPhone Duo (outer) and iPhone Duo (inner) join the mobile device
   registry, derived from Apple's published pixel specs at 3x.

Verified in Chromium: flat, a dialog stays centred at 313 of a 626pt
viewport; in book pose it centres at 153 inside the 0-305 leading segment
with its right edge at 293, while the backdrop still spans all 626. The
3-term calc on the offline overlay resolves to 367px in tabletop pose and
20px flat.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 15:57:08 +02:00
Codeman maintainer a5cf1f6005 docs(cli-registry): name the real tests and fields the catalogue docs point at
Three instructions a future contributor would follow literally were stale after
the last review round: the "Adding a CLI" checklist sent the agent-image reason to
AGENT_IMAGE_SPECIAL_CASES, a constant that no longer exists (it is
discovery.install.agentImageLayer on the entry in stock.ts), the trust-boundary
paragraph credited the embedded-commands pin to the invariants test when it is
test/cli-catalog-sync.test.ts, and install.sh claimed "the parity test" pinned the
DeepSeek Harness banner when no test did. That pin now exists: the invariants test
asserts the script's grep literal and the registry's discovery.identity.regex agree
on "DeepSeek Harness", and the comment names it.

docs/docker-cases.md separated the two reasons a CLI stays out of the shared npm
layer (no npmPackage at all versus an agentImageLayer entry), which it had folded
into one, and architecture-invariants no longer lists the agent image's CLI set by
hand.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:55:20 +02:00
Codeman maintainer 3566e8b5ff fix(install): let Skip in the AI CLI menu continue instead of aborting the install
Choosing "s" (Skip) in the new catalogue-driven install menu warned, printed the
install hints and then fell into the shared "The selected AI CLI failed to install"
gate one line below, because CLI_FOUND_COUNT is 0 by construction inside that block
and skipping does not change it. The AI CLI check runs before the clone and the
build, so a user who picked the documented skip option ended up with nothing
installed. The code this replaced guarded the gate with an elif on the skip choice.

The menu moves out of main() into offer_ai_cli_install() and the gate moves inside
the install branch: skipping continues to the clone, a chosen install that leaves
nothing behind is still fatal. Being a function, the interactive path can now be
driven with a stubbed read_reply, which is what nothing reached before: two
behavioural tests in test/install-sh-invariants.test.ts run the real function in a
real bash (skip continues with exit 0, a failed install dies with exit 1), and the
bash 3.2 CI step drives the skip path in the container as well.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:55:20 +02:00
Codeman maintainer c9c8ffddde test(mobile): read PHONE_MAX as an exclusive bound everywhere, drop the stale 430px baselines
Follow-up to #390. PHONE_MAX had become 599, an inclusive bound, while
three of its four consumers still read it as exclusive (width < PHONE_MAX
for phone); the one site that switched to <= disagreed with
getDeviceType(). It is 600 again with < at every site. The breakpoint
table in docs/mobile-testing-report.md says 600, and the three 430px
visual baselines are removed: they depict the tablet tier now, and the
visual suite recreates a missing baseline on its next run on the machine
that owns them. device-matrix.test.ts is also run through Prettier, which
the commit hook demanded and the format gate (src/ only) never did.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:53:22 +02:00
Codeman maintainer c1b4b440f4 chore(plugin): add npm run check:plugin with explicit manifest paths
Runs the mirror/version drift check and both strict validations. The paths
are spelled out because the documented pair ended in a bare `.` that reads as
a full stop when copied out of prose, which surfaced as `missing required
argument 'path'` on first use. Needs the `claude` CLI, so it is a local check
rather than a CI step.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:40:39 +02:00
Codeman maintainer c2d019d956 chore: move the maintainer PR bot out of this repository
`scripts/pr-bot/` was maintainer tooling, not part of the server, the CLI or the
npm package: a Telegram bot that reviews open pull requests in Codeman sessions
and reports to the maintainer. It now lives in its own private repository and
keeps running unchanged, as a client of Codeman's HTTP API like any other.

It moved because it grew a second watcher, for GitHub Discussions, and shipping
that here would mean publishing the briefs it hands its review sessions, the
judgement calls in them and its safety model. None of that helps anyone
installing Codeman, and all of it is easier to change when it is not a public
interface. The move cost nothing structurally: the whole tree depended on one
external package plus Node builtins.

What this removes from the repo, and nothing else: the sources, their three test
files, `config/tsconfig.pr-bot.json`, `docs/pr-bot.md`, the `pr-bot` npm script,
the bot's globs in the typecheck/lint/format scripts, and its knip entry. CLAUDE.md
keeps a short pointer in place of the section, because the bot still constrains
work in here: it takes the `prbot-<n>` and `dscbot-<n>` session names on the local
Codeman, holds clones under `~/.codeman/pr-bot/`, and fetches pull-request heads
into `refs/pr-bot/*` of this checkout, which it must never check out or reset.

The CHANGELOG entries from 1.25.0 and earlier still describe it. That is history
rather than drift, and is left alone.

Verified after the removal: typecheck, lint and format:check clean, and the suite
passes 6843 tests across 357 files, which is the previous run minus exactly the
70 tests that moved out with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 15:24:58 +02:00
Codeman maintainer fc098aaab2 docs(plugin): say to pick one install route, since plugin and user-level skill list twice
Measured with both installed: a fresh Claude Code lists `codeman` (the
user-level or per-case copy) and `codeman:codeman` (the plugin). Neither
shadows the other and both work, so this is noise rather than breakage, but
the README, the wiki and the plugin README now say to choose one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:31:26 +02:00
Codeman maintainer 49ab8bc2f1 fix(plugin): move the Claude Code plugin into plugins/codeman so an install no longer runs npm install
With the repo root as the plugin root, `claude plugin install codeman@codeman`
copied the whole checkout into its cache and, because that root carries a
package.json, ran an npm install there: 832 MB, 511 packages and this repo's
postinstall build on every installer's machine (measured from a clean worktree
of the previous commit). A plugin root must be a directory without one.

The plugin is now `plugins/codeman/`: its manifest, a README, and a MIRROR of
`skills/codeman/`. A mirror rather than a symlink because the install copies
the plugin directory and a link pointing outside it would dangle; a mirror
rather than the source because every install path, injector and doc already
names `skills/codeman/`. `scripts/sync-plugin.mjs` (replacing
sync-plugin-version.mjs) mirrors the skill and syncs both manifest versions
inside `version-packages`; `test/plugin-manifest.test.ts` pins byte-identity,
the versions, the absence of a package.json in the plugin root and that the
repo root `.claude-plugin/` holds only the marketplace manifest.

`claude plugin validate --strict` now passes for both the plugin and the repo
root. Install commands are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:28:57 +02:00
Codeman maintainer f6c08118dc feat(skill): ship the codeman agent skill as a Claude Code plugin from the repo's own marketplace
`.claude-plugin/marketplace.json` at the repo root makes
`/plugin marketplace add Ark0N/Codeman` work, and the one plugin it lists is
the repo itself (`source: "./"`), whose one component is `skills/codeman/`.
So `/plugin install codeman@codeman` is a third install route next to
`npx skills add` and `codeman skill install`, and the skill shows up in the
plugin directories that index Claude Code marketplaces.

Both manifests carry package.json's version: `scripts/sync-plugin-version.mjs`
rewrites them inside `version-packages`, right after `changeset version`, and
`test/plugin-manifest.test.ts` pins the equality, the skill's frontmatter name
(without it the installed skill would be named after a versioned cache dir),
and that no other plugin component (`commands/`, `agents/`, `hooks/`,
`.mcp.json`, `settings.json`) appears at the repo root, since an install would
silently ship it.

Verified with `claude plugin validate` (one expected warning: CLAUDE.md at a
plugin root is not plugin context) and a local marketplace add, install,
details, uninstall cycle against a clean checkout of this commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:24:35 +02:00
Codeman maintainer 7df2dc5955 docs: make a Discussions announcement step 8 of the COM release flow
Every release now gets an Announcements post shaped like #418 and #302:
features first, contributor mentions inline, Thanks at the end. The
step records the GraphQL command and ids, and why it exists: posts
stopped at 1.18 while ten releases shipped unannounced.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 14:16:44 +02:00
Codeman maintainer edeaa15986 feat(terminal): configurable normal and bold font weight (#403)
Bold text on the theme's default foreground carries exactly ONE cue, the
weight step. Claude Code marks its markdown bold with a bare ESC[1m and
changes no colour, and xterm substitutes a bright colour for bold only
when the foreground is a palette index 0-7, so the substitution never
fires for default-foreground text. A family shipping only a regular and
a bold face keeps that step small (measured on Consolas: glyph ink rises
from 14.25% to 16.57%), and picking a different family does not help,
because 400 stays 400 whatever the family. Lowering the NORMAL weight is
the only way to widen the gap.

Two per-device settings beside "Terminal font" in the Font group, each
defaulting to xterm's own value for its slot, so an untouched install
renders exactly as it did before. Both thread into the main terminal and
the Agent Teams panes, and apply on save without a reload.

The bundled face had to be unclamped in the same change or the settings
would look broken on a stock install. fonts/jetbrains-mono-variable.woff2
carries a wght axis of 100 to 800, but styles.css declared the face
`400 700`, and the descriptor is what the browser synthesizes from: at
that range 100, 200 and 300 rendered identically to 400 and 800
identically to 700 (measured in headless Chromium, both directions).
The two families ahead of it in the default stack, Fira Code and Cascadia
Code, exist only if the user installed them, so for most installs
"normal = 300" would have been a no-op. Declared `100 800`, every step is
distinct: 61%, 77% and 90% of the ink at 400, and 800 adds ~14% over 700.
Nothing in the stylesheets asks for a monospace weight outside 400-700,
so widening it changes nothing that rendered before.

Details that are easy to get wrong and are pinned by tests:

- Each slot falls back to its OWN xterm default, so an unset bold weight
  can never inherit `normal` and become a visible change.
- A live save refreshes both echo overlays. They cache
  terminal.options.fontWeight and paint it into their spans, so without
  it the characters being typed keep the old weight while the rest of the
  screen changes. Most visible on a phone, where local echo is on by
  default.
- A live save reaches open Agent Teams panes, which read their options at
  construction, exactly as applyTerminalSkin() propagates its own.
- A stored weight the picker does not list (a hand-set 350) is added to
  the select rather than dropped, so merely opening App Settings cannot
  reset it.
- _awaitTerminalFont() is untouched. CharSizeService measures through the
  CSS `font` shorthand, which resets the weight, so the measured face is
  always the 400 one and a weighted descriptor would request nothing new.

Verified end to end in a headless browser against a live server: the save
reaches the running terminal with no reload, the settings PUT stays 200
(both keys are display keys and are stripped before it, since
SettingsUpdateSchema is strict), the value survives a reload, and the
painted terminal really changes weight with the bundled font (lit-pixel
ink 0.83 / 0.95 / 1.00 / 1.13 / 1.21 at 100 / 300 / default / 700 / 800).

Proposed and analysed by @irisitymichaelgrundberg in discussion #403.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 14:09:13 +02:00
Codeman maintainer d2ff1814ed docs: close the last two Thanks gaps, 1.23.0 and 1.22.0
Auditing every release after the previous backfill turned up two more. 1.23.0
had no Thanks in either artifact; its three PRs (#337, #341, #338) are authored
by the maintainer, so like the others it credits the release it follows.
1.22.0 had the section on its GitHub release but never in CHANGELOG.md, which
is the drift that happens whenever the block is added post-hoc instead of in
the changeset.

Every release from 1.21.0 forward now carries a Thanks section in both places.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:45:42 +02:00
Codeman maintainer 9e2091255b docs: backfill Thanks sections for 1.26.0, 1.24.4 and 1.24.2
Those three shipped with maintainer-only commits and no Thanks section, on the
reasoning that a release with no contributor PRs has nobody to credit. That is
the wrong test: the newest tag is what GitHub marks Latest, so a contributor
who shipped in the release next door lands on a page acknowledging nobody.

Each now credits the release it follows and says so, rather than claiming work
its contributors did not do: 1.24.2 the hotfix on 1.24.1, 1.24.4 the same-day
follow-on to 1.24.3, 1.26.0 the day after 1.25.0. Wording is carried over
verbatim from those releases. The matching GitHub release bodies were edited to
match, since the two are separate artifacts once version-packages has run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:45:04 +02:00
Codeman maintainer 6030a520bd docs: add the Thanks section to the 1.28.1 changelog entry too
1.28.1 is a same-day follow-on to 1.28.0 and is the release people land on as
"Latest", so it credits the same three contributors rather than showing no
acknowledgement at all. Matches the section just added to its GitHub release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:43:40 +02:00
Codeman maintainer 465b842e97 docs: add the missing Thanks section to the 1.28.0 changelog entry
Every release credits its contributors in both places: a "### Thanks" block
and a comment on each merged PR. The PR comments went out, this did not.
Past releases carry it because the block was written INTO the changeset, which
is what feeds both CHANGELOG.md and the GitHub release body; mine went only on
the GitHub release, so the changelog was short a section. Put it in the
changeset next time rather than patching both by hand afterwards.

1.28.1 gets none on purpose: every commit in it is a maintainer commit, the
same call as 1.26.0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:35:33 +02:00
Codeman maintainer c4b74415ee chore: sync CLAUDE.md version to 1.28.1
COM step 4. Staged as a single hunk: the shared checkout also holds another
session's in-progress pr-bot discussions work in this file, which is left
untouched and uncommitted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:17:53 +02:00
Codeman maintainer 708cb2cbf0 fix(tabs): let a wrapped desktop tab strip grow the header instead of clipping itself
The fixed 120px/96px caps on the two wrapped layouts were row counts in disguise: a
third row was clipped into a ~4px scroller, hiding tabs inside a container nothing
invites you to scroll, while the header had the page below it to grow into. Both
layouts now share one rule capped at var(--tab-strip-max-height, 40vh), a safety net
for an absurd session count rather than a row limit.

Verified before shipping: .header is min-height + flex-shrink: 0 so it can grow, and
terminal-ui's ResizeObserver refits the terminal when it does; updateTabOverflowMode()
returns early for any non-desktop viewport, and below 1024px mobile.css pins the header
to max-height: 48px, so this is desktop-only in effect; the selector is comma-grouped
rather than :is(), so each arm keeps (0,2,0) and mobile.css's overrides still win on
source order. PostCSS parses the file cleanly (prettier ignores styles.css).

Authored in a parallel session against this shared checkout.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:08:57 +02:00
Codeman maintainer 48f30f3055 style: drop em-dashes from the text added in c2114615
House style, and these land in the changelog. Only the sentences added in the
previous commit are touched; the em-dashes in contributor text and in the
pre-existing COD-54/COD-115 comments are left alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:46:24 +02:00
Codeman maintainer c211461500 fix: merge-time follow-ups for #409, #404 and #399
#409 (Claude truecolor). The changeset becomes the changelog, and its premise
does not hold on tmux 3.2 or newer. Measured here on tmux 3.4: `default-terminal`
sits at its compiled default of `tmux-256color`, a live claude pane reports
`TERM=tmux-256color`, and supports-color reads that as 256 colors, where
rgb(55,55,55) lands on ESC[48;5;237m — visible, just not the color the theme
named. The invisible block the PR describes needs TERM to resolve to a 16-color
entry: tmux older than 3.2, or a ~/.tmux.conf setting `default-terminal screen`,
which Codeman's own tmux server does read (it passes no -f). Both the changeset
and the invariants paragraph now say that, so the next report here gets paired
with the reporter's tmux -V instead of being read as universal. The change itself
stands on the simpler argument: claude was one of two entries not asking for
truecolor while twelve do.

Also reorders buildClaudeEnv(). It applied the registry's unset/exports AFTER the
whole env was built, so a clis.json entry naming CODEMAN_HOOK_SECRET_FILE or PATH
would strip it on the direct-PTY path while the tmux pane kept it — buildEnvExports()
emits `...cliEnv` ahead of `export CODEMAN_MUX=1` and cannot. The block now runs
first and Codeman's own keys are assigned on top, matching the pane.

#404 (Ctrl+Z trap). Adds the missing changeset, and records what the trap does
not cover: an agent CLI already holds its tty with ISIG off (verified on three
live panes: `susp = ^Z -isig -icanon`), so this is defence for the startup window
rather than a fix for the steady state, and two input paths still reach the PTY
unfiltered — the mobile accessory bar's one-shot Ctrl and the CJK textarea.

#399 (path picker sort). The server sorts by name and cuts at 500, so the client
sorting those 500 by date gives "the newest of the first 500 by name", which is
wrong in exactly the >500-entry folder the date sort exists for. The status line
now says "(first 500 by name)" so the cut is legible, with the reasoning parked
on _sortEntries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:45:41 +02:00
Codeman maintainer a017e9a8e0 chore: version packages 2026-09-12 06:10:14 +02:00
Codeman maintainer 65ddedd1d4 fix: act on the 1.27.0 pre-release review
A Fable 5.1 reviewer read the whole release diff against 1.26.2 and returned
SHIP WITH FIXES. These are its findings, verified before acting on each.

**The changelog advertised a feature the code refuses (major).** The #401
changeset and docs/web-tabs.md both listed `*.localhost` in the loopback set.
The follow-up in 02b0e278 moved it out of the auto-route set on security
grounds and updated CLAUDE.md but neither of those, and that changeset becomes
the 1.27.0 CHANGELOG entry: a user would have read the release notes, tapped
`http://app.localhost:3000/` on a phone and got a connection error from a
documented feature. Both corrected, and the user guide now says why it is
excluded and that adding such a dashboard by hand still works.

**Dictation delivered its text twice (minor, #388).** `keydownSnapshot` started
`null`, so `keydownSnapshot ?? canonicalCount` at the input event read a counter
xterm had ALREADY bumped: on a fresh page load with no keydown yet, xterm's own
capture listener forwards the `insertText` itself (it is not gated behind a
keydown), then the snapshot equals the bumped count, `count > snapshot` is
false, and the controller emits the same text again. Reproduced directly
against the module: it emitted `hello` for input xterm had already delivered.
A `0` baseline restores that file's own invariant, that a missed recovery is
acceptable and a duplicated keystroke is not. Two regression tests, covering
both the xterm-already-delivered and genuinely-dropped halves.

**The sorted rail's arrow-key walk followed the DOM (minor).** `_tabKeydownHandler`
steps `querySelectorAll` order, which is `sessionOrder`, while a sorted rail
paints its rows with the flex `order` property, so ArrowDown from the top card
landed wherever that session happened to sit in the tab order. It now sorts its
node list by the COMPUTED order first: computed rather than inline, because web
tabs take their `order: 9999` from CSS and would otherwise read as 0 and lead
the walk. This is the one place that follows the paint; the Alt+N badge, the
drag model and the filter all still deliberately read the DOM.

**A trusted dashboard was auto-reused by a tapped link (minor, #401).** The
reuse loop skipped `managed` and direct-mode records but not `trusted`. A
trusted frame is mounted with `allow-same-origin`, i.e. on Codeman's origin
with the user's cookie, and these links come from agent output, which is the
threat model the loopback allowlist was just narrowed for. An agent that can
write into the dev server's tree could print a path that one tap opens inside
that privileged frame. Excluded from auto-reuse, with a test; opening it from
the Run dropdown is still an explicit action and unchanged.

**Two documentation claims that were no longer true.** CLAUDE.md said
test/location-overlay-commands.test.ts pins every remote pane command, but
remote claude and remote omp now have their own arm in `buildRemoteLaunchCommand`
and never reach `defaultRemoteCommandForMode`, which is what that test asserts,
so it pins nothing for them and changing either arm will not fail it. Named the
real pins instead. Also documented the arrow-key-walk exception in the rail
paragraph.

Left as follow-ups, deliberately: `POST /api/webviews` does not dedupe by URL
server-side, so two devices tapping one link concurrently can still save two
dashboards for one origin (pre-existing endpoint behaviour that #401 makes
reachable by a tap), and the location-overlay golden should assert the real
remote claude/omp commands rather than a branch neither reaches.

Full gate green: 359 files, 6869 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 06:09:43 +02:00
Codeman maintainer 8b23f3e260 feat(rail): sort the vertical tab rail by activity, and give its rows the home screen's card
The vertical rail lists exactly the sessions both home screens list, so it now
answers their question the same way instead of showing the raw tab order.

Order: new per-device `tabRailSort` (App Settings -> Appearance -> Tabs ->
Vertical Rail Order, default "By activity"). It runs `CodemanSessionOrder` over
rows classified by `_mobileOverviewState`, i.e. literally the home screens'
comparator, `lastSubmitAt`-anchored running group included.

It is applied as the flex `order` property, never by reordering the DOM.
`#sessionTabs` stays in `sessionOrder`, which is what keeps the Alt+N badge
honest (it names a shortcut, not a row position, so it deliberately does NOT
run 1,2,3 down a sorted rail), and keeps drag-and-drop, the arrow-key walk, the
sidebar filter and `_scrollActiveTabIntoView()` all reading the list they
always read. A session changing state then moves one inline style instead of
forcing the full rebuild that would restart every card's animation on every SSE
tick. The incremental render path re-applies it, since a state flip adds no tab
and never reaches the full rebuild, and an empty string is what clears it when
sorting stops. Web tabs are pinned past the cards by a CSS `order: 9999`, since
`renderWebviewTabs()` emits the same markup for every layout and the flex
default of 0 would interleave them. Drag is switched off while sorting (the
drop rewrites `sessionOrder` correctly and the sort puts the card straight
back, so the affordance would be a lie); 'manual' is the way back.

Cards: detailed rail rows become bordered cards on `--bg-card`, with the stamps
line on its own full-width row and the pill at its right end. The state dot
goes 6px to 9px, keeps its orbiting ring while working and gains the green
halo; idle mutes toward `--text-muted` as the home rail does. Needs/error/
waiting reuse `home-sessions-blink-red`/`-yellow` rather than a second copy.

These card rules are RAIL-SCOPED and deliberately absent from the comma-grouped
selectors that carry both vertical surfaces: the rail is an occasional,
resizable list you scan, while the sidebar is a permanently-docked nav column
where 20 stacked cards read as a wall. Every state-dot rule also excludes
`.tab-alert-action`/`.tab-alert-idle` by hand, because those alert rules are
only (0,3,0) and these are (0,5,1)+.

Lines: the lineage bracket already drew in the rail, but its track sat 6px from
the left edge, so half of its 11px outer glow was clipped by the window frame
and it read as a thread pinned to the frame. It now runs at 10px, mid-channel
in the gutter the rail already reserves.

Tests: test/tab-rail-order.test.ts (17) drives the real `isTabRailSorted()` and
`_tabRailSortOrder()` out of app.js, covering the row model (a WORKING row
ranked by `lastSubmitAt`, which would otherwise rank every running turn as
freshly started and fail no rendering test), the Alt+N badge, and the opt-out.

Verified in Chromium across sorted/manual/simple/header-strip/sidebar with the
setting flipped at runtime: no page errors, and the header strip and sidebar
render byte-identically to before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 06:09:28 +02:00
Codeman maintainer 02b0e27898 fix: merge-time follow-ups for #400, #401, #362 and #388
Each item is from the pre-merge review of the PR it names, applied on master
rather than by pushing to a contributor branch.

#400 (response viewer, shenlvkang-collab)
- The brief view opened at `scrollTop = 0`, right when it was a single card
  holding the last row. Now that it renders the whole turn, the top is the
  turn's first narration line and the answer can be screens below it, while
  loadFullContext already scrolls to the bottom of the same turn. A multi-row
  turn now opens at its newest text; a single card still opens at the top.

#401 (loopback links as web tabs, shenlvkang-collab)
- Drop `*.localhost` from the auto-route set. Every other member is an address
  literal that can only mean this box; a `*.localhost` DNS name is not one, and
  a resolver with a search domain retries `evil.localhost` as
  `evil.localhost.<search domain>`. The link source is agent-written terminal
  output, so that set is the whole confinement on a tap that makes Codeman
  fetch a URL server-side and persist it. The page-side test stays broader
  (`isOnBoxHostname`), where a false positive only declines to proxy.
- A link to the origin root navigated nothing: the path was flattened to '',
  which openWebview reads as "no deep link", leaving an open frame where it was.
- `this.webviews` being set does not mean it is loaded. initWebviews() assigns a
  truthy empty map and only then awaits the list, so a tap during page load
  found nothing to reuse and POSTed a duplicate record. Join the in-flight
  refresh instead.
- One dashboard per dev server rather than per host spelling, which is what the
  method's own comment already promised.
- Toast on the auto-create: it writes webviews.json, broadcasts over SSE and
  adds a Run-dropdown row on every signed-in device, with a new tab as its only
  previous signal.

#362 (remote omp continuation, timkjr)
- Accept the allowlisted `mode === 'omp'` arm as-is; a blanket registry render
  would hand deepseek a locally-resolved --profile and bypass claude's own
  overlay. A registry-declared switch is the follow-up if a third mode needs it.
- Revert the whole-file Prettier reformat of docs/remote-sessions.md (docs/ is
  hand-formatted and outside `npm run format`), keeping only the two new
  sections.
- Correct three stale passages: architecture-invariants' `exec claude
  --dangerously-skip-permissions`, the `exec <cli>` paragraph (claude and omp
  now have their own arms, and the claude pane's PID is the login shell), and
  omp-integration's `-c 'omp'`. RemoteCommandMode gains deepseek and omp.
- Add the missing `_maybeCaptureOmpSessionId` remote-guard test; the sibling
  guard in `_pinOmpRespawnId` had one and this path runs earlier, on the first
  idle turn.

#388 (keyCode 229 recovery, aakhter)
- Gate notifyCanonicalData on shouldSuppressTerminalQueryResponse and
  isTerminalFocusOrMouseReport. onData also carries the DA/DSR/CPR/OSC replies
  xterm answers during Ink redraws and its SGR mouse and focus reports; any of
  those landing between the keydown and the candidate's resolution was read as
  "xterm spoke for this keystroke", standing the recovery down and leaving the
  character dropped, worst on a busy agent pane. Reached through
  window.CodemanTerminalInput: the predicates live in a module IIFE that closes
  long before this call site, so bare references would throw into the
  surrounding try/catch and stop the notify from ever running.

Every fix has a test that fails without it (verified by reverting each).
Full gate green on the combined tree: 358 files, 6849 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 05:15:41 +02:00
Codeman maintainer 5b667264b4 chore: version packages 2026-09-10 03:20:58 +02:00
Codeman maintainer 57899f879e feat(ui): add a Blur entrance animation on all four surfaces
An iOS-style focus pull: the thing arrives out of focus and the blur fades
off it as the opacity comes up. Opacity leads the blur (full opacity around
45%, blur still lifting), which is what separates it from a cross-fade.
Ships on tabs (440ms), agent windows (560ms), the terminal pane (520ms) and
connection lines (380ms), plus a `Soft focus` theme that sets all four.
Default stays `legacy`, so an untouched install is unchanged.

The terminal pane is the one surface that cannot blur itself the documented
way, and `blur` takes a deliberate exception to the "never a filter on
.terminal-container" rule. Every alternative was measured against a live
xterm and does not work: a backdrop-filter veil on ::before blurs perfectly
while STATIC, and Chrome silently drops the backdrop the moment ANY
animation runs on that pseudo-element (the veil computes blur(15.3px) while
the text behind it stays razor sharp); driving the radius from rAF buys the
same full-screen blur per frame plus main-thread work. The cost the rule
exists to avoid is inherent to blurring a terminal, so the style buys it
knowingly: opt-in, off by default, one ~520ms run per session open, class
straight back off, will-change still unset. Worst-case price, headless
SwiftShader with no GPU: frame deltas 16.7ms -> 33.3ms for the run, against
16.7ms flat for `fade`. cols x rows measured unchanged at 178x38 before,
during and after, so FitAddon never sees it.

The line entrance animates `filter` too, where each line already carried
its glow. Both kinds now hold it in --line-glow and both keyframes say
`blur(N) var(--line-glow)`, so the function lists match and interpolate
instead of the glow vanishing for the run and popping back (a lineage
line's glow is a different colour, set per element). Its 100% frame omits
`opacity` on purpose so the endpoint comes from the element's own resting
value: 0.9 subagent, 0.72 lineage, 0.95 working.

test/entrance-animations.test.ts is a new static guard over the whole
feature, not just this style: the rule -> keyframes -> theme-option chain a
style silently does nothing without, the terminal's paint-only property
allowlist (the FitAddon rule), the --line-glow contract, and reduced-motion
coverage. Mutation-checked both ways.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 02:57:40 +02:00
Codeman maintainer d4fe3afc9d feat(files): raise the download cap to 2GB and stream /api/download
The 50MB cap on file-raw, the attachment /raw route and /api/download was
memory protection for a `readFile()` that no longer exists: file-raw and
/raw were rewritten to stream through `sendFileBody()` and answer Range
requests, so size costs a read stream rather than RSS (measured: a 600MB
download moved peak RSS by ~37MB). All the cap still did was refuse
legitimate downloads of build artifacts, videos and archives.

It is now MAX_FILE_DOWNLOAD_BYTES in config/buffer-limits.ts, default 2GB,
env CODEMAN_MAX_DOWNLOAD_BYTES, 0 = unlimited. `parseByteLimitEnv()` is
separate from the `parseInt(...) || default` idiom used elsewhere in that
file precisely because that idiom reads 0 as falsy and would silently
restore the default for the one value that means "no limit".

/api/download was the last route that really did buffer the whole file. It
now shares sendFileBody() with the other two, so it streams, advertises
Accept-Ranges, and is resumable. Its Content-Disposition also goes through
buildContentDisposition() rather than raw interpolation.

Refusals move from 400 to 413 across all three, which is the correct status
for the case; with the cap at 2GB it is a path almost nothing reaches now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 02:57:07 +02:00
Codeman maintainer 5130ca6633 fix(pr-bot): fail fast when the review model's budget is spent
Claude Code answers an exhausted model budget INSIDE the turn ("You've
reached your Fable limit. Run /usage-credits to continue or switch models
with /model.") and then sits there with nothing to write. The reviewer never
produces a report, so `runTurn` waited out its full 40-minute deadline and
reported a bare "timed out after 40 min without a report", which reads as a
hung reviewer rather than an account that needs attention.

Measured on 2026-09-08: #388, #393, #394 and #377 each lost 40 minutes this
way, and because every attempt counted, all four reached MAX_AUTO_RETRIES and
would NOT have been picked up again once the budget returned. One spent
afternoon quietly took the whole queue out of service.

`findModelLimitNotice()` reads the notice off the pane and `runTurn` returns
a new `limit` outcome instead of waiting. It is consulted in exactly two
places, both of which mean "the turn produced nothing": on a stop where
`isDone()` is still false, and on each timed-out wait slice. A review that
merely discusses usage limits in its own findings therefore cannot be
mistaken for one that hit the wall, and the pattern matches neither the model
name nor a straight apostrophe, since the pane renders a typographic one and
every model prints the same sentence.

A spent budget is an account condition, not a bad PR, so it no longer spends
the per-head retry budget: the queue resumes by itself when the budget does.
Telegram now names the cause and the file to change.

Tests use the pane captured verbatim off the run that lost the 40 minutes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 19:27:43 +02:00
Codeman maintainer a164c07f92 chore: version packages 2026-09-07 22:54:36 +02:00
Codeman maintainer 4f2dfb4e6d fix(mobile): carry resumeId through the phone overview's past rows
#386 made Codex conversations resumable from Past Sessions, and
resumeMobileOverviewSession() correctly passes row.resumeId on to
resumeHistorySession(). The phone's own row projection never copied the
field off the unified-list item though, so row.resumeId was always
undefined there and a tapped Codex row started a FRESH session on a thread
that was already on disk. The desktop path worked; only the phone was blind.

The test fails without the projection line, and pins the other half too: a
claude row must not grow a resumeId, since the field is what distinguishes
"resume this conversation" from "start a new one".

Docs: CLAUDE.md and architecture-invariants both still described the unified
list as merging Claude transcript files. It has been three stores since this
PR (Claude's ~/.claude/projects, omp's ~/.omp/agent/sessions, codex's
~/.codex/sessions), the alias field keeps its Claude-era name without being
Claude-only, and the scanner-only rule behind resumeId was written down
nowhere.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 22:44:25 +02:00
Codeman maintainer f1b7283393 fix(cli-registry): guard workDetect.workingLine like every other config regex
#385 made the composer glyph and the working status line per-CLI registry
data, which is right, but `workingLine` arrived as a config-supplied regex
validated with a bare `new RegExp()`. That skips `compileVersionRegex()`,
the helper the registry uses for exactly this: a `~/.codeman/clis.json`
override can set the field, the compiled pattern is run against every
accumulated PTY chunk and every pane capture, and a nested quantifier there
backtracks on the event loop for the whole server rather than one session.

Route it through the helper in both places, which are not redundant: the
schema refine rejects the entry at LOAD time so a bad pattern never reaches
a session, and `_workingLinePattern()` compiles through the same helper so
the runtime cannot hold a pattern the schema would have refused. The helper
returns null instead of throwing, so the Claude-pattern fallback stops being
a try/catch and becomes structural. Both shipped patterns compile unchanged,
and Claude's is behaviourally identical to CLAUDE_WORKING_LINE_PATTERN.

Also match the Codex footer case-insensitively on the E. It was
characterised against codex-cli 0.152.1, which prints a lowercase `esc`;
a version capitalising it would make the whole fix silently inert, since
the pane would simply never look like it was working.

Docs: CLAUDE.md, architecture-invariants and cli-registry.md all still
stated the Claude-mode-only rule this PR retires, and none of them named
the new capability or the regex guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 22:42:20 +02:00
Codeman maintainer 7fde978ce8 chore: version packages 2026-09-07 19:11:56 +02:00
Codeman maintainer 8ee7926e27 feat(agent-cases): tag agent-spawned case dirs and sweep their leftovers
A long orchestration creates one case directory per worker and deleting the
sessions never removed them, so ~/codeman-cases accumulated scratch folders
that were indistinguishable from real projects. They are now labelled and
have a cleanup path.

- src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven
  spawn gets a .codeman-agent-case.json marker (when, by whom, parent session,
  mode). Only the create branch writes it, so a linked case, a cloned repo or
  any pre-existing path is never labelled; reading is total, so a malformed
  marker means "not agent-created" rather than a half-trusted entry.
- The signal is the new X-Codeman-Agent-Origin header the skill preamble sets
  on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body
  field, falling back to a resolved parentSessionId so a worker spawned by a
  stale skill copy is still labelled.
- GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is
  a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage
  badges each case and offers a review-then-delete sweep that names every
  directory in its confirm and skips any case a live session is working in.
  Removal stays on the existing DELETE /api/cases/:name.
- Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was
  written per claude session and never removed (236 leftovers measured on a
  working machine). Now deleted with the session and swept at boot, guarded by
  a live-session keep set plus a 7-day age floor.

Verified end to end on an isolated instance: marker written for header, body
and lineage-only spawns, absent with no agent signal and for a pre-existing
directory; inUse flipping on session end; badge, sticky bar, confirm and sweep
driven in a browser; preamble seeded on create, removed on delete, boot sweep
taking only the aged orphans.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 19:09:24 +02:00
Codeman maintainer 61d22eee1c chore: version packages
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 00:06:00 +02:00
Codeman maintainer 92af855ce4 fix(base-path): keep the crash beacon under the mount, strip CODEMAN_BASE_URL in tests, add the wiring test
The merge-time items from the #381 review. navigator.sendBeacon is not fetch,
so the base-aware wrapper never saw the two crash-diag beacons and a sub-path
install posted them to the origin root every two seconds. The test suite now
strips CODEMAN_BASE_URL like CODEMAN_GESTURE, since the constructor reads it
as a fallback and an operator who exports it would see the root-install
byte-identity assertions fail. test/base-path-server.test.ts boots a real
WebServer under /codeman and checks the ingress strip, the base injection,
the rebased redirects, the 404 envelope and a prefixed WebSocket upgrade.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 23:11:01 +02:00
Codeman maintainer 80397fe140 fix(hooks,mobile): the merge-time items from the #367 and #368 reviews
#367 (UserPromptSubmit hook): `hook:prompt_submitted` went on the wire
unregistered; it is now in both SSE registries (158 = 158), and the hook only
lands in the run summary when the conversation actually moved, since one row
per prompt would evict useful rows from the 1000-event FIFO and clutter the
Summary timeline and /api/search.

#368 (Add Case header submit): the pending-state dimming targeted the footer
button, which the <=860px layout hides, so on a phone the only visible submit
control stayed at full brightness while a clone ran. The header button now
dims too, and a static test pins the header-submit contract so it cannot
silently disappear again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 23:05:26 +02:00
Codeman maintainer a2aaea3c0e docs(file-picker): state the Home/cases nesting the right way round, and document the new fallback chain
The two merge-time edits the #383 review asked for. The comment above the
picker's fallback chain said Home is nested under Codeman Cases; on the
native default it is the other way round (~/codeman-cases sits inside ~).
And the "Filesystem path picker" paragraph in architecture-invariants still
said the picker falls back to /mnt/d, which #383 changed to: the session's
Current Folder, then the Codeman Cases root, then /mnt/d, then the first
root. No code behaviour changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 21:20:45 +02:00
Codeman maintainer 9f5010aa51 fix(pr-bot): announce a bot-made merge once, cap automatic retries, show why a review failed
Observed on the first live merge (#383 via the Telegram button): runConfirmed
announced the merge and the scan five seconds later announced it again as a
closed PR. The scan now stays quiet for PRs the bot itself merged or closed,
and a merge of a `merge-with-fixes` verdict reminds that merging applies none
of the listed fixes.

A failed review used to be re-queued on every scan with no limit (two PRs
failed once each and were retried fine, but a head that keeps failing would
cost a session every ten minutes forever): three failures on one head now
stop the automatic retries until /review N or a new push. The failure notice
carries the reviewer's last message, so "finished without writing
report.json" says what it wrote instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 21:09:47 +02:00
Codeman maintainer f33b37c008 feat(pr-bot): review open PRs in Codeman sessions and report over Telegram
Maintainer tooling in scripts/pr-bot/: a daemon (systemd user unit
codeman-pr-bot) that lists open PRs with gh, reviews each head commit once in
a Codeman claude session (`prbot-<n>`) running in a private `git clone
--shared`, and sends the verdict, ranked findings, checks and a recommendation
to Telegram with action buttons. Merge, close, post-comment and approve-CI
happen only from a Telegram command or button plus a confirmation tap; the
bot never writes to GitHub on its own. The Telegram token and chat id come
from the existing notifier bot's env file.

Verified live: three PRs reviewed end to end (383, 363, 368), reports
delivered with buttons, reviewer sessions on the pinned model. Findings
along the way, each fixed and documented: a linked worktree inherits the
main checkout's model pin (hence the shared clone), undici's 5-minute header
timeout cut off the first review, gh was missing from the service PATH, and
the periodic scan orphaned an in-flight review's record.

typecheck/lint/format now cover scripts/pr-bot; tests in
test/pr-bot-{report,state,commands}.test.ts; guide in docs/pr-bot.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 21:05:42 +02:00
Codeman maintainer 8ad2215118 fix(docker): close the three adoption gaps the negative guarantee missed
Review follow-ups to #357. Each is a path that still touched, or still hid, a
container Codeman does not own.

**Export still mutated it.** The four fail-closed layers cover create/start/
stop/remove, but `POST /api/docker-cases/:name/export` reaches the container
twice through neither: a full export `docker commit`s it, and even a
workspace-only export `docker pause`s it first for snapshot consistency. Pause
freezes the owner's processes for as long as the tar takes, on a container we
promised not to touch. Full export is refused for an adopted case (it packages
someone else's container, with their logins, into a bundle Codeman hands out);
workspace-only keeps working and no longer pauses, accepting a live filesystem
the way `tar` does on any running host directory.

**A freshly linked OWNED case became unusable.** The run menu now probes the
container for its CLIs, and a failed probe hides every agent mode behind the
reason. For an adopted case that is right. For an owned one the container does
not exist until the first session launches it, so every newly linked Docker case
answered `container "codeman-case-x" not found (adoption never creates a
container — start it yourself first)` and offered nothing but Shell, for a
container the launch chain was about to create itself. A failed probe is
recorded only when the case is adopted; `CaseInfo.docker.owned` is on the wire
so the frontend can tell them apart. Verified in a browser: owned-with-no-
container offers all ten modes and no notice, adopted-but-stopped offers Shell
and says why.

**Multi-user gating.** Adoption is admin-only, unlike `docker-link` beside it.
Linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has
already confined to the caller's space; an adopted container's mounts are
whatever its owner gave it, so one mounting `/` hands the adopter a shell over
the whole host — exactly the workspace scoping multi-user mode exists to
enforce. Listing the engine's containers and browsing directories inside an
arbitrary one are machine-level reads and follow the docker-HOST policy for the
same reason. The preflight is deliberately not admin-only: the run menu fires it
for every docker case, so it admits a non-admin for a container already linked
to a case they can access, and nothing else.

Verified end to end against a real pre-existing root container (alpine + tmux,
no bind mounts): adopt, claude session inside it, workspace export, session
close and case unlink all left `StartedAt`, `RestartCount`, `Pid` and `Paused`
untouched; the pane ran the CONTAINER's claude, without
`--dangerously-skip-permissions`; a stopped container was refused at both
preflight and launch and was never started.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
2026-09-05 16:22:54 +02:00
Codeman maintainer 3d8ffcb9a2 Merge pull request #357 from dignfei/feat/docker-adopt-existing-container
feat(docker): attach a case to an already-running container

Conflicts came from work that landed after the PR was opened, and each is
resolved onto the newer abstraction rather than by keeping the older code:

- `defaultDockerCommandForMode` is registry-driven since #347, so the PR's
  `runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude
  Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the
  refusal is visible only inside the container, so an adopted root container
  otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and
  `test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch.

- The probe's mode list and its mode -> binary table both duplicated the
  registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is
  also what fixes the merge's silent regression: the hand-written list predates
  `omp`, and the run menu gates every docker case on this probe, so owned
  containers would have lost that mode. `shell` needs no arm — it declares no
  binary, so it is dropped from the lookup and reported available regardless.

- The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one
  `missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto
  it. Its test now pins the single gate instead of counting seven arms.

- The create arm keeps #349's swap-limit warning filter, which the adopted arm
  never reaches; the run-mode list gains `omp` from #353.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
2026-09-05 16:21:51 +02:00
Codeman maintainer 6f7add7ce4 chore: version packages 2026-09-04 20:46:19 +02:00
Codeman maintainer eeb5f9d0b2 docs: web-tab egress guard, capability revocation and referrer policy
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:15 +02:00
Codeman maintainer 2ab21c1b32 fix(webview): revoke proxy capabilities on logout and stamp Referrer-Policy
WebviewCapabilityStore.revokeOwner() shipped for two releases with a docstring
claiming logout called it and no caller at all. The capability is a bearer
credential exempt from cookie auth with a rolling TTL refreshed on every use, so
a proxy URL that leaked (browser history, a screenshot, a dashboard with a loose
referrer policy) stayed valid for as long as anything kept polling it.

- POST /api/logout revokes the caller's capabilities (all of them in single-user
  mode), the admin forced logout revokes the target user's, and user deletion
  revokes whatever that user had open. revokeOwner returns the count for the
  admin audit line.
- Proxied responses carry `Referrer-Policy: same-origin` and the upstream's own
  policy is dropped: every URL inside the frame carries the capability, and a
  dashboard on no-referrer-when-downgrade or unsafe-url handed it to any
  third-party host it linked. Verified with Playwright that a sandboxed frame
  under an upstream `unsafe-url` sends no Referer to a third party while the
  root-absolute fetch and the CSS-triggered 404 fallback still reach the
  dashboard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:13 +02:00
Codeman maintainer 550e08a791 fix(webview): refuse link-local and cloud-metadata targets on the resolved address
The web-tab proxy, its Test probe and its WebSocket relay accepted any http(s)
host. A live PoC relayed an IMDSv2-shaped PUT with custom headers to a loopback
echo server through a capability and no cookie, and 169.254.169.254 (decimal,
hex, IPv6-mapped, or via a DNS name) was as valid a dashboard as any other.

Loopback and RFC1918 stay allowed on purpose: a localhost Grafana is the feature.
Only link-local and the fixed cloud-metadata addresses are refused
(169.254.0.0/16, fe80::/10, fd00:ec2::254, 168.63.129.16, 100.100.100.200,
metadata.google.internal), at three stages that are each load-bearing:

- the Zod schema, so a save gets a clear refusal;
- a synchronous hostname check at every connect site, because net.connect skips
  DNS for an IP literal and a lookup hook never sees one;
- a `lookup` hook on an undici Agent (webviewFetch) and on the ws client, which
  judges the RESOLVED addresses of a name and refuses when any is blocked. This
  is what closes DNS rebinding, which a hostname-string check cannot.

Adds undici@^6 so the proxy runs the package's own fetch with the package's own
Agent; a package Agent handed to Node's bundled fetch can mismatch protocols.

Verified live on an isolated beta: 169.254.169.254.nip.io (a real name resolving
to the metadata address) is refused by probe, proxy (403) and WS relay (4003),
while 127.0.0.1.nip.io still passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:12 +02:00
Codeman maintainer 99ad9cb236 fix(docker): never exit the server unless something is known to restart it
#373 restarts the Compose container by exiting the server, which is right for
the shipped deployment: `restart: unless-stopped` relaunches it. The updater
verified that policy through the Docker socket and, when it could not (no
socket mounted), failed open and exited anyway. Failing open is the correct
choice for the GATE, where refusing would block every install without a
socket, but not for the kill: a container the daemon does not restart goes
down for good, with no UI left to recover it from. That is exactly the case a
plain `docker run` of this image without `--restart` produces, and the image
sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path.

The decision now happens server-side, where both the socket and the Compose
env are reachable, and rides down to the script as `--restart-by-exit 0|1`.
It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there
and only there, since that file is what sets the restart policy; the image ENV
deliberately does not) or when the daemon confirmed an auto-restart policy.
Otherwise the build still lands, the status becomes
`completed-needs-manual-restart` with the `docker restart` hint, and the
server keeps running. The shipped deployment is unchanged in effect: with the
socket it was already confirmed, and without it the declaration now covers it.

Also: a root-run `Start-Codeman.sh` (common on Unraid) created the
fingerprint baseline's `.codeman` directory before the container's first start
and left it root-owned, which the unprivileged server could then never write
its own state into. It is chowned to PUID:PGID when running as root.

Verified with a real image build of the merged tree (classic builder; this
box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain
present, the four CLIs at their pins, docker/.env absent, and `docker inspect
$HOSTNAME` returns the restart policy through the mounted socket as that user.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 14:36:35 +02:00
Codeman maintainer 823f56a243 Merge pull request #373 from opticon454/feature/docker-self-update
feat(docker): restore in-app self-update in the Compose dep
2026-09-04 14:25:56 +02:00
Codeman maintainer 72fd231d11 test(setup): one answer for CODEMAN_DATA_DIR, the strip from #371
#356 and #371 fixed the same leak two ways. #356 pointed CODEMAN_DATA_DIR at a
second throwaway directory and cleaned it up in afterAll and on exit; #371
deletes the variable along with CODEMAN_INSTANCE and CODEMAN_TMUX_SOCKET, so
`getDataDir()` falls back to `homedir()`, which the temp HOME already redirects.
Merged as they were, setup.ts set the variable and deleted it a few lines
later, and the second directory was created for nothing.

The strip wins: same protection, one tree to clean up, and the isolation test
#371 adds pins the list statically. The extra directory, its restore and its
two rmSync calls go, the vitest config `env` entries that set the same variable
go (they were documented as inert and would now be contradicted by the setup
file either way), the two test comments that described the old mechanism are
reworded, and CLAUDE.md's testing paragraph names the three stripped variables
and why CODEMAN_INSTANCE has to be stripped in the setup file rather than a hook.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 14:20:11 +02:00
Codeman maintainer 65d19c725e Merge pull request #371 from opticon454/fix/test-env-instance-isolation
fix(test): strip the instance-selection env vars in test/setup.ts
2026-09-04 14:20:04 +02:00
Codeman maintainer 80626567b2 chore: version packages 2026-09-04 14:01:20 +02:00
Codeman maintainer a81e87f440 fix(cli-registry): log why clis.json was ignored, and say 0600 when that is the rule
The loader refuses a `clis.json` with any group/world permission bit, read bits
included, so a file created with a normal umask (0644) is ignored. That is a
defensible posture for a file that chooses the binaries Codeman spawns, but two
things around it made the override feature look dead: the warning said
"group/world-writable", which a 0644 file is not, and `LoadResult.warnings` was
returned to a caller nobody wired up, so nothing anywhere printed it. A user
following the docs got silence.

The message now names the rule and the command that satisfies it, the loader
logs every warning once on first load (the result is memoized, so once per
process), the module header stops claiming that nothing ever writes (the
quarantine rename of a malformed file is a write, on first use) and the
registry doc gains a short section on the override file with the 0600
requirement in it. Whether the check should relax to writable bits only is a
separate decision; this keeps the shipped behaviour and makes it visible.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:51:20 +02:00
Codeman maintainer 2e0129f1f8 docs(test): name the real reason the suite could reach ~/.codeman
#356 stopped a bare suite run from overwriting the production
`remote-hosts.json` by pointing `CODEMAN_DATA_DIR` at a throwaway dir, and it
gated every case-tree delete on the temp HOME. Both changes are right; the
explanation written next to them is not. It says `os.homedir()` reads
/etc/passwd rather than `$HOME` on Linux, which would mean the temp HOME in
test/setup.ts never worked. It does: libuv checks the env var before the passwd
entry (measured: `HOME=/tmp/x node -e 'console.log(os.homedir())'` prints
/tmp/x), and CLAUDE.md's testing section relies on exactly that.

What bypasses the temp HOME is `CODEMAN_DATA_DIR` itself. `getDataDir()` reads
it as an absolute override before it looks at `homedir()`, so one inherited from
the shell (a second instance, a beta run) sends the whole suite at the real data
dir. That is the case setup.ts now closes, and #371 names the same variable from
the other direction.

The comments in setup.ts, the `safeRmHomeTree` helper, the voice-routes and
case-clone tests now say that, and the containment gate is described as what it
is: defense in depth. CLAUDE.md's testing paragraph gets the same note so the
next reader does not chase a homedir() bug that does not exist.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:50:22 +02:00
Codeman maintainer 28b44237ae fix(remote): classify the has-session probe by exit status, and forget it once the pane is back
#355 made the remote auto-reconnect watcher revive a dead pane only when the
durable remote tmux session is verifiably still alive, which is the right rule:
a clean Ctrl-C / Ctrl-D / exit tears that session down and must never relaunch
a fresh agent. Its probe, though, read `has-session`'s stdout and treated an
empty string as "gone". `tmux has-session` prints NOTHING on success (measured
on a scratch socket: exit 0, empty stdout, the failure message goes to stderr),
so every live remote session classified as gone and transport-drop reconnects
were silently disabled along with the clean-exit revives.

The probe now goes by exit status through a pure, unit-tested mapping
(`classifyRemoteAliveExit`): 0 is alive; ssh's own 255, a timeout (`killed`,
no numeric code) and a spawn failure are unknown, which the watcher already
treats as do-not-revive; any other status is the remote command's and means
gone (tmux's 1 for a missing session, 127 when tmux is not installed there).

Two smaller things in the same area:

- The cached answer was never invalidated, so after one successful reattach a
  stale `true` would have revived the NEXT clean exit (the original bug back
  after the first transport drop), and a cached `false` from a clean exit would
  have left a manually restarted session with auto-reconnect permanently off.
  The tick now forgets the cache entry whenever the pane is seen alive.
- The fire-and-forget probe has a 15s timeout against a 5s tick, so an
  unreachable host stacked up to three ssh processes per dead session. An
  in-flight set caps it at one.

The probe command is pinned as a literal string, and the reattach-then-clean-exit
sequence is driven through the watcher in the tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:50:22 +02:00
Codeman maintainer 1e24817b51 chore: version packages 2026-09-02 10:49:36 +02:00
Codeman maintainer 71ffbf18e4 chore: version packages 2026-09-01 21:55:00 +02:00
Codeman maintainer 826ddaa9aa build(docker): ship the Docker CLI in the Compose image, not the whole engine
docker/server.Dockerfile installed Debian's `docker.io` to get a client for the
socket mounted by Compose. That package is the full ENGINE: even with
--no-install-recommends it pulls 15 packages including containerd, runc, dmsetup
and iptables, none of which a container that only talks to a mounted socket can
use, and it ships Docker 20.10.24 (2023).

Copy the CLI and the buildx plugin from the official docker:29-cli image
instead. Measured on the same node:22-bookworm-slim base: 266 MB -> 108 MB, so
158 MB smaller with a current CLI (29.7.2) in place of a two-year-old one.

Three things verified rather than assumed, by building the real image and
running it:

- docker:cli is an ALPINE image, so copying a binary into this Debian one is
  only safe because the binaries are static Go builds (ldd: "Not a valid dynamic
  program"). In the built image, `docker --version`, `docker ps` and
  `docker build` all work against a mounted host socket as the unprivileged
  runtime user.
- buildx is copied on purpose. scripts/build-agent-image.mjs shells out to
  `docker build` and Codeman auto-builds the agent image on the first Docker
  case. Without the plugin that still works today — CLI 29 falls back to the
  classic builder, tested — but that builder is deprecated and will be dropped,
  so the plugin keeps the path supported.
- docker-compose is NOT copied: Codeman never shells out to it.

Pinned to the 29 major, matching how the base images here are pinned.
2026-09-01 21:54:56 +02:00
Codeman maintainer e2b72aafd7 chore: version packages 2026-09-01 11:32:50 +02:00
Codeman maintainer b15cc0eb1a fix(ui): keep the plan-usage chip's 5h slot when no session window is open
The header chip silently shrank from "5h 4% · 7d 52%" to a lone "7d 52%", which
reads as half the feature breaking rather than as an idle window.

Nothing was broken. Claude Code documents `rate_limits.five_hour` as "present
only while the API reports it and its resets_at has not passed", so between
5-hour session windows the key simply leaves the statusline payload. Codeman's
snapshot replaces the Claude half wholesale on every sample, so the segment
disappeared until usage opened a new window. Confirmed against a live 2.1.252
session by capturing real statusline payloads on an isolated tmux socket: the
boot render carries no `rate_limits` at all, and the first post-response render
carries both windows.

The slot now stays, with a dimmed em dash. Claude only: a missing CODEX bucket
means that plan has no such limit rather than an idle window, so those stay
omitted (pinned by the existing test). The placeholder can never stand alone
either — hasWindows() still gates the row, so a provider reporting nothing
renders nothing rather than a row of dashes. The tooltip says "5-hour limit: no
active session window" instead of dropping the line.

Verified in a real browser against a dev instance: the idle chip renders
"5h — · 7d 52%" with the dash at opacity 0.55 in --text-dim while the live
value keeps its green, and the chip holds its shape (100px idle vs 107px with
both windows).
2026-09-01 11:32:18 +02:00
Codeman maintainer 2a32b5064a Merge pull request #349 from opticon454/feature/docker-compose
Docker Compose deployment: Codeman runs in a container and spawns Docker cases
as SIBLING containers through the mounted host socket (Docker-outside-of-Docker).

Resolved the README conflict (master had grown to eight CLIs since the branch
was cut) and moved the Compose blurb out of the feature bullets into Quick
Start, next to the other ways of starting Codeman.

Three review findings from the PR discussion are fixed here rather than left
for a follow-up, because two of them are shipped-image problems:

- `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against
  the whole context-relative path, so `docker/.env` — which the deployment's own
  README tells the user to fill with CODEMAN_PASSWORD and provider API keys —
  was picked up by `COPY . .` and baked into the image at
  /opt/codeman/docker/.env. Verified in both directions against a real build
  context: with a canary secret in docker/.env, the unfixed ignore file lets
  /ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for
  the checked-in .env.example) leaves only the example behind.
- `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which
  still hardcoded ~/codeman-cases, so `codeman skill install --case <name>`
  reported "Case not found" on exactly the deployment the override exists for.
  Both now resolve through config/cases-dir.ts. state-store.ts keeps its own
  literal on purpose: that one migrates the historical ~/claudeman-cases
  directory by name and is about the old default, not the active location.
- CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the
  three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore
  joins the documented list of files that genuinely belong in the repo root.

The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug
fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/
deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's
lastClaudeSessionId was handed to every non-claude CLI.

Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax,
public assets, lockfile.
2026-09-01 11:32:03 +02:00
Codeman maintainer 0da0c8219d chore: version packages
1.24.2. Also corrects two numbers in the CLAUDE.md trust-dialog paragraph that
was written while the fix was still uncommitted: the keystroke cap is 6, not 3,
and the scan now schedules its own follow-up read rather than waiting on PTY
output that a static dialog never produces.
2026-09-01 02:31:31 +02:00
Codeman maintainer aaa93d4252 fix(session): answer Claude Code 2.1.252's reversed folder-trust dialog
Every claude session in a directory claude had not seen before died about six
seconds after it started (`Pane is dead (status 1)`), before the agent drew a
composer. Reproduced on a fresh case and measured.

Claude Code 2.1.252 rewrote the dialog. It used to be

  ❯ 1. Yes, I trust this folder
    2. No, exit

and is now unnumbered, reversed, and highlights the option that quits:

  ❯ No, exit
    Yes, I trust this folder

Detection still worked (the confirm affordance carries the match once the
numbered option text is gone), so the failure was entirely in the answer: the
auto-accept pressed Enter on the highlighted default, which is now exit.

trustDialogNextKey() reads the ❯ marker off the rendered pane and returns ONE
keystroke at a time: an arrow while the cursor is on the wrong option, Enter
only once the screen shows it on the trust option, and null for a frame that
does not say. Both layouts are handled, and which way the trust option lies is
read from the frame rather than assumed, so a further reordering costs a
repaint instead of a session. The last marked option wins, because the
direct-PTY fallback reads an append-only buffer where an older frame must not
out-vote the freshest one.

Two things only a live pane showed:

- The scan ran solely from the PTY onData handler. The arrow that moves the
  cursor is the last output the pane produces, so the first fix parked every
  session with the cursor sitting on the right option and no Enter ever sent.
  It now schedules its own follow-up read (_trustDialogTimer, cleared in
  _clearAllTimers()), offset past the scan throttle so the chain cannot break
  on a boundary.
- The keystroke cap goes 3 -> 6, since answering is no longer one press.

The bundled codeman skill had the same blind \r as its bounded fallback, so
preamble 1.21.0 replaces it with _trust_key/_accept_trust: read
terminal?full=1, steer onto the trust option, re-read, then confirm. Those
keystrokes go out under their own clientId, because input sequence numbers are
monotonic per client and spending prompt numbers on dialog keys would make the
next send-and-wait look like a stale duplicate and vanish while reporting
success. The readiness recipes in docs/extending-codeman.md,
docs/api-reference.md and the skill's own reference carry the corrected answer,
plus a symptom-table entry for a worker whose pane is dead seconds after spawn.

Verified live on an isolated instance (own data dir and tmux socket): fresh
case -> arrow at 5 s -> Enter at 7 s -> composer, with hasTrustDialogAccepted
recorded. With the server-side auto-accept disabled in a throwaway copy, the
skill's fallback cleared a genuinely parked dialog in 1.1 s and spawn_worker
took a brand-new case to a live composer in 7.2 s; spawn_workers + sendwait +
last_text then ran end to end.
2026-09-01 02:31:26 +02:00
Codeman maintainer 3518af3a9f docs: correct CLAUDE.md drift and document four undocumented subsystems
Audit of CLAUDE.md against the tree. Verified still accurate: the 31-module
frontend load order (matches index.html exactly), SSE registry parity at
157 = 157 (confirmed by running the parity test), config/ 21 files, types/ 22
domain files, 136 mobile device profiles, the version line, and every Quick
Reference command.

Drift corrected: 24 route modules to 25, ~220 handlers to ~227, system-routes
51 to 56, app.js ~5K lines to ~6.7K and 30 modules to 31, install.sh 92KB to
104KB. Completed the CLI resolver inventory, which was missing
deepseek-cli-resolver and omp-cli-resolver even though both modes are
documented, and named the shared cli-executable-resolver lookup chain.

Filled the gaps found by sweeping every src module against the file:

- Owner tab layouts (COD-359) had 6 source modules, 7 test files, 2 routes, an
  SSE event and a state.json key, with zero mentions anywhere in CLAUDE.md or
  docs/. The paragraph records the four things a reader would otherwise get
  wrong: it is backend-only as of 1.24.1 with no frontend consumer, the service
  is the sole mutation boundary, it projects onto PUT /api/session-order rather
  than replacing it, and reconciliation is gated on a successful restore.
- codeman doctor and codeman users were undocumented top-level CLI commands.
- Four subsystems whose invariants lived only in their @fileoverview:
  the workspace-trust dialog recognizer, proc-tree's bounded walk (the
  2026-07-30 incident that took a machine down), deepseek-web-server (one
  child process, deliberately not a shell session), and the Files panel
  search matcher (globs are never compiled to a RegExp).

Also fixes a stale "156 event types" comment in constants.js (actual: 157) and
a contradiction in AGENTS.md, which still carried the retired "never run the
full suite inside a managed tmux session" rule against CLAUDE.md's current
"npm test is the gate and is safe to run bare".

Note: the trust-dialog paragraph documents trustDialogNextKey(), which is part
of a sibling session's in-flight fix for the Claude Code 2.1.252 layout change
(unnumbered, reversed options with "No, exit" highlighted, so a blind carriage
return picks exit and kills the pane). That fix was uncommitted in the shared
tree when this landed, so the doc leads the code until it is committed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NMN8UuvdBim3iM87reuQ9Z
2026-09-01 02:10:09 +02:00
Codeman maintainer e5c5d890aa chore: version packages 2026-08-31 22:34:57 +02:00
Codeman maintainer d5b5f8f618 fix(docker): make the dsh profile install survive pnpm's build-script gate
Follow-up to #350, which fixed the actual blocker (issue #352): `dsh plugin` is
a thin forwarder that `spawnSync`s a literal `pnpm` with no npm fallback, so an
image without pnpm dies at exit 127 and takes the whole build with it.

That PR also pinned an allowlist of the two packages whose lifecycle scripts
pnpm blocked at the time. Replace it with a policy that cannot go stale: pnpm,
unlike npm, refuses dependency build scripts by default and FAILS the install
over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on pnpm 11.24), and the
names to allow move between rebuilds because `@deepseek-harness-tui/dsh-tui` is
resolved by dist-tag, not pinned: 0.9.3 pulled `@google/genai` (whose script is
a literal `preinstall: no-op`), 0.10.0-beta.x does not. An allowlist of two
names would have let the next tree break the build the same way. Allowing them
wholesale is also the exposure this image already accepts three layers up,
where `npm install -g` runs the install scripts of every transitive dep of the
five CLIs above with no gate at all.

Also correct a comment in the `/api/deepseek/install-profile` route that
asserted the opposite of what #352 proved ("dsh bundles its own package
manager, so no system pnpm is required"). The route's behavior is already
right: dsh's own "pnpm not found on PATH" stderr reaches the caller as the
OPERATION_FAILED detail, so the UI's "add a terminal profile" button names the
fix. Documented the prerequisite in docs/deepseek-integration.md, and taught
the docker-cases image smoke test about `dsh`/`omp` plus the profile check that
`dsh --version` does NOT cover.
2026-08-31 22:26:09 +02:00
Codeman maintainer 7762809202 Merge pull request #350 from opticon454/bugfix/dsh-pnpm
fix(docker): install pnpm for DeepSeek profile
2026-08-31 22:22:30 +02:00
Codeman maintainer 02bbf13b3c chore: version packages 2026-08-30 16:29:14 +02:00
Codeman maintainer da91b4353b Merge pull request #353 from timkjr/omp-mode
feat: add OMP (Oh My Pi) as a new session backend
2026-08-30 16:16:26 +02:00
Codeman maintainer d8688dc143 fix(web): drop the provider label from the plan-usage chip when there is only one
The chip prefixes every row with the provider name, so a machine that only
has Claude limits renders "CLAUDE 5H 60% 7D 23%" — a 46px label naming the
only thing it could possibly be. The name exists to tell two rows apart, so
it should only appear when there are two.

updatePlanUsageChip() now checks whether both Claude and Codex actually have
windows before building the rows, and emits the .pu-provider span only in
that case. The tooltip keeps naming the provider in both cases: it has the
room, and the chip no longer does.

Verified in a browser on an isolated beta instance: Claude-only renders bare
windows with no .pu-provider in the DOM, Codex-only the same, and the
two-provider chip is byte-identical to before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 13:59:40 +02:00
Codeman maintainer 23fae0c5af chore: version packages
Codex plan usage in the header chip (#346), a visible inline rename in
the session sidebar (#345), and the install.sh Tailscale re-run fix plus
the README network-access prompt description.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:25:32 +02:00
Codeman maintainer 23e32b22d5 docs(readme): describe the actual three-way network-access prompt
The installer bullet still described a two-way choice with 0.0.0.0 as "the
default", which predates the Tailscale option. The prompt has offered three
choices for a while (Tailscale / any device on your network / this machine
only), and the default is computed from what is already on the machine rather
than being fixed at 0.0.0.0.

Now states all three options, that the Tailscale one is a loopback bind
fronted by `tailscale serve` with the tailnet as the login, and how the
highlighted default is chosen. Line 220 already documented the Tailscale
option correctly; this was the only stale spot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:02:33 +02:00
Codeman maintainer 7dfb4acf24 fix(install): offer Tailscale setup on re-run instead of losing it to a failed build
The network-access prompt, where Tailscale serve is configured, runs AFTER
the build step. A build failure therefore exits before the question is ever
asked, and a user who then finishes the build by hand (rather than re-running
install.sh) ends up with a healthy loopback-only Codeman, a connected
Tailscale, and no serve mapping — with nothing anywhere pointing at
`install.sh tailscale`, the command that fixes it. Reported from a fresh
Ubuntu 24 install that died on the node-pty compile.

- maybe_offer_tailscale_repair(): on the update/re-run path, detect exactly
  that state (loopback bind + tailscale Running + no serve mapping fronting
  Codeman) and offer the retrofit. Silent for a deliberate non-loopback bind,
  silent once a mapping exists, silent when tailscale is absent, and prints
  the command instead of prompting when non-interactive. Returns 0 even when
  setup fails so it can never abort an update.
- print_security_notice(): the loopback branch now names
  `install.sh tailscale` when Tailscale is installed on the box, rather than
  the generic "tailscale serve / cloudflared tunnel" advice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 18:29:15 +02:00
Codeman maintainer d3f851a5e5 chore: version packages
install.sh installs a build toolchain on Linux (node-pty has no Linux
prebuild, so a stock Ubuntu 24 server died inside node-gyp with
"not found: make"), plus review hardening for #339: the write-queue
reset paths now release the one-chunk-in-flight gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 18:04:44 +02:00
Codeman maintainer a51563ce1f chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 19:14:41 +02:00
Codeman maintainer 93a1042bb3 Merge remote-tracking branch 'origin/feat/deepseek-harness' into feat/deepseek-agent-workers
# Conflicts:
#	CLAUDE.md
2026-08-25 19:02:34 +02:00
Codeman maintainer a628737d1f fix(deepseek): review-driven hardening across the harness integration
Fifteen review findings on the dsh mode, the serious ones first:

- Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys.
  _configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into
  every dsh pane and applyEnvOverrides() lands after it, so a non-granted
  owner who could redirect the base URL would have the operator's key sent
  as a bearer credential to a host of their choosing.
- Wait registry: until=stop/blocked is refused on docker and remote-SSH
  dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions).
  The HERDR triple is set via LOCAL tmux setenv, which crosses neither
  docker exec nor ssh, so such a session can never post a hook event and
  the wait burned its whole timeout on every turn.
- Approvals: a dsh item is an ALERT, not an answerable card. The answer
  route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the
  option parser cannot read a third-party TUI's frames, so an answer was a
  blind keystroke into a foreign composer), and the push notification
  carries no Approve/Deny actions for dsh sessions.
- Status shim (v3): --seq is forwarded and the server drops stale retried
  reports inside a 60s window (the TUI retries with backoff, so a retried
  'working' could land after 'blocked' and resolve an approval whose
  dialog was still on screen); 4xx responses exit 0 instead of retrying,
  so one misconfigured session cannot feed the auth rate-limit bucket
  until the hook endpoint 429s for the whole instance.
- Web-UI server: concurrent starts are serialized through a lock (two
  racing POSTs used to pick the same port and orphan the winner), and the
  readiness poll / timeout paths only clear or stop the singleton while it
  is still theirs. First click actually opens the tab now
  (refreshWebviews, not the nonexistent loadWebviews). DELETE
  /api/deepseek/web requires the privileged grant in multi-user mode.
- Cron: deepseek jobs run the same two-part launch gate as the HTTP
  create paths (impl moved into the resolver so all three share it) and no
  longer stamp a Claude default model on the session.
- Parity sweeps: quick-start's docker branch rejects deepSeekConfig like
  the remote branch; the Ralph auto-enable list gained deepseek;
  HookEventType gained agent_working; the phone overview run menu filters
  managed webview records like the desktop menu.
- install.sh: the dsh identity probe closes stdin (under curl|bash a
  child that reads stdin eats the rest of the script), bounds the exec
  with timeout where available, and is memoized to one scan per install.
- Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand
  identity (it rendered as an unstyled UA-grey button); stale markup
  comment about the web shortcut rewritten; clamp docs updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 19:01:29 +02:00
Codeman maintainer 015b865f56 fix(deepseek): review fixes for the transcript reader — docker/remote gate, poll memo, honest pairing docs
Four review findings on the worker-transcript feature:

- Docker and remote-SSH dsh sessions now keep the pane segmenter: their
  transcripts live in the container's / remote host's own ~/.dsh, which
  the local reader can never see, so the transcript path returned
  'nothing said yet' forever and an agent polling such a worker starved
  on an answer that existed. Gated on !session.docker && !session.remote
  (statically pinned) and documented in the integration guide.

- last-response reads are memoized on (path, mtime, size, blocks): the
  skill's last_text polls once per second, and each poll decompressed and
  reparsed the whole file on the event loop even when nothing had been
  appended. An unchanged poll now costs one stat.

- The pairing ladder's comment claimed /new is served by step 2; in truth
  the boot-window transcript wins for as long as it exists (deliberately:
  preferring newest-eligible would hand a worker its busier sibling's
  reply). The comment now states the real tradeoff instead of the
  aspirational one. Same for decodeZstdFrames' 'skipped' wording — a
  corrupt frame truncates the decode there, which is the safe behavior.

- stripReasoningPrefix no longer runs on user prompt text, so a prompt
  containing a literal </think> renders whole in blocks view.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 18:30:41 +02:00
Codeman maintainer 33f77c4680 fix(tabs): review fixes for the detailed rail — width-dialog default, compact wrap pass, rich-aware resets
Three review findings on the detailed-rows feature, all in its edge cases:

- The App Settings width select consulted the handheld defaults blob
  (tabRailWidth: 256) BEFORE the rich-aware default, which the renderer
  never reads — so a tablet's unsized rich rail rendered 320 while the
  dialog said 256, and a routine Save persisted the 256 (below the 288px
  tight threshold, permanently). The chain now mirrors
  applyTabRailWidth()'s actual resolution.

- _setTabRailWidth() re-rendered on a compact flip but never re-ran
  applyTabWrapSettings(), the one owner of the folder line, whose railRich
  input reads the compact class this function just toggled. A rich rail
  dragged below 240px kept emitting folder rows — persistently, for a
  stored width < 240, since the boot wrap pass runs before the class is
  first applied. The wrap pass now re-runs on the flip, with exactly one
  render either way.

- Both reset affordances (handle dblclick, Enter on the handle) reset to
  the hardcoded 256 even on a rich rail, landing it below the tight
  threshold; both now resolve the rich-aware default (320), via a new
  optional defaultWidth input on resolveTabRailKeyboardWidth().

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 18:26:07 +02:00
Codeman maintainer 6261b6f655 feat(skill): spawn and drive DeepSeek Harness workers
The agent skill could spawn a worker in any mode, but it could only
DRIVE a claude one: every other CLI has neither a real end-of-turn
signal nor an answer to read, so the recipes route them through output
markers.

dsh has both halves now -- its harness reports idle/working/blocked to
Codeman, and the previous commit reads its transcript -- so it joins
claude as a mode the four verbs work on unchanged. `spawn_workers alpha
beta:deepseek` is a mixed fleet in one call, and `sendwait` / `last_text`
/ `delete_session` need no per-mode variant.

Preamble 1.20.0 (SKILL.md's §0 heredoc regenerated from it):

- `spawn_worker` grows a deepseek branch that gates on the harness
  composer. ⚠️ Readiness there is NOT the stop signal: the harness
  reports idle at BOOT ~300 ms before its composer paints (measured
  2.26 s vs 2.56 s after spawn), so a send-and-wait fired straight after
  quick-start resolves on the boot edge, reports a turn that never ran,
  and strands the prompt in a pane not yet taking input. Waiting for the
  composer also spends that edge, since signals are edge-triggered.
- `spawn_workers` takes `name[:mode]`, so a mixed fleet stays one
  concurrent call. Case names still have to be unique -- the mode never
  disambiguates two workers that would share a directory.
- `sendwait` asks for `wait:"stop,exit"` instead of the `wait:true`
  default set. That set also carries `idle`, which for an external CLI is
  inferred from output stabilization: on a dsh worker whose TUI repaints
  rarely, the re-wait resolved in 0 ms with `signal:"idle"` on a turn
  with three minutes left to run. It also makes a wrong mode loud -- the
  modes that cannot deliver `stop` answer 400 before writing anything,
  instead of resolving on a flap.
- The self-heal resend carries `delivered:true` forward. The resend is a
  tagged duplicate, so the server truthfully reports `delivered:false`
  about a write it skipped, and §1's cleanup then read a completed turn
  as an undelivered one and kept a finished worker forever.
- dsh workers spawn with the permission posture the Run button sends,
  because the harness default still asks and a worker parked on an
  approval row cannot finish a fan-out. The multi-user clamp still
  applies.

Docs: a worked dsh flow in recipes.md, readiness and the signal rules in
verbs.md, and the corrections this makes necessary -- `stop`/`blocked`
are no longer claude-only, and `last-response` is no longer permanently
empty for deepseek. The integration guide gains a section on reading a
session back and driving one as a worker; its web-UI section was also
stale (that server moved out of a shell session).

The static guard that keeps those lists from naming some external CLIs but
not others is extended rather than exempted: it now knows the three real
classes inside that family (no transcript, no hook signals, and the
positive twin -- the modes whose answers can be read), with the hook class
derived from `hooksAvailableForMode()` so the predicate and the prose
cannot drift apart. Any other partial list still fails, and a new backend
belongs to none of the classes until someone says so.
2026-08-25 04:17:48 +02:00
Codeman maintainer d1bc0c517d feat(deepseek): read dsh session transcripts for last-response
`GET /api/sessions/:id/last-response` is how an agent (and the Response
Viewer) reads what a worker said. DeepSeek was falling through to the
pane segmenter with the other external CLIs, which for this mode is not
merely coarse but wrong: dsh-TUI paints a full-screen splash, so a
`last-response` call on a fresh dsh session answered with its ASCII-art
logo -- and anything polling for a worker's first reply reads that as a
reply.

dsh does not belong in that group. It writes a structured JSONL
transcript per session, so read it. Four things in that file shaped the
reader, all measured against real transcripts on disk:

1. dsh appends ONE ZSTD FRAME PER WRITE, and Node's zlib zstd decoder
   (one-shot and streaming alike) stops at the first frame end: a real
   56-line transcript decoded as 1 line / 158 bytes -- the session header
   alone, i.e. a silent truncation that reads as "nothing said yet"
   forever. `zstdFrameRanges()` walks frame and block headers to find
   exact boundaries; splitting on the 4-byte magic would corrupt
   everything after a magic sequence occurring inside compressed data.
   zstd is resolved at RUNTIME because it landed in Node 22.15 while the
   project floor is 22.0, so an older Node keeps the pane behaviour.
2. Every turn also records a plugin-sourced `user/message` (the runtime
   context snapshot), which must not render as the user's own words.
3. A turn that ends in an error carries the provider's message; it is
   surfaced as `Turn error: …` (and a non-error early stop as
   `Turn ended: …`) rather than as an empty string, which an agent reads
   as "still thinking" through fifteen polls.
4. Reply text is assembled per (turn, step): a finalized message wins and
   the streamed deltas fill in only for a step that never finalized, so a
   partial answer is readable mid-turn and never doubled. "Finalized" is
   tracked as a set of steps rather than as non-empty text, because a
   step whose whole reply was reasoning strips to '' at the `</think>`
   boundary and would otherwise resurrect the raw deltas in its place.

Session-to-transcript pairing is by the transcript's own header `cwd`
plus a boot window against the session's createdAt, never by
reproducing dsh's directory mangling (already two forms on disk) and
never by newest-mtime alone -- mtime alone handed a freshly spawned
worker its predecessor's answer in the same case directory.

An empty result still wins over the pane; only a Node that cannot decode
zstd falls back to it.
2026-08-25 04:14:26 +02:00
Codeman maintainer c30dfaf0e7 fix(deepseek): run the web UI server in the background, not in a shell tab
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a
shell tab running the server next to it. The shell was deliberate - the server
lived in an ordinary session so it was visible, scrollable, killable and died
with its tab, and nothing new had to supervise a long-lived HTTP server. That
reasoning was sound and the result was still wrong in use: opening a dashboard
should open one tab, and after the first launch the terminal is pure noise.

The server moves to a background child process owned by a new
`src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session
gave away for free is now explicit, which is most of the module:

- Exactly one server. A second click reuses the running one instead of racing
  it for a port; the session flow could not do this at all, because two clicks
  were simply two sessions.
- Restarted when the requested authority changes. `--trusted-host` fences dsh's
  own /api against the browser authority, and a Codeman reachable at both
  loopback and a tailnet name has two. Reusing a server fenced for the other
  origin renders a page whose every call 403s, which reads as a broken
  dashboard rather than a misconfigured one, so a mismatch restarts instead.
- Killed on shutdown. The child is detached so its whole plugin tree can be
  signalled at once, which also means it would outlive Codeman and hold its
  port against the next start - the exact EADDRINUSE this feature already got
  wrong once.
- Boot output captured and returned. With no shell tab there is nowhere else
  for a stack trace to land, so a failed spawn reports its own tail.

The endpoint is fenced at the same bar as the profile installer and for the
same reason: booting a dsh profile executes the plugin code in it, so this is a
privileged action even though it reads as "open a page". `authority` comes from
the client (`location.host`) because only the browser knows which origin is in
play, and it is regex-confined at the schema boundary - defence in depth behind
the argv-array spawn, admitting host:port in the shapes a browser authority can
take and nothing readable as a second argument.

`GET /api/deepseek/web-port` is gone; port selection moved into the supervisor,
which is the thing that knows whether a server is already running. The two
client-side probe helpers went with it, since the server now owns the wait.

Verified over the tailnet authority end to end: no session is created (session
count unchanged, one tab), the server runs on 3081 beside the user's own dsh
web on 3080, status reports the tailnet authority, and the proxied dashboard
renders with zero 4xx. Full gate green (6148 passed, +6).
2026-08-25 03:08:15 +02:00
Codeman maintainer 15ae5f5d81 fix(deepseek): make the web-UI shortcut pick a free port, verify it, and trust its frame
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.

1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
   precisely the port a DeepSeek user is most likely to be serving on already,
   so the launch died with EADDRINUSE against the user's own server. The port
   now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
   free loopback port by BINDING it (a connect probe cannot tell "free" from
   "listening but not answering yet").

2. It opened the tab unconditionally. The crashed server left a saved dashboard
   pointing at nothing, with the failure only visible in a shell tab nobody had
   a reason to look at. The launch now polls the existing webview probe until
   the URL answers, and on timeout reports the error naming the shell tab
   instead of persisting a dead dashboard.

3. The saved tab was untrusted, so the frame was sandboxed without
   `allow-same-origin` and the dashboard was broken twice over: the dsh
   client-runtime reads `localStorage` while loading its plugins and died there
   ("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
   opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
   every `/api` call no matter which authority `--trusted-host` named. Passing
   `location.host` only means anything once the frame actually carries that
   origin, so `--trusted-host` had never once done its job. The managed tab is
   now created `trusted: true`.

   That trade is real and deliberate: a trusted proxied frame is same-origin
   with Codeman and can reach Codeman's API. It is defensible only because this
   dashboard is an agent harness Codeman just started itself, on loopback, which
   can already run code as the user. It is not a precedent for trusting
   third-party dashboards, which is why it is set at this one call site rather
   than defaulted.

Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.

`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.

Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
2026-08-25 02:39:57 +02:00
Codeman maintainer 14de2b7012 fix(tabs): keep the created stamp reachable on a tight rail, and do not skip the first render
Two review nits on the vertical rail's detailed rows.

1. The tab-rail-tight rule (below 288px) hides `.tab-meta-created`, and its
   comment claimed the value "survives in the row's title attribute either way".
   It did not: the only title carrying it lived ON that element, and a
   `display: none` element has no hover target, so the created stamp was not
   shrunk but gone with no way to ask for it. Rather than just correcting the
   comment, `_sidebarRichMetaHTML()` now puts BOTH absolute stamps on the
   `.tab-meta` line itself, so the pill and the gaps around the stamps remain as
   hover targets. An item's own title still wins where the item is visible.

2. applyTabOrientation() decided whether applyTabWrapSettings() had already
   re-rendered by comparing `_tallTabsEnabled` before and after. That reads an
   UNDEFINED previous value as "it rendered", but applyTabWrapSettings()
   deliberately renders nothing on its first call ever (it only establishes the
   baseline: `prevTallTabs !== undefined && prevTallTabs !== showFolder`). So on
   a first call that also flips the folder row, neither function rendered and the
   rows stayed stale. Reachable when the pre-paint script throws and leaves the
   layout attributes on their catch-branch fallbacks for applyTabOrientation() to
   correct. The guard now mirrors applyTabWrapSettings()'s own condition.

Both new tests were run against the unfixed code first and fail there, which is
the only thing that makes them regression tests. (The third, "does not render
twice", passes either way by design: it pins that fix 2 did not introduce a
double rebuild.)

Verified in a real browser against a live server with two sessions, driving the
narrowing through _setTabRailWidth() the way the resize drag does: at the 320
default the row reads "CREATED 2m ago · IDLE <1m" with the created element
displayed; at 256 the tight class is on, the created element computes to
display:none, the visible text drops to "IDLE <1m", and the meta line's title
still reads "First created: ...". At 220 the compact threshold drops rich rows
entirely. Screenshots confirm no truncation artifacts in either state.

Full gate green (6104 passed), typecheck, lint, format, frontend-syntax and
public-assets all clean.
2026-08-24 22:58:28 +02:00
Codeman maintainer cdceede33d fix(deepseek): atomic shim write, honest attribution comment, name-fallback profile classifier
The three smaller review nits, plus the first real test coverage for the status
shim (it had none: it is emitted as a STRING, so tsc never sees it).

1. The shim was written with a plain writeFileSync. The TUI can be exec'ing that
   exact path while an upgraded Codeman refreshes it, and a reader catching a
   half-written file gets a syntax error, exits non-zero, and is retried four
   times per state change for a file that will never parse. Now temp + rename
   (atomic within the directory), with the temp chmod'ed before the rename since
   writeFileSync's mode only applies on create, and removed if the write throws.
   SHIM_VERSION bumped to 2, because SHIM_SOURCE changed and an existing v1 shim
   would otherwise keep matching the embedded marker and never be refreshed.

2. The pane-id comment claimed the ambient env "cannot be spoofed by an argument
   the agent itself could influence". The agent runs IN that pane and can invoke
   the shim with CODEMAN_SESSION_ID unset and any argv it likes. It buys nothing
   it did not already have (the hook-secret file is readable from the same pane,
   so it can POST /api/hook-event directly), but the comment read like a security
   boundary. Rewritten to say what the preference actually buys: correct
   attribution when a TUI mangles or re-uses the pane argument. Accidents, not
   adversaries.

3. classifyProfile() folded the directory name into the same haystack as the
   bundles, but only the TUI arm could match a bare name, so a stock profile
   whose package.json has no dsh.profile.bundles (hand-edited, older layout,
   mid-install) classified as `unknown` -> launchable -> eligible as the DEFAULT
   pick, which is exactly the pane-dies-on-arrival failure the two-part
   availability gate exists to prevent. The stock names are now a LAST-resort
   fallback consulted after the bundle patterns, so real bundle evidence still
   wins over a name the user chose. The loose `tui` arm gained word boundaries:
   it decides which profile boots by default, and matching the middle of
   `intuition` is not a rule anyone could predict.

New test/deepseek-status-shim.test.ts runs the generated script the way the
harness does -- real node process, real argv, real env, real listener -- and
covers the exit-code contract that makes the retry behaviour safe: mapped states
post and exit 0, an unknown verb or unmapped state exits 0 WITHOUT posting (a
non-zero there would be four HTTP requests per state change forever), a rejecting
server or an unreachable one exits non-zero so the caller retries, the hook secret
is read at execution time, and `node --check` parses the file (a template-literal
typo in SHIM_SOURCE is invisible to tsc).

Trap worth recording, hit while writing it: the tests must spawn the shim
ASYNCHRONOUSLY. The listener lives in the test process, so spawnSync blocks the
event loop that has to accept the connection, the shim waits out its own 1500ms
socket timeout and exits 1, and it reads exactly like a broken shim (measured:
Socket._onTimeout in its --trace-exit output, server logging nothing).

Verified: full gate green (6142 passed, +10), typecheck/lint/format clean.
2026-08-24 18:00:52 +02:00
Codeman maintainer 2034719d61 fix(deepseek): close the env-var clamp hole, bound the profile install, make the hook gate per-session
Three review findings on the DeepSeek Harness mode, plus one the third exposed.

1. The multi-user clamp was bypassable by a sibling field on the same request.
   clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but
   DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER
   _configureDeepSeek(), so a non-granted owner sending
   envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated
   instance: a session created with permissionMode "read-only" and that override
   ran with DSH_PERMISSION_MODE=danger-full-access in its pane.

   Every other CLI's bypass is a command-line flag reachable only through the
   per-CLI config, which is why the config clamp alone is the whole gate for
   them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted
   owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to
   what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on
   that list because it aims the launcher at a profile tree whose plugin code
   runs at boot, before any approval row can apply. Verified end to end in real
   multi-user mode: a non-granted user sending both now gets workspace-write and
   no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched.

2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout`
   signals only the direct child, and a plugin install fans out into
   package-manager children that keep the inherited stdio pipes open, so `close`
   never fires and the held-open request leaks with no route-level deadline.
   Reproduced: with a 1.5s built-in timeout the promise was still unsettled after
   6s and both fan-out children were alive. Now detached: true plus negative-pid
   SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a
   last-resort reap for a grandchild that escaped the group. Same probe after the
   change: close fires, direct child and both grandchildren dead.

3. hooksAvailableForMode() promised more than a dsh session can deliver.
   deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that
   triple is the only reason a dsh session posts hook events, so `until=stop` was
   accepted and then blocked for the caller's whole timeout: the exact
   infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now
   takes HookCapabilityOptions and every call site passes sessionHookOptions(),
   with the deepseek arm reading `!== false` so a forgotten one degrades to the
   old behaviour. The refusal names the setting rather than saying "no Claude
   Code hooks", which would send the caller hunting a bug that is really a
   setting they chose. Profile conformance stays unknowable at request time and
   is documented as such. The stale "True for `claude` and nothing else" docblock
   is corrected.

4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a
   claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent
   capture read Claude's own transcript, and adding deepseek silently widened
   both to a mode that has none. They compare mode === 'claude' directly now, and
   a static check pins them there.

Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the
wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 --
bridge off plus explicit until=stop is a 400 naming the setting, bridge off with
no `until` still 200s on idle/exit, bridge on accepts stop.
2026-08-24 16:01:02 +02:00
Codeman maintainer b330f1d9e8 feat(tabs): give the vertical rail the home screen's per-session detail
The vertical tab rail (tabOrientation 'vertical') listed names and nothing
else, while the rich sidebar and both home screens already answered the
question a docked column exists to answer: which of these sessions wants me
next, and how long has it been like that. The rail is a docked column too, so
it now draws the same row.

- New per-device setting tabRailDetail ('rich' | 'simple', default rich),
  App Settings -> Appearance -> Tabs, in SettingsUpdateSchema + displayKeys and
  stamped as data-tab-rail-detail by the pre-paint script, so a detailed rail
  does not flash through simple rows on every load.
- ONE gate for both vertical surfaces: isRichTabRows() =
  isSessionSidebarRich() || isTabRailRich(). The row model, the markup and the
  20s in-place clock are the existing rich-sidebar ones, classified by
  _mobileOverviewState/_mobileOverviewSince, so the rail, the sidebar, the
  desktop home rail and the phone overview cannot disagree about what
  "working" means or which stamp measures it.
- Detail rides on its OWN attribute, exactly as the sidebar's does, so every
  existing [data-tab-orientation='vertical'] rule keeps matching both variants
  untouched. A flip of detail ALONE still forces a full render (the stamps line
  is emitted by the row template, not toggled by CSS) and re-runs
  applyTabWrapSettings(), which owns the folder line and is now rail-aware.
- CSS: every rich paint rule gains a rail twin as a COMMA-GROUPED selector,
  never :is() - an :is() list takes its most specific argument, which would
  lift the sidebar arm from (0,3,1) to the rail's (0,5,1) and let these rules
  outrank things they never used to.
- Width is why there are thresholds. At 256px the stamps line ellipsizes
  mid-word, the same reason the rich sidebar is 300px, so a rail that has never
  been sized defaults to 320 (RICH_DEFAULT_WIDTH, the existing Wide preset,
  which also keeps the settings select on a named choice). A width the user has
  chosen is never overridden: below 288px the created stamp is dropped rather
  than truncated (tab-rail-tight, CSS only) and below 240px the rows go back to
  simple (tab-rail-compact, which re-renders).
- The rich clock is armed and disarmed by applyTabOrientation() as well as
  applySessionListLayout(); a leaked interval would rewrite stamps in a list
  that no longer has any.

Also fixes a data-loss bug in the inline tab rename that predates the rail and
reproduces in every layout, header strip included: Escape set the input to ''
and blurred it, and the blur handler commits - so cancelling a rename PUT an
empty name, and the tab fell back to its folder label (measured against a live
server: ["rail-alpha","","rail-gamma"]). Escape now calls cancelRename(), which
invalidates the edit so the blur that follows the input's removal is a no-op.

Tests: rail-detail gate, the three ways it turns back off (simple, compact,
horizontal), sidebar-wins, render-on-detail-flip and the plumbing/CSS guards in
test/session-list-layout.test.ts; the rename cancel in test/inline-rename.test.ts
(browser suite), pinned by running it against the old code first. Verified live
against a real server on an isolated instance: detailed/simple/compact/header/
sidebar variants, click-select, the ... menu, inline rename, Alt+N, the in-place
stamp tick and a full settings-picker round-trip including reload.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 04:45:02 +02:00
Codeman maintainer 4cda150493 feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.

DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:

1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
   $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
   `base` -- the interactive terminal front door is always a third-party
   plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
   and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
   button gates on the latter, because reporting only the binary would spawn a
   pane that dies on arrival. When the binary is present but no profile is,
   the run menu offers to install one (POST /api/deepseek/install-profile).

2. The permission switch is an env var, not a flag. The harness has no
   command-line permission option; its sandbox/approval rows read
   DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
   Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
   own workspace-write, which still asks, so the multi-user clamp is the
   only-if-sent branch and clamps to workspace-write, never read-only.

3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
   earned that. The terminal front door reports idle/working/blocked to a
   supervising process over a generic env-gated contract; a generated shim
   (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
   report to /api/hook-event as stop / agent_working / permission_prompt. So a
   dsh session gets definitive respawn triggers, real wait-endpoint signals and
   real Approvals Inbox items instead of output-stabilization guesswork.
   `agent_working` is new (157th SSE constant) and joins
   APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
   alert at once.

The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.

Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.

Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.

Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 03:37:56 +02:00
Codeman maintainer 9cfd8e8989 fix(docker): survive xAI installer's own /usr/local/bin/grok symlink
The agent-image grok step copied /root/.grok/bin/grok onto /usr/local/bin/grok
with cp -L. Newer versions of xAI's install.sh already create
/usr/local/bin/grok as a symlink to that same binary, so the copy failed with
'same file' and the --no-cache rebuild died at the grok layer (2026-08-24).
Stage the copy under a temp name, drop whatever the installer left at the
destination, then move into place - correct against both old and new
installers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 01:12:38 +02:00
Codeman maintainer 8fe393826b chore: version packages 2026-08-24 01:00:51 +02:00
Codeman maintainer 7a340fe7bc fix(tabs): review-driven hardening for the tab-layout foundation and vertical rail
Post-merge follow-ups from the deep review of #334 and #335, so they ship in
the same release as the features.

Tab-layout foundation (#335):
- PUT /api/session-order drops unknown/foreign ids again instead of 400ing
  the whole write, in both the owner and the admin path (single-user requests
  are the synthetic admin, so that path is the one the browser hits). The
  frontend debounces its reorder push and swallows errors, so a session
  deleted inside the debounce window silently cost the user the entire
  reorder - and the endpoint sits on the stable /api/v1 surface, where the
  pre-layout server merged leniently.
- A failed mux restore no longer locks explicit deletions into 500s for the
  process lifetime: runSessionDeletion and webviewDeleted degrade to
  best-effort without layout coordination, while the automated stale sweep
  (runStaleSessionCleanup) stays fail-closed.
- sse-events doc comment: no 'suppressed' hook event exists; hooks stay 8.
- registerSessionWithLayout resolves its owner through ownerLayoutKey()
  instead of a hardcoded '@single'.

Vertical rail (#334) - all rail-awareness gaps in sidebar-only predicates,
unified behind the new _isVerticalTabList() (sidebar OR rail):
- Drag-reorder read the insertion side from clientX in the rail, so
  before/after was effectively arbitrary on vertical rows; the drag-over
  indicators now draw as top/bottom edges there like the sidebar's.
- The active tab is scrolled into view in the rail (Alt+N/palette selection
  used to leave the row below the fold).
- Floating subagent/ultracode windows anchor to the RIGHT of rail tabs, and
  the connector redraw gates (render tail + strip scroll) cover the rail.
- Server-seeded tabOrientation is applied when the async settings load
  resolves, not only at boot, so a fresh device shows the rail immediately.
- The pre-paint script stamps data-tab-orientation and --tab-rail-width
  (sidebar-wins and solo carve-outs included), removing the flash of the
  header strip on every vertical-mode load.
- The session name font defaults to 12px, the sidebar's historical 0.75rem
  size, so installs that never touch the new slider are not restyled.

Also documents the rail in CLAUDE.md (second #sessionTabs host, mover
ordering, the axis-predicate rule) and gives tab-rail-resize.js its
@dependency/@loadorder header. Full gate green (6093 tests); the excluded
browser suite was run by hand - only the known environmental failures
(opencode/codex binaries) remain, identical to pristine master.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 00:59:48 +02:00
Codeman maintainer f3c615b669 fix(files): give the file preview a working detach button
The button next to the file preview's close icon was Copy Content, whose
overlapping-pages glyph reads as a pop-out control - and for a PDF or any
media/binary preview it was completely dead: those branches never fill
filePreviewContent, so the click hit an empty-content guard and did nothing,
with no feedback.

There is now a real detach button that opens the previewed file in a browser
tab (raw route for PDFs/images/media/text, the server-converted PDF preview
for docx/pptx), severs window.opener by hand so a blocked pop-up stays
detectable, closes the overlay on success (which also stops any playing
media), and disarms on close so it can never open a stale file. The copy
button now toasts 'Nothing to copy in this preview' instead of staying
silent.

Verified live with Playwright against an isolated instance: button visible
and armed on a PDF preview, file-raw answers 200, clicking opens the URL and
tears the overlay down, text previews keep a working copy buffer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 00:59:27 +02:00
Codeman maintainer c173ae0264 Merge remote-tracking branch 'origin/master' into worktree-grok-mode
# Conflicts:
#	src/web/public/app.js
2026-08-24 00:32:25 +02:00