Commit Graph
261 Commits
Author SHA1 Message Date
Codeman maintainer 1105d646f3 fix(terminal): #555 landing fixes
Comment and doc corrections that #555 made stale, no behaviour change.

- stock.ts: the grok and omp altScreen comments compared their strip to
  opencode's, which is now strip-mux-and-mouse rather than the narrow
  strip. Grok now says it shares antigravity's strip until measured, and
  omp drops opencode from its comparison.
- terminal-ui.js: the touch-tap comment named Claude/Codex/Gemini as the
  stripped modes, but the gate is now the cliMouseTracking flag alone and
  covers opencode too, so it names the two stripping flavours instead.
- src/types/session.ts: the cliMouseTracking JSDoc (the flag the browser
  now gates on exclusively) listed only claude/codex/gemini; it now names
  the strip-full and strip-mux-and-mouse modes, including opencode under
  tmux.
- src/session.ts: the usesMux getter doc now names isMuxMouseStripMode,
  since the replay strip passes usesMux to it as well.
- docs/architecture-invariants.md: the narrow-strip list gains
  grok/deepseek/omp (matching the PR's own CLAUDE.md line), the
  "must REMEMBER" heading covers both DECSET-stripping flavours, and the
  cliMouseTracking writer is described as the full-or-mouse branch it
  really is.
- docs/wiki/The-Dashboard.md: the user manual said every non-Claude CLI
  scrolls locally; opencode's wheel and swipes now page its conversation.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 05:30:20 +02:00
Codeman maintainer ff2f81541a Merge #555: opencode drags select text and the wheel pages its transcript 2026-10-09 05:28:51 +02:00
Codeman maintainer 7d8c188f83 fix(gemini): read Gemini CLI's composer bar and spinner line so a turn can end
A gemini session stayed "working" for good after its first turn. Its braille
spinner trips the generic SPINNER_PATTERN and marks the pane working, but only
a composer glyph arms the idle confirmation and gemini declared none, so it
fell back to Claude's `❯`, which gemini never draws.

Measured on live Gemini CLI 0.63.0 panes (capture-pane every 300 ms through
real turns with a shell call at 40, 120 and 200 columns, YOLO and default
approval mode, plus the raw PTY stream). The turns ran against a local
stand-in for the Gemini API (GOOGLE_GEMINI_BASE_URL, which Codeman's custom
endpoint support already sets), since the CLI's TUI does not depend on the
backend and no account is needed for it:

- The TUI repaints its whole bottom region every frame, composer included,
  and the composer sits between a `▄` bar and a `▀` bar; the submitted prompt
  is echoed between the same bars. The `▀` bar arms the idle check: every
  repaint carries it, tmux's reattach repaint too. The composer's prompt
  character is no good: it follows the approval mode (`*` in YOLO), and its
  `>` also starts the echoed prompt, which would make the submit verifier
  read a submitted prompt as stranded and press Enter again.
- While a turn runs a line `⠦ Thinking... (esc to cancel, 6s)` animates about
  every 80 ms (largest gap mid-turn: 214 ms). The label can be any loading
  phrase, so the working line is the `(esc to cancel, <n>` suffix, or a
  spinner frame opening a line for when a long phrase wraps that suffix.
  Nothing at rest matches either.
- A tool confirmation (default mode) replaces the composer, stops the
  spinner and the pane goes silent (3.9 s gap), so it reads as idle.

Verified on an isolated instance from this branch: a YOLO turn emitted one
session:working (+170 ms) and one session:idle (2.5 s after the last output);
a default-mode turn went working -> idle while the confirmation waited ->
working once allowed -> idle at the end; after a server restart four restored
gemini panes went busy -> idle in about 4 s; a fresh launch settled in 3 s.

The launch-settle and uncharacterised-CLI tests that used gemini as their
example of a CLI without work detection now use grok and deepseek.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 03:37:41 +02:00
Codeman maintainer 112c533ac7 fix(omp): read omp's status bar so an omp turn can end
An omp session stayed "working" for good once a turn started, the same
latch pi had: omp's braille spinner trips the SPINNER_PATTERN fast path,
and only a composer glyph arms the idle confirmation. omp declared none,
so it fell back to Claude's `❯`, which omp never draws once its setup
wizard is done.

Measured on live omp 18.8.6 and 18.0.11 panes, holding a turn open
against a local endpoint that never answers: the input row is `╰─ <text>`
and is redrawn at submit, at the end of a turn, at launch and on
reattach. While a turn runs, the status bar's leading `π` becomes a
braille spinner plus the elapsed time (` ⠼ 14s > ⬢ model > ...`; 18.0.11
pads it with two spaces, past a minute it reads `1m`), with a
`⎋ Working…` row above it. The registry entry now names the input row as
the glyph and either working signal as the working line.

The glyph also switches the submit verifier on for omp, which reads the
input row the way it reads Claude's composer. A prompt sent mid-turn goes
to omp's Steering queue and clears the row, so the verifier stands down.
Text left in the row after an Enter is the one case it re-presses.

Verified end to end on a sandboxed instance (own HOME and PATH, omp
18.8.6): session:idle at launch, session:working during a turn,
session:idle about 3 s after it ended, and a restored pane settled idle
about 3 s after a server restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 02:43:22 +02:00
Codeman maintainer e7d661158b fix(opencode): read opencode's composer bar and footer spinner so a turn can end
An opencode session that ran a tool stayed "working" for good. The running
tool row draws a braille spinner (`⠋ Sleep for 12 seconds...`), which trips
the generic SPINNER_PATTERN and marks the pane working, but only a composer
glyph arms the idle confirmation and opencode declared none, so it fell back
to Claude's `❯`, which opencode never draws. Measured on an isolated
instance: a 16 s turn latched busy/isWorking for the rest of the session.
A text-only turn had the opposite problem and never showed as working.

Measured on live opencode 1.3.0 panes (capture-pane every 250-300 ms through
real turns at 40, 60, 120 and 200 columns, plus the raw PTY stream):

- Every composer row starts with a `┃` bar, and the submitted prompt lands in
  the transcript with the same bar, so a turn's first repaint arms the idle
  check, and tmux's reattach repaint does the same for a restored pane.
- While a turn runs the footer row starts with an 8-cell knight-rider
  spinner, `⬝■■■■■■⬝  esc interrupt`, redrawn about every 40 ms (largest
  gap mid-turn: 121 ms). At rest the TUI is silent and nothing on screen
  draws a `⬝`/`■` run, the 200-column sidebar included.
- The working line is the spinner run, `[⬝■]{8}`, not the label: tmux ships
  `esc` and `interrupt` as separate words joined by cursor moves, so the
  label never reaches the stream detector, and at 40 columns the footer
  wraps it to `esc` / `interr` / `upt`. All 344 spinner chunks of a turn
  match the run after Codeman's ANSI strip.
- A pending permission prompt replaces the composer and stops the spinner,
  so it reads as idle (waiting on the user).
- The last `┃` row on screen is the composer's agent/model row, or the
  permission box's closing bar, never the prompt text, so the submit
  verifier stands down and can never press Enter into a dialog.

Verified on an isolated instance from this branch: a 15 s tool turn emitted
exactly one session:working (+271 ms) and one session:idle (3 s after the
spinner stopped); a permission prompt read idle and the allowed turn went
working -> idle; after a server restart both restored opencode panes (one
at rest, one on a permission prompt) went busy -> idle in about 4 s; a
fresh launch reached an open page as idle in 3 s.

The launch-settle tests that used opencode as their example of a CLI
without work detection now use gemini and antigravity.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 02:40:04 +02:00
Codeman maintainer db168cdbe5 fix(session): never settle a just-prompted pane idle at launch
40560ace made the 3 s launch settle announce its idle. For a CLI with
work detection that edge could now land in the middle of a turn: a
prompt sent ~2.9 s after launch has not been marked working yet (the
working line goes through the deferred parsers), so the settle called
the pane idle and a send-and-wait registered for that prompt resolved
before the turn even started. Before 40560ace the settle was silent, so
this edge is new.

The settle now also leaves a pane to `_confirmIdle()` when a prompt was
submitted since the timer was armed, the same way it already does for a
pane marked working; that confirmation reads the screen before it ends
the turn. A CLI without work detection still settles: nothing else ever
would.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 02:18:26 +02:00
Codeman maintainer b539780f33 fix(pi): read pi's composer rule so a pi turn can end
A pi session stayed "working" for good once a turn started. pi's braille
spinner trips the SPINNER_PATTERN fast path, which marks the pane working,
but only a composer glyph arms the idle confirmation and pi declared none,
so it fell back to Claude's `❯`, which pi never draws. Measured on beta136:
an errored turn stayed busy/isWorking for 3+ minutes after pi was back at
rest.

pi has no composer glyph. Measured on a live pi 1.1.0 pane (capture-pane
every 250 ms through a turn): its composer sits between two `─` rules, and
while a turn runs it rewrites the top rule as `── ⠏ Working ───` on every
frame. The registry entry now names the rule as the glyph that arms the
check and a spinner frame inside it as the working line.

The same glyph settles a reattached pi pane (a restored pane gets no launch
timer): tmux's reattach repaint carries `─`, which arms the confirmation.
The submit verifier reads the last rule, finds no prompt text and stands
down, so it can never press Enter on a pi pane.

Verified on an isolated instance from this branch: a real pi turn emitted
session:working then session:idle, and after a server restart the restored
pi pane went busy -> idle in about 3 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 01:49:19 +02:00
Codeman maintainer 40560aced1 fix(session): announce when a fresh or re-attached agent pane goes idle
A fresh codex, pi or opencode tile could spin "working" forever while the
pane sat at its composer. The external-CLI launch timer set the status to
idle WITHOUT an event, only `needsRefresh`. When the launch paint never
marked the pane working, the later idle confirmation found the status
already idle and announced nothing either, so every open browser kept the
`busy` from the spawn broadcast. A reload fixed it, which is why only open
pages were stuck. Measured on the 1.36.0 beta: w8 (codex) had
`lastPromptTime: 0`, i.e. no idle edge ever, and a page held a fresh pi
session at `busy` while GET /api/sessions said `idle`. pi and opencode hit
it on every launch (no work detection, so the timer is all they have);
codex only when its launch paint lost the race against the timer.

- `_concludeIdle()` is now the one place a pane is concluded idle: status,
  working flag and prompt stamp change together and `idle` is emitted
  (session:idle + state broadcast). `_confirmIdle()` uses it too.
- The launch settle (`_settlePaneStartup`) concludes a pane still in its
  spawn-time `busy`. A pane already working is left to `_confirmIdle()`
  only when its CLI declares `capabilities.workDetect`; for the rest the
  timer is the only thing that can settle it, so it also clears a working
  flag a launch spinner glyph latched.
- `_armPaneSettle()` also arms it for a RESTORED pane (Codeman restart,
  auto-reattach, tile Attach) of a CLI without work detection (at this
  commit shell, opencode, gemini, antigravity, pi, grok, deepseek and omp),
  which used to stay `busy` server-side with nothing to clear it. Restored claude and
  codex panes stay on their composer glyph, so a restart mid-turn is never
  called idle. No `needsRefresh` there: an attach refetches by itself.

Pre-existing since the OpenCode integration, not a regression of this
release. Restore-path gap reported by the opencode and pi sessions working
the same symptom.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 01:49:04 +02:00
RandalixandClaude Opus 5.5 5ba729fcbb fix(terminal): strip opencode's mouse DECSETs so a drag selects text again
opencode's TUI enables mouse tracking. tmux runs with `mouse off`, so it passes
the PANE's DECSETs straight through to the tmux client, and the browser's xterm
obeyed them: `mouseTrackingMode` flipped to 'any' (measured 62 none / 18 any over
16s) and xterm then reported DRAGS to the TUI instead of selecting locally.

In that state marking text produced no selection at all, so copy-on-select
silently did nothing (5/5 dead drags while `any`), and the obvious fallback —
Ctrl+C — is opencode's `app_exit`, which ended the session. Both were hit here.

opencode needs the middle strip: alt-screen toggles AND mouse DECSETs, but NOT
`3J` (a TUI is not a `clear` consumer). That is `altScreen: 'strip-mux-and-mouse'`
+ `isMuxMouseStripMode`, applied to the live stream (session.ts) and the replay
of a stored buffer, now the exported `stripReplayBuffer()` (session-routes.ts).

The browser's mouse-report gate keeps no mode list any more:
`_shouldReportMouseToCli()` reads only `cliMouseTracking`. The server sets that
flag solely in the mouse-strip branch (`_recordStrippedMouseMode`, one caller),
so it can only be true for a mode whose DECSETs are stripped, and whichever modes
the registry strips, the browser follows. Clicks still reach opencode through the
hand-encoded SGR tap it gates.

The `altScreen` JSDoc gets the decision table its three independent choices need
(alt-screen / `3J` / mouse DECSETs), written from the predicates, including that
`preserve` and `strip-mux-only` take the same runtime row. The table is pinned for
every stock CLI, with and without tmux, on both the live strip and the replay
strip, plus the published flag (test/claude-scrollback-strip.test.ts), so the two
halves cannot drift and a mis-ordered replay branch fails.

Docs and comments that still said opencode keeps its mouse reporting or gets the
narrow strip are updated (CLAUDE.md, architecture-invariants, scrollback and
copy-shortcut plans, session.ts, terminal-ui.js).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-08 14:55:02 +02:00
Codeman maintainer 9209922ea4 fix(deepseek): reject the dsh footer's non-model words instead of requiring a digit
The digit rule from 21ae48a5 hid the official DeepSeek ids (`deepseek-chat`,
`deepseek-reasoner` carry no digit), so a session on the official route with
the model field on showed the logo alone (its bundle row pins a provider
alone, so the config had nothing either). It also still misread a folder name
with a digit when every field before it was off.

Now the captured field is rejected when it is what the field can be when it is
NOT the model, and read otherwise:

- capabilities.modelDetect.rejectWords (registry data, single tokens, compared
  ignoring case; the schema bounds them and requires a screenLine). dsh lists
  every effort id its adapters offer (pi-ai THINKING_LEVELS plus the DeepSeek
  adapter's off/low/high/max) and the shipped mode ids, from dsh 0.1.1-rc.2 /
  dsh-TUI 0.10.0-beta.1. A mode's drawn label (`plan mode`, `full access`,
  CJK) can never be one captured field.
- In the shared screen reader, for every CLI: a field equal to the session's
  own working-directory basename is the folder, never the model.

Fixtures: `deepseek-chat` and `deepseek-reasoner` with the model field on are
read; every effort id, `default`, `plan mode`, and the folder name first (with
and without a digit) are not; the live qwen footer still reads `qwen3.8-27b`.
Known gaps, all off by default, are named in stock.ts: a custom mode id drawn
raw, a git branch or a one-word session title first, and the non-compact
footer layout (nothing read there; the route config applies).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 21:16:53 +02:00
Codeman maintainer 661fe3dc13 feat(sessions): a dsh session shows its route config's model while its screen names none
displayModel gains a `config` source, ranked below any report from the running
CLI and above the launch model: custom endpoint, then statusline or screen,
then config, then launch, then nothing. The screen still wins whenever it
names a model, since that is what the running TUI uses.

- Registry data: capabilities.modelDetect gains `configResolver`, a NAMED
  reader (src/model-config-resolvers.ts), like a launcher profile; dsh names
  'deepseek-route' (the reader from the previous commit). `screenLine` becomes
  optional; the schema refuses a modelDetect naming nothing, an unknown
  reader, or screenLines without a screenLine.
- Session: the reader runs from _withPaneLifecycle's finally, so at every pane
  start, attach and relaunch, with the session's own launch config
  (legacyConfigForMode) and env (its clamped overrides, then the server's), so
  a per-session DSH_HOME is the home read. Async; a read that lands after a
  newer one or after the session stopped is dropped; a remote or docker
  session reads nothing locally. A change emits displayModelChanged
  (broadcast and persist). Not restored after a restart: the next attach
  reads it again, and a restored screen value outranks it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 19:41:56 +02:00
Codeman maintainer 8392854619 feat(sessions): publish the model each session runs (displayModel)
A session header can only name the model a session runs if the server knows
it, so SessionState gains `displayModel: { model, source }`, resolved in a pure
module (src/session-display-model.ts), strongest first:

- custom-endpoint: a Custom Model Endpoint Profile's modelId answers the
  session, whatever alias the CLI prints;
- statusline / screen: the newest report from the running CLI itself.
  Claude's statusLine exporter already posts model.display_name on every
  render; the status-telemetry route now records it (only for a CLI with
  capabilities.statusLineTelemetry). A CLI whose registry entry declares the
  new capabilities.modelDetect has its footer read off the pane capture the
  idle/working probe already takes (no extra tmux call), so an in-session
  /model switch is followed at the next transition;
- launch: the model the session was launched with (claude's --model or the
  app-wide default, another CLI's <cli>Config.model), read where the registry
  says the model param lives;
- nothing known: no field, never a placeholder.

modelDetect is registry data, measured on live panes: dsh-TUI's status line
on the row under its composer (qwen3.8-27b on the owner's route) and codex's
`<model> <effort> ·` footer on its last row. Both anchor on chrome only that
CLI draws, over the last rows of the screen only; a transcript line shaped like
the footer is never taken (fixture tests). The pattern goes through
compileVersionRegex() with exactly one capture group, checked at load time.

An unreadable or covered footer keeps the last model (unlike the watching
label: a model does not stop running when something covers its row). Model
text is untrusted: escape sequences and control characters are stripped and it
is capped at 64 characters. A change emits displayModelChanged, broadcast
(session:updated) and persisted; a restart restores a CLI-reported model until
the next report.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 13:31:59 +02:00
Codeman maintainer 06c4c7da16 fix(sessions): scope the launch model to claude and pin it with the advisor (#514, #515, #530 landing)
Maintainer merge-time fixes for the three PRs that landed together on the
session create / launch / persistence path.

#514 findings (bot verdict merge-with-fixes):
- minor, fixed: SessionState.model was published and persisted for every
  mode, so a codex/opencode cron session reported the app-wide Claude
  default it never ran on. toState() now emits it only where the new
  cliTakesSessionModel() holds (registry capability model.source ===
  'claude-settings-file', no CLI id branch). POST /api/sessions uses the
  same helper for its non-claude refusal, so refusal and publication cannot
  drift. Recovery then hands back undefined for other modes on its own.
- nit, fixed: the `model` schema admitted a leading dash (and '.', '[').
  The first character must now be a letter or digit; still a subset of the
  registry's model-claude pattern, so nothing accepted is refused at launch.
- nit, fixed (reject, the consistent choice): `model` with
  attachRemoteSession was silently dropped. Now a 400 INVALID_INPUT, as
  #514 does for non-claude CLIs and quick-start does for remote cases.
  advisorModel (#530) gets the same refusal there. effort and envOverrides
  keep their older silent ignore on that branch so no existing caller breaks.

#515 finding (bot verdict merge, one nit):
- nit, fixed: the types/session.ts @fileoverview described CodexConfig as
  (model, resumeSessionId); it now lists reasoningEffort, bypass,
  animations and renderMode too.

Audit of the merged combination (not reviewed before):
- The conflict resolutions in session.ts (toState), types/session.ts,
  reboot-restore-routes.ts, server.ts (restoreMuxSessions), CLAUDE.md and
  skills/codeman/reference/endpoints.md (+ plugin mirror) keep both sides
  correctly; nothing was lost or doubled.
- A claude session with both `model` and `advisorModel` launches with
  `--model <id>` and ONE merged `--settings` JSON (ultracode + advisorModel,
  or advisorModel beside `--effort <level>`), on the tmux template
  (including the resume || new variant and with the statusLine exporter)
  and on the direct-PTY fallback. Both values (and effort) survive
  restoreMuxSessions onto a dead pane, a reboot restore into a fresh pane,
  and restartCli/dead-pane respawn via _buildRespawnPaneOptions.
- quick-start and ralph-loop take no per-session `model` (matching #514's
  scope, POST /api/sessions only) and launch on the app-wide default, which
  toState now persists for claude, so recovery stays consistent.
- No defect found in the combination beyond the findings above. Noted, not
  changed: advisorModel is still published for any mode a caller sends it
  with (launch-inert there; the UI and skill send it for claude only).

Tests: test/session-model-recovery.test.ts pins the pair through both
recovery shapes for effort ultracode/high/none, the recovery constructors'
fields, the tmux-manager builder hop, and the codex/opencode/shell
non-publication; test/advisor-model.test.ts pins the launch lines and a
real direct-PTY Session's pty.spawn argv; the route test covers flag-shaped
models, attach refusals and the published fields. Docs: SessionState.model
docstring, the reboot-restore-registry header, the golden test comment and
the CLAUDE.md model/advisor bullets.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-04 23:52:41 +02:00
Codeman maintainer b8038a592c Merge pull request #530 from Ark0N/feat/advisor-model
feat: Claude advisor tool support (per session, App Settings default, skill workers)

# Conflicts:
#	CLAUDE.md
#	plugins/codeman/skills/codeman/reference/endpoints.md
#	skills/codeman/reference/endpoints.md
#	src/session.ts
#	src/types/session.ts
#	src/web/routes/reboot-restore-routes.ts
#	src/web/server.ts
2026-10-04 23:25:01 +02:00
Codeman maintainer 01f403dc1b feat(claude): advisor tool support, per session and as an App Settings default
Claude Code's advisor tool (code.claude.com/docs/en/advisor) lets the session's
main model consult a second, stronger model at decision points: before
committing to an approach, on a recurring error, and before declaring a task
done. Codeman can now start claude sessions with one.

- `advisorModel` field on POST /api/sessions, /api/quick-start and
  /api/ralph-loop/start (fable, opus, sonnet or a full model id in those
  families; haiku cannot advise and is refused). Stored on the session and
  persisted, so respawn, boot restore and reboot restore keep it. Remote and
  docker quick-starts refuse it, as they refuse effort.
- App Settings, Models, "Advisor" segment (Default / Sonnet / Opus / Fable),
  synced as `claudeAdvisorModel`. Run, resume and the Ralph wizard send it.
  Default sends nothing, leaving the CLI's own /advisor choice in charge.
- Carried as the `advisorModel` key in the launch's single --settings JSON,
  merged with ultracode and the statusLine exporter, never the --advisor
  flag: `claude --advisor haiku` exits 1 at launch, which would leave a dead
  pane on every respawn, while the settings key degrades to no advisor. A
  launch without an advisor is byte-identical to before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-04 19:05:43 +02:00
Michael GrundbergandClaude Opus 5.5 3df113fc54 fix(sessions): keep a session's model through recovery, refuse it off claude
SessionState now carries the model a session launched with, and both
recovery constructors (mux recovery and reboot restore) pass it back, so a
recovered session relaunches on the same --model rather than the account
default. A top-level `model` sent with any other CLI is refused, since
those take their model in their own config object, and an empty string
means no per-session model, as it does for modelOverride.

CLAUDE.md now describes both routes for a Claude model. The tests pin
which of `model` and `modelOverride` reaches the launch and which the
case file, and that a model opening with a dash renders as --model's value.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 17:24:12 +02:00
Codeman maintainer 2eece4f8f9 fix(cron): send a paste-mode prompt's Enter as its own write
A cron job in "Paste (direct)" input mode wrote `<text>\r` into the pane
in one piece. Claude Code (measured on 2.1.283) takes a burst of about a
hundred characters as a paste, so the `\r` landed as a newline and the
prompt sat unsent on the composer while the run reported `prompt_sent`.

Delivery now lives in `deliverCronPrompt()`. Paste mode writes the text
raw, waits CRON_PASTE_ENTER_DELAY_MS (300 ms), sends `\r` as a separate
write down the same PTY (so it cannot overtake the text), and arms the
session's composer check through the new public
`Session.verifySubmitted()`, which re-presses Enter while the prompt is
still visibly unsent. A session with nothing to write to now fails the
run instead of reporting the prompt as sent. Typed mode is unchanged.

Verified on an isolated instance: a paste-mode job with a 104-character
prompt submitted on the first Enter and Claude answered.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 17:10:12 +02:00
Codeman maintainer 7659ca8b44 fix(session): keep a tab working while Claude waits for its own workers
When Claude hands work to an ultracode workflow or background agents, it
ends its own turn and closes it with `✻ Waiting for 1 dynamic workflow to
finish` instead of `✻ Brewed for 1m 18s`, then resumes by itself when the
workers report back. The pane sits quiet with the composer up, so the idle
probe called the session idle for the whole wait. At phone width the
workflow's progress row also drops its ticking timer, so nothing on screen
changes for minutes.

A new optional registry field, `capabilities.workDetect.awaitingLine`,
names that closing row, and `_probePaneWorking()` counts it as work.
Claude renders the row once from a snapshot and never redraws it, so the
same words stay on screen after the workers finish. `isAwaitingWorkers()`
therefore tests only the newest column-0 row directly above the composer,
never the whole pane and never the PTY stream; a follow-up turn always
puts rows of its own there. The column-0 anchor also keeps an agent from
holding its own tab busy by printing the sentence.

Verified against the live Mac mini pane that reported the bug (2.1.283),
and end to end on an isolated instance: an ultracode session running a
90 s workflow at 46 columns stayed busy through the wait and the
follow-up turn, then went idle 6 s after that turn closed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:18:27 +02:00
Michael GrundbergandClaude Opus 5.5 b80d47aff8 feat(session): close sessions whose agent exited cleanly (#486)
* fix(cleanup): keep .claude-images while a sibling session uses the same dir

cleanupSession() recursively removes {workingDir}/.claude-images. That
directory belongs to the working directory rather than to the session, and
several sessions routinely share one case directory, so closing one session
deleted the pasted images a live sibling still referred to.

The removal now runs only when no other live session has the same working
directory. A session that is itself being cleaned up does not count as live,
so two sessions of one case closed together still remove the dir.

Split out ahead of the exited-agent sweep for Ark0N/Codeman#446, which closes
sessions unattended and would otherwise make the loss routine.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(session): close sessions whose agent exited cleanly (#446)

Part 2 of Ark0N/Codeman#446. Part 1 records an exited agent as
SessionState.paneExit. A session whose agent the user ended with /exit is
now closed through cleanupSession(), the same path the X button takes, so
finished sessions stop piling up on the board. The lifecycle log records
the reason as "agent exited cleanly (status 0)", and the conversation stays
resumable from the Resume list.

shouldCloseCleanlyExitedSession() in the new pure module pane-exit-sweep.ts
holds the rule. It closes a session only when all of these hold:

- The exit status is an explicit numeric 0 with no signal. An absent status
  is how a SIGKILL presents on tmux 3.2a, so it counts as unknown and the
  row stays. A non-zero status or any signal also keeps the row, with the
  exit code on the tab.
- Two authoritative pane reads agreed on that exit.
  TmuxManager.getPaneExitReadCount() counts them, and a failed, empty or
  skipped read neither confirms nor resets the count.
- No start, attach or relaunch is running for the pane.
  Session.paneLifecycleInFlight covers _setupOrAttachMuxSession(), whose
  dead-pane branch revives an exited pane on purpose, and restartCli().

setPaneExit() already scopes paneExit to local mux-backed sessions, so
remote, docker and direct-PTY sessions are never closed.

planRebootRestore() now refuses a record whose persisted paneExit is a
clean exit. That covers an agent that exited just before a reboot, before
the sweep reached it. A crashed agent's record stays eligible, like its row.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(web): show "exited" on the phone overview and desktop home rail (#446)

Part 1 of Ark0N/Codeman#446 taught the tab strip and the rich rail rows
to say that a session's agent has exited. The phone overview and the
desktop home rail still said "idle", beside a green or pulsing dot.

_mobileOverviewExit() in mobile-overview.js is now the one rule for all
three surfaces, and _sidebarRichRow() uses it as well. It changes what a
row shows and leaves the row's state alone, because the state still picks
the section and the sort order. An exited row gets an "exited" pill, a
neutral dot and row accent, and a duration measured from when the server
first saw the pane dead. A pending permission prompt or question still
wins, as it does on the tab.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cleanup): close the gaps review found in the #446 sweep and image guard

Four fixes from a dual review of Ark0N/Codeman#446 part 2.

- The .claude-images guard compares canonical paths, so a sibling that
  reaches the same directory through a symlink keeps it. Its comment used to
  say that case only missed a deletion; it caused one.
- A detached session counts as a live sibling. DELETE ?killMux=false removes
  it from the server's map while its pane keeps running, so the guard now
  reads persisted records too, and exempts only sessions being killed rather
  than every session in cleaningUp.
- A session being closed refuses startInteractive() and startShell(). The
  /interactive route awaits listener setup before the start, and a start
  that raced the close could launch a CLI in a tmux session whose record was
  then deleted. A failed close clears the mark again.
- The clean-exit sweep tries each exit once, keyed by session id and the
  exit's at stamp, so a close that fails is not retried and logged every
  two seconds.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(session): keep a clean exit that lands within 10 s of a pane start (#446)

A CLI that prints a startup error ("not logged in", a bad profile, a config
error) and exits 0 used to lose its tab, and the error with it, about 4 s
after launch. The sweep now keeps any clean exit that lands within
CLEAN_EXIT_MIN_PANE_LIFETIME_MS (10 s) of the last start, attach or relaunch
finishing (Session.paneStartedAt, stamped when _withPaneLifecycle ends). The
row stays as "exited (0)" for the user to read and close.

Verified on an isolated instance: a shell that ran `exit 0` 2 s after start
kept its row, one that exited after 13 s was closed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 22:18:07 +02:00
69a71287e6 fix(sessions): stop pinning the w1-myapp placeholder as Claude's /resume title (#457)
* fix(sessions): stop pinning the w1-myapp placeholder as Claude's /resume title

Local claude spawns passed the tab name as `--name`. That flag is not only the
cross-session peer name: it is also the prompt-box label, the `/resume` picker
entry and the terminal title, and a pinned title stops Claude generating its own
(`customTitle ?? aiTitle`). So every conversation of a case was listed in
`/resume` as the same `w1-myapp`, and none of them got a generated title. On one
workspace, 34 of 34 conversations spawned with `--name` had no ai-title, while
every conversation spawned without it had one.

Only a name the user chose is pinned now: `Session.cliPinnedName` is the name
when `nameSource === 'manual'`, carried to the builders as a separate `cliName`
so the tab/mux name is untouched. Placeholder and auto names let Claude title
the conversation again.

A rename in Codeman also reaches `/resume`: the new name is appended to the
conversation's transcript as the `custom-title` row `/rename` writes (never
creating the file, never writing an empty title). For a pane spawned without
`--name` this holds immediately; a pane spawned with one re-appends its own
title each turn, so there the new name holds from the next spawn.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(sessions): skip no-op renames and docker sessions when syncing the /resume title

A same-name PUT (the Session Options field saves on blur and recomposes the
unchanged placeholder) no longer flips nameSource to manual or appends a
custom-title row, and docker sessions skip the host transcript scan since their
transcript lives in the container. The skill pages no longer use a w<N>- name
as the peer-name example, and the changeset notes the re-append caveat.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: record that nameSource decides --name and renames reach /resume

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: codeman-local <codeman@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-24 01:48:33 +02:00
Codeman maintainer 43d4be8eeb fix(session): merge-time fixes for the dead-pane resume pin (#467)
- test/setup.ts strips CLAUDE_CONFIG_DIR (pinned in test-env-isolation), so
  transcript-fixture tests such as session-custom-model-restart no longer go
  red on a machine that exports it for a separate Claude account (#255).
- The vanished-tmux-session branch of _setupOrAttachMuxSession() relaunches
  the CLI through createSession() just like a failed respawn, so it now takes
  the same resume pin. A genuinely new session is unaffected.
- After a dead-pane respawn of a fallback-chain CLI, _claudeSessionId names
  the conversation the walk actually pinned instead of the chain tail, which
  the walk may have passed over for lack of a transcript.
- _claudeConfigDir() trims the override like claudeProjectsDir() does.
- The remote-reattach test is labelled as documentation, since the pin
  builder's own remote guard would make it pass either way.
- CLAUDE.md: the create-path pin persists through toState() as
  resumeSessionId, and the end of the walk adds no pin rather than clearing
  the launch seed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:27 +02:00
Codeman maintainer 7a30a31430 fix(terminal): merge-time fixes for the silent-failure paths (#431)
- While another device holds the pane width (_paneWidthRefused), a resize
  now asks for the container's width without applying it locally
  (_geometryForResizeRequest: rows follow the container, columns stay at
  the PTY's). Fitting first re-wrapped the whole buffer to the container
  and back on every 30s mobile retry, and throttledResize ran the
  scrollback clear for a resize that brings no redraw. selectSession
  clears the flag, since it belongs to the previous pane. New unit tests
  run the real mixin against a fake terminal and fail without the fix.
- Session seeds _ptyCols/_ptyRows at spawn (_notePtySpawnGeometry), so a
  reattached pane reports its tmux window's real size through ptyGeometry.
- Session.resize's declined-branch comment names ptyGeometry, not the
  deleted ptyCols/ptyRows getters.
- Delete the dead terminalGeometryAgrees() and its window export.
- test/xterm-private-api.test.ts header: it pins the exact locked version,
  so any bump fails, not only a major.
- The main-terminal fit sweep also matches fitAddon?.fit?.(), and
  CLAUDE.md names the modules it actually covers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer 0462a5d5a0 fix(approvals): merge-time fixes for the watching badge (#473)
- session.ts: a pane capture that fails now CLEARS the watching label
  (and emits watchingChanged so pages drop the badge) instead of keeping
  the last one, so a failed capture degrades toward an alert rather than
  pre-acknowledging the next real idle prompt. Test updated; invariant
  noted in architecture-invariants.
- approvals-ui.js: the header bell counts only unacknowledged items
  (pendingApprovalsCount), matching codeman tui's pendingApprovalCount();
  pinned in watching-no-alert.test.ts.
- mobile-overview.js: move the orphaned "Pill copy per state" JSDoc back
  onto MOBILE_OVERVIEW_PILL_LABEL.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer de4b1db490 Merge pull request #431 from rounakdatta/feat/mobile-terminal-resilience
fix(terminal): four silent-failure paths — renderer freeze, replay race, reconnect gap, unbounded fetches
2026-09-23 11:32:14 +02:00
Codeman maintainer 8536aaef7b Merge pull request #473 from irisitymichaelgrundberg/feat/session-watching-badge
feat(approvals): let a session watching its own background work keep quiet (#468)

# Conflicts:
#	src/config/cli-registry/stock.ts
2026-09-23 11:32:11 +02:00
Codeman maintainer 6a01412af9 Merge pull request #466 from irisitymichaelgrundberg/feat/pane-exit-reporting
feat(tmux): report that a pane's agent has exited (#446, part 1)
2026-09-23 11:32:02 +02:00
Michael GrundbergandClaude Opus 5 9a48c43aa1 docs(watching): a restart is not a gap, and here is the measurement
Claimed after a manual test that a session comes back from a server restart
without its badge until it next produces output. Measured instead of assumed,
and it is wrong: a codex session with a background terminal still running had
its label back within about 20 seconds of the restart, with no input from
anyone. Reconciliation re-attaches the pane, the attach repaint carries the
composer glyph, the idle confirmation arms on it, and the probe re-reads the
label — the ordinary path, doing the ordinary thing.

What produced the false claim was a session whose monitor had simply expired
while it sat there. Its footer carries no chip, so `watching: null` was the
right answer and there was nothing missing to restore.

Recorded at the field and in the invariants, because the shape of this invites
exactly one wrong fix: a polling timer to keep a value fresh that the pane
already refreshes by itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 08:00:16 +02:00
Michael GrundbergandClaude Opus 5 b13672596f fix(watching): tell the page when the badge goes away
Reported from a manual test: a Codex session went on showing the watching badge
after its background terminal had finished. The server was right and the page
was stale — `Session.watching` changes while the session's status does not, and
nothing broadcast it.

The label is usually SET on the idle transition, which broadcasts anyway, so the
badge always appeared correctly. It CLEARS when the work ends, and a CLI can end
background work without taking a turn: codex repaints its background-terminal
row away and stays idle, so `_confirmIdle()` concludes without emitting `idle`
(that emit is guarded by `wasWorking || isInitialReady`) and no other event
fires. Every open page kept drawing a badge the server had already dropped.

`_readWatching()` now emits `watchingChanged` when, and only when, the label
really changes, and the wiring pushes the session state on it. No new SSE event:
the badge reads off the session payload every surface already has.

A/B measured on an isolated beta with the page loaded and then left untouched.
Without this commit the server dropped the label at t+50s and the page still
showed the badge at t+100s; with it, page and server cleared in the same
ten-second window.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 17:59:45 +02:00
Michael GrundbergandClaude Opus 5 05c788ce9d fix(watching): close the review findings on the label and its window
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief)
found the trust boundary weaker than the comments around it claimed. Eleven
findings, all applied.

The two blockers were both about who can write the row the label is read from.
Claude's window covered two rows, and the second one is the status line, whose
command a session running with permissions bypassed can write into its own
`.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of
its own and silence its own idle alert. The default window is one row now, which
is the footer and nothing else, and the constant says why. Separately, the label
reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML,
which is an injection sink for any config-supplied pattern whose capture group is
permissive; it goes through escapeHtml() like every other untrusted string in
that file.

The Codex entry could not be fixed the same way, and now says so. Its row is
third from the bottom only while a terminal runs; with none running that slot
holds the last row of the transcript, so matching the complete row (with the
`/stop to close` tail, window narrowed to three) raises the bar without closing
it. What contains it is `hooks: 'none'`: no hook event from a codex session
reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an
alert. The registry comment, `docs/cli-registry.md` and the test all state that
rather than claiming a guarantee the code does not have.

Also from the review: the TUI header badge no longer counts an acknowledged
item, which was the same gate the classifier fix already went through and was
wrong for human acknowledgement too; the TUI approval card reads the quiet
reason and drops to a new `info` tone instead of asking for a reply; the badge
carries an aria-label, because the phone it was built for has no hover target;
the schema refuses `watchingLines` without a `watchingLine`; and the pattern and
its window are resolved together rather than one memoized and one not.

Documentation moved with it. The mechanism now lives in
`docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer,
`docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped
buzzing, and both that page and the changeset name the limitation neither did
before: a question asked in plain prose is not a dialog, so it is silenced along
with the false alarms while background work runs.

Verified live again after the narrowing, on an isolated beta: a Claude session
reported `1 monitor` and took its idle prompt acknowledged, and a Codex session
reported `1 background terminal` against the full-row anchor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 17:11:33 +02:00
Michael GrundbergandClaude Opus 5 9c286eeddf fix(session): persist an exit retraction, and let tests reach the watcher
Ten findings from a two-model review of this branch. Both reviewers cleared the
detection logic itself; everything here is a gap around it.

A route that starts a command in a pane now PERSISTS as well as broadcasts.
`/interactive` and `/shell` did neither before, and the pane-exit watcher cannot
cover for them: its next tick finds `paneExit` already cleared in memory,
reports no change and writes nothing, so `state.json` kept saying the agent had
exited for as long as the session stayed quiet. Nothing reads that record for a
decision yet, which is exactly why it had to be fixed now — part 2 is designed
to read it. The `clearPaneExitForNewPane()` docstring claimed its callers
already persisted; that claim was false for these two, and now says what the
caller owes instead.

The watcher's four guards were unreachable by any test. `refreshPaneExits()`
opened with `if (IS_TEST_MODE) return;`, so the read gate, the in-flight
suppression, the generation counter and the empty-read rule could each be
deleted with the whole suite green. The tmux call moves into `readPaneRows()`,
which a test subclass overrides — the shape `runRemoteReconnectTick` already
uses in this file for the same reason — and the test-mode gate moves with it, so
what a test cannot do is spawn a process rather than exercise the bookkeeping.
Each of the four guards now has a test that fails when it is deleted.

The muted status dot turned out to be a specificity fight on three surfaces, not
two. `.tab-status.error` was not excluded, so a session whose agent exited and
whose PTY-exit breaker then tripped lost its red dot to the mute — the state the
browser answers with a "restart it?" confirm, and a needs-you colour by the same
argument that protects the two alert classes. And mobile.css gives a `busy` dot
a 9px size and a green glow with `!important`, while `status` stays `busy` for a
pane whose agent died mid-turn, so a phone rendered a grey dot still wearing the
green halo beside a badge reading "exited". Both measured against the real
stylesheets, both now excluded, and the CSS test reads mobile.css too instead of
being structurally blind to half the problem.

Six comments said things that were not true. Two named the stats collector as
what replaces a restored reading, which is the opposite of the design. The
interval constant argued that 2000 ms keeps a read inside a tick, when the
5000 ms exec timeout means it cannot — which is why the in-flight guard exists.
`MuxSession.discovered` did not say the flag is permanent, though `saveSessions()`
serializes it. The empty-read docstring claimed a distinction that `|| true`
makes impossible. The invariants doc promised more than its drift test delivers.
And CLAUDE.md had no pointer at all, leaving its two hardest prohibitions
("never set `status: 'error'`", "never null the pid") only in the file it is
meant to route people to.

Refs Ark0N/Codeman#446.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 15:24:09 +02:00
Rounak DattaandClaude Opus 5 e1e7dc5bd8 fix(terminal): Ark0N's read of the #464 geometry work
Five items, two of which he could only see by running it, plus six smaller
ones. Taking the two blockers first, because both were wrong in ways the
existing tests could not catch.

**Adopting the PTY's rows put the CLI's input line off-screen.** A phone that
took a desktop's 43 rows into a viewport with room for 18 painted an
`.xterm-screen` far taller than its container; xterm's own viewport then had
nothing to scroll, so the bottom of the frame sat below the container with no
gesture able to reach it. Output visible, typing invisible, for as long as the
desktop kept the claim hot. `reconcilePtyGeometry` adopts COLUMNS ONLY now:
width is the axis Ink's wrap and `eraseLines` arithmetic depend on, and keeping
the local row count keeps the composer at the bottom of a viewport that
scrolls. Measured at his geometry — a 360x300 container against a 198x43 pane
now keeps 13 rows, takes 198 columns, paints 202px into a 210px container, and
the input line is inside the box.

**`capture-geometry-retry.browser.test.ts` failed, and CI could not see it**
because the file is in `BROWSER_TEST_GLOBS`. Its premise WAS the clamp —
`getTerminalDimensions()` floored while `fitAddon.fit()` did not — which this
work removes at the source, so it can never hold again at any viewport. The
case survives on its own terms: a pane already drawing at the requested size
must not be replayed. Its premise is now the #464 invariant itself, that the
floored report and the terminal agree, which is a stronger guard because the
clamp coming back fails it here rather than silently restoring the replay loop.
The helper docblock that repeated the old premise is corrected too.

**A session with no pane reported 120x40 and the client adopted it.**
`resize()` writes `_ptyCols`/`_ptyRows` only when `ptyProcess` is set and
nothing seeds them from the spawn geometry, so a dead-pane session still held
the constructor defaults — clicking that tab resized the browser terminal to
120x40 and, on anything narrower, claimed another device owned the pane when
none existed. `Session.ptyGeometry` returns null without a pane, the HTTP route
answers `{}` and the socket sends no frame at all. The raw `ptyCols`/`ptyRows`
getters are deleted rather than left available to be misused again.

**The 40-column floor clipped the pane with nothing able to reach it.** The
affordance keyed on a PTY mismatch, and the floor produces no mismatch — xterm
and the PTY agree throughout, the terminal is simply wider than the box. It
keys on what does not FIT now, MEASURED (`.xterm-screen` against the container,
on the next frame, because the screen takes its width with the render) rather
than derived from cell arithmetic. Measured at 360px: font 24 applies 40
columns and paints 560px, and all 200px of the overhang is reachable.
`.pty-oversized` is renamed `.term-overflows-x`, because after this the old
name describes only one of the two causes.

**"Scroll sideways" did not work on touch for the sessions it targets.**
`touch-action: pan-x` is cancelled before it starts by the `preventDefault()`
`touchstart` calls on every 'content' tap. The terminal's own touchmove handler
pans the container now, with the axis locked once per gesture so a diagonal
cannot pan and scroll at once, and the CSS grants no `touch-action` at all —
handing the browser a pan AS WELL would move the pane twice for one finger on
the taps where that preventDefault does not run. Measured under real touch
dispatch: a 140px swipe reaches `scrollLeft` 140 where it reached 0 before, the
buffer does not move with it, and a vertical swipe still scrolls the scrollback.

Three defects in the above, found while checking it rather than by being told:

- `canPanHorizontally` first tested `scrollWidth > clientWidth` alone, which is
  true of a container that is not a scroller — a sideways swipe would have
  locked the axis, done nothing, AND suppressed the vertical scroll it should
  have been. Gated on the class as well.
- The notice advised scrolling sideways whenever the PTY was wider, including
  when it still fitted and nothing scrolled. It is gated on measured overflow,
  and on a comparison against the width this container WOULD request rather
  than the one it currently holds — once adopted those are equal, so the second
  question answers itself false while the condition is still true.
- `_syncTerminalOverflowAffordance` could throw out of `document.getElementById`
  before reaching its try block. It runs off every geometry change, so a
  cosmetic affordance could have taken the resize down with it.

The smaller items:

- `docs/architecture-invariants.md` no longer explains the equality guard as a
  clamp signature; it records what the clamp used to do and why it cannot any
  more. Edited by hand — that file is outside the Prettier glob, and letting
  Prettier near it rewrote eleven unrelated emphasis markers.
- `throttledResize`'s HTTP fallback reads the reply. It is the path where a
  declined resize is least likely to be noticed, because no socket means no
  `{"t":"zc"}` frame either.
- The changeset covers the whole release: the geometry work, the queued replay
  clear, the renderer watchdog, the body-covering fetch deadline, the WebSocket
  output-gap reconcile, the build-generated service-worker precache and
  per-build cache key, and the crash-trail hygiene.
- `@xterm/headless` is declared in the root devDependencies instead of being
  reached through workspace hoisting.
- The output-gap marker is cleared after any response arrives, not only when
  the capture was non-empty: a server that answers with an empty capture HAS
  reconciled us, and leaving the marker set refetched on every reconnect.
- `e587d845`'s message claimed a test asserted the failed-load copy against the
  built asset. It did not — that assertion lived in a probe deleted with the
  other scratch scripts, so the claim was false when it was written. There is a
  real test now, and it reads the source rather than `dist/`, because `dist/` is
  not committed and a test that skips when it is absent would pass for the wrong
  reason in CI.

`Session.ptyGeometry` gets behavioural coverage against the real class in
`session-resize-arbitration.test.ts` rather than a source guard, including the
contrast — a pane that does exist still reports, and still follows a resize —
so "always null" would fail it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 18:53:46 +05:30
Rounak DattaandClaude Opus 5 abf1d1f1ca fix(terminal): the PTY and the browser terminal must never disagree about size
Issue #464, "text gets muffled sometimes, in both TUI default and fullscreen".
The screenshot is not a dropped frame or a frozen renderer — it is arithmetic.
Claude Code's TUI wraps its frame at the width the PTY reported and erases the
previous frame by walking the cursor up the rows it believes that frame took. A
browser terminal of a different width makes each logical line occupy more
physical rows than Ink counted, so `eraseLines(n)` clears too few and the new
frame paints over rows nothing erased: doubled lines, and short tool summaries
sitting inside longer prose rows with the prose's tail still visible.

Reproduced against this repo's own xterm before changing anything — a 120-column
PTY against a 62-column terminal renders every wrapped line twice. `test/
terminal-pty-geometry.test.ts` pins that, and pins the clean render at matching
widths beside it, so the assertion cannot be satisfied by code that fixes
nothing.

Four ways the two drifted apart, none of them observable from either end:

1. `fitAddon.fit()` resizes xterm to `proposeDimensions()` RAW while every
   server-facing path reported those floored at 40x10. Measured in Chrome at
   430px: font size 44 proposed 13 columns, the server was told 40, and xterm
   stayed at 13. Three call sites each did their own fit-then-floor, and two
   re-read the proposal after the fit — `_shrinkPaddingToFit()` runs exactly
   there, so the container had moved.
2. `throttledResize` (keyboard up) and `sendResize` (session detached into its
   own window) reflowed locally and withheld only the SIGWINCH. That is the one
   combination that cannot be right: a reflow nothing is rendering for buys
   nothing and costs correctness. Both now withhold everything, and the
   keyboard's settle timer still sends the one resize that stops the PTY going
   stale.
3. `setFontSize`/`setFontFamily`/`setFontWeight` move the cell size — a geometry
   change — and told the server nothing at all, so raising the font on a phone
   left the CLI wrapping at the old column count.
4. `Session.resize` DECLINES a small-viewport request while a desktop connection
   holds an active sizing claim, and said nothing, because resize was write-only.

`syncTerminalGeometry()` is now the one function that may change the terminal's
size: it fits, floors and applies as a single step, so the numbers xterm holds
are the numbers the server is told. A test sweeps every module for a bare
`fit()` on the main terminal, and finds exactly one — the owner's own.

For (4) the client cannot win, so it is told the truth instead: both transports
answer a resize with `session.ptyCols`/`ptyRows` (`{"t":"zc"}` on the socket,
the body of the resize POST) and `_onPtyGeometryReport` adopts them. A terminal
that keeps a shape the PTY refused does not render "too narrow", it renders
garbled. Adopting can leave the pane wider than the screen and the container is
`overflow: hidden`, so `.pty-oversized` grants horizontal reach for exactly as
long as the mismatch lasts: correct-and-reachable beats correct-and-clipped
beats garbled. That rule sets both overflow axes and its own `touch-action`
because mobile.css loads later and sets `.terminal-container { overflow:
visible; touch-action: none }` — a bare `overflow-x` would leave overflow-y
computing to `auto` and hand the browser a vertical scroll container the
terminal's touch handler knows nothing about.

Verified in Chrome at 430px against a live server, with a desktop client holding
the claim: the phone adopts 198x43, gets `overflow-x: auto` / `overflow-y:
hidden` / `touch-action: pan-x`, 758px of reach to the right, and keeps its own
vertical scrolling. The pre-fix build was measured in the same harness for the
control.

Two things this deliberately does not do. It does not change who owns the pane
size — the desktop still wins, and `_startMobileResizeRetry` still takes it back
once that goes idle. And `throttledResize` still holds the PTY's shape for the
whole keyboard animation rather than sending a SIGWINCH per step; that decision
predates this and was not re-tested here.

Also in this commit, Ark0N's third-pass review items on #431:

- The response viewer's byte-buffer fallback and `_onSessionClearTerminal` both
  used the no-param `/terminal` form, capped only by `terminalBufferMaxBytes`
  (32MB) — the largest body the frontend asks for anywhere. One carried no
  deadline at all and the other got the 15s tail budget. Both now take the
  full-history budget.
- A `?full=1` capture that outruns its deadline falls back to the bounded tail.
  The pane is blanked before that fetch, so an abort used to leave a black
  rectangle, discard the queued live output and never reach `_connectWs`. A
  failed load now still opens the socket, says one dim line where the content
  would have been, and clears the tab's spinner — which nothing did, so a failed
  select left `aria-busy="true"` set forever.
- `_wsOutputGapSession` is cleared at the repaint that settles it, not in a
  `finally` that also ran on the catch. A reconcile that threw, or hit the new
  deadline — the flaky link the marker exists for — dropped the gap with nothing
  to retry it. `ws.onopen` no longer clears it up front either.
- The replay-clear invariant is pinned in the gate, which is the drift this PR
  exists to fix: `_resetTerminalForReplay` must be a queued write and nothing
  else, and no module may blank the terminal with a `clear()+reset()` pair.
- `DIAG_ENTRY_MAX_CHARS` replaces the hardcoded 300, bound through a local
  first: `CodemanDiag?.x` still throws a ReferenceError when the identifier was
  never declared, and that is the one function in the app that must not throw.
- panels-ui's two kill-all clears route through the same helper, and the
  xterm-version guard's comment says "resolved lockfile version" rather than
  "dependency RANGE", which is what it has pinned since the last round.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 13:07:33 +05:30
Michael GrundbergandClaude Opus 5 64c288a683 feat(codex): read Codex's own background-terminal row
Codex states background work too, and it says so in a different place. Claude
writes `· 1 monitor ·` on the last row of the screen; Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its
composer, which puts that row third from the bottom once the status line and
the composer are counted.

So how far up the screen to look is now per-CLI data as well:
`capabilities.workDetect.watchingLines`, bounded to 1..8 by the schema, and
defaulting to Claude's two. That bound is the point. The window is half the
injection guard, since every row it adds is another row the agent itself may be
able to write, and the label is what silences an idle alert. The other half is
the anchor, and Codex's is ` · /ps to view`: chrome naming a slash command only
the CLI can offer, so a session that writes "I left 1 background terminal
running for you" into its own output matches nothing.

Measured against a live codex-cli 0.154.0 pane rather than read out of a
binary. The row appears when the terminal starts, follows the composer down as
the conversation grows, and is gone after `/stop`. Verified end to end on an
isolated beta: the session payload carried `watching: "1 background terminal"`
and the badge rendered with it, and both cleared when the terminal stopped. The
fixtures in the tests are that capture verbatim.

Codex has no hook signals, so no idle prompt and no false NEEDS YOU row: for a
Codex session this is the badge alone, which is the case the maintainer said a
registry field could cover and a hook never could. Cross-CLI tests pin that
neither pattern fires on the other's screen, and that a CLI declaring nothing
still reports nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 08:58:52 +02:00
Michael GrundbergandClaude Opus 5 74884a20eb feat(approvals): let a watching session keep quiet, and fix the tui gate
The badge alone left the row in NEEDS YOU, which is the thing the issue was
about. The fix is the alert that does not fire.

An idle prompt from a session that is watching its own background work now
opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to
`notePrompt()`, which sets `acknowledgedAt` and records why in a new
`acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has
always meant "the alert this prompt armed is spent", and the prompt itself
stays pending, answerable and available as Read My Mind context. A wrong label
therefore costs a card that does not blink, never an alert that was never
created.

Every surface follows from that. The broadcast carries the reason, so a live
page declines to arm the tab alert and raises no desktop notification. The push
is skipped, since a false alarm is hardest to ignore on a phone. A reloading
page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And
`classifySession()` now reads it too, which is a pre-existing bug fixed here:
acknowledging on one device cleared the alert everywhere except `codeman tui`.
It re-arms for free, because the next idle prompt supersedes the item and is
built fresh. Only `idle` is eligible, so a dialog that blocks the agent still
goes red whatever else it started.

The label is pane-derived and therefore prompt-injectable, so it is now read
from the last two rows of the screen only, with Claude's pattern anchored on
the `·` its footer joins items with, ANSI-stripped and length-capped at the
source. An agent that prints `· 1 monitor ·` into its own output finds no
match.

Verified on an isolated beta: a session that armed a monitor took its idle
prompt acknowledged with no alert on any surface, wore the badge, and showed
"quiet, watching 1 monitor" on its still-answerable card; the same session with
the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts`
pins both directions across all four surfaces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 08:10:05 +02:00
Michael GrundbergandClaude Opus 5 1cb0441bd8 fix(session): degrade the resume pin to the session id, not to nothing
A single pin that failed its transcript gate returned the options untouched,
so `resumeSessionId` fell back to `_resumeSessionId` — undefined for an
ordinary session — and the renderer emitted the bare
`claude --dangerously-skip-permissions --session-id "<this.id>"`. Every
session prompted before its first `/clear` owns a transcript under that id,
so the dropped pin handed back exactly the refusal this branch removes, with
no `||` branch to catch it. It was also a regression against master on the
`restartCli()` path, which pinned `_claudeSessionId ?? this.id` and, since the
constructor seeds that field, could never land unpinned.

The pin now walks three candidates in priority order — the conversation
chain's tail, the launch seed, then the session's own id — and takes the first
one a transcript backs. A candidate that misses is passed over rather than
ending the walk.

Falling off the end pins nothing, which also settles the second half of the
problem: the old code skipped the transcript check whenever the pin was the
session's own id, so a genuinely new pane rendered the two-branch form after
all. That costs a brand-new session claude's "No conversation found" line in
its scrollback, and `wrapWithNice()` prefixes only the first branch of the
rendered `a || b`, so the branch that actually runs loses its priority for the
life of the session. With no transcript anywhere the bare `--session-id` is
the correct command, so the comment claiming an unchanged shape is now true.

The transcript lookup reads the server process's own `CLAUDE_CONFIG_DIR` when
a session declares none. A pane inherits the server environment through tmux,
so on an install that exports it the CLI writes its transcripts there and
every lookup under `~/.claude` was a false negative — which under the old code
meant the colliding command. `claudeCredentialsPath()` and
`realClaudeConfigDir()` resolve the same directory the same way. The header
sentence calling a skipped resume "the safe direction" described the opposite
of what happens at this call site, and says so now.

The create-path fallback writes `_resumeSessionId` alongside the create
options. That branch leaves `isRestored` false, so `_claudeSessionId` is
recomputed from the launch fields and settled on `this.id` while the CLI
resumed the chain tail; the response viewer, Read My Mind and the unified-list
alias map read that field until the next first-hand hook.

Four new tests: a chain tail with no transcript while the session id has one,
no transcript anywhere, the create path's alias, and the process-env lookup.
All four fail against the previous commit. Two existing tests move with the
gate — the guess-refusal test now backs the session's own id, and the
custom-model restart test gives its working pane the transcript that makes
`--session-id` collide in the first place, alongside a new one pinning the
no-transcript case.

CLAUDE.md described the pin as a `restartCli()`-only thing sourced from the
live conversation id. All three halves of that moved here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 18:21:21 +02:00
Michael GrundbergandClaude Opus 5 3f2cde2db7 feat(session): say when a session is watching its own background work
An agent that arms a monitor, backgrounds a shell or hands a task to a
cloud session is told to end its turn. The pane then falls quiet, Claude
Code's idle_prompt notification arrives a minute later, and every surface
files the session under NEEDS YOU with nothing for a human to answer.

Claude states what it is still running on the last row of its screen
(`⏵⏵ bypass permissions on · 1 monitor · ← for agents`). That row is now
`capabilities.workDetect.watchingLine` in the CLI registry, guarded by
compileVersionRegex() like every other config regex, and the idle probe
reads it off the capture it already takes: `watchingLabel()` in
session-activity.ts searches the last five lines only, so a session that
PRINTS "1 monitor" is not mistaken for one running it.

The label lands on Session.watching and rides toLightDetailedState() out
to every surface. The phone overview, the desktop home rail and the rich
sidebar rows wear it as a `watching` badge in the accent colour, beside
the state pill and never in place of it: an agent can arm a monitor and
ask a question in the same breath, and only the pill says which.

Verified end to end against a throwaway session on an isolated beta
instance: the payload carried `watching: "1 monitor"` once the turn
ended, the badge rendered next to a yellow `waiting` pill, and both
cleared when the monitor died.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 18:00:30 +02:00
Michael GrundbergandClaude Opus 5 47e7935274 fix(session): resume the conversation when respawning a dead pane
A CLI that launches with `--session-id <id>` refuses an id that is already in
use (claude: `Error: Session ID ... is already in use.`), and every session
whose agent has been prompted owns a transcript under that id. The dead-pane
respawn in `_setupOrAttachMuxSession()` passed the bare launch line, so
recovering such a session relaunched a CLI that died on startup, the pane went
dead again at once, and the conversation was stranded behind a tab that looked
merely idle.

`restartCli()` has pinned a resume id against this since the custom-model work,
and its comment states the assumption that made the other path look safe:
"Unlike the dead-pane respawn, this one kills a WORKING pane whose conversation
already has a transcript". A pane whose agent exited has a transcript too.

Both relaunch paths now build options through
`_buildRespawnPaneOptionsWithResumePin()`, and so does the create-path fallback
after a failed respawn, which otherwise met the same refusal that made it the
fallback. Four gates guard the pin, each standing for a way of resuming the
WRONG conversation or of making a working relaunch fail.

A remote or docker session is never pinned. Unlike `restartCli()`, whose route
refuses both, the dead-pane respawn is reached by every session shape. Their
pane commands already render a self-healing `--session-id || --resume`, and
both flip to resume-first once the resume id differs; the conversation lives on
the far side, so a local id resolves to nothing there and the `--session-id`
fallback then collides with the transcript the far side does hold.

The id comes from the conversation CHAIN rather than `_claudeSessionId`, which
also holds history-correlated guesses keyed on the working directory.
`_recordClaudeSessionInChain()` refuses those so they cannot "write a foreign
conversation into this pane's permanent record", and launching from one is
worse than the display bug that rule prevents. The chain tail also outranks the
launch seed, which is written once at construction and never moves off a
`/clear`.

A pin no transcript backs is dropped, because the fallback branch keeps
`--session-id <this.id>` and would collide. A synthetic `restored-<fragment>`
id from socket discovery is dropped too, and logged: it fails claude's `uuid`
token pattern, so the renderer would emit the unpinned command while the caller
believed otherwise.

Tests cover each gate and the rendered command. Four of them fail against the
unfixed source; the remote and docker ones were separately checked against a
build with only that guard removed, since they pass on master for the wrong
reason.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 15:36:17 +02:00
Michael GrundbergandClaude Opus 5 90a95f562b feat(session): publish and persist a local pane's agent exit
The mux layer now knows a pane's agent has exited. This puts it on the session
record, where the board and, later, the reboot restore can see it.

`SessionState.paneExit` carries `{ status?, signal?, at }` and rides the
existing `session:updated` broadcast through `toState()`. No new SSE event. The
server pulls each answer from `mux.getPaneExit()` rather than off a broadcast
payload, so the raw reading never reaches a browser: for a remote or docker
session that reading is the death of an ssh client or a `docker exec`, not of
the agent.

The field is tri-state, and the third state is its absence: `undefined` means
Codeman does not know, and it never reads as alive. `Session.setPaneExit()`
forces that unknown for every shape a dead local pane does not describe. A
direct-PTY session owns no pane. A remote SSH session's local pane holds the
ssh client, whose death means a transport drop OR an exit, which is the
ambiguity PR #355 settled by not guessing. A docker case's local pane holds a
`docker exec` into the container's own tmux. And a session rebuilt from the
socket has no provenance at all: `reconcileSessions()` gives it a synthetic
`restored-<fragment>` id that matches no `state.json` entry, so a remote
session rediscovered after `mux-sessions.json` was lost arrives with no
`remote` field and looks local — `MuxSession.discovered` marks it, and absent
metadata there counts as unproven rather than as proof. The scoping lives on
`Session` rather than in `TmuxManager` so there is one copy of the rule.

`status` and `pid` are untouched. `status: 'error'` belongs to the PTY-exit
circuit breaker and makes the browser offer a restart, and a null `pid` is what
makes the browser re-attach and launch a fresh CLI. A reading that repeats the
previous answer writes nothing and broadcasts nothing.

An unknown answer never reads as alive, but a stale KNOWN one would keep
reading as exited, so `clearPaneExitForNewPane()` retracts it on every path
that puts a new command in the pane: the start/attach path, the `restartCli()`
relaunch behind a custom-model switch, and the remote reattach. Without the
second of those, switching an endpoint on an exited session launched a new
command and then persisted and broadcast the old exit straight back onto it.

`toState()` is also what `state.json` persists, so the record survives a
reboot, which is the only thing that does: a reboot takes the tmux server, and
with it every live signal and every `mux-sessions.json` entry. Nothing reads it
there yet — making the restore refuse such a session is a behavior change that
belongs with the part that closes them. Recovery threads the saved value back
through the constructor so the first persist after boot cannot blank it.

Refs Ark0N/Codeman#446.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 10:01:09 +02:00
Codeman maintainer 4205f6930f fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine.

**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.

**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.

**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.

**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.

The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:58:33 +02:00
Codeman maintainer 19ffe9b7a8 fix(input): make sure a prompt sent through the API actually leaves the composer
Claude Code 2.1.277 takes typed text the moment its composer paints but
ignores Enter for the first 30 to 50 seconds after it (measured 2026-09-19
through the input route: an Enter at 28 s stranded the prompt, one at 51 s
submitted it). The text+Enter pair `sendInput` sends 50 ms apart therefore
left every programmatic prompt sitting unsent, and every waiter burned its
timeout on a turn that never started.

Server: `SubmitVerifier` (session-submit-verifier.ts), armed from
`writeViaMux` for every mux write that carried a carriage return, reads the
pane on a 2 s to 60 s schedule and re-sends Enter only while the last
composer line (the CLI's own prompt glyph) still holds the head of what was
sent. An empty composer, other text, or no composer line at all ends it; a
newer write replaces the schedule.

Skill: `sendwait` gets the same loop (`_composer_text`, no-break space
stripped by its bytes for BSD sed) for servers that predate this, and the
preamble version moves to 1.30.1 so seeded agents pick up the fresh copy.
SKILL.md's heredoc and the plugin mirror are regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:38:59 +02:00
Michael GrundbergandClaude Opus 5 fa52753e8b fix(sessions): undo a failed rebuild without deleting the user's data
A second review of the previous commit found that its own repair for the
session leak introduced three defects, all from reaching for
cleanupSession() to undo a half-built session. That function is the
user-initiated delete, not an undo.

It banked the session's historical token and cost totals into the lifetime
figures, and a reboot never runs cleanup, so those totals had never been
counted before; every failed rebuild added them again. It saw the pin that
had just been restored and demoted the record to `stopped`, which this pass
reads as the durable marker of a deliberate kill, so a pinned session whose
rebuild failed became permanently unrestorable. And it recursively removed
`.claude-images` from the working directory, which belongs to the workspace
rather than to the session, so a failed rebuild destroyed the pasted images
of any other live session in that repo.

discardPartiallyBuiltSession() now undoes only what the construction did:
the map entry, the tab-layout slot, the listeners and any pane the launch
created before throwing. The persisted record, the lifetime totals, the
Ralph state and the workspace's files are left alone.

Re-applying the persisted state also splits in two, which removes the first
two defects at the root rather than only at the call site. The half that
shapes the pane, the custom-model environment and the nice priority, still
runs before the spawn. The half that is the session's own history now runs
after it, so a session whose pane never started carries no totals and no pin
for anything downstream to misread.

The rest of that review. The multi-user workspace confinement re-check read
the requesting user's grant, and returns true for an admin, so the case its
own comment described was the one it missed; it now resolves the entry
owner's grant through isWorkingDirAllowedForUsername, the way cron does. A
forbidden workspace goes back on offer, matching both the registry's stated
contract and the API reference. The client re-reads the plan after a restore
instead of blanking the banner, so entries the server put back stay
reachable, and a 409 now says a restore is already running rather than
reporting a failure. A dismiss arriving mid-restore wins, through a
generation counter the route carries across its take. The re-application
also restores the tab colour, the image-watcher flag and the original
pinnedAt, via a new Session.restorePin that does not re-stamp the pin time.
The phone breakpoint gains min-width: 0, without which a nowrap flex item
never shrinks and the buttons still overflow, and it folds into the existing
phone block.

Ralph's loop configuration still does not survive a restore, because
toState() reads it off a live tracker and there is no way to keep it without
arming the loop. The method now says so rather than leaving it implied.

Tests. The capacity test could not fail on the property it existed for: it
filled the board past the cap before the loop, so a single pre-loop check
would have passed it. It now leaves one seat, so only a per-iteration check
restores exactly one entry. New tests cover the ordering around the spawn,
a throw before the loop returning the whole plan and releasing the flight,
the dismiss-during-restore race, and that the failure path calls the narrow
discard rather than the delete. The shared mock context gains the port
method it was missing, which is what made the first run of these tests fail
for the wrong reason.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:04:35 +02:00
Codeman maintainer 5b920cb43d feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream
and named tabs wrong five ways (every prompt, every write path, a bare Esc
eating the next prompt's first character, pasted newlines as Enter, any CSI
clearing the draft) and replaced the whole name, which dropped the case from
the tab and reset the w<n> counter. This lands the feature with each of those
closed:

- First prompt means the first: applyAutoName() flips a placeholder to
  `auto` whether or not the string changed. nameSource is now the tri-state
  placeholder | auto | manual; the name setter is the only manual path.
- Only user-originated input counts: write()/writeViaMux() take
  SessionWriteOptions.fromUser, set by the browser WS path and POST /input
  only, so Ralph, respawn, cron, approvals and the trust-dialog keys can
  never name a tab. A startMode 'shell' CLI never feeds the tracker (a
  capability, not an id check); the send-key route feeds trackUserInput()
  because its line feed bypasses the session.
- Prefix form `w3-case: title`: parseSessionPrefix() already renders it as
  the title with the prefix in the tooltip and the next-session counter
  still matches it. Composed within MAX_SESSION_NAME_LENGTH.
- Tracker rules per key: bare Esc resolves at chunk end; mouse/focus
  reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R
  taint the draft so Enter submits nothing rather than a fragment;
  bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space;
  the draft keeps its head past 8192 code points; an escape past 64 bytes
  is abandoned.
- Title: slash commands by shape (a path is a prompt), `!` escapes
  refused, first sentence only past 8 code points ("e.g." is not a title),
  72 code points on a word boundary.
- Synced `autoNameSessions` setting, default OFF (the prompt reaches
  mux-sessions.json, session:updated and /api/search), App Settings ->
  Appearance -> Tabs, read fresh per prompt after the eligibility check.

Tests: test/session-auto-name.test.ts (tracker, title, composition,
ownership, emit gating), the wiring test (once, prefix, setting off,
manual protected), test/routes/session-name-routes.test.ts (PUT /name
flips to manual and persists). Verified live on an isolated instance: API
and browser-typed prompts name the tab, a second prompt does not, shells
and renamed tabs are untouched, nameSource survives a restart.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 17:59:16 +02:00
Codeman maintainer c4322513d9 Merge pull request #376 from shenlvkang-collab/feat/auto-session-names-upstream 2026-09-15 17:19:57 +02:00
d fei 01da577053 fix(input): recover when the seq counter falls behind the server watermark
Browser input is delivered exactly once by (clientId, seq). The server records a
watermark per clientId and discards anything not above it as a duplicate — but
acknowledged it with an ACK indistinguishable from "applied". The client then
dropped the record from its queue, the UI looked perfectly normal, and the
terminal received nothing at all.

The counter is persisted to localStorage through a debounced write. Kill the page
between "sent" and "persisted" and the restored counter is below the server's
watermark, after which every keystroke lands under it, is discarded, and is
ACKed. Reloading does not help: the clientId is restored from localStorage
alongside that stale counter. Measured on a real session — typing into the same
session from a fresh browser (new clientId, no watermark on the server) worked
perfectly, which is what localised the fault to client state.

Three changes:
- on rejection the server replies {"t":"ia",seq,"dup":true,"last":<watermark>}.
  It still ACKs, so the client can drop the record from its queue, but it now
  says the input was not applied and supplies the number needed to climb out.
- on `dup` the client lifts its counter above the watermark and re-queues.
  ⚠️ Only records whose FIRST delivery is being retried are re-sent: a retry
  judged duplicate means the mechanism is working (the original did arrive), and
  re-sending would type the same text twice.
- the counter is now persisted synchronously. The queue payload can stay
  debounced, but the counter is the thing that has to survive a crash, and
  leaving it on the lossiest path cancels the only guarantee there is.

⚠️ Reading the watermark is defensive: the session arrives through a structured
port, and a port missing that method must not take the whole input path down —
a throw inside the handler means the ACK is never sent and the record is stuck in
the client queue forever, which is worse than the ambiguity being fixed. A mock
port's test timeout is what exposed this.

(cherry picked from commit 05bb7081cc)
2026-09-14 23:56:19 +02:00
Codeman maintainer 942bf37e48 fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a
custom OpenAI-compatible endpoint by injecting env vars or a config file
and restarting the CLI in place. Review of the apply path found four
things, two of them destructive. This lands all four plus the smaller
items from the same review.

1. Clearing a selection did not clear it. The injected vars reach the CLI
   via `tmux setenv`, which persists at the tmux-session level and is
   inherited by `respawn-pane` (measured: `setenv FOO bar` survived two
   successive `respawn-pane -k`), so deleting the keys from the session's
   envOverrides relaunched the CLI still pointed at the old endpoint, and
   for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just
   been deleted. `Session.setCustomModel()` now reports the removed keys,
   queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys`
   carries them into `applyEnvOverrides()`, which `setenv -u`s them before
   re-applying the live overrides, on the same path that already unsets
   the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket
   that `setenv -u HOME` hands the next respawn the global HOME back.

2. Applying a model to a local claude session killed the pane. The
   relaunch was `claude --session-id <id>` and Claude refuses an id that
   already has a transcript, and unlike the dead-pane respawn this one
   kills a working pane first. `restartCli()` now pins the live
   conversation id as the resume id for that respawn when the CLI's launch
   declares a `fallback` chain, which renders the same
   `--resume <id> || --session-id <id>` shape the docker and remote pane
   commands use. Gated on the registry shape, not the CLI id: an entry
   whose resume id is minted by the CLI itself never declares that chain.

3. pi, omp and grok wrote their config file and then launched without the
   `--model` that selects it, so the file was ignored. The registry entry
   now declares `customModelInjection.launchModel` (`custom/{modelId}` for
   pi and omp, grok's `[model.codeman-custom]` block name), the builder
   renders it, and `_withCustomModelLaunchModel()` applies it onto the
   respawn options through `legacyConfigField`, leaving the stored
   <Mode>Config untouched so a clear falls back to the user's own model.
   A model id the CLI's `model` token pattern cannot carry is refused
   with a 400 rather than silently dropped by the argv engine.

4. Remote (SSH) and Docker sessions reported `restarted: true` and changed
   nothing: their `restartCli()` reattaches the durable tmux rather than
   relaunching the agent, and the env lands on the local pane. Both are
   refused with a 400 until those paths are plumbed.

Smaller items from the same review:

- The selection survives a Codeman restart as the disk-only `__customModel`
  bookkeeping (endpoint, model, injected key NAMES, config dir, launch
  model; never the values, which carry the API key). Recovery re-derives
  the values from the endpoint store through the same apply path the route
  uses and keeps the bookkeeping even when the endpoint is gone, so a
  later clear still has keys to unset.
- Discovery goes through `webviewFetch()`, so the RESOLVED address is
  judged by the same egress guard the web-tab proxy uses, and `baseUrl`
  reuses `webviewUrlSchema` (http(s) only, no embedded credentials,
  link-local and cloud-metadata addresses refused). undici's `fetch failed`
  wrapper is unwrapped so the user sees the ECONNREFUSED underneath.
- `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session
  config dir 0700/0600 (pi and omp embed the key literally), and that dir
  is removed with the session.
- `PR.md` is gone from the repo root and the design doc moved to
  `docs/custom-model-endpoints-plan.md` with the LAN address and the
  personal name scrubbed; every reference follows. The guide's `authStyle`
  text matches the shipped schema (`bearer | api-key`, default `bearer`)
  and says that `customModelEndpointsEnabled` is read by nothing until
  the picker lands.
- `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts`
  (four real type errors fixed). It is not yet wired into `npm run typecheck`
  because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json`
  there is the one-line follow-up.

Tests: `test/session-custom-model-restart.test.ts` drives a real Session and
fails on the unfixed code for items 1 to 3; the route suite covers item 4
and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets
run before the overrides and that a shell-metachar key never reaches tmux.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:46:28 +02:00
Codeman maintainer 1e42cb4e2d Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses)
2026-09-14 23:46:27 +02:00
Codeman maintainer fa1ea8d9fe fix(statusline): unset a stale user statusline var, write the exporter script atomically
Three small follow-ups from the #361 review.

A tmux setenv survives respawn-pane, so _configureStatusLineUserCommand
returning early when the user has no statusline left a previously
exported CODEMAN_USER_STATUSLINE_CMD in place: a user who deleted their
own statusline kept getting the stale one wrapped, and lost Codeman's
footer print-through, until the tmux session was recreated. It now
issues `setenv -u` in that case, the same shape as the effort-level
cleanup in applyEnvOverrides.

ensureStatusLineExporterScript truncated and rewrote a script that live
sessions execute on every statusline render, and chmod'd it after the
write. It now writes a temp file next to the target, chmods that, and
rename()s it into place.

The non-tmux direct-PTY fallback carries no exporter; that is now stated
at the spawn site and in the architecture-invariants paragraph rather
than left as a silent gap.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:59:00 +02:00
DevvynandClaude Sonnet 5 41416566aa feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).

- Registry: capabilities.customModelInjection per CLI entry (env /
  configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
  model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
  custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
  (POST /api/sessions/:id/custom-model), reusing the existing
  respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
  CLI's privilegedEnvKeys, closing a pre-existing gap where several were
  already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
  against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
  every CLI's injected values through a real HTTP shape

Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
  a top-level model string + [model_providers.custom]); fixing it then
  surfaced a genuine, documented protocol incompatibility (Codex only
  speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
  don't implement)
- Claude Code's async session-title-generation call validates
  ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
  hangs the whole -p invocation on an unrecognized name; documented for
  chunk 6, worked around in the standalone script only (--bare is NOT
  safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
  and api-key headers) reliably hung a real server; removed the option
  entirely rather than just changing the default

Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00
timkjrandClaude Sonnet 5 797f0d387c fix(remote): address review feedback on omp/claude respawn continuity
- Remote omp command now renders through buildSpawnCommandFromRegistry
  (the mode-agnostic engine local/docker spawns use) instead of the
  buildOmpCommand() the CLI-registry refactor deleted.
- Session._pinOmpRespawnId()/_maybeCaptureOmpSessionId() now skip
  host-local ~/.omp resolution entirely for a remote session and fall
  back to --continue: that resolver only ever reads THIS host's
  filesystem, which is meaningless (and could wrongly alias an
  unrelated local conversation) for a conversation that lives on the
  remote host.
- Remote-claude launch now honors an explicit resumeSessionId distinct
  from sessionId (mirrors claudeDockerPaneCommand's shape), and
  validates sessionId the same way that sibling does before
  interpolating it into the remote shell command.
- Add the still-missing header-cwd half of the trailing-slash test,
  and document respawn/reattach continuation + auto-reconnect-vs-
  clean-exit in docs/remote-sessions.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 22:11:54 -05:00
Ark0N 344e93c824 Merge pull request #386 from irisitymichaelgrundberg/feat/codex-resume
Merging with the phone-overview resumeId fix and the two unified-list doc passages applied on master.
2026-09-07 22:43:46 +02:00