Commit Graph
63 Commits
Author SHA1 Message Date
Codeman maintainer ecd577157b fix(cli-registry): codex launch defaults as registry data, ultra footer, schema doc defaults
- Codex footer model detection (c28): the modelDetect.screenLine effort
  alternation is now built from CODEX_REASONING_EFFORTS plus 'default', so
  'ultra' (offered by the codexReasoningEffort App Setting and codex's own
  /model picker) is read and the launch enum and the footer reader cannot
  drift again. Still one capture group, 125 characters, no new quantifier.
  New session-display-model case loops every effort level, ultra included.

- No CLI-id branching for launch defaults (c27): the two mode === 'codex'
  branches the synced codex model/effort defaults added to the create and
  quick-start routes are replaced by a registry capability,
  capabilities.launchDefaults (launch param -> settings key, values from a
  closed enum), declared on the codex entry only. The resolver moved from
  web/codex-launch-defaults.ts to web/launch-defaults.ts as
  applyLaunchDefaults(mode, configs, customEndpoint), filling the entry's
  legacyConfigField object through legacyConfigAliases, still re-validating
  with SettingsUpdateSchema and never overwriting a caller's value. The
  route exclusions are unchanged (create: not remote; quick-start: not
  remote, not Docker, not a custom model endpoint), and quick-start still
  derives the session model from a bag without ompConfig, as before.
  schema.ts refuses an undeclared param, an unknown settings key, an empty
  map, and launchDefaults on an entry with no legacyConfigField.

- The no-id-branching guard now carries an exact occurrence count per
  allowlisted key, so a new copy of an already approved expression fails
  instead of riding the old approval, with a synthetic anti-vacuity case.

- SettingsUpdateSchema JSDoc (c21/c29): 'classic' is the tabArrangement
  default and 'compact' the headerStatsStyle default, matching the
  resolvers and the pre-paint script; state/case/ledger are marked opt-in.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 09:35:56 +02:00
Codeman maintainer 1105d646f3 fix(terminal): #555 landing fixes
Comment and doc corrections that #555 made stale, no behaviour change.

- stock.ts: the grok and omp altScreen comments compared their strip to
  opencode's, which is now strip-mux-and-mouse rather than the narrow
  strip. Grok now says it shares antigravity's strip until measured, and
  omp drops opencode from its comparison.
- terminal-ui.js: the touch-tap comment named Claude/Codex/Gemini as the
  stripped modes, but the gate is now the cliMouseTracking flag alone and
  covers opencode too, so it names the two stripping flavours instead.
- src/types/session.ts: the cliMouseTracking JSDoc (the flag the browser
  now gates on exclusively) listed only claude/codex/gemini; it now names
  the strip-full and strip-mux-and-mouse modes, including opencode under
  tmux.
- src/session.ts: the usesMux getter doc now names isMuxMouseStripMode,
  since the replay strip passes usesMux to it as well.
- docs/architecture-invariants.md: the narrow-strip list gains
  grok/deepseek/omp (matching the PR's own CLAUDE.md line), the
  "must REMEMBER" heading covers both DECSET-stripping flavours, and the
  cliMouseTracking writer is described as the full-or-mouse branch it
  really is.
- docs/wiki/The-Dashboard.md: the user manual said every non-Claude CLI
  scrolls locally; opencode's wheel and swipes now page its conversation.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 05:30:20 +02:00
Codeman maintainer ff2f81541a Merge #555: opencode drags select text and the wheel pages its transcript 2026-10-09 05:28:51 +02:00
Codeman maintainer 7e2b9ed8f6 feat(opencode): show the model an opencode session runs in its tile and split headers
An opencode tile showed only the session name, while Claude Code, codex and
DeepSeek tiles add `· <model>`: opencode declared no `modelDetect`, so its
screen was never read for a model and it is launched without a model param.

opencode draws the model on its composer's agent row, directly above the
box's bottom edge: `┃  Build  Big Pickle OpenCode Zen`. Read from its own
1.3.0 source, the row is the agent, the model's name, the provider's name and
`· <variant>` when the model has one, and only colour tells model from
provider. So the field is all of it, exactly what opencode itself shows (the
owner's choice over a short id that only appears after the first reply).

- The pattern anchors on that row sitting directly above the `╹` edge, ends
  the field at a double space (where the 200-column layout's sidebar shares
  the row), skips the `No provider selected` placeholder, and takes the LAST
  such row in the window through a lookahead, so a composer-shaped row the
  agent prints higher up can never stand in for it. A test with a forged pair
  inside the window fails without the lookahead.
- It reads 8 rows: the home screen puts up to five rows of opencode's own
  chrome under the composer (key hints, a tip, the cwd/version row). The
  schema's `screenLines` bound goes from 4 to 8, the reader's own cap; the
  comment there records why a taller window is only safe with such a pattern.
- A permission prompt or shell mode hides the row; the last model is kept.

Measured against every captured opencode 1.3.0 frame (home screen and in
session, 40/60/120/200 columns, mid-turn and at rest, permission prompt):
the model was read everywhere it is drawn and nowhere else. Live on an
isolated instance from this branch, a restored opencode session published
`displayModel: Big Pickle OpenCode Zen` (source: screen) and its tile header
rendered `oc-home · Big Pickle OpenCode Zen`. The owner's own home-screen pane
on the 1.36.0 beta reads the same.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 03:37:42 +02:00
Codeman maintainer 7d8c188f83 fix(gemini): read Gemini CLI's composer bar and spinner line so a turn can end
A gemini session stayed "working" for good after its first turn. Its braille
spinner trips the generic SPINNER_PATTERN and marks the pane working, but only
a composer glyph arms the idle confirmation and gemini declared none, so it
fell back to Claude's `❯`, which gemini never draws.

Measured on live Gemini CLI 0.63.0 panes (capture-pane every 300 ms through
real turns with a shell call at 40, 120 and 200 columns, YOLO and default
approval mode, plus the raw PTY stream). The turns ran against a local
stand-in for the Gemini API (GOOGLE_GEMINI_BASE_URL, which Codeman's custom
endpoint support already sets), since the CLI's TUI does not depend on the
backend and no account is needed for it:

- The TUI repaints its whole bottom region every frame, composer included,
  and the composer sits between a `▄` bar and a `▀` bar; the submitted prompt
  is echoed between the same bars. The `▀` bar arms the idle check: every
  repaint carries it, tmux's reattach repaint too. The composer's prompt
  character is no good: it follows the approval mode (`*` in YOLO), and its
  `>` also starts the echoed prompt, which would make the submit verifier
  read a submitted prompt as stranded and press Enter again.
- While a turn runs a line `⠦ Thinking... (esc to cancel, 6s)` animates about
  every 80 ms (largest gap mid-turn: 214 ms). The label can be any loading
  phrase, so the working line is the `(esc to cancel, <n>` suffix, or a
  spinner frame opening a line for when a long phrase wraps that suffix.
  Nothing at rest matches either.
- A tool confirmation (default mode) replaces the composer, stops the
  spinner and the pane goes silent (3.9 s gap), so it reads as idle.

Verified on an isolated instance from this branch: a YOLO turn emitted one
session:working (+170 ms) and one session:idle (2.5 s after the last output);
a default-mode turn went working -> idle while the confirmation waited ->
working once allowed -> idle at the end; after a server restart four restored
gemini panes went busy -> idle in about 4 s; a fresh launch settled in 3 s.

The launch-settle and uncharacterised-CLI tests that used gemini as their
example of a CLI without work detection now use grok and deepseek.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 03:37:41 +02:00
Codeman maintainer f3b2080e69 fix(codex): see the background-terminal row under codex 0.162's hint row
A codex session waiting on a background terminal read as plainly idle,
with no "1 background terminal" badge. Codex pins that row above its
composer, and the registry looked for it in the last three non-blank
rows. Codex 0.162.0 added a hint row under the status line at rest
(`  ← for agents · ? for shortcuts`), which pushes the chip to FOURTH from
the bottom exactly when the idle probe reads it. Measured live on the
1.36.0 beta with `sleep 600` started as a background terminal: chip,
composer, status line, hint row; the server reported `watching: null`.
While a prompt is typed the hint goes away and the chip is third again.

The codex entry now declares `watchingLines: 4`. The trade is stated in
the entry: with no terminal running, the fourth row from the bottom is
the last transcript row (the last two while typing), which the agent
writes. As before, this is contained by codex having no hooks: a forged
row costs a wrong badge, never a silenced alert.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 03:29:37 +02:00
Codeman maintainer a58991e6d8 fix(pi): read the model off pi's footer so pi tabs and tiles name it
A pi session showed no model in its tab or tile unless one was passed at
launch, and none at all for the default route. pi draws its model in its
own footer, but its registry entry declared no modelDetect, so the pane
probe never read it.

The footer, from pi 1.1.0's footer code (0.84.4's is the same) and a live
pane (`0.8%/253k (auto)       qwen3.8-27b-pi • xhigh`): usage and context
on the left, then at least two spaces and `[(provider) ]<model>`, with
` • <thinking>` for a reasoning model and ` → <routed model>` when
routed. The last two rows are read (an extension status row can sit
below), and the context field picks the stats row out of them.

pi truncates the right side to fit a narrow pane with no ellipsis,
leaving exactly two spaces of padding. A name with nothing after it is
therefore read only with three or more spaces in front, and with two only
when a following ` •`/` →` proves it whole, so a cut-off name is never
shown. `no-model`, pi's placeholder, is rejected.

The read rides the idle confirmation pi gained with its workDetect entry.
Verified on an isolated instance: a fresh pi session published
qwen3.8-27b-pi (source screen) about 8 s after launch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 03:10:43 +02:00
Codeman maintainer 3fe5d1b278 fix(codex): read the model above codex 0.162's new footer hint row
Every codex tile and tab showed no model, unlike claude and deepseek.
Codex reports its model only in the footer under the composer
(`  GPT-6-Luna default · ~/codeman-cases/testcase`), and the registry read
the pane's LAST row for it. Codex 0.162.0 added a hint row under the
footer at rest (`  ← for agents · ? for shortcuts`, or `  ? for
shortcuts`), so the last row was always the hint and `displayModel`
stayed null. Measured live on the 1.36.0 beta: the hint is there at rest
and after a turn, and gone while a prompt is being typed (the footer is
the last row again then).

The codex modelDetect window is now two rows, and the footer must be
either the last row or followed by exactly one more two-space-indented
row. Anchoring to the end of the window keeps the guard the one-row rule
had: with the footer hidden, the last two rows are a transcript line and
the `›` composer, so a footer-shaped line the agent printed is not read.
Tests cover both 0.162 hint variants, the typing layout, the hint never
read as a model, and a forged footer-plus-indented pair above the
composer.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 03:10:36 +02:00
Codeman maintainer 112c533ac7 fix(omp): read omp's status bar so an omp turn can end
An omp session stayed "working" for good once a turn started, the same
latch pi had: omp's braille spinner trips the SPINNER_PATTERN fast path,
and only a composer glyph arms the idle confirmation. omp declared none,
so it fell back to Claude's `❯`, which omp never draws once its setup
wizard is done.

Measured on live omp 18.8.6 and 18.0.11 panes, holding a turn open
against a local endpoint that never answers: the input row is `╰─ <text>`
and is redrawn at submit, at the end of a turn, at launch and on
reattach. While a turn runs, the status bar's leading `π` becomes a
braille spinner plus the elapsed time (` ⠼ 14s > ⬢ model > ...`; 18.0.11
pads it with two spaces, past a minute it reads `1m`), with a
`⎋ Working…` row above it. The registry entry now names the input row as
the glyph and either working signal as the working line.

The glyph also switches the submit verifier on for omp, which reads the
input row the way it reads Claude's composer. A prompt sent mid-turn goes
to omp's Steering queue and clears the row, so the verifier stands down.
Text left in the row after an Enter is the one case it re-presses.

Verified end to end on a sandboxed instance (own HOME and PATH, omp
18.8.6): session:idle at launch, session:working during a turn,
session:idle about 3 s after it ended, and a restored pane settled idle
about 3 s after a server restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 02:43:22 +02:00
Codeman maintainer e7d661158b fix(opencode): read opencode's composer bar and footer spinner so a turn can end
An opencode session that ran a tool stayed "working" for good. The running
tool row draws a braille spinner (`⠋ Sleep for 12 seconds...`), which trips
the generic SPINNER_PATTERN and marks the pane working, but only a composer
glyph arms the idle confirmation and opencode declared none, so it fell back
to Claude's `❯`, which opencode never draws. Measured on an isolated
instance: a 16 s turn latched busy/isWorking for the rest of the session.
A text-only turn had the opposite problem and never showed as working.

Measured on live opencode 1.3.0 panes (capture-pane every 250-300 ms through
real turns at 40, 60, 120 and 200 columns, plus the raw PTY stream):

- Every composer row starts with a `┃` bar, and the submitted prompt lands in
  the transcript with the same bar, so a turn's first repaint arms the idle
  check, and tmux's reattach repaint does the same for a restored pane.
- While a turn runs the footer row starts with an 8-cell knight-rider
  spinner, `⬝■■■■■■⬝  esc interrupt`, redrawn about every 40 ms (largest
  gap mid-turn: 121 ms). At rest the TUI is silent and nothing on screen
  draws a `⬝`/`■` run, the 200-column sidebar included.
- The working line is the spinner run, `[⬝■]{8}`, not the label: tmux ships
  `esc` and `interrupt` as separate words joined by cursor moves, so the
  label never reaches the stream detector, and at 40 columns the footer
  wraps it to `esc` / `interr` / `upt`. All 344 spinner chunks of a turn
  match the run after Codeman's ANSI strip.
- A pending permission prompt replaces the composer and stops the spinner,
  so it reads as idle (waiting on the user).
- The last `┃` row on screen is the composer's agent/model row, or the
  permission box's closing bar, never the prompt text, so the submit
  verifier stands down and can never press Enter into a dialog.

Verified on an isolated instance from this branch: a 15 s tool turn emitted
exactly one session:working (+271 ms) and one session:idle (3 s after the
spinner stopped); a permission prompt read idle and the allowed turn went
working -> idle; after a server restart both restored opencode panes (one
at rest, one on a permission prompt) went busy -> idle in about 4 s; a
fresh launch reached an open page as idle in 3 s.

The launch-settle tests that used opencode as their example of a CLI
without work detection now use gemini and antigravity.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 02:40:04 +02:00
Codeman maintainer b539780f33 fix(pi): read pi's composer rule so a pi turn can end
A pi session stayed "working" for good once a turn started. pi's braille
spinner trips the SPINNER_PATTERN fast path, which marks the pane working,
but only a composer glyph arms the idle confirmation and pi declared none,
so it fell back to Claude's `❯`, which pi never draws. Measured on beta136:
an errored turn stayed busy/isWorking for 3+ minutes after pi was back at
rest.

pi has no composer glyph. Measured on a live pi 1.1.0 pane (capture-pane
every 250 ms through a turn): its composer sits between two `─` rules, and
while a turn runs it rewrites the top rule as `── ⠏ Working ───` on every
frame. The registry entry now names the rule as the glyph that arms the
check and a spinner frame inside it as the working line.

The same glyph settles a reattached pi pane (a restored pane gets no launch
timer): tmux's reattach repaint carries `─`, which arms the confirmation.
The submit verifier reads the last rule, finds no prompt text and stands
down, so it can never press Enter on a pi pane.

Verified on an isolated instance from this branch: a real pi turn emitted
session:working then session:idle, and after a server restart the restored
pi pane went busy -> idle in about 3 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-09 01:49:19 +02:00
RandalixandClaude Opus 5.5 5ba729fcbb fix(terminal): strip opencode's mouse DECSETs so a drag selects text again
opencode's TUI enables mouse tracking. tmux runs with `mouse off`, so it passes
the PANE's DECSETs straight through to the tmux client, and the browser's xterm
obeyed them: `mouseTrackingMode` flipped to 'any' (measured 62 none / 18 any over
16s) and xterm then reported DRAGS to the TUI instead of selecting locally.

In that state marking text produced no selection at all, so copy-on-select
silently did nothing (5/5 dead drags while `any`), and the obvious fallback —
Ctrl+C — is opencode's `app_exit`, which ended the session. Both were hit here.

opencode needs the middle strip: alt-screen toggles AND mouse DECSETs, but NOT
`3J` (a TUI is not a `clear` consumer). That is `altScreen: 'strip-mux-and-mouse'`
+ `isMuxMouseStripMode`, applied to the live stream (session.ts) and the replay
of a stored buffer, now the exported `stripReplayBuffer()` (session-routes.ts).

The browser's mouse-report gate keeps no mode list any more:
`_shouldReportMouseToCli()` reads only `cliMouseTracking`. The server sets that
flag solely in the mouse-strip branch (`_recordStrippedMouseMode`, one caller),
so it can only be true for a mode whose DECSETs are stripped, and whichever modes
the registry strips, the browser follows. Clicks still reach opencode through the
hand-encoded SGR tap it gates.

The `altScreen` JSDoc gets the decision table its three independent choices need
(alt-screen / `3J` / mouse DECSETs), written from the predicates, including that
`preserve` and `strip-mux-only` take the same runtime row. The table is pinned for
every stock CLI, with and without tmux, on both the live strip and the replay
strip, plus the published flag (test/claude-scrollback-strip.test.ts), so the two
halves cannot drift and a mis-ordered replay branch fails.

Docs and comments that still said opencode keeps its mouse reporting or gets the
narrow strip are updated (CLAUDE.md, architecture-invariants, scrollback and
copy-shortcut plans, session.ts, terminal-ui.js).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-08 14:55:02 +02:00
Codeman maintainer 9209922ea4 fix(deepseek): reject the dsh footer's non-model words instead of requiring a digit
The digit rule from 21ae48a5 hid the official DeepSeek ids (`deepseek-chat`,
`deepseek-reasoner` carry no digit), so a session on the official route with
the model field on showed the logo alone (its bundle row pins a provider
alone, so the config had nothing either). It also still misread a folder name
with a digit when every field before it was off.

Now the captured field is rejected when it is what the field can be when it is
NOT the model, and read otherwise:

- capabilities.modelDetect.rejectWords (registry data, single tokens, compared
  ignoring case; the schema bounds them and requires a screenLine). dsh lists
  every effort id its adapters offer (pi-ai THINKING_LEVELS plus the DeepSeek
  adapter's off/low/high/max) and the shipped mode ids, from dsh 0.1.1-rc.2 /
  dsh-TUI 0.10.0-beta.1. A mode's drawn label (`plan mode`, `full access`,
  CJK) can never be one captured field.
- In the shared screen reader, for every CLI: a field equal to the session's
  own working-directory basename is the folder, never the model.

Fixtures: `deepseek-chat` and `deepseek-reasoner` with the model field on are
read; every effort id, `default`, `plan mode`, and the folder name first (with
and without a digit) are not; the live qwen footer still reads `qwen3.8-27b`.
Known gaps, all off by default, are named in stock.ts: a custom mode id drawn
raw, a git branch or a one-word session title first, and the non-compact
footer layout (nothing read there; the route config applies).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 21:16:53 +02:00
Codeman maintainer 21ae48a5c8 fix(deepseek): the dsh footer field after a switched-off model is not the model
dsh-TUI draws the model as its status line's first field only while the
status bar's model field is on (the default). Switched off, the first field
is the reasoning effort (` medium · th-scratch`), else the mode, else the
cwd's basename, and the footer pattern read that as the model, which would
also outrank the route config added in the previous commit.

The captured field must now carry a digit, as a model id does (a version) and
an effort word, a mode name or most folder names do not. A model id without
one (`deepseek-chat`) is not read from the screen and the session falls back
to its route config: silent, never wrong.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 19:54:17 +02:00
Codeman maintainer 661fe3dc13 feat(sessions): a dsh session shows its route config's model while its screen names none
displayModel gains a `config` source, ranked below any report from the running
CLI and above the launch model: custom endpoint, then statusline or screen,
then config, then launch, then nothing. The screen still wins whenever it
names a model, since that is what the running TUI uses.

- Registry data: capabilities.modelDetect gains `configResolver`, a NAMED
  reader (src/model-config-resolvers.ts), like a launcher profile; dsh names
  'deepseek-route' (the reader from the previous commit). `screenLine` becomes
  optional; the schema refuses a modelDetect naming nothing, an unknown
  reader, or screenLines without a screenLine.
- Session: the reader runs from _withPaneLifecycle's finally, so at every pane
  start, attach and relaunch, with the session's own launch config
  (legacyConfigForMode) and env (its clamped overrides, then the server's), so
  a per-session DSH_HOME is the home read. Async; a read that lands after a
  newer one or after the session stopped is dropped; a remote or docker
  session reads nothing locally. A change emits displayModelChanged
  (broadcast and persist). Not restored after a restart: the next attach
  reads it again, and a restored screen value outranks it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 19:41:56 +02:00
Codeman maintainer 8392854619 feat(sessions): publish the model each session runs (displayModel)
A session header can only name the model a session runs if the server knows
it, so SessionState gains `displayModel: { model, source }`, resolved in a pure
module (src/session-display-model.ts), strongest first:

- custom-endpoint: a Custom Model Endpoint Profile's modelId answers the
  session, whatever alias the CLI prints;
- statusline / screen: the newest report from the running CLI itself.
  Claude's statusLine exporter already posts model.display_name on every
  render; the status-telemetry route now records it (only for a CLI with
  capabilities.statusLineTelemetry). A CLI whose registry entry declares the
  new capabilities.modelDetect has its footer read off the pane capture the
  idle/working probe already takes (no extra tmux call), so an in-session
  /model switch is followed at the next transition;
- launch: the model the session was launched with (claude's --model or the
  app-wide default, another CLI's <cli>Config.model), read where the registry
  says the model param lives;
- nothing known: no field, never a placeholder.

modelDetect is registry data, measured on live panes: dsh-TUI's status line
on the row under its composer (qwen3.8-27b on the owner's route) and codex's
`<model> <effort> ·` footer on its last row. Both anchor on chrome only that
CLI draws, over the last rows of the screen only; a transcript line shaped like
the footer is never taken (fixture tests). The pattern goes through
compileVersionRegex() with exactly one capture group, checked at load time.

An unreadable or covered footer keeps the last model (unlike the watching
label: a model does not stop running when something covers its row). Model
text is untrusted: escape sequences and control characters are stripped and it
is capped at 64 characters. A change emits displayModelChanged, broadcast
(session:updated) and persisted; a restart restores a CLI-reported model until
the next report.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 13:31:59 +02:00
Codeman maintainer b451b3851e fix(mcp): no file text in sync errors, follow relocated config dirs, docs and Settings polish (#521 review)
Maintainer merge-time fixes for the MCP server sync (opt-in mcpSyncEnabled, synced, default OFF).

M1, parse errors echoed config text (secrets included) into the HTTP response and Settings:
smol-toml's TomlError carries a code frame of the offending lines and V8's JSON "Unexpected
token" errors quote source. Both catch sites now go through describeMcpSyncError(): a parse
failure is reported by line/column only ("not valid TOML (line 3, column 21)", "not valid
JSON"), an errno failure by Node's own message (code, syscall, path), the module's own
messages via a McpConfigError class, anything else as "unexpected error". Tests put a secret
on the broken line (TOML, both JSON message shapes, and a write refused at the re-parse that
would have quoted a copied server's env) and assert it is absent from the result and from the
route's response body; they fail against the old code.

M2, CODEX_HOME / CLAUDE_CONFIG_DIR / XDG_CONFIG_HOME were ignored, so a sync could create a
file the CLI never reads and report success: new optional registry field
capabilities.mcpConfig.relocation { envVar, path } (registry data, no id branch; schema
reuses the env-name and no-traversal path rules). Declared for claude (CLAUDE_CONFIG_DIR,
checked in the 2.1.289 binary), codex (CODEX_HOME), opencode (XDG_CONFIG_HOME) and gemini
(GEMINI_CLI_HOME, gemini-cli paths.ts); antigravity follows $HOME only (agy 1.1.12 has no
relocation var). Resolved from the server process env at call time: absolute moves the file,
empty means unset, anything else reports the target with the new status "skipped" plus the
reason and writes nothing. Dedupe is now by resolved file. When a caller overrides `home`
without passing `env`, process.env is not consulted, and the route tests clear those vars so
a CI runner's XDG_CONFIG_HOME can never aim a write outside the temp HOME.

M3, feature undocumented: CLAUDE.md Key Patterns paragraph (opt-in, admin-only, additive
only, backups, re-parse validation, 0600 for copied secrets, names-only responses with
position-only parse errors, capabilities.mcpConfig and relocation), a Settings-Reference row
in the wiki, and docs/cli-registry.md + docs/api-reference.md updated for relocation, the
"skipped" status and the error policy.

Nits:
- N1 Preview/Sync before Save: the UI remembers the saved value on open and says "Save
  settings to turn MCP sync on first" instead of calling the routes; the 403 message also
  says to turn it on and save.
- N2 non-admins in multi-user mode: _applyMcpSyncAdminGate() hides the whole MCP group, called
  from applyMcpSyncVisibility() and the codeman:me event like the CLI-management gate.
- N3 scope chip says "synced".
- N4 "(1 servers)" pluralised; the unsupported list only names installed CLIs (route test
  pins it with a per-test installed set).
- N5 McpSyncResult / McpSyncTargetResult moved to src/types/mcp-sync.ts (barrel export); only
  the route imported them, so no churn.

Verified with an isolated instance (throwaway HOME, own instance and tmux socket) and
Playwright: chip, save-first message, preview rendering and the admin gate.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-04 23:52:41 +02:00
Codeman maintainer e082d8e438 Merge pull request #515 from irisitymichaelgrundberg/feat/codex-reasoning-effort
feat(codex): start a codex session at a chosen reasoning effort
2026-10-04 23:24:39 +02:00
Codeman maintainer 9649b5019b Merge pull request #521 from opticon454/feat/mcp-sync
feat(mcp): sync MCP servers across enabled CLIs

# Conflicts:
#	src/config/cli-registry/schema.ts
#	src/config/cli-registry/types.ts
#	src/web/public/settings-ui.js
2026-10-04 23:24:26 +02:00
DevvynandClaude Sonnet 5.5 2cf37529e9 fix(terminal): address review: Key tester isolates shortcuts, Codex stays on line feed
- app.js: the shortcut dispatcher returns early for events aimed at a data-raw-keys
  field, so Ctrl+W / Ctrl+L / Escape / Alt+1 / Ctrl+K pressed in the Key tester no
  longer kill the session, clear the terminal or close Settings
- stock.ts: drop Codex's esc-enter (a line feed works); no stock CLI declares a chord.
  The esc-enter path is tested through a clis.json override
- tests: unused port (3194), Ctrl+Enter asserts no keypress, shortcut-isolation test
  (verified to fail without the guard)
- docs/comments point at capabilities.newline; set-input class, trailing whitespace

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-03 11:19:42 +08:00
DevvynandClaude Sonnet 5.5 f39c66e4e8 feat(terminal): newline chord as registry data, plus a Key tester in Settings
capabilities.newline replaces choosing the Shift+Enter bytes in the send-key
route. Key tester shows the keydown/keypress/keyup a browser reports.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 21:53:31 +08:00
DevvynandClaude Sonnet 5.5 4398dbfad0 feat(mcp): make sync opt-in and address review
Opt-in (mcpSyncEnabled, default OFF; routes 403 until on). Review fixes:
- codex TOML read/validated with smol-toml: CRLF, inline tables and
  command-less tables no longer yield a duplicate [mcp_servers.x]; the new
  text is re-parsed before writing
- null-prototype tables and own-key checks; unsafe names ignored at every level
- servers switched off in their own CLI (codex/opencode/antigravity) are not copied
- only CLIs that are installed or already have a config file take part
- files receiving env/headers are left 0600; symlinked configs are written through
- one apply at a time (409), unique tmp files cleaned on failure, failed status
- routes set real HTTP status codes; api-reference section; format type single-sourced

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 21:15:31 +08:00
DevvynandClaude Sonnet 5.5 41a10b159e feat(mcp): add Antigravity, fix Gemini http/sse shape, report unsupported CLIs
Formats verified against real agy/gemini/codex mcp add output.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 18:27:07 +08:00
DevvynandClaude Sonnet 5.5 e6b258fc44 feat(mcp): sync MCP servers across enabled CLIs
Adds capabilities.mcpConfig to the CLI registry (Claude, Gemini, Codex,
OpenCode), an additive src/mcp-sync.ts, GET/POST /api/mcp-sync and a
Settings > Agents & CLIs control. Never edits or removes an existing
server; backs up each file it changes; reports conflicts.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-02 18:27:06 +08:00
Michael GrundbergandClaude Opus 5.5 45db24bacf feat(codex): start a codex session at a chosen reasoning effort
codexConfig takes a `reasoningEffort`, one of the levels codex accepts,
and the session starts with `--config model_reasoning_effort=<level>`.
The registry declares one literal per level, gated on the enum, because
an argv token cannot splice a value into a literal.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 16:59:46 +02:00
Codeman maintainer a0fbd1d28d docs(registry): note the cliMouseTracking half of claude's wheel rule (#498)
- claude's declared-for-later wheelForward says the live rule in
  _shouldForwardWheelToApp is the version AND the server-published
  cliMouseTracking flag, so whoever wires the field up needs both

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:28:29 +02:00
Codeman maintainer 272b56d47b fix(session): merge-time fixes for #491
- claude watchingLine: the lookahead keys on "Artifact" alone, so a
  footer truncated mid-chip ("1 Artifact…", "1 Artifact comm…") is still
  refused instead of reporting the shell beside it; comment follows
- test: both truncations return no watching label
- invariants: a chip that waits on a human never counts as watching, and
  the ^ anchor is what stops the retry past the chip

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:28:29 +02:00
Codeman maintainer 92921b9107 Merge pull request #491 from irisitymichaelgrundberg/fix/artifact-comment-monitor-needs-you
fix(session): alert for an agent waiting on artifact comments

# Conflicts:
#	src/config/cli-registry/stock.ts
2026-09-28 16:20:52 +02:00
Codeman maintainer 7659ca8b44 fix(session): keep a tab working while Claude waits for its own workers
When Claude hands work to an ultracode workflow or background agents, it
ends its own turn and closes it with `✻ Waiting for 1 dynamic workflow to
finish` instead of `✻ Brewed for 1m 18s`, then resumes by itself when the
workers report back. The pane sits quiet with the composer up, so the idle
probe called the session idle for the whole wait. At phone width the
workflow's progress row also drops its ticking timer, so nothing on screen
changes for minutes.

A new optional registry field, `capabilities.workDetect.awaitingLine`,
names that closing row, and `_probePaneWorking()` counts it as work.
Claude renders the row once from a snapshot and never redraws it, so the
same words stay on screen after the workers finish. `isAwaitingWorkers()`
therefore tests only the newest column-0 row directly above the composer,
never the whole pane and never the PTY stream; a follow-up turn always
puts rows of its own there. The column-0 anchor also keeps an agent from
holding its own tab busy by printing the sentence.

Verified against the live Mac mini pane that reported the bug (2.1.283),
and end to end on an isolated instance: an ultracode session running a
90 s workflow at 46 columns stayed busy through the wait and the
follow-up turn, then went idle 6 s after that turn closed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:18:27 +02:00
Michael GrundbergandClaude Opus 5.5 a9b48320a3 fix(session): alert for an agent waiting on artifact comments
An agent that publishes an artifact arms a monitor for its comments and
ends its turn. Claude Code shows that on the footer as `1 Artifact
comment monitor`, and #473 put that chip on the list of background work,
so the session counted as watching and its idle prompt opened already
acknowledged. Unlike every other chip on the list, that monitor waits
on the user: the agent hears nothing until somebody comments.

Claude's `watchingLine` now refuses any footer that carries the chip,
through a lookahead over the whole row, so a shell running beside the
monitor cannot report the session as watching either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 07:50:25 +02:00
DevvynandClaude Opus 5.5 0a52a99ca9 feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis

Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI

Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).

Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.

Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).

Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.

Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.

Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.

27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else

window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.

Fixed in two places:

- server.ts: after building `available`, intersect the nine real
  SessionMode ids against `enabledClis()`. git/cloudflared (utility
  binaries, not CLI registry entries) and deepseekBinary (a secondary
  installed-only flag for the "add a profile" affordance) are deliberately
  left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
  `window.__codemanCliAvailable` in place and refreshes the welcome screen,
  the mobile overview and an already-open Run menu, mirroring the existing
  `installDeepSeekProfile()` pattern for the same "injected once, needs an
  explicit patch" reason — without this half, the server-side fix alone
  still left every surface stale until the next reload.

New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.

Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap

Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:

- A new "Status as of 2026-09-22" section up top: what's actually
  implemented (verified by grepping the routes/schema/UI, not just trusting
  the commit messages), the availability-flag staleness bug found and fixed
  in this session (commit 0c77dd0a) with its devbox verification record, and
  one real outstanding gap.

- The outstanding gap: a custom CLI created via Phase 5's write API has no
  way to actually be launched. The Run menu is static per-mode markup with
  no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
  custom entry, confirm it can be launched" verify step was never actually
  exercised against this. Documented with two candidate fixes, neither
  started.

- Each phase's checklist flipped to [x] where confirmed present in the tree,
  Status lines updated from TODO to DONE, and the two originally-open
  questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
  marked resolved against what actually shipped.

No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* docs: add the CLI-registry deployment plan and the parked Copilot plan

Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:

- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
  series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
  is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
  name in docs/cli-enable-disable-plan.md's own header as a sibling plan
  tracked separately — kept for continuity, not active on this branch.

The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* fix(cli-registry): render enabled CLIs in launch surfaces

* test(cli-registry): update frontend branch guard

* fix(test): isolate suite from deployment environment

* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent

shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.

Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): hide shell's toggle entirely instead of greying it out

A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical

renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* style: prettier fixes from the master merge

* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels

Four gaps found verifying #476 against the #343 review trail:

- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
  (the nine per-CLI resolvers and the generic registry one) caches in its own
  closure, with a negative-cache backoff of up to 5 minutes, and nothing
  cleared them. invalidateCliExecutableResolvers(binaries) now drops those
  caches per binary; install (success or failure), create, edit and delete
  call it plus invalidateCliResolverCache(id). Before this, a CLI installed
  from Settings could fail to launch for minutes, and an edited custom entry
  kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
  ignoring the entry's searchDirs and the login-shell lookup that spawn and
  the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
  sit behind an explicit confirm; the confirm now names the exact command,
  which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
  of the word ("Claude", "Codex"). It uses the registry label again, which is
  identical to the old static table for every stock CLI (now pinned).

14 new tests; 9 of them fail against the previous head and pass here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs

Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
  does not parse or has group/world permission bits instead of overwriting it
  (isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
  existence/duplicate checks inside the serialized step, plus a unique tmp
  name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
  section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md

Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
  body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
  map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
  label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
  "Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* test(cli-registry): pin the #343/#347 findings #476 makes reachable

A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 01:48:26 +02:00
Codeman maintainer da6fa663e7 fix(terminal): merge-time fixes for the copy gutter strip (#469)
- stock.ts: claude is no longer the only entry declaring transcriptGutter;
  codex declares it too.
- architecture-invariants: the strip applies when the session's CLI declares
  a margin (not detection), and a note that it keys on the session's launch
  mode, not on what is running in the pane (a claude pane dropped to a shell
  still loses up to two columns; copyStripMargin is the escape hatch).
- render-index-html test: the gutter map is injected for a solo
  /session/:id render as well.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer 10f87428c3 fix(cli-registry): merge-time fixes for the run-button accents (#463)
- mobile.css: gemini and antigravity run/gear rules get `!important` like
  pi/omp/grok/deepseek, so the gear half no longer keeps the skin accent
  while the body takes the mode colour (two-tone button on the default skin).
- test/skin-themes.test.ts: static guard that every run mode with a base
  `.btn-toolbar.btn-run.mode-<id>` rule also has a resting rule inside the
  `html:not([data-skin="og"])` block; ids are derived from the stylesheet.
- stock.ts: grok's accent comment names zinc-300 (border/badge colour);
  gemini's accent is #8ab4f8 to match its tab badge and run-mode dot, noted
  as the one exception to the border-colour method.
- types.ts: "(below)" -> "(above)".
- docs/cli-registry.md, CLAUDE.md: `accent` is now measured, not transcribed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer 8536aaef7b Merge pull request #473 from irisitymichaelgrundberg/feat/session-watching-badge
feat(approvals): let a session watching its own background work keep quiet (#468)

# Conflicts:
#	src/config/cli-registry/stock.ts
2026-09-23 11:32:11 +02:00
Codeman maintainer 94b093b617 Merge pull request #469 from irisitymichaelgrundberg/feat/copy-dedent-pane-margin
feat(terminal): take the transcript gutter off a copy, at the width the CLI declares
2026-09-23 11:32:02 +02:00
Michael GrundbergandClaude Opus 5 8d45b92eba docs(codex): record that a sub-agent leaves no row to read
Measured on the same beta, codex-cli 0.154.0: a sub-agent started without
waiting outlives the turn exactly as a background terminal does — the sandboxed
process was still running — and codex shows nothing for it. The last rows of the
pane are the composer and the status line, and `Sub-agents running` belongs to
the on-demand `/subagents` panel rather than to the row above the composer.

So there is no second codex label to add. A codex session waiting on a sub-agent
reads as plainly idle, which misfiles nothing (codex raises no idle prompts) and
simply leaves that one kind of quiet unexplained until codex pins a row for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 17:35:56 +02:00
Michael GrundbergandClaude Opus 5 05c788ce9d fix(watching): close the review findings on the label and its window
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief)
found the trust boundary weaker than the comments around it claimed. Eleven
findings, all applied.

The two blockers were both about who can write the row the label is read from.
Claude's window covered two rows, and the second one is the status line, whose
command a session running with permissions bypassed can write into its own
`.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of
its own and silence its own idle alert. The default window is one row now, which
is the footer and nothing else, and the constant says why. Separately, the label
reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML,
which is an injection sink for any config-supplied pattern whose capture group is
permissive; it goes through escapeHtml() like every other untrusted string in
that file.

The Codex entry could not be fixed the same way, and now says so. Its row is
third from the bottom only while a terminal runs; with none running that slot
holds the last row of the transcript, so matching the complete row (with the
`/stop to close` tail, window narrowed to three) raises the bar without closing
it. What contains it is `hooks: 'none'`: no hook event from a codex session
reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an
alert. The registry comment, `docs/cli-registry.md` and the test all state that
rather than claiming a guarantee the code does not have.

Also from the review: the TUI header badge no longer counts an acknowledged
item, which was the same gate the classifier fix already went through and was
wrong for human acknowledgement too; the TUI approval card reads the quiet
reason and drops to a new `info` tone instead of asking for a reply; the badge
carries an aria-label, because the phone it was built for has no hover target;
the schema refuses `watchingLines` without a `watchingLine`; and the pattern and
its window are resolved together rather than one memoized and one not.

Documentation moved with it. The mechanism now lives in
`docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer,
`docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped
buzzing, and both that page and the changeset name the limitation neither did
before: a question asked in plain prose is not a dialog, so it is silenced along
with the false alarms while background work runs.

Verified live again after the narrowing, on an isolated beta: a Claude session
reported `1 monitor` and took its idle prompt acknowledged, and a Codex session
reported `1 background terminal` against the full-row anchor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 17:11:33 +02:00
Michael GrundbergandClaude Opus 5 ce80b7a212 feat(terminal): take the transcript gutter off a copy, at the width the CLI declares
Copying a paragraph out of a Claude Code or Codex pane puts that pane's own
two-column transcript gutter on the clipboard, so every pasted line arrives
indented. #451 shipped the trailing half of the copy clean and left the leading
half out, because deriving the width from the selection fires on 73% of ordinary
indented text and cannot tell a margin from content.

The width is DECLARED rather than derived. `capabilities.transcriptGutter` on
the CLI registry is a bounded integer; claude and codex each declare 2, measured
on live panes, and no other stock entry declares any, so a CLI whose transcript
layout nobody has measured is never touched. The server publishes the map as
`window.__codemanTranscriptGutter`, built by filtering `enabledClis()` on the
capability rather than by listing ids, and `_activeCliGutterColumns()` looks the
active session's mode up in it. The copy path reads no terminal buffer at all.

The declared width is a CEILING, not the answer: `clean()` strips the lesser of
it and the run every selected line shares. A block can therefore only shift as a
unit, the structure inside a selection survives by construction, and a selection
reaching column 0 loses nothing. That is what keeps a `git log` body at its own
four-space indent inside an agent's two-column gutter.

Codex was measured separately, because it renders nothing like Claude: it draws
boxes narrower than the pane and pushes its transcript into ordinary scrollback.
On a live 0.154.0 answer its `•`/`›`/`⚠` markers sit in the gutter, prose
continuations sit at 2, and a nested YAML block the model wrote rendered at
2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120, 160, 198, 235 and 282 columns
its indents were 0, 2, 4, 6 and 8 at every one, never 1. Copying that YAML out
of a live Codex pane now yields 0/2/4/6: gutter gone, nesting intact.

Two derived versions were built and measured first, and both are recorded in the
code because both looked correct:

- Painted trailing padding — a full-screen TUI writes real spaces across the
  unused part of a row, a shell leaves them never-written for xterm to trim —
  has no false positives and never over-stripped. It is also a function of pane
  WIDTH: the padding exists only while a rendered line stops short of the CLI's
  own layout width, and Claude's prose wraps to fill it. Dragging the same two
  prose rows of one live transcript at five window sizes, the share of padded
  rows ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298 columns, so the
  strip silently did nothing at every ordinary size while a corpus captured
  entirely at 282 columns said it worked.
- Taking the narrowest indent on the rows around the selection fires at every
  width and over-strips about 1% of selections, because a file listing inside
  the transcript can be the narrowest thing on screen.

Measured over 1,392,281 selections — every 1, 2, 3, 5, 10 and 20-row window of
real Claude screens replayed from live PTY streams at 100, 120, 160, 198, 235
and 282 columns — the declared width over-strips none, breaks no relative indent
and alters no text, and serves 100% of the selections whose own indent covers
the gutter. Verified end to end in a browser with a real mouse drag and a real
Ctrl+C: Claude and Codex panes paste flush at 123, 198 and 298 columns, a shell
pane is untouched at every one.

The strip sits behind `copyStripMargin` (App Settings, Selection & clipboard),
per-device and default ON: a display key, absent from the .strict()
SettingsUpdateSchema, read as `!== false` because the desktop branch of
getDefaultSettings() returns {}. The toggle is checked before the map.

Two review findings from #451, handled:

- The mid-row flag governs ONE line now. `range.start.x > 0` excludes only the
  first selected line, the one whose margin the mousedown genuinely cut off, so
  the same three rows no longer produce three different clipboard results.
- The reversed-drag finding does not reproduce on the pinned xterm.
  `getSelectionPosition()` reads `_selectionService.selectionStart`, whose
  getter returns `SelectionModel.finalSelectionStart`, and that swaps the pair
  when `areSelectionValuesReversed()` says so. A real upward mouse drag through
  chromium against xterm 6.0 reports the same range as the downward drag.
  `_normalisedSelectionRange()` keeps the ordering as a guard, because the model
  one layer down exposes the unnormalised fields under the same two names.

Tests: test/terminal-copy-clean.test.ts (64, up from 31), plus the injected
script stripped in test/server-index-title.test.ts. Every guard is pinned:
removing any one of seven reds at least one test, including declaring the wrong
gutter width. Full suite green, 7,861 passed, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 13:33:50 +02:00
Michael GrundbergandClaude Opus 5 64c288a683 feat(codex): read Codex's own background-terminal row
Codex states background work too, and it says so in a different place. Claude
writes `· 1 monitor ·` on the last row of the screen; Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its
composer, which puts that row third from the bottom once the status line and
the composer are counted.

So how far up the screen to look is now per-CLI data as well:
`capabilities.workDetect.watchingLines`, bounded to 1..8 by the schema, and
defaulting to Claude's two. That bound is the point. The window is half the
injection guard, since every row it adds is another row the agent itself may be
able to write, and the label is what silences an idle alert. The other half is
the anchor, and Codex's is ` · /ps to view`: chrome naming a slash command only
the CLI can offer, so a session that writes "I left 1 background terminal
running for you" into its own output matches nothing.

Measured against a live codex-cli 0.154.0 pane rather than read out of a
binary. The row appears when the terminal starts, follows the composer down as
the conversation grows, and is gone after `/stop`. Verified end to end on an
isolated beta: the session payload carried `watching: "1 background terminal"`
and the badge rendered with it, and both cleared when the terminal stopped. The
fixtures in the tests are that capture verbatim.

Codex has no hook signals, so no idle prompt and no false NEEDS YOU row: for a
Codex session this is the badge alone, which is the case the maintainer said a
registry field could cover and a hook never could. Cross-CLI tests pin that
neither pattern fires on the other's screen, and that a CLI declaring nothing
still reports nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 08:58:52 +02:00
Michael GrundbergandClaude Opus 5 74884a20eb feat(approvals): let a watching session keep quiet, and fix the tui gate
The badge alone left the row in NEEDS YOU, which is the thing the issue was
about. The fix is the alert that does not fire.

An idle prompt from a session that is watching its own background work now
opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to
`notePrompt()`, which sets `acknowledgedAt` and records why in a new
`acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has
always meant "the alert this prompt armed is spent", and the prompt itself
stays pending, answerable and available as Read My Mind context. A wrong label
therefore costs a card that does not blink, never an alert that was never
created.

Every surface follows from that. The broadcast carries the reason, so a live
page declines to arm the tab alert and raises no desktop notification. The push
is skipped, since a false alarm is hardest to ignore on a phone. A reloading
page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And
`classifySession()` now reads it too, which is a pre-existing bug fixed here:
acknowledging on one device cleared the alert everywhere except `codeman tui`.
It re-arms for free, because the next idle prompt supersedes the item and is
built fresh. Only `idle` is eligible, so a dialog that blocks the agent still
goes red whatever else it started.

The label is pane-derived and therefore prompt-injectable, so it is now read
from the last two rows of the screen only, with Claude's pattern anchored on
the `·` its footer joins items with, ANSI-stripped and length-capped at the
source. An agent that prints `· 1 monitor ·` into its own output finds no
match.

Verified on an isolated beta: a session that armed a monitor took its idle
prompt acknowledged with no alert on any surface, wore the badge, and showed
"quiet, watching 1 monitor" on its still-answerable card; the same session with
the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts`
pins both directions across all four surfaces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 08:10:05 +02:00
Michael GrundbergandClaude Opus 5 3f2cde2db7 feat(session): say when a session is watching its own background work
An agent that arms a monitor, backgrounds a shell or hands a task to a
cloud session is told to end its turn. The pane then falls quiet, Claude
Code's idle_prompt notification arrives a minute later, and every surface
files the session under NEEDS YOU with nothing for a human to answer.

Claude states what it is still running on the last row of its screen
(`⏵⏵ bypass permissions on · 1 monitor · ← for agents`). That row is now
`capabilities.workDetect.watchingLine` in the CLI registry, guarded by
compileVersionRegex() like every other config regex, and the idle probe
reads it off the capture it already takes: `watchingLabel()` in
session-activity.ts searches the last five lines only, so a session that
PRINTS "1 monitor" is not mistaken for one running it.

The label lands on Session.watching and rides toLightDetailedState() out
to every surface. The phone overview, the desktop home rail and the rich
sidebar rows wear it as a `watching` badge in the accent colour, beside
the state pill and never in place of it: an agent can arm a monitor and
ask a question in the same breath, and only the pill says which.

Verified end to end against a throwaway session on an isolated beta
instance: the payload carried `watching: "1 monitor"` once the turn
ended, the badge rendered next to a yellow `waiting` pill, and both
cleared when the monitor died.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 18:00:30 +02:00
DevvynandClaude Sonnet 5 d3ee9f23c2 fix(cli-registry): correct accent colours, and a real gemini/antigravity/omp rendering bug
Two related fixes, found while re-measuring stock.ts's `accent` field
against the actual rendered UI (docs/cli-registry.md flags this field as
"transcribed, not authoritative — re-measure before wiring one up"):

1. A real, user-visible bug: `.btn-toolbar.btn-run.mode-gemini`,
   `.mode-antigravity` and `.mode-omp` had no override rule inside the
   `html:not([data-skin="og"])` block, unlike codex/pi/grok/deepseek, which
   do. The generic `.btn-toolbar.btn-run` rule in that block resolves at
   higher specificity than the base sheet's per-mode pair, so all three
   rendered as plain claude-blue on every skin except `og` — including
   `daylight-blue`, which is the actual DEFAULT skin for a fresh install
   (index.html's pre-paint script), not an edge case. Added the three
   missing rules, sourced from each CLI's own already-designed og-skin
   colours (no new colours invented), mirroring the exact pattern
   pi/grok/deepseek already use. Also corrected the stale comment on the
   pi rule, which claimed this was still broken for gemini/antigravity.

2. `stock.ts`'s `accent` field was simply wrong for most CLIs — e.g. claude
   was registered as Anthropic's brand orange (#d97757) while its button
   renders blue, antigravity was registered purple while it renders cyan,
   pi was registered green while it renders pink. Measured each CLI's real
   `border-color` from its own `.mode-<id>` rule on the og skin (the
   cleanest single representative hex each entry's gradient resolves
   around) and corrected all 9 non-shell entries to match. `accent` has no
   reader yet (confirmed via the DECLARED_FOR_LATER guard test), so this
   changes no rendered output — it's a data-accuracy fix, matching the
   registry's own "transcribed, not authoritative" warning taken literally.
   Also fixed a false claim in types.ts's doc comment for the field
   ("CSS derives every per-CLI gradient from it via --cli-accent") — no
   such CSS variable exists anywhere in the codebase.

Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format:check/
check:public-assets/check:frontend-syntax all clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-21 12:23:19 +08:00
DevvynandClaude Sonnet 5 5fc391a47c fix(custom-model): address fourth pre-merge review + merge upstream master (Ark0N)
Merged upstream/master (22 commits: reboot-restore recovery feature,
terminal keycode229 recovery work, install.sh/CLI-catalog generator
changes, CHANGELOG/version bump to 1.30.0) into this branch. No
conflicts; git auto-merged every overlapping file (CLAUDE.md,
docs/api-reference.md, app.js, index.html, styles.css, routes/index.ts,
session-routes.ts, schemas.ts, server.ts).

Two required fixes from the latest review:

1. privilegedEnvKeys widening (stock.ts) changes behaviour outside this
   feature. The reviewer decided to keep both CLAUDE_CODE_MAX_CONTEXT_TOKENS
   and CLAUDE_CONFIG_DIR listed (types.ts's rule that every traffic-
   redirecting var this feature introduces must appear there stays
   literally true), and asked for the real consequences documented
   instead of hidden:
   - Corrected session-env-clamp.ts's fileoverview, which stated the
     opposite of what the code now does (reboot-restore's clamp call
     used to be able to strip nothing for claude; it now strips a
     persisted CLAUDE_CONFIG_DIR for a non-granted owner).
   - Corrected the rationale comments in stock.ts: privilegedEnvKeys
     has exactly one consumer (ownerClampedEnvKeys, feeding the
     generic envOverrides clamp on create/quick-start/reboot-restore),
     not the custom-model routes.
   - Added a CLAUDE.md line to the CLAUDE_CONFIG_DIR gotcha covering
     the admin-only-in-multi-user-mode and reboot-restore-strips-it
     consequences.
   - Added a "Claude multi-user clamp" test next to the existing
     DeepSeek/OMP ones, pinning the new stripping behaviour.

2. GET .../running-status (custom-model-routes.ts) no longer passes
   the raw llama-swap `cmd` field (the literal launch line, which can
   carry model paths and --api-key) to the browser -- the frontend
   only ever reads model/state, cmd exists solely for server-side
   parseCtxFromCmd() during discovery. Added a test asserting the
   response never contains cmd or a planted secret.

Also regenerated config/clis.stock.json and install.sh's catalogue
block (npm run generate:cli-catalog) to clear drift introduced by the
upstream merge, since it was failing the sync check.

Left to the reviewer, as they said they'd take at merge: the two
"comments pointing at removed code" cleanups, the two stale CLAUDE.md
counts, and the small items list (mode==='claude' frontend branch,
isCliAvailable() unknown-id gap, shared confirmed flag ordering,
one-shot cancel toast severity, pumpLlamaSwapLogTail buffer cap).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-19 18:05:51 +08:00
DevvynandClaude Sonnet 5 e2034177c5 fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.

Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
  published dependencies) into a scratch dir purely to read
  dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
  completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
  the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
  completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
  same endpoint. dsh's own error template ("DeepSeek API error (HTTP
  ${status})") reproduces the originally-reported
  "dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.

- New registry field `appendV1Suffix` (env kind only, deepseek's entry
  alone — claude/gemini must NOT get it, since claude was already
  confirmed working against the unmodified baseUrl). When set,
  buildCustomModelInjection runs endpoint.baseUrl through the same
  withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
  use, instead of writing it verbatim.

Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.

2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:41:19 +08:00
DevvynandClaude Sonnet 5 211b872335 feat(custom-model): skip Claude Code's first-run wizard on custom-model launches
A fresh, isolated CLAUDE_CONFIG_DIR (used to keep an injected API key
away from a stored claude.ai OAuth login) looks like a brand-new Claude
Code profile to the CLI, so it replays its ENTIRE first-run sequence on
every single launch: the theme picker, the security-notes screen, the
per-project "trust this folder?" dialog, and (running with
--dangerously-skip-permissions) a one-time bypass-permissions warning —
confirmed live, none of which a real, already-onboarded profile shows
again.

- New registry-declared env-kind field `skipFirstRunPrompts` (alongside
  apiKeyTrustFile, which it reuses) — claude's entry only, carried
  through buildCustomModelInjection (pure) into
  applyCustomModelInjection (IO).
- seedFirstRunOnboardingState(): merges hasCompletedOnboarding: true and
  this session's own projects[workingDir].hasTrustDialogAccepted: true
  into the same <configDir>/.claude.json the API-key trust file already
  writes to — other projects and other fields on this session's own
  entry are left untouched.
- seedSkipBypassPermissionsPrompt(): merges
  skipDangerousModePermissionPrompt: true into <configDir>/settings.json,
  a separate file, same corrupt-tolerant merge behavior.
- applyCustomModelInjection() gains an optional workingDir parameter,
  threaded from session.workingDir (dedicated apply route) /
  resolvedCasePath (quick-start route) — boot recovery omits it
  (a dialog already answered once needs no re-seed on the same,
  persisted isolated directory).

Tests added at the pure-builder, IO-wrapper (including merge-preserves-
other-fields and corrupt-file-tolerance cases), and existing directory-
listing assertions updated for the new settings.json file. Typecheck/
lint/format clean; full suite shows no new regressions (baseline
pre-existing Windows-environment failures unchanged, 8 more passing
tests than before — the ones added here).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 09:21:36 +08:00
DevvynandClaude Sonnet 5 97464bfa27 fix(custom-model): pre-approve the injected API key in the isolated Claude config dir
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.

Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 13:08:07 +08:00
DevvynandClaude Sonnet 5 0e8b1981af fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker:

1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still
   coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same
   config directory and warns about it (confirmed cosmetic - the API key
   wins for actual requests, verified via a real session's own API Usage
   Billing line). A custom-model claude session now gets an isolated
   CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty,
   no files written into it) so there is nothing to conflict with. projects
   is symlinked (junction on Windows) back into the real config dir so the
   response viewer, subagent windows and Read My Mind keep working for that
   session, best-effort.

2. Context-window overflow. Claude Code assumes a large default context
   window for a model id it doesn't recognise and never compacts, so a
   custom endpoint's real, much smaller context (verified live: a 400
   exceeding a 16384-token llama-swap model with a stock ~33.7K-token system
   prompt) silently overflows. Discovery now also learns each model's real
   context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx),
   but ONLY for a model llama-swap's own /v1/models response already marks
   status.value === 'loaded' - never an unloaded one, since llama-swap
   treats ?model= as a routing hint and probing an unloaded model risks
   triggering an actual, slow, GPU-swapping load as a side effect of
   read-only discovery. A server with no status field at all gets no
   enrichment rather than a guess; a model not probed this round keeps its
   previously-learned value until it disappears from the list entirely.
   Stored per model (CustomModelHost.modelContextLengths) and applied via a
   new contextLengthVar registry field, set to
   CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude.

Both new fields live on the existing env-kind customModelInjection
capability shape, declared only on claude's registry entry - every other
CLI's injection is unaffected (pinned by test).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 12:08:26 +08:00
Codeman maintainer 942bf37e48 fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a
custom OpenAI-compatible endpoint by injecting env vars or a config file
and restarting the CLI in place. Review of the apply path found four
things, two of them destructive. This lands all four plus the smaller
items from the same review.

1. Clearing a selection did not clear it. The injected vars reach the CLI
   via `tmux setenv`, which persists at the tmux-session level and is
   inherited by `respawn-pane` (measured: `setenv FOO bar` survived two
   successive `respawn-pane -k`), so deleting the keys from the session's
   envOverrides relaunched the CLI still pointed at the old endpoint, and
   for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just
   been deleted. `Session.setCustomModel()` now reports the removed keys,
   queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys`
   carries them into `applyEnvOverrides()`, which `setenv -u`s them before
   re-applying the live overrides, on the same path that already unsets
   the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket
   that `setenv -u HOME` hands the next respawn the global HOME back.

2. Applying a model to a local claude session killed the pane. The
   relaunch was `claude --session-id <id>` and Claude refuses an id that
   already has a transcript, and unlike the dead-pane respawn this one
   kills a working pane first. `restartCli()` now pins the live
   conversation id as the resume id for that respawn when the CLI's launch
   declares a `fallback` chain, which renders the same
   `--resume <id> || --session-id <id>` shape the docker and remote pane
   commands use. Gated on the registry shape, not the CLI id: an entry
   whose resume id is minted by the CLI itself never declares that chain.

3. pi, omp and grok wrote their config file and then launched without the
   `--model` that selects it, so the file was ignored. The registry entry
   now declares `customModelInjection.launchModel` (`custom/{modelId}` for
   pi and omp, grok's `[model.codeman-custom]` block name), the builder
   renders it, and `_withCustomModelLaunchModel()` applies it onto the
   respawn options through `legacyConfigField`, leaving the stored
   <Mode>Config untouched so a clear falls back to the user's own model.
   A model id the CLI's `model` token pattern cannot carry is refused
   with a 400 rather than silently dropped by the argv engine.

4. Remote (SSH) and Docker sessions reported `restarted: true` and changed
   nothing: their `restartCli()` reattaches the durable tmux rather than
   relaunching the agent, and the env lands on the local pane. Both are
   refused with a 400 until those paths are plumbed.

Smaller items from the same review:

- The selection survives a Codeman restart as the disk-only `__customModel`
  bookkeeping (endpoint, model, injected key NAMES, config dir, launch
  model; never the values, which carry the API key). Recovery re-derives
  the values from the endpoint store through the same apply path the route
  uses and keeps the bookkeeping even when the endpoint is gone, so a
  later clear still has keys to unset.
- Discovery goes through `webviewFetch()`, so the RESOLVED address is
  judged by the same egress guard the web-tab proxy uses, and `baseUrl`
  reuses `webviewUrlSchema` (http(s) only, no embedded credentials,
  link-local and cloud-metadata addresses refused). undici's `fetch failed`
  wrapper is unwrapped so the user sees the ECONNREFUSED underneath.
- `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session
  config dir 0700/0600 (pi and omp embed the key literally), and that dir
  is removed with the session.
- `PR.md` is gone from the repo root and the design doc moved to
  `docs/custom-model-endpoints-plan.md` with the LAN address and the
  personal name scrubbed; every reference follows. The guide's `authStyle`
  text matches the shipped schema (`bearer | api-key`, default `bearer`)
  and says that `customModelEndpointsEnabled` is read by nothing until
  the picker lands.
- `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts`
  (four real type errors fixed). It is not yet wired into `npm run typecheck`
  because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json`
  there is the one-line follow-up.

Tests: `test/session-custom-model-restart.test.ts` drives a real Session and
fails on the unfixed code for items 1 to 3; the route suite covers item 4
and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets
run before the overrides and that a shell-metachar key never reaches tmux.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:46:28 +02:00
Codeman maintainer 1e42cb4e2d Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses)
2026-09-14 23:46:27 +02:00
Ark0N e6e5a62d9b Merge pull request #380 from opticon454/feature/cli-catalog-consumers
feat(cli-registry): drive install.sh and the Docker agent image from the CLI catalogue
2026-09-14 15:55:08 +02:00