mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-05 06:59:42 +02:00
b5d8122ea75d78f1b0fcccb23d3c32be49298817
94
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
54c591c84d | Merge origin/master (cron paste-mode Enter fix) into the landing branch | ||
|
|
2eece4f8f9 |
fix(cron): send a paste-mode prompt's Enter as its own write
A cron job in "Paste (direct)" input mode wrote `<text>\r` into the pane in one piece. Claude Code (measured on 2.1.283) takes a burst of about a hundred characters as a paste, so the `\r` landed as a newline and the prompt sat unsent on the composer while the run reported `prompt_sent`. Delivery now lives in `deliverCronPrompt()`. Paste mode writes the text raw, waits CRON_PASTE_ENTER_DELAY_MS (300 ms), sends `\r` as a separate write down the same PTY (so it cannot overtake the text), and arms the session's composer check through the new public `Session.verifySubmitted()`, which re-presses Enter while the prompt is still visibly unsent. A session with nothing to write to now fails the run instead of reporting the prompt as sent. Typed mode is unchanged. Verified on an isolated instance: a paste-mode job with a 104-character prompt submitted on the first Enter and Claude answered. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
a0fbd1d28d |
docs(registry): note the cliMouseTracking half of claude's wheel rule (#498)
- claude's declared-for-later wheelForward says the live rule in _shouldForwardWheelToApp is the version AND the server-published cliMouseTracking flag, so whoever wires the field up needs both Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
272b56d47b |
fix(session): merge-time fixes for #491
- claude watchingLine: the lookahead keys on "Artifact" alone, so a
footer truncated mid-chip ("1 Artifact…", "1 Artifact comm…") is still
refused instead of reporting the shell beside it; comment follows
- test: both truncations return no watching label
- invariants: a chip that waits on a human never counts as watching, and
the ^ anchor is what stops the retry past the chip
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
||
|
|
92921b9107 |
Merge pull request #491 from irisitymichaelgrundberg/fix/artifact-comment-monitor-needs-you
fix(session): alert for an agent waiting on artifact comments # Conflicts: # src/config/cli-registry/stock.ts |
||
|
|
7659ca8b44 |
fix(session): keep a tab working while Claude waits for its own workers
When Claude hands work to an ultracode workflow or background agents, it ends its own turn and closes it with `✻ Waiting for 1 dynamic workflow to finish` instead of `✻ Brewed for 1m 18s`, then resumes by itself when the workers report back. The pane sits quiet with the composer up, so the idle probe called the session idle for the whole wait. At phone width the workflow's progress row also drops its ticking timer, so nothing on screen changes for minutes. A new optional registry field, `capabilities.workDetect.awaitingLine`, names that closing row, and `_probePaneWorking()` counts it as work. Claude renders the row once from a snapshot and never redraws it, so the same words stay on screen after the workers finish. `isAwaitingWorkers()` therefore tests only the newest column-0 row directly above the composer, never the whole pane and never the PTY stream; a follow-up turn always puts rows of its own there. The column-0 anchor also keeps an agent from holding its own tab busy by printing the sentence. Verified against the live Mac mini pane that reported the bug (2.1.283), and end to end on an isolated instance: an ultracode session running a 90 s workflow at 46 columns stayed busy through the wait and the follow-up turn, then went idle 6 s after that turn closed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
a9b48320a3 |
fix(session): alert for an agent waiting on artifact comments
An agent that publishes an artifact arms a monitor for its comments and ends its turn. Claude Code shows that on the footer as `1 Artifact comment monitor`, and #473 put that chip on the list of background work, so the session counted as watching and its idle prompt opened already acknowledged. Unlike every other chip on the list, that monitor waits on the user: the agent hears nothing until somebody comments. Claude's `watchingLine` now refuses any footer that carries the chip, through a lookahead over the whole row, so a shell running beside the monitor cannot report the session as watching either. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
0a52a99ca9 |
feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (
|
||
|
|
da6fa663e7 |
fix(terminal): merge-time fixes for the copy gutter strip (#469)
- stock.ts: claude is no longer the only entry declaring transcriptGutter; codex declares it too. - architecture-invariants: the strip applies when the session's CLI declares a margin (not detection), and a note that it keys on the session's launch mode, not on what is running in the pane (a claude pane dropped to a shell still loses up to two columns; copyStripMargin is the escape hatch). - render-index-html test: the gutter map is injected for a solo /session/:id render as well. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
10f87428c3 |
fix(cli-registry): merge-time fixes for the run-button accents (#463)
- mobile.css: gemini and antigravity run/gear rules get `!important` like pi/omp/grok/deepseek, so the gear half no longer keeps the skin accent while the body takes the mode colour (two-tone button on the default skin). - test/skin-themes.test.ts: static guard that every run mode with a base `.btn-toolbar.btn-run.mode-<id>` rule also has a resting rule inside the `html:not([data-skin="og"])` block; ids are derived from the stylesheet. - stock.ts: grok's accent comment names zinc-300 (border/badge colour); gemini's accent is #8ab4f8 to match its tab badge and run-mode dot, noted as the one exception to the border-colour method. - types.ts: "(below)" -> "(above)". - docs/cli-registry.md, CLAUDE.md: `accent` is now measured, not transcribed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
8536aaef7b |
Merge pull request #473 from irisitymichaelgrundberg/feat/session-watching-badge
feat(approvals): let a session watching its own background work keep quiet (#468) # Conflicts: # src/config/cli-registry/stock.ts |
||
|
|
94b093b617 |
Merge pull request #469 from irisitymichaelgrundberg/feat/copy-dedent-pane-margin
feat(terminal): take the transcript gutter off a copy, at the width the CLI declares |
||
|
|
8d45b92eba |
docs(codex): record that a sub-agent leaves no row to read
Measured on the same beta, codex-cli 0.154.0: a sub-agent started without waiting outlives the turn exactly as a background terminal does — the sandboxed process was still running — and codex shows nothing for it. The last rows of the pane are the composer and the status line, and `Sub-agents running` belongs to the on-demand `/subagents` panel rather than to the row above the composer. So there is no second codex label to add. A codex session waiting on a sub-agent reads as plainly idle, which misfiles nothing (codex raises no idle prompts) and simply leaves that one kind of quiet unexplained until codex pins a row for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
05c788ce9d |
fix(watching): close the review findings on the label and its window
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief) found the trust boundary weaker than the comments around it claimed. Eleven findings, all applied. The two blockers were both about who can write the row the label is read from. Claude's window covered two rows, and the second one is the status line, whose command a session running with permissions bypassed can write into its own `.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of its own and silence its own idle alert. The default window is one row now, which is the footer and nothing else, and the constant says why. Separately, the label reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML, which is an injection sink for any config-supplied pattern whose capture group is permissive; it goes through escapeHtml() like every other untrusted string in that file. The Codex entry could not be fixed the same way, and now says so. Its row is third from the bottom only while a terminal runs; with none running that slot holds the last row of the transcript, so matching the complete row (with the `/stop to close` tail, window narrowed to three) raises the bar without closing it. What contains it is `hooks: 'none'`: no hook event from a codex session reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an alert. The registry comment, `docs/cli-registry.md` and the test all state that rather than claiming a guarantee the code does not have. Also from the review: the TUI header badge no longer counts an acknowledged item, which was the same gate the classifier fix already went through and was wrong for human acknowledgement too; the TUI approval card reads the quiet reason and drops to a new `info` tone instead of asking for a reply; the badge carries an aria-label, because the phone it was built for has no hover target; the schema refuses `watchingLines` without a `watchingLine`; and the pattern and its window are resolved together rather than one memoized and one not. Documentation moved with it. The mechanism now lives in `docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer, `docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped buzzing, and both that page and the changeset name the limitation neither did before: a question asked in plain prose is not a dialog, so it is silenced along with the false alarms while background work runs. Verified live again after the narrowing, on an isolated beta: a Claude session reported `1 monitor` and took its idle prompt acknowledged, and a Codex session reported `1 background terminal` against the full-row anchor. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ce80b7a212 |
feat(terminal): take the transcript gutter off a copy, at the width the CLI declares
Copying a paragraph out of a Claude Code or Codex pane puts that pane's own two-column transcript gutter on the clipboard, so every pasted line arrives indented. #451 shipped the trailing half of the copy clean and left the leading half out, because deriving the width from the selection fires on 73% of ordinary indented text and cannot tell a margin from content. The width is DECLARED rather than derived. `capabilities.transcriptGutter` on the CLI registry is a bounded integer; claude and codex each declare 2, measured on live panes, and no other stock entry declares any, so a CLI whose transcript layout nobody has measured is never touched. The server publishes the map as `window.__codemanTranscriptGutter`, built by filtering `enabledClis()` on the capability rather than by listing ids, and `_activeCliGutterColumns()` looks the active session's mode up in it. The copy path reads no terminal buffer at all. The declared width is a CEILING, not the answer: `clean()` strips the lesser of it and the run every selected line shares. A block can therefore only shift as a unit, the structure inside a selection survives by construction, and a selection reaching column 0 loses nothing. That is what keeps a `git log` body at its own four-space indent inside an agent's two-column gutter. Codex was measured separately, because it renders nothing like Claude: it draws boxes narrower than the pane and pushes its transcript into ordinary scrollback. On a live 0.154.0 answer its `•`/`›`/`⚠` markers sit in the gutter, prose continuations sit at 2, and a nested YAML block the model wrote rendered at 2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120, 160, 198, 235 and 282 columns its indents were 0, 2, 4, 6 and 8 at every one, never 1. Copying that YAML out of a live Codex pane now yields 0/2/4/6: gutter gone, nesting intact. Two derived versions were built and measured first, and both are recorded in the code because both looked correct: - Painted trailing padding — a full-screen TUI writes real spaces across the unused part of a row, a shell leaves them never-written for xterm to trim — has no false positives and never over-stripped. It is also a function of pane WIDTH: the padding exists only while a rendered line stops short of the CLI's own layout width, and Claude's prose wraps to fill it. Dragging the same two prose rows of one live transcript at five window sizes, the share of padded rows ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298 columns, so the strip silently did nothing at every ordinary size while a corpus captured entirely at 282 columns said it worked. - Taking the narrowest indent on the rows around the selection fires at every width and over-strips about 1% of selections, because a file listing inside the transcript can be the narrowest thing on screen. Measured over 1,392,281 selections — every 1, 2, 3, 5, 10 and 20-row window of real Claude screens replayed from live PTY streams at 100, 120, 160, 198, 235 and 282 columns — the declared width over-strips none, breaks no relative indent and alters no text, and serves 100% of the selections whose own indent covers the gutter. Verified end to end in a browser with a real mouse drag and a real Ctrl+C: Claude and Codex panes paste flush at 123, 198 and 298 columns, a shell pane is untouched at every one. The strip sits behind `copyStripMargin` (App Settings, Selection & clipboard), per-device and default ON: a display key, absent from the .strict() SettingsUpdateSchema, read as `!== false` because the desktop branch of getDefaultSettings() returns {}. The toggle is checked before the map. Two review findings from #451, handled: - The mid-row flag governs ONE line now. `range.start.x > 0` excludes only the first selected line, the one whose margin the mousedown genuinely cut off, so the same three rows no longer produce three different clipboard results. - The reversed-drag finding does not reproduce on the pinned xterm. `getSelectionPosition()` reads `_selectionService.selectionStart`, whose getter returns `SelectionModel.finalSelectionStart`, and that swaps the pair when `areSelectionValuesReversed()` says so. A real upward mouse drag through chromium against xterm 6.0 reports the same range as the downward drag. `_normalisedSelectionRange()` keeps the ordering as a guard, because the model one layer down exposes the unnormalised fields under the same two names. Tests: test/terminal-copy-clean.test.ts (64, up from 31), plus the injected script stripped in test/server-index-title.test.ts. Every guard is pinned: removing any one of seven reds at least one test, including declaring the wrong gutter width. Full suite green, 7,861 passed, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
64c288a683 |
feat(codex): read Codex's own background-terminal row
Codex states background work too, and it says so in a different place. Claude writes `· 1 monitor ·` on the last row of the screen; Codex pins `1 background terminal running · /ps to view · /stop to close` ABOVE its composer, which puts that row third from the bottom once the status line and the composer are counted. So how far up the screen to look is now per-CLI data as well: `capabilities.workDetect.watchingLines`, bounded to 1..8 by the schema, and defaulting to Claude's two. That bound is the point. The window is half the injection guard, since every row it adds is another row the agent itself may be able to write, and the label is what silences an idle alert. The other half is the anchor, and Codex's is ` · /ps to view`: chrome naming a slash command only the CLI can offer, so a session that writes "I left 1 background terminal running for you" into its own output matches nothing. Measured against a live codex-cli 0.154.0 pane rather than read out of a binary. The row appears when the terminal starts, follows the composer down as the conversation grows, and is gone after `/stop`. Verified end to end on an isolated beta: the session payload carried `watching: "1 background terminal"` and the badge rendered with it, and both cleared when the terminal stopped. The fixtures in the tests are that capture verbatim. Codex has no hook signals, so no idle prompt and no false NEEDS YOU row: for a Codex session this is the badge alone, which is the case the maintainer said a registry field could cover and a hook never could. Cross-CLI tests pin that neither pattern fires on the other's screen, and that a CLI declaring nothing still reports nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
74884a20eb |
feat(approvals): let a watching session keep quiet, and fix the tui gate
The badge alone left the row in NEEDS YOU, which is the thing the issue was about. The fix is the alert that does not fire. An idle prompt from a session that is watching its own background work now opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to `notePrompt()`, which sets `acknowledgedAt` and records why in a new `acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has always meant "the alert this prompt armed is spent", and the prompt itself stays pending, answerable and available as Read My Mind context. A wrong label therefore costs a card that does not blink, never an alert that was never created. Every surface follows from that. The broadcast carries the reason, so a live page declines to arm the tab alert and raises no desktop notification. The push is skipped, since a false alarm is hardest to ignore on a phone. A reloading page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And `classifySession()` now reads it too, which is a pre-existing bug fixed here: acknowledging on one device cleared the alert everywhere except `codeman tui`. It re-arms for free, because the next idle prompt supersedes the item and is built fresh. Only `idle` is eligible, so a dialog that blocks the agent still goes red whatever else it started. The label is pane-derived and therefore prompt-injectable, so it is now read from the last two rows of the screen only, with Claude's pattern anchored on the `·` its footer joins items with, ANSI-stripped and length-capped at the source. An agent that prints `· 1 monitor ·` into its own output finds no match. Verified on an isolated beta: a session that armed a monitor took its idle prompt acknowledged with no alert on any surface, wore the badge, and showed "quiet, watching 1 monitor" on its still-answerable card; the same session with the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts` pins both directions across all four surfaces. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f2cde2db7 |
feat(session): say when a session is watching its own background work
An agent that arms a monitor, backgrounds a shell or hands a task to a cloud session is told to end its turn. The pane then falls quiet, Claude Code's idle_prompt notification arrives a minute later, and every surface files the session under NEEDS YOU with nothing for a human to answer. Claude states what it is still running on the last row of its screen (`⏵⏵ bypass permissions on · 1 monitor · ← for agents`). That row is now `capabilities.workDetect.watchingLine` in the CLI registry, guarded by compileVersionRegex() like every other config regex, and the idle probe reads it off the capture it already takes: `watchingLabel()` in session-activity.ts searches the last five lines only, so a session that PRINTS "1 monitor" is not mistaken for one running it. The label lands on Session.watching and rides toLightDetailedState() out to every surface. The phone overview, the desktop home rail and the rich sidebar rows wear it as a `watching` badge in the accent colour, beside the state pill and never in place of it: an agent can arm a monitor and ask a question in the same breath, and only the pill says which. Verified end to end against a throwaway session on an isolated beta instance: the payload carried `watching: "1 monitor"` once the turn ended, the badge rendered next to a yellow `waiting` pill, and both cleared when the monitor died. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d3ee9f23c2 |
fix(cli-registry): correct accent colours, and a real gemini/antigravity/omp rendering bug
Two related fixes, found while re-measuring stock.ts's `accent` field
against the actual rendered UI (docs/cli-registry.md flags this field as
"transcribed, not authoritative — re-measure before wiring one up"):
1. A real, user-visible bug: `.btn-toolbar.btn-run.mode-gemini`,
`.mode-antigravity` and `.mode-omp` had no override rule inside the
`html:not([data-skin="og"])` block, unlike codex/pi/grok/deepseek, which
do. The generic `.btn-toolbar.btn-run` rule in that block resolves at
higher specificity than the base sheet's per-mode pair, so all three
rendered as plain claude-blue on every skin except `og` — including
`daylight-blue`, which is the actual DEFAULT skin for a fresh install
(index.html's pre-paint script), not an edge case. Added the three
missing rules, sourced from each CLI's own already-designed og-skin
colours (no new colours invented), mirroring the exact pattern
pi/grok/deepseek already use. Also corrected the stale comment on the
pi rule, which claimed this was still broken for gemini/antigravity.
2. `stock.ts`'s `accent` field was simply wrong for most CLIs — e.g. claude
was registered as Anthropic's brand orange (#d97757) while its button
renders blue, antigravity was registered purple while it renders cyan,
pi was registered green while it renders pink. Measured each CLI's real
`border-color` from its own `.mode-<id>` rule on the og skin (the
cleanest single representative hex each entry's gradient resolves
around) and corrected all 9 non-shell entries to match. `accent` has no
reader yet (confirmed via the DECLARED_FOR_LATER guard test), so this
changes no rendered output — it's a data-accuracy fix, matching the
registry's own "transcribed, not authoritative" warning taken literally.
Also fixed a false claim in types.ts's doc comment for the field
("CSS derives every per-CLI gradient from it via --cli-accent") — no
such CSS variable exists anywhere in the codebase.
Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format:check/
check:public-assets/check:frontend-syntax all clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
|
||
|
|
1a99b5836c | Merge pull request #430 from opticon454/custom-model-run-menu | ||
|
|
035bfbc2fe |
fix(remote): merge-time fixes for Wake-on-LAN
The MAC-count limit lived in two places that disagreed. RemoteHostSchema.wakeMac's 128-character cap admits seven comma-separated MACs while parseMacList takes at most four, all-or-nothing, so a five-MAC value validated, was written to remote-hosts.json, and then resolved to NO wake target: POST /api/sessions/:id/wake answered "No wake-on-LAN target configured for this host" and the banner offered "Configure WoL" for a host the user had just configured. MAX_WAKE_MACS now lives in src/config/remote-wake-limits.ts and both sides refine against it. Its own module because src/remote-wake.ts is import-fenced to session-routes.ts and server.ts (the wiring guard that stops a watcher waking a host), and because schemas.ts must not drag dgram/net/child_process into every request-validating module. The documented 40 s request budget also omitted the wake's own cost. A `command` target is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS and runs BEFORE the readiness poll, so a slow one pushed a wakeCommand host's worst case to ~68 s, past the 60 s proxy_read_timeout the budget exists to stay under. _wakeAndWait now subtracts the wake's measured elapsed time from the readiness budget, floored at one poll interval so a wake that ate the whole budget still gets one probe. A magic packet is effectively instant and is unaffected, which is why live testing never saw it. Also: the two new endpoints are documented in docs/api-reference.md with the import fence stated as the rule it is, CLAUDE.md's frontend module count moves to 34, and the release changesets carry the Thanks section. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5fc391a47c |
fix(custom-model): address fourth pre-merge review + merge upstream master (Ark0N)
Merged upstream/master (22 commits: reboot-restore recovery feature,
terminal keycode229 recovery work, install.sh/CLI-catalog generator
changes, CHANGELOG/version bump to 1.30.0) into this branch. No
conflicts; git auto-merged every overlapping file (CLAUDE.md,
docs/api-reference.md, app.js, index.html, styles.css, routes/index.ts,
session-routes.ts, schemas.ts, server.ts).
Two required fixes from the latest review:
1. privilegedEnvKeys widening (stock.ts) changes behaviour outside this
feature. The reviewer decided to keep both CLAUDE_CODE_MAX_CONTEXT_TOKENS
and CLAUDE_CONFIG_DIR listed (types.ts's rule that every traffic-
redirecting var this feature introduces must appear there stays
literally true), and asked for the real consequences documented
instead of hidden:
- Corrected session-env-clamp.ts's fileoverview, which stated the
opposite of what the code now does (reboot-restore's clamp call
used to be able to strip nothing for claude; it now strips a
persisted CLAUDE_CONFIG_DIR for a non-granted owner).
- Corrected the rationale comments in stock.ts: privilegedEnvKeys
has exactly one consumer (ownerClampedEnvKeys, feeding the
generic envOverrides clamp on create/quick-start/reboot-restore),
not the custom-model routes.
- Added a CLAUDE.md line to the CLAUDE_CONFIG_DIR gotcha covering
the admin-only-in-multi-user-mode and reboot-restore-strips-it
consequences.
- Added a "Claude multi-user clamp" test next to the existing
DeepSeek/OMP ones, pinning the new stripping behaviour.
2. GET .../running-status (custom-model-routes.ts) no longer passes
the raw llama-swap `cmd` field (the literal launch line, which can
carry model paths and --api-key) to the browser -- the frontend
only ever reads model/state, cmd exists solely for server-side
parseCtxFromCmd() during discovery. Added a test asserting the
response never contains cmd or a planted secret.
Also regenerated config/clis.stock.json and install.sh's catalogue
block (npm run generate:cli-catalog) to clear drift introduced by the
upstream merge, since it was failing the sync check.
Left to the reviewer, as they said they'd take at merge: the two
"comments pointing at removed code" cleanups, the two stale CLAUDE.md
counts, and the small items list (mode==='claude' frontend branch,
isCliAvailable() unknown-id gap, shared confirmed flag ordering,
one-shot cancel toast severity, pumpLlamaSwapLogTail buffer cap).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
|
||
|
|
e2034177c5 |
fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.
Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
published dependencies) into a scratch dir purely to read
dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
same endpoint. dsh's own error template ("DeepSeek API error (HTTP
${status})") reproduces the originally-reported
"dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.
- New registry field `appendV1Suffix` (env kind only, deepseek's entry
alone — claude/gemini must NOT get it, since claude was already
confirmed working against the unmodified baseUrl). When set,
buildCustomModelInjection runs endpoint.baseUrl through the same
withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
use, instead of writing it verbatim.
Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.
2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
211b872335 |
feat(custom-model): skip Claude Code's first-run wizard on custom-model launches
A fresh, isolated CLAUDE_CONFIG_DIR (used to keep an injected API key away from a stored claude.ai OAuth login) looks like a brand-new Claude Code profile to the CLI, so it replays its ENTIRE first-run sequence on every single launch: the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running with --dangerously-skip-permissions) a one-time bypass-permissions warning — confirmed live, none of which a real, already-onboarded profile shows again. - New registry-declared env-kind field `skipFirstRunPrompts` (alongside apiKeyTrustFile, which it reuses) — claude's entry only, carried through buildCustomModelInjection (pure) into applyCustomModelInjection (IO). - seedFirstRunOnboardingState(): merges hasCompletedOnboarding: true and this session's own projects[workingDir].hasTrustDialogAccepted: true into the same <configDir>/.claude.json the API-key trust file already writes to — other projects and other fields on this session's own entry are left untouched. - seedSkipBypassPermissionsPrompt(): merges skipDangerousModePermissionPrompt: true into <configDir>/settings.json, a separate file, same corrupt-tolerant merge behavior. - applyCustomModelInjection() gains an optional workingDir parameter, threaded from session.workingDir (dedicated apply route) / resolvedCasePath (quick-start route) — boot recovery omits it (a dialog already answered once needs no re-seed on the same, persisted isolated directory). Tests added at the pure-builder, IO-wrapper (including merge-preserves- other-fields and corrupt-file-tolerance cases), and existing directory- listing assertions updated for the new settings.json file. Typecheck/ lint/format clean; full suite shows no new regressions (baseline pre-existing Windows-environment failures unchanged, 8 more passing tests than before — the ones added here). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
97464bfa27 |
fix(custom-model): pre-approve the injected API key in the isolated Claude config dir
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.
Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
|
||
|
|
0e8b1981af |
fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker: 1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same config directory and warns about it (confirmed cosmetic - the API key wins for actual requests, verified via a real session's own API Usage Billing line). A custom-model claude session now gets an isolated CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty, no files written into it) so there is nothing to conflict with. projects is symlinked (junction on Windows) back into the real config dir so the response viewer, subagent windows and Read My Mind keep working for that session, best-effort. 2. Context-window overflow. Claude Code assumes a large default context window for a model id it doesn't recognise and never compacts, so a custom endpoint's real, much smaller context (verified live: a 400 exceeding a 16384-token llama-swap model with a stock ~33.7K-token system prompt) silently overflows. Discovery now also learns each model's real context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx), but ONLY for a model llama-swap's own /v1/models response already marks status.value === 'loaded' - never an unloaded one, since llama-swap treats ?model= as a routing hint and probing an unloaded model risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery. A server with no status field at all gets no enrichment rather than a guess; a model not probed this round keeps its previously-learned value until it disappears from the list entirely. Stored per model (CustomModelHost.modelContextLengths) and applied via a new contextLengthVar registry field, set to CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude. Both new fields live on the existing env-kind customModelInjection capability shape, declared only on claude's registry entry - every other CLI's injection is unaffected (pinned by test). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG |
||
|
|
942bf37e48 |
fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a custom OpenAI-compatible endpoint by injecting env vars or a config file and restarting the CLI in place. Review of the apply path found four things, two of them destructive. This lands all four plus the smaller items from the same review. 1. Clearing a selection did not clear it. The injected vars reach the CLI via `tmux setenv`, which persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so deleting the keys from the session's envOverrides relaunched the CLI still pointed at the old endpoint, and for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just been deleted. `Session.setCustomModel()` now reports the removed keys, queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys` carries them into `applyEnvOverrides()`, which `setenv -u`s them before re-applying the live overrides, on the same path that already unsets the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket that `setenv -u HOME` hands the next respawn the global HOME back. 2. Applying a model to a local claude session killed the pane. The relaunch was `claude --session-id <id>` and Claude refuses an id that already has a transcript, and unlike the dead-pane respawn this one kills a working pane first. `restartCli()` now pins the live conversation id as the resume id for that respawn when the CLI's launch declares a `fallback` chain, which renders the same `--resume <id> || --session-id <id>` shape the docker and remote pane commands use. Gated on the registry shape, not the CLI id: an entry whose resume id is minted by the CLI itself never declares that chain. 3. pi, omp and grok wrote their config file and then launched without the `--model` that selects it, so the file was ignored. The registry entry now declares `customModelInjection.launchModel` (`custom/{modelId}` for pi and omp, grok's `[model.codeman-custom]` block name), the builder renders it, and `_withCustomModelLaunchModel()` applies it onto the respawn options through `legacyConfigField`, leaving the stored <Mode>Config untouched so a clear falls back to the user's own model. A model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. 4. Remote (SSH) and Docker sessions reported `restarted: true` and changed nothing: their `restartCli()` reattaches the durable tmux rather than relaunching the agent, and the env lands on the local pane. Both are refused with a 400 until those paths are plumbed. Smaller items from the same review: - The selection survives a Codeman restart as the disk-only `__customModel` bookkeeping (endpoint, model, injected key NAMES, config dir, launch model; never the values, which carry the API key). Recovery re-derives the values from the endpoint store through the same apply path the route uses and keeps the bookkeeping even when the endpoint is gone, so a later clear still has keys to unset. - Discovery goes through `webviewFetch()`, so the RESOLVED address is judged by the same egress guard the web-tab proxy uses, and `baseUrl` reuses `webviewUrlSchema` (http(s) only, no embedded credentials, link-local and cloud-metadata addresses refused). undici's `fetch failed` wrapper is unwrapped so the user sees the ECONNREFUSED underneath. - `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session config dir 0700/0600 (pi and omp embed the key literally), and that dir is removed with the session. - `PR.md` is gone from the repo root and the design doc moved to `docs/custom-model-endpoints-plan.md` with the LAN address and the personal name scrubbed; every reference follows. The guide's `authStyle` text matches the shipped schema (`bearer | api-key`, default `bearer`) and says that `customModelEndpointsEnabled` is read by nothing until the picker lands. - `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts` (four real type errors fixed). It is not yet wired into `npm run typecheck` because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json` there is the one-line follow-up. Tests: `test/session-custom-model-restart.test.ts` drives a real Session and fails on the unfixed code for items 1 to 3; the route suite covers item 4 and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets run before the overrides and that a shell-metachar key never reaches tmux. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1e42cb4e2d |
Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses) |
||
|
|
e6e5a62d9b |
Merge pull request #380 from opticon454/feature/cli-catalog-consumers
feat(cli-registry): drive install.sh and the Docker agent image from the CLI catalogue |
||
|
|
dae2ac580f |
fix(terminal): read the colour env from the registry on every local spawn path
buildClaudeEnv(), the direct-PTY fallback taken when mux creation fails, now
reads getCli('claude').env and applies its unset and exports lists. It used to
delete COLORTERM and CLAUDECODE from a hand-maintained list of its own, which
left it contradicting the registry entry that the tmux pane and the attach
client both read. An engine value needing a mux name has nothing to resolve
against on this path, so it is skipped rather than guessed.
Claude no longer unsets NO_COLOR. The invisible-background bug does not need
it, and unsetting it overrides a preference the user set deliberately, so a
user who exports NO_COLOR globally keeps monochrome panes. The other seven
truecolor CLIs still unset it; that inconsistency is intentional and the
comment on the entry says so.
The invariants doc gains a Terminal colour env paragraph under Session launch
modes, where a reader looking up Claude will find it — the previous sentence
sat under a heading that lists only the non-Claude CLIs. It now says the lists
are the stock catalog and a clis.json override replaces them wholesale, and
that the declarations reach the tmux pane, its attach client and the direct
PTY but not a remote pane, whose command carries no env exports at all. Docker
hands COLORTERM=truecolor to every mode, including the two the registry says
must unset it.
The changeset named six peer CLIs and there are seven: deepseek also exports
truecolor. A test beside the existing OpenCode assertion pins the new
behaviour, so a future registry edit cannot make the backgrounds vanish again
in silence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7767b16d4f |
fix(terminal): let Claude use truecolor so its themed backgrounds render
Claude draws the user's own messages as a block of background color, and inside a Codeman pane that block was invisible. tmux hands each pane TERM=screen, which supports-color reads as 16 colors, and Claude's registry entry deleted COLORTERM on top of that. Claude therefore quantized every RGB color its theme asked for down to the basic palette, where rgb(55, 55, 55) and every other dark background becomes ESC[40m, the terminal's own black. Changing the color in a custom Claude theme moved nothing on screen. Claude now exports COLORTERM=truecolor and unsets NO_COLOR, matching codex, gemini, antigravity, pi, grok and omp. CLAUDECODE stays unset, because Claude reads it as a signal that it is running nested inside itself. Both the tmux session and the attach client read this one registry entry, so they cannot disagree. PR #3 introduced the unset in February, citing xterm.js#484 for the claim that xterm.js mishandles truecolor. xterm.js closed that issue in April 2019, Codeman now depends on @xterm/xterm 6, and TmuxManager already sets terminal-overrides ",*:Tc" on its own tmux server, so 24-bit color reaches the browser today for every CLI that asks for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a0628a40e8 |
fix(cli-registry): address maintainer review on #380
Rebased onto current master (the one real conflict was the import line
in docker-hosts.ts Ark0N flagged; kept both), then addressed every
point from the review:
**1. Rebase.** Done — this branch now sits on current upstream/master.
**2. Agent-image special cases are data now, not an id-keyed table
outside stock.ts.** `AGENT_IMAGE_SPECIAL_CASE_IDS`/`AGENT_IMAGE_SPECIAL_CASES`
are gone. `CliDiscovery.install.agentImageLayer?: { kind: 'dedicated';
reason: string }` is a field on the registry entry itself (pi,
deepseek), `reason` is required by schema.ts, both producers
(docker-hosts.ts and cli-catalog.mjs) filter on its presence instead
of an id, and the coverage test reads it from the generated catalogue.
Also added the npm-package-name validation to the TS producer, which
only the .mjs one had — same SAFE_PACKAGE regex, duplicated
(necessarily, one side can't import the other) and now pinned
byte-identical by a new parity test.
**3. Changeset said five, it's eight.** (Not nine — see the DeepSeek
point below, which changes the true count.) Reworded to state it
structurally rather than pin a number that will go stale again.
Then the four behavior-changing findings:
- **DeepSeek was offered as a normal install option but can't actually
drive a pane.** `npm install -g @deepseek-ai/dsh` installs the
launcher only; DeepSeek ships no profile that can run standalone.
The generator now emits an empty install command for any
`launcherProfile` entry, so install.sh's menu (which requires a
non-empty command) skips it and falls through to its docs URL hint
instead — matching what the old hand-written code did before this
PR replaced it.
- **wget-only hosts lost every automatic install, including the npm
ones that never needed curl.** The menu-building loop now filters
PER ENTRY (only a command starting with `curl ` is held back) rather
than wiping the whole menu when DOWNLOADER != curl.
- **The DISPLAY/TRUSTED split and the catalogue refresh didn't hold up
under review** (refresh's only real write was the label; it ran
before the Node existence check; its own eval-detection test was
tripped by the word "eval'd" in a comment). Dropped entirely per
your own recommendation — embedded catalogue only, no network
fetch, no second array. install-sh-invariants.test.ts now asserts
the refresh/DISPLAY machinery does not exist rather than testing its
internals.
The three take-or-leave items, applied:
- `dsh_banner_probe`'s bash 3.2 empty-array bug: `${runner[@]}` →
`${runner[@]+"${runner[@]}"}`. Verified live in a real `bash:3.2.57`
container with `timeout` removed from PATH — crashed before, clean
now, full `detect_all_clis` path exercised end to end.
- `docker-agent-image-coverage.test.ts` now anchors on each layer's
`<binary> --version` proof line instead of `Dockerfile.includes(binary)`,
which stayed true if a layer were deleted but its comment survived.
- Doc drift: docs/docker-cases.md (four → five, and now describes the
data field), docker/agent.Dockerfile's "other four CLIs" comment (no
longer a magic number — CLI_NPM_PACKAGES is generated and can grow),
CLAUDE.md's install.sh size (104KB → ~112KB) and its stale mention of
the now-dropped refresh.
Verified: tsc clean, prettier clean, the full targeted suite (142
tests across the 8 affected files) green, and the full `npm test` gate
diffed BY TEST NAME against a clean upstream/master baseline run on
this same machine — identical 201-name failure set both sides (168
tests / 67 files, all pre-existing Windows-environment noise: symlinks,
PTY spawning, POSIX permission bits — none of it touching anything
this PR changes), zero new failures either side of the diff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
|
||
|
|
b6f75b87f5 |
fix(custom-model): don't clamp DEEPSEEK_API_KEY as a privileged env key
CI caught a real regression: DEEPSEEK_API_KEY was added to deepseek's privilegedEnvKeys alongside DEEPSEEK_BASE_URL on the theory that "the pair travels together," but that contradicts the documented and tested design (clampEnvOverridesForOwner()'s own docstring in session-routes.ts) — a non-granted owner supplying their OWN DeepSeek key removes privilege rather than granting it, since the exfiltration vector is the BASE URL (which redirects the server's own forwarded key to a foreign host), not the key itself. Removed it from the list; test/deepseek-mode.test.ts's existing two clamp tests now pass again. Also swapped that test's "unrelated override" example off CODEX_HOME, which the earlier commit in this same PR legitimately made privileged (closing a real pre-existing gap, documented in PR.md) — so it stopped being a valid "unrelated" example the moment that fix landed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3 |
||
|
|
61779745aa |
test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
|
||
|
|
41416566aa |
feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. Covers local hardware (llama.cpp, Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter). Off by default (customModelEndpointsEnabled, synced, default OFF). - Registry: capabilities.customModelInjection per CLI entry (env / configContentEnv / configDir / unsupported kinds) - Pure injection builder (custom-model-injection.ts) turning an endpoint + model id into the real env vars / config content per CLI - Endpoint store + CRUD routes (custom-model-hosts.ts, custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded - Session integration: Session.setCustomModel()/restartCli() (POST /api/sessions/:id/custom-model), reusing the existing respawn-pane -k primitive to restart the CLI process with new env - Multi-user hardening: every new redirect-capable env var added to its CLI's privilegedEnvKeys, closing a pre-existing gap where several were already reachable via the generic envOverrides field's prefix allowlist - Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries against a real endpoint outside the web UI, independent of tmux/sessions - Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying every CLI's injected values through a real HTTP shape Real end-to-end validation against a live llama-swap server (inside a codeman/agent:llm-test Docker image with all 9 CLI binaries) found and fixed three real bugs before they shipped: - Codex's config.toml schema was wrong ([model].default table instead of a top-level model string + [model_providers.custom]); fixing it then surfaced a genuine, documented protocol incompatibility (Codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement) - Claude Code's async session-title-generation call validates ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and hangs the whole -p invocation on an unrecognized name; documented for chunk 6, worked around in the standalone script only (--bare is NOT safe for a real interactive session, which needs hooks) - The discovery route's authStyle: 'both' option (send both Authorization and api-key headers) reliably hung a real server; removed the option entirely rather than just changing the default Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see PR.md and deployment_plan.md for the full chunk breakdown and confidence table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3 |
||
|
|
d4fe3afc9d |
feat(files): raise the download cap to 2GB and stream /api/download
The 50MB cap on file-raw, the attachment /raw route and /api/download was memory protection for a `readFile()` that no longer exists: file-raw and /raw were rewritten to stream through `sendFileBody()` and answer Range requests, so size costs a read stream rather than RSS (measured: a 600MB download moved peak RSS by ~37MB). All the cap still did was refuse legitimate downloads of build artifacts, videos and archives. It is now MAX_FILE_DOWNLOAD_BYTES in config/buffer-limits.ts, default 2GB, env CODEMAN_MAX_DOWNLOAD_BYTES, 0 = unlimited. `parseByteLimitEnv()` is separate from the `parseInt(...) || default` idiom used elsewhere in that file precisely because that idiom reads 0 as falsy and would silently restore the default for the one value that means "no limit". /api/download was the last route that really did buffer the whole file. It now shares sendFileBody() with the other two, so it streams, advertises Accept-Ranges, and is resumable. Its Content-Disposition also goes through buildContentDisposition() rather than raw interpolation. Refusals move from 400 to 413 across all three, which is the correct status for the case; with the cap at 2GB it is a path almost nothing reaches now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f1b7283393 |
fix(cli-registry): guard workDetect.workingLine like every other config regex
#385 made the composer glyph and the working status line per-CLI registry data, which is right, but `workingLine` arrived as a config-supplied regex validated with a bare `new RegExp()`. That skips `compileVersionRegex()`, the helper the registry uses for exactly this: a `~/.codeman/clis.json` override can set the field, the compiled pattern is run against every accumulated PTY chunk and every pane capture, and a nested quantifier there backtracks on the event loop for the whole server rather than one session. Route it through the helper in both places, which are not redundant: the schema refine rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` compiles through the same helper so the runtime cannot hold a pattern the schema would have refused. The helper returns null instead of throwing, so the Claude-pattern fallback stops being a try/catch and becomes structural. Both shipped patterns compile unchanged, and Claude's is behaviourally identical to CLAUDE_WORKING_LINE_PATTERN. Also match the Codex footer case-insensitively on the E. It was characterised against codex-cli 0.152.1, which prints a lowercase `esc`; a version capitalising it would make the whole fix silently inert, since the pane would simply never look like it was working. Docs: CLAUDE.md, architecture-invariants and cli-registry.md all still stated the Claude-mode-only rule this PR retires, and none of them named the new capability or the regex guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a49be03f96 |
Merge pull request #385 from irisitymichaelgrundberg/fix/work-detection-external-clis
Merging with follow-up fixes applied on master: workingLine routed through compileVersionRegex() in both the schema refine and _workingLinePattern(), the Codex footer matched case-insensitively on the E, plus the doc passages that stated the retired Claude-mode-only rule. |
||
|
|
cfc8fe7e41 |
Merge pull request #381 from mtiller/feat/reverse-proxy-base-url
feat(web): support a reverse-proxy base URL |
||
|
|
51957e2ed4 |
fix(session): let each CLI declare how its own pane shows work
A Codex session reported `isWorking: false` for its entire life, including
mid-turn. Codeman has four paths that mark a session working, and all four were
inert for Codex:
- The spinner fast path tests eight braille frames, and Codex animates none.
- The activity-streak fallback was wrapped in `!isExternalCliMode(mode)`.
- The pane probe inside `_confirmIdle` would have matched, since Codex prints
`esc to interrupt`, but arming it required the literal glyph `❯` and Codex
draws `›` on its composer row.
- The text detector sat inside `_processExpensiveParsers`, whose first statement
returns early for an external CLI.
Add an optional `workDetect: { promptGlyph, workingLine }` to CliCapabilities,
so the two strings that differ per CLI are registry data rather than constants
in the detector. Claude declares its existing pair and behaves as before. Codex
declares `›` and `esc to interrupt`. The text detector moves above the
external-CLI early return, guarded on the descriptor so a CLI without one still
skips the ANSI strip that the early return used to save it.
A CLI that declares no descriptor falls back to Claude's pair, and the
activity-streak gate now reads "has a descriptor, or is not external", so the
plain shell mode keeps the behaviour it had.
Rewrite the test that asserted the old premise in its own comment, so it makes
the same guarantee for a genuinely uncharacterised CLI, and add Codex coverage
built from verbatim pane captures on Codex CLI 0.152.1.
|
||
|
|
3d8ffcb9a2 |
Merge pull request #357 from dignfei/feat/docker-adopt-existing-container
feat(docker): attach a case to an already-running container Conflicts came from work that landed after the PR was opened, and each is resolved onto the newer abstraction rather than by keeping the older code: - `defaultDockerCommandForMode` is registry-driven since #347, so the PR's `runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the refusal is visible only inside the container, so an adopted root container otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and `test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch. - The probe's mode list and its mode -> binary table both duplicated the registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is also what fixes the merge's silent regression: the hand-written list predates `omp`, and the run menu gates every docker case on this probe, so owned containers would have lost that mode. `shell` needs no arm — it declares no binary, so it is dropped from the lookup and reported available regardless. - The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one `missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto it. Its test now pins the single gate instead of counting seven arms. - The create arm keeps #349's swap-limit warning filter, which the adopted arm never reaches; the run-mode list gains `omp` from #353. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1 |
||
|
|
7e4914d991 |
feat(web): support a reverse-proxy base URL (--base-url / CODEMAN_BASE_URL)
Codeman can now be mounted under a sub-path behind a reverse proxy that forwards the prefix unchanged (e.g. https://host/codeman/). Default is `/` (root), which is byte-identical to the historical behavior. Design — few choke points, mirrored ingress/egress: - src/config/base-path.ts: pure single-source normalize/validate/join/strip. - Server ingress: stripBasePath() inside Fastify rewriteUrl, so routes stay declared prefix-agnostic; un-prefixed requests (hooks, health, docker bridge hitting the raw port) pass through unchanged. - Server egress: one onSend hook prepends the base to root-absolute Location headers (covers all redirects). - HTML: renderIndexHtml points <base href> at the mount and injects window.__CODEMAN_BASE__ — ONLY when a base is set (inert at root). - Frontend runtime URLs: CodemanBase.url() route builder in constants.js, applied transparently by a fetch wrapper and explicitly at the EventSource/WebSocket/window.open/<img|iframe|a>-src sites. - sw.js derives its base from self.location; manifest uses relative start_url/scope. - Web-tab proxy: proxyPrefixFor(cap, basePath) is the single base-aware root that cascades to the injected <base>, HTML/attr rewrites, runtimeUrlShim, Set-Cookie Path and Location; capabilityFromReferer strips the base off the browser Referer, while the ingress parsers stay base-agnostic (rewriteUrl already stripped it). --base-url rides the daemon relaunch (buildWebArgs) and the service unit (resolveServicePlan). constants.js is guarded against a missing `window` for isolated unit-test contexts. Tests: test/base-path.test.ts (pure helpers), base-path coverage in webview-proxy/render-index-html/daemon-control; CodemanBase stubbed in the vm-isolated panels-ui test contexts. Docs: Remote-Access.md (sub-path section + nginx example), security-architecture.md env table, CLAUDE.md pattern. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XUkPBxbumnct6qSrx4JDju |
||
|
|
a81e87f440 |
fix(cli-registry): log why clis.json was ignored, and say 0600 when that is the rule
The loader refuses a `clis.json` with any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored. That is a defensible posture for a file that chooses the binaries Codeman spawns, but two things around it made the override feature look dead: the warning said "group/world-writable", which a 0644 file is not, and `LoadResult.warnings` was returned to a caller nobody wired up, so nothing anywhere printed it. A user following the docs got silence. The message now names the rule and the command that satisfies it, the loader logs every warning once on first load (the result is memoized, so once per process), the module header stops claiming that nothing ever writes (the quarantine rename of a malformed file is a write, on first use) and the registry doc gains a short section on the override file with the 0600 requirement in it. Whether the check should relax to writable bits only is a separate decision; this keeps the shipped behaviour and makes it visible. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu |
||
|
|
c5b84fb5f4 |
docs(cli-registry): annotate overlays.credStore as declared-for-later
Review item 4 named THREE live tables duplicating registry data. Two are now read from the entry (`defaultRemoteCommandForMode`, `defaultDockerCommandForMode`); the third, `resolveDockerCredentialArtifacts`, is not — and it was left neither wired nor annotated, which is the state that item explicitly rules out. It is not wired because the shape cannot express the live table: `credStore` is ONE store per CLI, and `CRED_STORES` needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek's entry declares none at all even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is at once the worst thing in that file to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest. It belongs in its own change, measured against a real container. So it is annotated instead, at the field, in the type's declared-for-later header, in docs/cli-registry.md, and in the pinned DECLARED_FOR_LATER list — the last of which means wiring it later makes a test fail rather than leaving a stale comment behind. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
4830e662f9 |
refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. Code that used to ask "which CLI is this?" reads the entry instead. Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every spawn command as a literal string, captured from the hand-written builders before they were deleted, and `test/location-overlay-commands.test.ts` does the same for all 20 remote and in-container pane commands. Config can never contain shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only in this release. OMP is included as a registry entry rather than a tenth hand-written builder, so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of `buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen and doctor ladders all drop out. Guard rails: - `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id branching reappears outside `stock.ts`, in any of its four shapes (`===`, `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the negated forms, which is how 36 of them survived an earlier pass. Every allowlisted branch carries its reason. - `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities; deriving one from another shipped the `until=stop`-hangs-on-shell bug. - `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config` wire field is separate, bridged only by `legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. - Registry data resolves AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks). A module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. - Six fields are annotated DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/ `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured. A test pins the list so it cannot quietly grow. Three user-visible changes, all deliberate and named: - `probeDockerCliVersion()` derives the in-container binary from the registry rather than assuming it equals the mode name (`antigravity` runs `agy`). - The remote CLI version probe now covers grok and deepseek, which the hardcoded map it replaces omitted while its own comment said the rule was "every mode except shell". - `codeman doctor`'s CLI rows are generated from the entries, so Claude's install hint is the install command rather than a docs URL, five CLIs gain hints they never had, and the row order follows the catalog. Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars (matching the `cliId` pattern) before its failure message quotes the value back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading the hand-editable `clis.json`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
2a32b5064a |
Merge pull request #349 from opticon454/feature/docker-compose
Docker Compose deployment: Codeman runs in a container and spawns Docker cases as SIBLING containers through the mounted host socket (Docker-outside-of-Docker). Resolved the README conflict (master had grown to eight CLIs since the branch was cut) and moved the Compose blurb out of the feature bullets into Quick Start, next to the other ways of starting Codeman. Three review findings from the PR discussion are fixed here rather than left for a follow-up, because two of them are shipped-image problems: - `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against the whole context-relative path, so `docker/.env` — which the deployment's own README tells the user to fill with CODEMAN_PASSWORD and provider API keys — was picked up by `COPY . .` and baked into the image at /opt/codeman/docker/.env. Verified in both directions against a real build context: with a canary secret in docker/.env, the unfixed ignore file lets /ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for the checked-in .env.example) leaves only the example behind. - `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which still hardcoded ~/codeman-cases, so `codeman skill install --case <name>` reported "Case not found" on exactly the deployment the override exists for. Both now resolve through config/cases-dir.ts. state-store.ts keeps its own literal on purpose: that one migrates the historical ~/claudeman-cases directory by name and is about the old default, not the active location. - CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore joins the documented list of files that genuinely belong in the repo root. The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/ deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's lastClaudeSessionId was handed to every non-claude CLI. Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax, public assets, lockfile. |
||
|
|
9841f4ffb9 |
refactor(omp): align omp resolver + doctor with upstream shared CLI resolver
- omp-cli-resolver.ts already uses createCliExecutableResolver; add dedicated test/omp-cli-resolver.test.ts mirroring pi's (version-probe accept/reject, negative-cache backoff, VITEST hermeticity gate) - dependency-registry omp entry now requires OMP_VERSION_REGEX match like pi, so codeman doctor and the run-mode resolver agree on what counts as installed - system-routes /api/omp/status surfaces version |
||
|
|
4f5678fac4 | feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) | ||
|
|
4cda150493 |
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/ antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI as a Codeman web tab. DeepSeek is wired unlike its siblings in three ways, each of which is the reason for a design decision rather than an accident: 1. The agent is a PROFILE, not the binary. `dsh` is a launcher over $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and `base` -- the interactive terminal front door is always a third-party plugin. So availability is two questions: `isDeepSeekAvailable()` (binary) and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run button gates on the latter, because reporting only the binary would spawn a pane that dies on arrival. When the binary is present but no profile is, the run menu offers to install one (POST /api/deepseek/install-profile). 2. The permission switch is an env var, not a flag. The harness has no command-line permission option; its sandbox/approval rows read DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access). Exported via `tmux setenv`, never on the spawn line. Absent = the harness's own workspace-write, which still asks, so the multi-user clamp is the only-if-sent branch and clamps to workspace-write, never read-only. 3. It is the only non-claude mode that passes hooksAvailableForMode(), and it earned that. The terminal front door reports idle/working/blocked to a supervising process over a generic env-gated contract; a generated shim (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each report to /api/hook-event as stop / agent_working / permission_prompt. So a dsh session gets definitive respawn triggers, real wait-endpoint signals and real Approvals Inbox items instead of output-stabilization guesswork. `agent_working` is new (157th SSE constant) and joins APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its alert at once. The resolver needs the strictest identity probe of the family: `dsh` is not merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell), so `dsh --help` must print the harness's own banner before a candidate is handed a spawn line. Model is deliberately not a session field -- it is a composition entry in the profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider keys named by a settings-file `apiKeyEnv` stay out, which is pi's 34-provider-key problem in a new shape. Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the status endpoint's two-part answer, the no-profile refusal, the profile bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the permission mode injected via setenv, and the full status bridge -- a send-and-wait returned signal "stop" from a real turn, and blocked/working created and cleared an Approvals Inbox item. Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md (decisions + honest gaps). Tests: test/deepseek-mode.test.ts, test/deepseek-cli-resolver.test.ts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f8c8e99d1 |
feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|