mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
0a52a99ca918cf8341446672eaa22b2f655fd413
14
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0a52a99ca9 |
feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (
|
||
|
|
05c788ce9d |
fix(watching): close the review findings on the label and its window
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief) found the trust boundary weaker than the comments around it claimed. Eleven findings, all applied. The two blockers were both about who can write the row the label is read from. Claude's window covered two rows, and the second one is the status line, whose command a session running with permissions bypassed can write into its own `.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of its own and silence its own idle alert. The default window is one row now, which is the footer and nothing else, and the constant says why. Separately, the label reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML, which is an injection sink for any config-supplied pattern whose capture group is permissive; it goes through escapeHtml() like every other untrusted string in that file. The Codex entry could not be fixed the same way, and now says so. Its row is third from the bottom only while a terminal runs; with none running that slot holds the last row of the transcript, so matching the complete row (with the `/stop to close` tail, window narrowed to three) raises the bar without closing it. What contains it is `hooks: 'none'`: no hook event from a codex session reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an alert. The registry comment, `docs/cli-registry.md` and the test all state that rather than claiming a guarantee the code does not have. Also from the review: the TUI header badge no longer counts an acknowledged item, which was the same gate the classifier fix already went through and was wrong for human acknowledgement too; the TUI approval card reads the quiet reason and drops to a new `info` tone instead of asking for a reply; the badge carries an aria-label, because the phone it was built for has no hover target; the schema refuses `watchingLines` without a `watchingLine`; and the pattern and its window are resolved together rather than one memoized and one not. Documentation moved with it. The mechanism now lives in `docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer, `docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped buzzing, and both that page and the changeset name the limitation neither did before: a question asked in plain prose is not a dialog, so it is silenced along with the false alarms while background work runs. Verified live again after the narrowing, on an isolated beta: a Claude session reported `1 monitor` and took its idle prompt acknowledged, and a Codex session reported `1 background terminal` against the full-row anchor. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f2cde2db7 |
feat(session): say when a session is watching its own background work
An agent that arms a monitor, backgrounds a shell or hands a task to a cloud session is told to end its turn. The pane then falls quiet, Claude Code's idle_prompt notification arrives a minute later, and every surface files the session under NEEDS YOU with nothing for a human to answer. Claude states what it is still running on the last row of its screen (`⏵⏵ bypass permissions on · 1 monitor · ← for agents`). That row is now `capabilities.workDetect.watchingLine` in the CLI registry, guarded by compileVersionRegex() like every other config regex, and the idle probe reads it off the capture it already takes: `watchingLabel()` in session-activity.ts searches the last five lines only, so a session that PRINTS "1 monitor" is not mistaken for one running it. The label lands on Session.watching and rides toLightDetailedState() out to every surface. The phone overview, the desktop home rail and the rich sidebar rows wear it as a `watching` badge in the accent colour, beside the state pill and never in place of it: an agent can arm a monitor and ask a question in the same breath, and only the pill says which. Verified end to end against a throwaway session on an isolated beta instance: the payload carried `watching: "1 monitor"` once the turn ended, the badge rendered next to a yellow `waiting` pill, and both cleared when the monitor died. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c087d0ae4d |
fix(mobile): raise the phone breakpoint from 430px to 600px
The phone tier stopped at innerWidth < 430 and @media (max-width: 430px), so every current large phone landed in the tablet layout: the 430pt iPhone 14 Pro Max, 15 Plus, 15 Pro Max and 16 Plus, the 440pt iPhone 16 Pro Max and 17 Pro Max, Pixel 6 Pro, 7 Pro and OnePlus 12 Pro, the 448pt Pixel 8 Pro and 9 Pro XL, and the Galaxy Z Fold 5 cover screen at 460. On those devices the header icon row replaced the session pill, the toolbar kept the desktop Run Shell button instead of Enter and the mic, the keyboard accessory bar could never become visible because its .visible rule lives inside the phone block, and the toolbar jumped to the top of the page when the keyboard opened. The new cutoff is 600, the line test/mobile/devices.ts already draws between large phones (430-599) and small tablets (600-767). No physical device sits between 480 and 600, but a phone zoomed out one or two steps in Safari does: a 440pt iPhone at 85% or 75% page zoom reports 518px or 587px and still needs the phone controls, which a 480 cutoff would have taken away. The phone block is max-width: 599px and the tablet block starts at min-width: 600px, so a 600px device is a tablet in CSS and in getDeviceType() alike instead of straddling the boundary the way 430pt phones did. The number changes everywhere it is encoded: JS, CSS, comments, CLAUDE.md, the CI tests that pin the phone block, and the test:mobile helpers. Measurement history that names 430px stays as written. |
||
|
|
4f2dfb4e6d |
fix(mobile): carry resumeId through the phone overview's past rows
#386 made Codex conversations resumable from Past Sessions, and resumeMobileOverviewSession() correctly passes row.resumeId on to resumeHistorySession(). The phone's own row projection never copied the field off the unified-list item though, so row.resumeId was always undefined there and a tapped Codex row started a FRESH session on a thread that was already on disk. The desktop path worked; only the phone was blind. The test fails without the projection line, and pins the other half too: a claude row must not grow a resumeId, since the field is what distinguishes "resume this conversation" from "start a new one". Docs: CLAUDE.md and architecture-invariants both still described the unified list as merging Claude transcript files. It has been three stores since this PR (Claude's ~/.claude/projects, omp's ~/.omp/agent/sessions, codex's ~/.codex/sessions), the alias field keeps its Claude-era name without being Claude-only, and the scanner-only rule behind resumeId was written down nowhere. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4f5678fac4 | feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) | ||
|
|
4cda150493 |
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/ antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI as a Codeman web tab. DeepSeek is wired unlike its siblings in three ways, each of which is the reason for a design decision rather than an accident: 1. The agent is a PROFILE, not the binary. `dsh` is a launcher over $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and `base` -- the interactive terminal front door is always a third-party plugin. So availability is two questions: `isDeepSeekAvailable()` (binary) and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run button gates on the latter, because reporting only the binary would spawn a pane that dies on arrival. When the binary is present but no profile is, the run menu offers to install one (POST /api/deepseek/install-profile). 2. The permission switch is an env var, not a flag. The harness has no command-line permission option; its sandbox/approval rows read DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access). Exported via `tmux setenv`, never on the spawn line. Absent = the harness's own workspace-write, which still asks, so the multi-user clamp is the only-if-sent branch and clamps to workspace-write, never read-only. 3. It is the only non-claude mode that passes hooksAvailableForMode(), and it earned that. The terminal front door reports idle/working/blocked to a supervising process over a generic env-gated contract; a generated shim (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each report to /api/hook-event as stop / agent_working / permission_prompt. So a dsh session gets definitive respawn triggers, real wait-endpoint signals and real Approvals Inbox items instead of output-stabilization guesswork. `agent_working` is new (157th SSE constant) and joins APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its alert at once. The resolver needs the strictest identity probe of the family: `dsh` is not merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell), so `dsh --help` must print the harness's own banner before a candidate is handed a spawn line. Model is deliberately not a session field -- it is a composition entry in the profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider keys named by a settings-file `apiKeyEnv` stay out, which is pi's 34-provider-key problem in a new shape. Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the status endpoint's two-part answer, the no-profile refusal, the profile bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the permission mode injected via setenv, and the full status bridge -- a send-and-wait returned signal "stop" from a real turn, and blocked/working created and cleared an Approvals Inbox item. Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md (decisions + honest gaps). Tests: test/deepseek-mode.test.ts, test/deepseek-cli-resolver.test.ts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f8c8e99d1 |
feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
cdbde9f36f |
home-order follow-ups: fresh stamps for the blocked group, sort/display agreement, live Alt+N projection
Three follow-ups from the 1.19.0 review of the activity-ordered home screens: - Hook events now ride the same debounced session state broadcast the working/idle handlers use. The blocked group ranks on lastActivityAt, and without this a permission prompt raised after page load kept ranking by whatever stamp the browser loaded with. - A working row with no submit stamp now shows the lastActivityAt fallback its sort anchor already uses: a row must never be ranked by a number it does not display. - Alt+digit resolves through the live-session projection the render paints (sessionOrder minus dead ids), so a stale id cannot shift every painted number off its target, web tabs included. New tests pin both surfaces to one shared order and the numbering to the live projection. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
19aabe34d2 |
feat: sort both home-screen session lists by activity, not tab order
The phone overview and the desktop tab rail list the same sessions, so
they now share one order (CodemanSessionOrder in constants.js, pure and
unit-tested): blocked on you first (longest-blocked at the top), then
running longest-turn-first, then quiet most-recently-quiet first.
The tiebreak flips direction halfway down on purpose: for a state a
session is still in, longer is more urgent; for a state it has stopped
in, more recent is more relevant. The running group keys off the pane's
last Enter (lastSubmitAt), never lastActivityAt, because a working pane
repaints about once a second and would rank every turn as freshly
started. A 0 stamp means "unknown" and sorts last within its state.
The desktop rail was previously in raw tab order. Its number badge stays
the Alt+1..9 index, so on a sorted rail it deliberately no longer runs
1,2,3 downward: it names a shortcut, not a row position. Its second
stamp changes from "active 3m ago" to the state duration the order is
computed from ("created 1d ago . working 40m"), since both working rows
otherwise read "active just now" and the order looked arbitrary.
The tab strip itself is untouched: still user-ordered and drag-sortable.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
c5b59633d8 |
feat(pi): add Pi (pi.dev) as a sixth CLI run mode (#206)
SessionMode gains 'pi', a first-class backend alongside Claude Code, OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose tab identity, welcome button, run-mode entry, cron agentType, Docker and remote-SSH command defaults, and clone-repo Brain option. Pi is a different shape of CLI from the other four, and three decisions follow from that: - It has NO permission prompts and no sandbox, so there is no --dangerously-skip-permissions analog and none was invented. The privilege-shaped knob is the tri-state approveProjectTrust, which makes pi load and EXECUTE repo-local .pi/extensions TypeScript and install missing project packages. clampExternalCliBypassForOwner() therefore puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets --no-approve even when no config was sent, because pi's own default is a prompt the session user could answer themselves. That helper had zero test coverage; it now has coverage for all four CLIs. - Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode context, so admitting them would widen the allowlist for every mode at once. Auth goes through pi's /login or the server's own environment. --api-key is deliberately never wired: it would put a provider secret on the spawn command line. - pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the main screen with terminal-owned scrollback, and its 0.84.0 fullscreen mode is runtime-switchable via /settings; that flip was measured to put the pane into the alt screen, which the strip would have corrupted. pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires semver-shaped output, because `pi` is a short generic name a stray binary can shadow; GET /api/pi/status surfaces path and version so a misresolution is diagnosable rather than presenting as a broken mode. Docker installs pi in its own --ignore-scripts step so that flag cannot affect the other four CLIs, and seeds its credentials per-file rather than whole-dir (~/.pi/agent also holds sessions, extensions and package trees). Verified end to end against pi 0.84.1 on an isolated instance: resolver search-dir fallback, flag construction, piConfig persistence across a full server restart, the trust prompt and its --no-approve suppression, the rose Run button on the default daylight-blue skin (the nested skin block eats per-mode gradients unless the rule lives inside it), and the buffer local-echo policy, which pi tolerates where codex did not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cf3183abf7 |
chore: version packages
Release 1.16.6: phone overview started/idle stamps, plus fixes for the selection-dialog keyboard lockout, the accessory bar arrows bypassing the local-echo overlay, and recovered sessions being restamped as newly created on every server restart. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1ea39de650 |
fix(mobile): gate the phone overview's run picker on CLI availability
MOBILE_OVERVIEW_RUN_MODES / _buildMobileOverviewRunMenu is a separate, hardcoded duplicate of the toolbar's #runModeMenu (mobile-overview.js is a newer feature that mirrors the toolbar menu's look/behavior rather than reusing its render), so it never picked up #201's isCliAvailable() gating and offered every backend regardless of what the server actually has installed. Gate it the same way: skip an entry unless isCliAvailable(mode), shell always exempt. Added functional + static regression tests mirroring the toolbar menu's own test pattern. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
23f258a85d |
chore: version packages
Release 1.9.8 (aicodeman) and 0.1.8 (xterm-zerolag-input). Fixes macOS session start (`posix_spawnp failed.`, issues #6 and #204): node-pty ships its macOS spawn-helper as mode 0644 and macOS launches every PTY through it. `scripts/fix-node-pty.mjs` (npm run fix:node-pty) chmods every helper, prebuilds/ included, then verifies by really opening a PTY; the blind Node-22+ rebuild is gone. `spawnPtyWithHelperRepair()` self-heals an already broken install on the first failed spawn. Adds the phone home screen (session overview under 430px, per-device `mobileOverviewEnabled`, default ON) and a guided Tailscale path in install.sh, plus `install.sh tailscale` to retrofit it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |