mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
3518af3a9f98e9e3d03c7fab024187dbfdc59c2a
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c0423bf560 | fix(omp): complete omp wiring in UI files, skill docs, and tests after rebase | ||
|
|
6261b6f655 |
feat(skill): spawn and drive DeepSeek Harness workers
The agent skill could spawn a worker in any mode, but it could only
DRIVE a claude one: every other CLI has neither a real end-of-turn
signal nor an answer to read, so the recipes route them through output
markers.
dsh has both halves now -- its harness reports idle/working/blocked to
Codeman, and the previous commit reads its transcript -- so it joins
claude as a mode the four verbs work on unchanged. `spawn_workers alpha
beta:deepseek` is a mixed fleet in one call, and `sendwait` / `last_text`
/ `delete_session` need no per-mode variant.
Preamble 1.20.0 (SKILL.md's §0 heredoc regenerated from it):
- `spawn_worker` grows a deepseek branch that gates on the harness
composer. ⚠️ Readiness there is NOT the stop signal: the harness
reports idle at BOOT ~300 ms before its composer paints (measured
2.26 s vs 2.56 s after spawn), so a send-and-wait fired straight after
quick-start resolves on the boot edge, reports a turn that never ran,
and strands the prompt in a pane not yet taking input. Waiting for the
composer also spends that edge, since signals are edge-triggered.
- `spawn_workers` takes `name[:mode]`, so a mixed fleet stays one
concurrent call. Case names still have to be unique -- the mode never
disambiguates two workers that would share a directory.
- `sendwait` asks for `wait:"stop,exit"` instead of the `wait:true`
default set. That set also carries `idle`, which for an external CLI is
inferred from output stabilization: on a dsh worker whose TUI repaints
rarely, the re-wait resolved in 0 ms with `signal:"idle"` on a turn
with three minutes left to run. It also makes a wrong mode loud -- the
modes that cannot deliver `stop` answer 400 before writing anything,
instead of resolving on a flap.
- The self-heal resend carries `delivered:true` forward. The resend is a
tagged duplicate, so the server truthfully reports `delivered:false`
about a write it skipped, and §1's cleanup then read a completed turn
as an undelivered one and kept a finished worker forever.
- dsh workers spawn with the permission posture the Run button sends,
because the harness default still asks and a worker parked on an
approval row cannot finish a fan-out. The multi-user clamp still
applies.
Docs: a worked dsh flow in recipes.md, readiness and the signal rules in
verbs.md, and the corrections this makes necessary -- `stop`/`blocked`
are no longer claude-only, and `last-response` is no longer permanently
empty for deepseek. The integration guide gains a section on reading a
session back and driving one as a worker; its web-UI section was also
stale (that server moved out of a shell session).
The static guard that keeps those lists from naming some external CLIs but
not others is extended rather than exempted: it now knows the three real
classes inside that family (no transcript, no hook signals, and the
positive twin -- the modes whose answers can be read), with the hook class
derived from `hooksAvailableForMode()` so the predicate and the prose
cannot drift apart. Any other partial list still fails, and a new backend
belongs to none of the classes until someone says so.
|
||
|
|
4cda150493 |
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/ antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI as a Codeman web tab. DeepSeek is wired unlike its siblings in three ways, each of which is the reason for a design decision rather than an accident: 1. The agent is a PROFILE, not the binary. `dsh` is a launcher over $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and `base` -- the interactive terminal front door is always a third-party plugin. So availability is two questions: `isDeepSeekAvailable()` (binary) and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run button gates on the latter, because reporting only the binary would spawn a pane that dies on arrival. When the binary is present but no profile is, the run menu offers to install one (POST /api/deepseek/install-profile). 2. The permission switch is an env var, not a flag. The harness has no command-line permission option; its sandbox/approval rows read DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access). Exported via `tmux setenv`, never on the spawn line. Absent = the harness's own workspace-write, which still asks, so the multi-user clamp is the only-if-sent branch and clamps to workspace-write, never read-only. 3. It is the only non-claude mode that passes hooksAvailableForMode(), and it earned that. The terminal front door reports idle/working/blocked to a supervising process over a generic env-gated contract; a generated shim (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each report to /api/hook-event as stop / agent_working / permission_prompt. So a dsh session gets definitive respawn triggers, real wait-endpoint signals and real Approvals Inbox items instead of output-stabilization guesswork. `agent_working` is new (157th SSE constant) and joins APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its alert at once. The resolver needs the strictest identity probe of the family: `dsh` is not merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell), so `dsh --help` must print the harness's own banner before a candidate is handed a spawn line. Model is deliberately not a session field -- it is a composition entry in the profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider keys named by a settings-file `apiKeyEnv` stay out, which is pi's 34-provider-key problem in a new shape. Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the status endpoint's two-part answer, the no-profile refusal, the profile bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the permission mode injected via setenv, and the full status bridge -- a send-and-wait returned signal "stop" from a real turn, and blocked/working created and cleared an Approvals Inbox item. Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md (decisions + honest gaps). Tests: test/deepseek-mode.test.ts, test/deepseek-cli-resolver.test.ts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f8c8e99d1 |
feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
1c94995290 |
docs(skill): hooks are a setting now, not who created the directory
The workspace-hooks install makes the skill's central hooks rule wrong in the cautious direction. Six places told a worker that a linked case or a raw workingDir has no `stop`/`blocked` and that send-and-wait cannot be trusted there, so an agent would hand-roll output-marker synchronization in exactly the workspaces where `wait:true` now works. Rewritten against the setting rather than directory provenance: - verbs.md §5.1: the where-to-spawn table, the rule paragraph (now naming `workspaceHooksEnabled`, default ON, the add-only merge, and the boot sweep of recovered sessions), and the silent-failure warning. The three cases that stay hook-less regardless are called out: remote SSH sessions, docker cases that opted out, and a workspace Codeman cannot write to. - verbs.md §5.3: the send-and-wait precondition is "the workspace has the hooks block", not "a case Codeman created". - endpoints.md: the Signals-by-mode table is now keyed on the setting, with rows for OFF, for remote/docker-opt-out, and for a session from an older server. The old create-path grep list becomes a "before 1.18.x" note. - SKILL.md §2 + the cost list, recipes.md Flow-1 contrast, messaging.md step 1. "Check, do not assume" is kept and promoted to the load-bearing habit, because the setting is not visible from the call and a session created by an older server that has not restarted still has nothing. The `spawn_worker` hooks grep STAYS: it guards the setting being off, remote sessions, and older servers. Only its diagnostic changes, since "pick an unused name" is no longer the fix. That text lives in both the §0 heredoc and `preamble.sh`, which `test/agent-skill.test.ts` pins byte-identical, so both are patched with the same bytes. Docs only, no behavior change. 23 skill tests green, full test:ci 5109 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f18097cb23 |
perf(skill): make the codeman skill spawn workers instead of deliberating
Measured against a live 1.18.1 server, the API does the whole job in about ten seconds: two cold claude workers spawned and ready in 6.3s, both tasked and both answers read in 4.0s more. The slowness users reported was agent-side. Three causes, all of them things the skill taught: - It taught serial spawning. Nothing in the main document showed `&`/`wait`, so "spawn two workers" read as "do the readiness ladder twice", which is one model turn per worker. - It had no spawn primitive. The happy path had to be reassembled on every run from where-to-spawn, a four-stage readiness ladder, send-and-wait, the fan-out caveats and a recipe with two variants. Each is a decision, and most carry a warning. - It cost ~16k tokens before the first call, at 3.6:1 prose to code, with 25 warning glyphs and 55 occurrences of "never". A document that is mostly failure modes teaches caution, and caution bills as thinking tokens. The preamble now defines the verbs rather than describing them: spawn_worker, spawn_workers (concurrent), sendwait, last_text. Section 1 composes them into the whole job in one Bash call and says to stop reading there. Two ceremonies the measurements retired: the pid poll (one iteration, 33ms, and wait-output already blocks on the composer) and reading settings.local.json to check hooks for a case quick-start creates, which always has them. That check stays required for linked cases and raw paths, where its absence silently breaks send-and-wait. The bootstrap's write condition now greps the version stamp, so a stale or truncated preamble self-heals rather than failing and asking for a manual rm. The stamp line is kept bare because the grep anchors on it with $; an inline comment there would rewrite the file on every bootstrap. Section 5 moved to reference/verbs.md behind an index, cutting the always-paid SKILL.md from ~16.4k to ~7.6k tokens. Section numbers and anchor slugs are unchanged, so existing references still resolve; all 201 anchors across the five files were checked, with the checker positive-controlled against an injected bad link. Verified by extracting the code blocks from the shipped file and running them against the live server: bootstrap plus full fast path, two workers resolving on the definitive stop signal, answers read and sessions deleted, in 6.8s. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |