diff --git a/CLAUDE.md b/CLAUDE.md index 8803833e..e5a4bf9b 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -159,7 +159,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph | **Search** | `src/search-service.ts` | Pure in-memory core for `GET /api/search` | | **Attachments** | `src/attachment-registry.ts`, `attachment-magic`, `generated-artifact-attachments`, `session-attachment-history`, `document-preview-cache`, `document-thumbnailer`, `document-conversion-limiter`, `config/attachment-guard` | See Key Patterns | | **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/` (`claude-md.ts` + `case-template.md`) | `templates/` holds the CLAUDE.md scaffold generated into new cases | -| **Web** | `src/web/server.ts` ★, `sse-events.ts`, `routes/*.ts` (20 modules + barrel; `session-routes.ts` ★), `route-helpers.ts`, `ports/*.ts`, `middleware/auth.ts`, `schemas.ts`, `self-update.ts`, `plan-usage-latest.ts`, `ws-connection-registry.ts`, `heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` | | +| **Web** | `src/web/server.ts` ★, `sse-events.ts`, `routes/*.ts` (21 modules + barrel; `session-routes.ts` ★), `route-helpers.ts`, `ports/*.ts`, `middleware/auth.ts`, `schemas.ts`, `self-update.ts`, `plan-usage-latest.ts`, `ws-connection-registry.ts`, `heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` | | | **Frontend** | `src/web/public/app.js` (~5K lines, core) + 28 modules + `sw.js` | See Frontend section for the load order, which is authoritative | | **Types** | `src/types/index.ts` (barrel) → 20 domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts | @@ -212,6 +212,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Read My Mind intent profiles** (phase 1 of `docs/readmymind-plan.md`; `readMyMindEnabled`, SYNCED, default OFF): per-CASE profiles (user-stated `goals` + the user's recent real prompts), keyed by owner + realpath(workingDir) so they survive `/clear`/respawns and multi-user scoping is structural. Capture rides the transcript (`transcript:user_prompt` from `transcript-watcher.ts`), NOT the input paths: `POST /input` sees only programmatic prompts and the WS channel is raw keystrokes. The listener lives inside `startTranscriptWatcher()`'s `if (!watcher)` block (outside it would duplicate per hook event) and is claude-only + gated on the setting per event. Store: `src/intent-store.ts` singleton, `intents.json` written 0600 tmp+rename (prompts can contain secrets; never fed to `/api/search`). Endpoints: GET/PUT/DELETE `/api/sessions/:id/intent` + POST `/api/sessions/:id/readmymind` (`readmymind-routes.ts`, ownership via `findSessionOrFail` WITH `req`; registrations stay the bare `app.('path')` shape, the endpoints.md drift scanner cannot see generics). **Phase 2 (predictor + 🧠 button)**: `readmymind-context.ts` is the PURE budgeted assembler (9 ranked sources, drop order siblings→away→workspace→tools, sections 1-4 truncate only); IO lives in `readmymind-collectors.ts` (transcript TAIL read — the live watcher keeps only a 500-char snippet — + git signals, skipped for remote-SSH cases) and the route; `readmymind-predictor.ts` reuses the AiCheckerBase spawn mechanics standalone (verdict-shaped base vs freeform JSON) as a mutable singleton routes call and tests stub. Claude-mode only (400), one in flight per session (409 CONFLICT), model = `readMyMindModel` setting defaulting to `AI_CHECK_MODEL` (opus, decided). Frontend `readmymind-ui.js`: header 🧠 marker-hidden (`btn-readmymind--hidden`) until the setting is ON; phones hide it in mobile.css and get a keyboard-accessory 🧠 key instead (ships in BOTH bar templates, revealed by the `rmm-enabled` class on the BAR element — setMode() rebuilds button innerHTML, so per-key state would be wiped; synced at init + every `applyHeaderVisibilitySettings()`). Alternate suggestions render as tappable rows that swap into the editable field without losing edits; Rethink rejects the whole shown set. Suggestions render via value/`textContent` ONLY and Send/Insert go through `POST /input` (server-side, so the sendEnterKey/local-echo trap does not apply) — nothing auto-sends, ever. User guide: `docs/readmymind.md`. +**Voice dictation via Claude** (`claudeVoiceEnabled`, SYNCED, default OFF): the mic button can transcribe through this machine's Claude Code login instead of a Deepgram key, using the same speech-to-text service the CLI's own `/voice` mode uses. ⚠️ **Claude Code's voice mode itself is unusable here**: it opens the HOST's microphone (`sox`/`arecord`), and the CLI runs in a headless tmux pane while the human is in a browser elsewhere. So Codeman captures in the browser and borrows only the backend. Audio goes browser → Codeman → Anthropic (`src/web/voice-stream.ts`): the OAuth token never reaches the page, and the browser only sends PCM and receives text. ⚠️ Credentials are **read-only** (`src/claude-credentials.ts`) and Codeman never refreshes them — a refresh rotates the refresh token and could sign the user out of their own CLI; an elapsed token reports `expired` instead. ⚠️ Capture MUST be linear16/16 kHz/mono, so it uses an **AudioWorklet**, not MediaRecorder (which cannot emit raw PCM); `voice-pcm-worklet.js` is fetched from JS, so it is invisible to `cacheBustAssets` and borrows voice-input.js's `?v=` token — **edit the two together**. ⚠️ Transcript frames carry the WHOLE running transcript, not deltas: the Claude path replaces where the Deepgram path appends. Provider choice is `voiceSettings.provider` (`auto` prefers Claude → Deepgram → Web Speech). → `docs/claude-voice-plan.md` + **Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `docs/agent-teams/`. **Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit) @@ -318,7 +320,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L ### API Routes -~200 handlers across 23 route files in `src/web/routes/`: system (45), sessions (34), cases (29), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), approvals (3), readmymind (4), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details. +~200 handlers across 24 route files in `src/web/routes/`: system (45), sessions (34), cases (29), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), approvals (3), readmymind (4), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), voice (1 + the `/ws/voice/stream` relay), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details. **HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`). diff --git a/docs/api-reference.md b/docs/api-reference.md index e3331855..17bce601 100644 --- a/docs/api-reference.md +++ b/docs/api-reference.md @@ -471,6 +471,28 @@ All four enforce session ownership in multi-user mode; a foreign session id answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the same directory are distinct by construction. +## Voice dictation + +Browser dictation transcribed through this server's Claude Code login, i.e. the +same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synced +`claudeVoiceEnabled` setting (default OFF). Design: +[`claude-voice-plan.md`](claude-voice-plan.md). + +- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?, + expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody + signed in to Claude Code on the server), `expired` (the access token elapsed; + running any Claude session refreshes it) or `malformed`. The OAuth token + itself is never returned by this or any other endpoint. +- `GET /ws/voice/stream?language=&keyterms=` (WebSocket, not under `/api`) + relays one dictation. Client sends binary frames of signed 16-bit + little-endian PCM, 16 kHz mono (<= 64 KB per frame), plus JSON control frames + `{"t":"finalize"}` (ask for the final transcript) and `{"t":"stop"}`. Server + sends `{"t":"ready"}`, `{"t":"transcript","text","final"}` (each frame is the + WHOLE running transcript, not a delta), `{"t":"error","message"}` and + `{"t":"closed"}`. Close codes: `4003` disallowed Host/Origin, `4004` + unavailable (reason in the close reason), `4008` too many concurrent streams. + Streams are capped in count and length (`src/config/voice.ts`). + ## Authentication Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque diff --git a/docs/claude-voice-plan.md b/docs/claude-voice-plan.md new file mode 100644 index 00000000..a0fd7d27 --- /dev/null +++ b/docs/claude-voice-plan.md @@ -0,0 +1,121 @@ +# Claude voice dictation in Codeman + +Wire Codeman's existing mic button to the same speech-to-text service Claude Code's own +`/voice` mode uses, so dictation works with **no third-party API key** for anyone already +signed in to Claude Code on the server. + +## Why the CLI's own voice mode cannot be reused directly + +Claude Code 2.1.x ships voice input: `/voice hold|tap|off` arms it, the CLI opens the +**host's** microphone (native `audio-capture-napi`, falling back to `sox`/`arecord` on Linux +after probing `/proc/asound/cards`), streams PCM upstream and types the transcript into its +own composer. + +Every part of that is on the wrong machine for Codeman. The CLI runs inside a tmux pane on +the server, which is typically headless and has no sound card at all, while the human is in +a browser on a phone somewhere else. Toggling `/voice` in the pane from Codeman would arm a +microphone nobody is sitting in front of. So Codeman keeps capturing audio in the browser, +where the user actually is, and only borrows the CLI's **transcription backend**. + +## The backend, as the CLI uses it + +Extracted from the 2.1.226 binary (`connectVoiceStream`): + +| | | +| --- | --- | +| URL | `wss://api.anthropic.com/api/ws/speech_to_text/voice_stream` | +| Query | `encoding=linear16`, `sample_rate=16000`, `channels=1`, `endpointing_ms=300`, `utterance_end_ms=1000`, `language=`, `use_conversation_engine=true`, `stt_provider=deepgram-nova3` | +| Headers | `Authorization: Bearer `, `User-Agent`, `x-app: cli`, `anthropic-client-platform`, optional `x-config-keyterms` | +| Audio | raw binary frames, PCM signed 16-bit little-endian, 16 kHz, mono | +| Keepalive | `{"type":"KeepAlive"}` on open, then every 8 s | +| Finalize | `{"type":"CloseStream"}`, then wait for the endpoint frame | +| Downstream | `{"type":"TranscriptText"\|"TranscriptInterim","data":"…"}` (running interim), `{"type":"TranscriptEndpoint"}` (promotes the pending interim to final), `{"type":"TranscriptError",…}`, `{"type":"error","message":…}` | + +Deepgram Nova-3 runs server-side, so the Deepgram-quality result arrives without a Deepgram +account. Verified against the live endpoint before this design was written: connect, stream +PCM, receive interims and an endpoint frame. + +## Architecture + +The browser cannot call that endpoint itself: it would need the OAuth bearer token in page +JavaScript (and CORS would refuse anyway). So the audio goes browser → Codeman → Anthropic, +and Codeman is the only thing that ever touches the token. + +``` +mic → AudioWorklet (Float32 → PCM16 @16 kHz) + → wss:///ws/voice/stream [cookie/basic auth, Origin+Host guarded] + → VoiceStreamRelay (reads ~/.claude/.credentials.json per connect) + → wss://api.anthropic.com/api/ws/speech_to_text/voice_stream + ← {"t":"transcript","text":…,"final":…} → existing _insertText() path +``` + +Nothing about the insert path changes: the transcript lands in the same preview overlay, +the same direct/compose insert modes, the same green Send button. + +### Server pieces + +- **`src/claude-credentials.ts`** — locate and parse the Claude Code OAuth credentials. + `parseClaudeCredentials()` is pure (JSON string + `now` → status) and unit-tested; + `readClaudeOAuthToken()` wraps it with IO: `$CLAUDE_CONFIG_DIR/.credentials.json` or + `~/.claude/.credentials.json`, and on macOS the login keychain + (`security find-generic-password -s "Claude Code-credentials"`). + **Read-only, always.** Codeman never writes credentials and never refreshes the token: a + refresh rotates the refresh token, and racing Claude Code's own refresh could sign the + user out of their CLI. An expired token surfaces as a plain "run a Claude session to + refresh" error instead. + The token is never logged, never returned by any endpoint, and never sent to the browser. + +- **`src/web/voice-stream.ts`** — pure `buildVoiceStreamUrl()` / `buildVoiceStreamHeaders()` / + `sanitizeKeyterms()` (ASCII-only, deduped, 1024-char cap, mirroring the CLI), plus + `VoiceStreamRelay`, which owns one upstream socket: keepalive timer, audio passthrough, + transcript translation, finalize, and the caps below. + +- **`src/web/routes/voice-routes.ts`** + - `GET /api/voice/status` → `{ available, reason, subscriptionType?, expiresAt? }`. Never + the token. `available:false` with a machine-readable `reason` (`disabled`, `no-credentials`, + `expired`) is what the settings row and the provider resolver read. + - `GET /ws/voice/stream?language=&keyterms=` → the relay. Same upgrade guard as + `/ws/sessions/:id/terminal`: allowed Host, same-site Origin, and the global auth hook has + already run on the handshake. + +Caps, because an open mic is an open pipe: one stream per connection, `MAX_VOICE_STREAMS` +concurrent server-wide, a hard `MAX_STREAM_MS` per stream, and a per-frame size cap. A tab +left recording cannot bill an unbounded amount of upstream audio. + +### Frontend pieces + +- **`voice-pcm-worklet.js`** — an `AudioWorkletProcessor` converting Float32 blocks to PCM16 + and posting ~256 ms frames back. `MediaRecorder` cannot produce raw PCM, which is why the + existing Deepgram path (container audio, auto-detected) cannot be reused as-is. Falls back + to `ScriptProcessorNode` where AudioWorklet is unavailable. +- **`ClaudeVoiceProvider`** in `voice-input.js` — mirrors `DeepgramProvider`'s shape + (`start({language, keyterms, onStream, onResult, onError, onEnd})`) so `VoiceInput` treats + the three providers uniformly. +- **Provider resolution** — new `voiceSettings.provider`: `auto` (default) | `claude` | + `deepgram` | `webspeech`. `auto` picks Claude when `/api/voice/status` reports it + available, else Deepgram when a key is set, else Web Speech. Pinning a provider always + wins, so an existing Deepgram user can keep exactly what they have. + +### Settings + +- `claudeVoiceEnabled` — synced, **default OFF**, gating the whole server side. Off is the + honest default: turning it on means this machine's Claude subscription starts paying for + transcription for whoever can reach the UI, and the audio goes to Anthropic rather than to + wherever it went before. One switch in Settings → Voice, and the mic works with no key. +- `voiceSettings.provider` — per the resolution table above; joins the existing synced + `voiceSettings` object. + +## Things worth knowing + +- **This uses an undocumented endpoint with subscription credentials.** It is the user's own + token, on the user's own machine, driving the user's own Claude Code install, but it is not + a published API and Anthropic can change or restrict it. Default-OFF is deliberate; the + Deepgram and Web Speech paths stay untouched as the supported fallbacks. +- **Multi-user mode**: every user's dictation would run on the server owner's Claude + credentials, exactly as every user's *sessions* already run on them. Consistent, but worth + stating out loud in the settings copy. +- **Token lifetime** is about 8 hours, refreshed by Claude Code itself whenever it runs. The + relay re-reads the file on every connect rather than caching, so a refresh is picked up on + the next press of the mic. +- **HTTPS or localhost**: `getUserMedia` needs a secure context. Prod is HTTPS behind + `tailscale serve`, so this is already satisfied; the existing error copy covers the rest. diff --git a/src/claude-credentials.ts b/src/claude-credentials.ts new file mode 100644 index 00000000..e2445375 --- /dev/null +++ b/src/claude-credentials.ts @@ -0,0 +1,126 @@ +/** + * @fileoverview Read-only access to the Claude Code OAuth credentials. + * + * Claude Code stores its subscription OAuth tokens in + * `$CLAUDE_CONFIG_DIR/.credentials.json` (default `~/.claude/.credentials.json`, + * mode 0600) on Linux/Windows, and in the login keychain on macOS. Codeman reads + * the access token to authenticate the voice-dictation relay + * (`src/web/voice-stream.ts`) against the same speech-to-text service the CLI's + * own `/voice` mode uses. + * + * ⚠️ READ-ONLY, deliberately. Codeman never writes this file and never performs + * an OAuth refresh: a refresh ROTATES the refresh token, so racing Claude Code's + * own refresh could invalidate the user's CLI login. An expired access token is + * reported as `expired` and the caller tells the user to run a Claude session + * (which refreshes it) instead. + * + * ⚠️ The token is a bearer secret: it is never logged, never persisted, never + * included in any API response, and never sent to the browser. + */ + +import { readFile } from 'fs/promises'; +import { execFile } from 'child_process'; +import { homedir, userInfo } from 'os'; +import { join } from 'path'; + +/** Result of inspecting the credential store. The token is present only on 'ok'. */ +export type ClaudeCredentialStatus = 'ok' | 'expired' | 'missing' | 'malformed'; + +export interface ClaudeOAuthCredentials { + status: ClaudeCredentialStatus; + /** Bearer token. Present only when status is 'ok'. Never log or serialize this. */ + accessToken?: string; + /** Epoch ms the access token expires at, when the store reports one. */ + expiresAt?: number; + /** e.g. 'max', 'pro'. Display-only, safe to surface. */ + subscriptionType?: string; +} + +/** Skew applied to the stored expiry so a token that dies mid-stream is refused up front. */ +const EXPIRY_SKEW_MS = 60_000; + +/** macOS keychain service holding the same JSON blob as `.credentials.json`. */ +const KEYCHAIN_SERVICE = 'Claude Code-credentials'; + +/** Keychain lookups shell out; keep them short so a locked keychain cannot hang a request. */ +const KEYCHAIN_TIMEOUT_MS = 3000; + +/** + * Parse a `.credentials.json` payload. Pure: no IO, no clock read (pass `now`), + * so the expiry and shape handling are unit-testable. + * + * Returns 'malformed' for anything that is not the expected `claudeAiOauth` + * shape rather than throwing — a hand-edited or half-written file must degrade + * to "voice unavailable", never to a 500. + */ +export function parseClaudeCredentials(raw: string, now: number): ClaudeOAuthCredentials { + let parsed: unknown; + try { + parsed = JSON.parse(raw); + } catch { + return { status: 'malformed' }; + } + if (!parsed || typeof parsed !== 'object') return { status: 'malformed' }; + + const oauth = (parsed as { claudeAiOauth?: unknown }).claudeAiOauth; + if (!oauth || typeof oauth !== 'object') return { status: 'malformed' }; + + const record = oauth as Record; + const accessToken = typeof record.accessToken === 'string' ? record.accessToken.trim() : ''; + if (!accessToken) return { status: 'malformed' }; + + const expiresAt = typeof record.expiresAt === 'number' ? record.expiresAt : undefined; + const subscriptionType = typeof record.subscriptionType === 'string' ? record.subscriptionType : undefined; + + // An expired token is a real state (the CLI refreshes on its next run), not a + // malformed store: report it separately so the UI can say something useful. + if (expiresAt !== undefined && expiresAt - EXPIRY_SKEW_MS <= now) { + return { status: 'expired', expiresAt, subscriptionType }; + } + return { status: 'ok', accessToken, expiresAt, subscriptionType }; +} + +/** Path of the credentials file, honoring CLAUDE_CONFIG_DIR like the CLI does. */ +export function claudeCredentialsPath(env: NodeJS.ProcessEnv = process.env): string { + const configDir = typeof env.CLAUDE_CONFIG_DIR === 'string' && env.CLAUDE_CONFIG_DIR.trim(); + return join(configDir || join(homedir(), '.claude'), '.credentials.json'); +} + +/** Read the macOS keychain entry. Resolves to null on any failure (locked, absent, non-mac). */ +function readKeychainCredentials(): Promise { + return new Promise((resolve) => { + execFile( + 'security', + ['find-generic-password', '-a', userInfo().username, '-w', '-s', KEYCHAIN_SERVICE], + { encoding: 'utf-8', timeout: KEYCHAIN_TIMEOUT_MS }, + (err, stdout) => resolve(err ? null : stdout.trim() || null) + ); + }); +} + +/** + * Locate and parse the Claude Code OAuth credentials. + * + * File first (present on every platform once the CLI has run there), keychain + * second on macOS. Never caches: Claude Code rewrites the store roughly every + * 8 hours, and a cached token would go stale inside a long-lived server. + */ +export async function readClaudeOAuthCredentials(now: number = Date.now()): Promise { + let fileResult: ClaudeOAuthCredentials | null = null; + try { + fileResult = parseClaudeCredentials(await readFile(claudeCredentialsPath(), 'utf-8'), now); + } catch { + fileResult = null; + } + if (fileResult && fileResult.status !== 'malformed') return fileResult; + + if (process.platform === 'darwin') { + const raw = await readKeychainCredentials(); + if (raw) { + const keychainResult = parseClaudeCredentials(raw, now); + if (keychainResult.status !== 'malformed') return keychainResult; + } + } + + return fileResult ?? { status: 'missing' }; +} diff --git a/src/config/voice.ts b/src/config/voice.ts new file mode 100644 index 00000000..3a666331 --- /dev/null +++ b/src/config/voice.ts @@ -0,0 +1,56 @@ +/** + * @fileoverview Bounds and endpoint config for Claude voice dictation. + * + * Backs the browser → Codeman → Anthropic dictation relay (`src/web/voice-stream.ts`, + * `src/web/routes/voice-routes.ts`; design in `docs/claude-voice-plan.md`). + * + * Why everything here is bounded: an open microphone is an open pipe. Each live + * stream holds a browser socket, an upstream socket and a keepalive timer, and + * every second of audio is billed against the server owner's Claude subscription. + * A tab left recording (phone in a pocket, forgotten laptop) must cost a bounded + * amount, so streams die on their own at `MAX_STREAM_MS` and the server refuses + * more than `MAX_CONCURRENT_STREAMS` at once. + * + * The audio frame cap is a memory guard on a socket that carries attacker-shaped + * binary data: PCM16 at 16 kHz mono is 32 KB/s, so a 256 ms frame is ~8 KB and + * anything near 64 KB is either a broken client or an attempt to make the relay + * buffer for someone else. + */ + +/** Upstream speech-to-text service (the one Claude Code's own `/voice` mode uses). */ +export const VOICE_STREAM_HOST = 'wss://api.anthropic.com'; + +/** Path of the streaming speech-to-text endpoint. */ +export const VOICE_STREAM_PATH = '/api/ws/speech_to_text/voice_stream'; + +/** + * Base override, for tests (point the relay at a local mock) and for users on an + * Anthropic-compatible gateway. Must be a ws:// or wss:// origin. + */ +export function voiceStreamBase(env: NodeJS.ProcessEnv = process.env): string { + const override = typeof env.CODEMAN_VOICE_STREAM_BASE === 'string' ? env.CODEMAN_VOICE_STREAM_BASE.trim() : ''; + if (override && /^wss?:\/\//.test(override)) return override.replace(/\/+$/, ''); + return VOICE_STREAM_HOST; +} + +/** Upstream drops an idle socket; the CLI pings at 8s and so do we. */ +export const KEEPALIVE_INTERVAL_MS = 8000; + +/** Hard ceiling on one dictation. Long enough for any real utterance, short enough to bound a forgotten mic. */ +export const MAX_STREAM_MS = 5 * 60_000; + +/** Concurrent relays server-wide. Dictation is a human-paced, one-at-a-time act. */ +export const MAX_CONCURRENT_STREAMS = 4; + +/** Largest single audio frame accepted from the browser (~2s of PCM16 @16 kHz mono). */ +export const MAX_AUDIO_FRAME_BYTES = 64 * 1024; + +/** How long to wait for the final transcript after the client asks to finalize. */ +export const FINALIZE_TIMEOUT_MS = 3000; + +/** Upstream caps the keyterms header; mirrors the CLI's own limit. */ +export const MAX_KEYTERMS_HEADER_CHARS = 1024; + +/** Audio format the endpoint is opened with. The browser worklet must match exactly. */ +export const AUDIO_SAMPLE_RATE = 16000; +export const AUDIO_CHANNELS = 1; diff --git a/src/web/ports/config-port.ts b/src/web/ports/config-port.ts index 88fb7f18..322a60ad 100644 --- a/src/web/ports/config-port.ts +++ b/src/web/ports/config-port.ts @@ -19,6 +19,8 @@ export interface ConfigPort { getTerminalHistoryConfig(): Promise; /** Synced `agentSkillEnabled` app setting (default OFF); gates per-case agent-skill injection. */ getAgentSkillEnabled(): Promise; + /** Synced `claudeVoiceEnabled` app setting (default OFF); gates the Claude voice dictation relay. */ + getClaudeVoiceEnabled(): Promise; getDefaultClaudeMdPath(): Promise; getLightState(identity?: { username: string; role: 'admin' | 'user' }): unknown; getLightSessionsState(): unknown[]; diff --git a/src/web/public/index.html b/src/web/public/index.html index 6bc54df3..d32b4e5e 100644 --- a/src/web/public/index.html +++ b/src/web/public/index.html @@ -1966,6 +1966,18 @@
Active provider
— +
+
+ Speech engine + Auto prefers Claude when this server can transcribe, then Deepgram, then the browser. +
+ +
Insert mode @@ -1979,6 +1991,23 @@
+
+

Claude

synced
+
+
+
+ Transcribe with this server's Claude login + Dictation with no API key, through the same service Claude Code's own /voice mode uses. Microphone audio goes to Anthropic and is billed to this machine's Claude subscription, for everyone who can reach this UI. +
+ +
+
+
Server status
+ — +
+
+
+

Deepgram Nova-3

device
diff --git a/src/web/public/settings-ui.js b/src/web/public/settings-ui.js index 2c5f1357..8175bfae 100644 --- a/src/web/public/settings-ui.js +++ b/src/web/public/settings-ui.js @@ -482,17 +482,18 @@ Object.assign(CodemanApp.prototype, { const voiceCfg = VoiceInput._getDeepgramConfig(); document.getElementById('voiceDeepgramKey').value = voiceCfg.apiKey || ''; document.getElementById('voiceLanguage').value = voiceCfg.language || 'en-US'; - document.getElementById('voiceKeyterms').value = voiceCfg.keyterms || 'refactor, endpoint, middleware, callback, async, regex, TypeScript, npm, API, deploy, config, linter, env, webhook, schema, CLI, JSON, CSS, DOM, SSE, backend, frontend, localhost, dependencies, repository, merge, rebase, diff, commit, com'; + document.getElementById('voiceKeyterms').value = voiceCfg.keyterms || DEFAULT_VOICE_KEYTERMS; document.getElementById('voiceInsertMode').value = voiceCfg.insertMode || 'direct'; + document.getElementById('voiceProvider').value = voiceCfg.provider || 'auto'; + document.getElementById('appSettingsClaudeVoice').checked = settings.claudeVoiceEnabled ?? false; // Reset key visibility to hidden const keyInput = document.getElementById('voiceDeepgramKey'); keyInput.type = 'password'; document.getElementById('voiceKeyToggleBtn').textContent = 'Show'; - // Update provider status - const providerName = VoiceInput.getActiveProviderName(); - const providerEl = document.getElementById('voiceProviderStatus'); - providerEl.textContent = providerName; - providerEl.className = 'voice-provider-status' + (providerName.startsWith('Deepgram') ? ' active' : ''); + // Update provider status. The Claude row needs a fresh server probe: the + // setting is synced, so another device may have flipped it since page load. + this._renderVoiceProviderStatus(); + VoiceInput.refreshClaudeStatus().then(() => this._renderVoiceProviderStatus()); // Updates section — show current version, reset transient result/progress UI. this._initUpdatesSection(); @@ -1830,6 +1831,37 @@ Object.assign(CodemanApp.prototype, { } }, + /** + * Paint both Voice status rows: which provider a mic press would use, and what + * the server reports about its Claude login. Called on open and again once the + * /api/voice/status probe resolves. + */ + _renderVoiceProviderStatus() { + const providerEl = document.getElementById('voiceProviderStatus'); + if (providerEl) { + const providerName = VoiceInput.getActiveProviderName(); + providerEl.textContent = providerName; + const live = providerName.startsWith('Deepgram Nova') || providerName.startsWith('Claude (this'); + providerEl.className = 'voice-provider-status' + (live ? ' active' : ''); + } + const claudeEl = document.getElementById('voiceClaudeStatus'); + if (!claudeEl) return; + const status = VoiceInput._claudeStatus; + const text = !status + ? 'Checking...' + : status.available + ? `Ready${status.subscriptionType ? ` (${status.subscriptionType})` : ''}` + : status.reason === 'expired' + ? 'Login expired - run a Claude session to refresh' + : status.reason === 'no-credentials' + ? 'No Claude Code login on the server' + : status.reason === 'malformed' + ? 'Claude credentials unreadable' + : 'Off - enable it above'; + claudeEl.textContent = text; + claudeEl.className = 'voice-provider-status' + (status?.available ? ' active' : ''); + }, + async saveAppSettings() { // Gesture overlay is injected at page render (server-side), so a change to it // only takes effect on reload — remember the prior value to decide below. @@ -1892,6 +1924,7 @@ Object.assign(CodemanApp.prototype, { // Claude Permissions settings agentTeamsEnabled: document.getElementById('appSettingsAgentTeams').checked, agentSkillEnabled: document.getElementById('appSettingsAgentSkill').checked, + claudeVoiceEnabled: document.getElementById('appSettingsClaudeVoice').checked, claudeModel: document.getElementById('appSettingsClaudeModel').value, opusContext1mEnabled: document.getElementById('appSettingsOpusContext1m').checked, remoteAutoReconnect: document.getElementById('appSettingsRemoteAutoReconnect').checked, @@ -1931,6 +1964,7 @@ Object.assign(CodemanApp.prototype, { // Save voice settings to localStorage + include in server payload for cross-device sync const voiceSettings = { + provider: document.getElementById('voiceProvider').value, apiKey: document.getElementById('voiceDeepgramKey').value.trim(), language: document.getElementById('voiceLanguage').value, keyterms: document.getElementById('voiceKeyterms').value.trim(), @@ -2110,6 +2144,10 @@ Object.assign(CodemanApp.prototype, { this.closeAppSettings(); + // Voice availability is a server-side answer, so re-probe after a save: + // otherwise the mic keeps using the pre-save provider until the next reload. + VoiceInput.refreshClaudeStatus(); + // The gesture overlay is injected at page render (server reads // gestureControlEnabled from settings.json), so a change only takes effect on // reload. Reload when it actually changed — the server PUT above already diff --git a/src/web/public/voice-input.js b/src/web/public/voice-input.js index 1bcfaec2..28c22985 100644 --- a/src/web/public/voice-input.js +++ b/src/web/public/voice-input.js @@ -1,7 +1,13 @@ /** - * @fileoverview Voice input with Deepgram Nova-3 (primary) and Web Speech API (fallback). + * @fileoverview Voice input with three providers: Claude (this server's Claude Code + * login), Deepgram Nova-3, and the Web Speech API. * - * Defines two singleton objects: + * Defines three singleton objects: + * + * - ClaudeVoiceProvider — Dictation through Codeman's own `/ws/voice/stream`, which + * relays to the speech-to-text service Claude Code's `/voice` mode uses. No API key: + * the server holds the OAuth token, the browser only sends PCM16 @16 kHz (AudioWorklet, + * since MediaRecorder cannot emit raw PCM) and receives text. See docs/claude-voice-plan.md. * * - DeepgramProvider — Direct browser-to-Deepgram WebSocket connection for speech-to-text. * Captures audio via MediaRecorder, streams chunks every 250ms, handles KeepAlive pings, @@ -14,6 +20,7 @@ * Includes a temporary green Send button that replaces the settings gear icon after voice input. * Web Speech API has auto-retry (up to 2x) for premature onend and iOS Safari stability check. * + * @globals {object} ClaudeVoiceProvider * @globals {object} DeepgramProvider * @globals {object} VoiceInput * @@ -22,9 +29,13 @@ * @loadorder 3 of 15 — loaded after mobile-handlers.js, before notification-manager.js */ -// Codeman — Voice input with Deepgram Nova-3 and Web Speech API fallback +// Codeman — Voice input with Claude, Deepgram Nova-3 and Web Speech API // Loaded after mobile-handlers.js, before app.js +/** Dev vocabulary sent to the recognizer as a hint. Shared by every provider and the settings form. */ +const DEFAULT_VOICE_KEYTERMS = + 'refactor, endpoint, middleware, callback, async, regex, TypeScript, npm, API, deploy, config, linter, env, webhook, schema, CLI, JSON, CSS, DOM, SSE, backend, frontend, localhost, dependencies, repository, merge, rebase, diff, commit, com'; + // ═══════════════════════════════════════════════════════════════ // Voice Input (Deepgram Nova-3 + Web Speech API fallback) // ═══════════════════════════════════════════════════════════════ @@ -245,7 +256,282 @@ const DeepgramProvider = { }; /** - * VoiceInput - Speech-to-text with Deepgram Nova-3 (primary) and Web Speech API (fallback). + * ClaudeVoiceProvider - Speech-to-text through this Codeman server's Claude Code + * login, i.e. the same service the CLI's own `/voice` mode uses. No API key. + * + * Audio goes browser -> Codeman -> Anthropic: the OAuth token never leaves the + * server, so the browser only ever sends PCM and receives text + * (docs/claude-voice-plan.md). + * + * ⚠️ The upstream endpoint is opened as linear16 / 16 kHz / mono, so capture MUST + * be raw PCM at that rate. MediaRecorder cannot emit raw PCM (container formats + * only), which is why this path uses an AudioWorklet rather than reusing + * DeepgramProvider's recorder. The AudioContext is constructed at 16000 Hz so the + * browser does the resampling. + * + * ⚠️ Transcript frames carry the WHOLE running transcript, not deltas. Callers + * must replace, never concatenate. + */ +const ClaudeVoiceProvider = { + _ws: null, + _stream: null, + _audioContext: null, + _workletNode: null, + _sourceNode: null, + _scriptNode: null, + _silenceTimeout: null, + _onResult: null, + _onError: null, + _onEnd: null, + _finalized: false, + + /** How long without any transcript before the recording gives up on its own. */ + SILENCE_MS: 6000, + + /** + * Start streaming. + * @param {object} opts - { language, keyterms[], onResult(text, isFinal), onError(msg), onEnd(), onStream(stream) } + */ + async start(opts) { + this._onResult = opts.onResult; + this._onError = opts.onError; + this._onEnd = opts.onEnd; + this._finalized = false; + + if (!navigator.mediaDevices?.getUserMedia) { + this._onError?.('Microphone requires a secure context (HTTPS). Use --https flag or access via localhost.'); + this._cleanup(); + return; + } + try { + this._stream = await navigator.mediaDevices.getUserMedia({ + audio: { noiseSuppression: true, echoCancellation: true, autoGainControl: true } + }); + } catch (err) { + const msg = err.name === 'NotAllowedError' + ? 'Microphone access denied. Check browser settings.' + : 'Microphone error: ' + err.message; + this._onError?.(msg); + this._cleanup(); + return; + } + opts.onStream?.(this._stream); + + const params = new URLSearchParams(); + if (opts.language) params.set('language', opts.language); + if (opts.keyterms?.length) params.set('keyterms', opts.keyterms.join(',')); + const proto = location.protocol === 'https:' ? 'wss:' : 'ws:'; + try { + this._ws = new WebSocket(`${proto}//${location.host}/ws/voice/stream?${params}`); + } catch (err) { + this._onError?.('Failed to open voice stream: ' + err.message); + this._cleanup(); + return; + } + this._ws.binaryType = 'arraybuffer'; + + this._ws.onopen = () => { + // Capture starts only once the socket is up: PCM buffered before that would + // be the oldest audio, and dropping it keeps the transcript aligned with what + // the user hears themselves saying. + this._startCapture().catch((err) => { + this._onError?.('Microphone capture failed: ' + err.message); + this.stop(); + }); + this._resetSilenceTimeout(); + }; + + this._ws.onmessage = (event) => { + let msg; + try { + msg = JSON.parse(event.data); + } catch (_e) { + return; + } + if (msg.t === 'transcript' && msg.text) { + this._resetSilenceTimeout(); + this._onResult?.(msg.text, msg.final === true); + } else if (msg.t === 'error') { + this._onError?.(msg.message || 'Voice transcription failed'); + } + }; + + this._ws.onerror = () => { + // onclose carries the actionable detail (close code); nothing useful here. + }; + + this._ws.onclose = (event) => { + if (event.code === 4004) { + this._onError?.(this._unavailableMessage(event.reason)); + } else if (event.code === 4008) { + this._onError?.('Too many voice streams are already running on this server.'); + } else if (event.code === 4003) { + this._onError?.('Voice stream refused (origin not allowed).'); + } else if (event.code !== 1000 && !this._finalized) { + this._onError?.('Voice stream closed: ' + (event.reason || `code ${event.code}`)); + } + this._stopCapture(); + const onEnd = this._onEnd; + this._onEnd = null; + onEnd?.(); + }; + }, + + /** Map the server's close reason onto something a user can act on. */ + _unavailableMessage(reason) { + if (reason === 'expired') return 'Claude login expired. Run a Claude session to refresh it, then try again.'; + if (reason === 'disabled') return 'Claude voice is off. Enable it in Settings > Voice.'; + return 'No Claude Code login found on the server. Sign in with `claude` there, or use Deepgram.'; + }, + + /** Wire mic -> 16 kHz PCM16 frames -> WebSocket. */ + async _startCapture() { + const Ctx = window.AudioContext || window.webkitAudioContext; + // Ask for 16 kHz directly so the browser resamples; Safari may hand back its + // own rate, which _pcmFromFloat32 then downsamples to match. + this._audioContext = new Ctx({ sampleRate: 16000 }); + if (this._audioContext.state === 'suspended') await this._audioContext.resume(); + this._sourceNode = this._audioContext.createMediaStreamSource(this._stream); + + if (this._audioContext.audioWorklet) { + await this._audioContext.audioWorklet.addModule(this._workletUrl()); + this._workletNode = new AudioWorkletNode(this._audioContext, 'pcm-frame-processor'); + this._workletNode.port.onmessage = (event) => this._sendAudio(event.data); + this._sourceNode.connect(this._workletNode); + // A worklet with no destination is not pulled in some engines; a zero-gain + // sink keeps the graph running without echoing the mic to the speakers. + const sink = this._audioContext.createGain(); + sink.gain.value = 0; + this._workletNode.connect(sink).connect(this._audioContext.destination); + return; + } + + // Fallback for engines without AudioWorklet (older Safari): deprecated, but + // it is this or no dictation at all there. + this._scriptNode = this._audioContext.createScriptProcessor(4096, 1, 1); + this._scriptNode.onaudioprocess = (event) => { + this._sendAudio(this._pcmFromFloat32(event.inputBuffer.getChannelData(0), this._audioContext.sampleRate)); + }; + this._sourceNode.connect(this._scriptNode); + this._scriptNode.connect(this._audioContext.destination); + }, + + /** + * Worklet URL carrying this page's cache-bust token. + * + * ⚠️ Static assets are served `immutable` for a year, and `cacheBustAssets` + * only rewrites `.js` refs in `