feat(voice): dictate through the server's Claude Code login, no API key

The mic button previously needed a Deepgram API key, or fell back to the
browser's Web Speech engine. It can now transcribe through the same
speech-to-text service Claude Code's own /voice mode uses, so anyone signed
in to Claude Code on the server gets dictation with no third-party account.

Claude Code's voice mode cannot be driven directly: it opens the HOST's
microphone (sox/arecord), and the CLI runs in a headless tmux pane while the
human is in a browser somewhere else. So capture stays in the browser and only
the transcription backend is borrowed.

Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the
page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since
MediaRecorder cannot emit raw PCM) and receives text.

- GET /api/voice/status reports readiness and never the token
- GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade
  guard as the terminal socket, plus caps on concurrency, stream length and
  frame size
- credentials are read-only: Codeman never refreshes them, since a refresh
  rotates the refresh token and could sign the user out of their own CLI
- claudeVoiceEnabled (synced, default OFF) gates the whole server side
- voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers
  Claude, then a configured Deepgram key, then the browser

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-08-10 11:19:51 +02:00
parent 752374abc7
commit 4b51ba306e
19 changed files with 1904 additions and 22 deletions
+12
View File
@@ -857,6 +857,16 @@ export const SettingsUpdateSchema = z
* add-only at create; a marker keeps user-authored copies untouched.
*/
agentSkillEnabled: z.boolean().optional(),
/**
* Let browser dictation transcribe through this machine's Claude Code login,
* the same speech-to-text service the CLI's own `/voice` mode uses
* (docs/claude-voice-plan.md). SYNCED, default OFF: enabling it spends the
* operator's Claude subscription on transcription for anyone who can reach
* the UI, and routes microphone audio to Anthropic rather than to whichever
* provider was configured before. The Deepgram and Web Speech paths are
* untouched by this flag.
*/
claudeVoiceEnabled: z.boolean().optional(),
/**
* Approvals Inbox (header bell + drawer, phone overview answer buttons,
* push Approve/Deny action buttons). SYNCED, default OFF (opt-in): even
@@ -970,6 +980,8 @@ export const SettingsUpdateSchema = z
// Voice settings (cross-device sync)
voiceSettings: z
.object({
/** 'auto' | 'claude' | 'deepgram' | 'webspeech'. Unknown values fall back to auto client-side. */
provider: z.string().max(20).optional(),
apiKey: z.string().max(200).optional(),
language: z.string().max(20).optional(),
keyterms: z.string().max(500).optional(),