feat(voice): dictate through the server's Claude Code login, no API key

The mic button previously needed a Deepgram API key, or fell back to the
browser's Web Speech engine. It can now transcribe through the same
speech-to-text service Claude Code's own /voice mode uses, so anyone signed
in to Claude Code on the server gets dictation with no third-party account.

Claude Code's voice mode cannot be driven directly: it opens the HOST's
microphone (sox/arecord), and the CLI runs in a headless tmux pane while the
human is in a browser somewhere else. So capture stays in the browser and only
the transcription backend is borrowed.

Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the
page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since
MediaRecorder cannot emit raw PCM) and receives text.

- GET /api/voice/status reports readiness and never the token
- GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade
  guard as the terminal socket, plus caps on concurrency, stream length and
  frame size
- credentials are read-only: Codeman never refreshes them, since a refresh
  rotates the refresh token and could sign the user out of their own CLI
- claudeVoiceEnabled (synced, default OFF) gates the whole server side
- voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers
  Claude, then a configured Deepgram key, then the browser

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-08-10 11:19:51 +02:00
parent 752374abc7
commit 4b51ba306e
19 changed files with 1904 additions and 22 deletions
+29
View File
@@ -1966,6 +1966,18 @@
<div class="set-row-text"><span class="set-row-label">Active provider</span></div>
<span class="voice-provider-status" id="voiceProviderStatus">&mdash;</span>
</div>
<div class="set-row has-field" data-search="voice provider claude deepgram web speech engine">
<div class="set-row-text">
<span class="set-row-label">Speech engine</span>
<span class="set-row-desc">Auto prefers Claude when this server can transcribe, then Deepgram, then the browser.</span>
</div>
<select id="voiceProvider" class="set-select">
<option value="auto">Auto</option>
<option value="claude">Claude</option>
<option value="deepgram">Deepgram</option>
<option value="webspeech">Browser (Web Speech)</option>
</select>
</div>
<div class="set-row has-field" data-search="voice insert mode compose">
<div class="set-row-text">
<span class="set-row-label">Insert mode</span>
@@ -1979,6 +1991,23 @@
</div>
</div>
<div class="set-group">
<div class="set-group-head"><h4>Claude</h4><span class="set-scope">synced</span></div>
<div class="set-group-body">
<div class="set-row" data-search="claude voice dictation subscription no api key">
<div class="set-row-text">
<span class="set-row-label">Transcribe with this server's Claude login</span>
<span class="set-row-desc">Dictation with no API key, through the same service Claude Code's own /voice mode uses. Microphone audio goes to Anthropic and is billed to this machine's Claude subscription, for everyone who can reach this UI.</span>
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsClaudeVoice"><span class="slider"></span></label>
</div>
<div class="set-row" data-search="claude voice status credentials">
<div class="set-row-text"><span class="set-row-label">Server status</span></div>
<span class="voice-provider-status" id="voiceClaudeStatus" data-i18n-skip>&mdash;</span>
</div>
</div>
</div>
<div class="set-group">
<div class="set-group-head"><h4>Deepgram Nova-3</h4><span class="set-scope">device</span></div>
<div class="set-group-body">