mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-06 23:49:41 +02:00
feat(voice): dictate through the server's Claude Code login, no API key
The mic button previously needed a Deepgram API key, or fell back to the browser's Web Speech engine. It can now transcribe through the same speech-to-text service Claude Code's own /voice mode uses, so anyone signed in to Claude Code on the server gets dictation with no third-party account. Claude Code's voice mode cannot be driven directly: it opens the HOST's microphone (sox/arecord), and the CLI runs in a headless tmux pane while the human is in a browser somewhere else. So capture stays in the browser and only the transcription backend is borrowed. Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since MediaRecorder cannot emit raw PCM) and receives text. - GET /api/voice/status reports readiness and never the token - GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade guard as the terminal socket, plus caps on concurrency, stream length and frame size - credentials are read-only: Codeman never refreshes them, since a refresh rotates the refresh token and could sign the user out of their own CLI - claudeVoiceEnabled (synced, default OFF) gates the whole server side - voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers Claude, then a configured Deepgram key, then the browser Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1966,6 +1966,18 @@
|
||||
<div class="set-row-text"><span class="set-row-label">Active provider</span></div>
|
||||
<span class="voice-provider-status" id="voiceProviderStatus">—</span>
|
||||
</div>
|
||||
<div class="set-row has-field" data-search="voice provider claude deepgram web speech engine">
|
||||
<div class="set-row-text">
|
||||
<span class="set-row-label">Speech engine</span>
|
||||
<span class="set-row-desc">Auto prefers Claude when this server can transcribe, then Deepgram, then the browser.</span>
|
||||
</div>
|
||||
<select id="voiceProvider" class="set-select">
|
||||
<option value="auto">Auto</option>
|
||||
<option value="claude">Claude</option>
|
||||
<option value="deepgram">Deepgram</option>
|
||||
<option value="webspeech">Browser (Web Speech)</option>
|
||||
</select>
|
||||
</div>
|
||||
<div class="set-row has-field" data-search="voice insert mode compose">
|
||||
<div class="set-row-text">
|
||||
<span class="set-row-label">Insert mode</span>
|
||||
@@ -1979,6 +1991,23 @@
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="set-group">
|
||||
<div class="set-group-head"><h4>Claude</h4><span class="set-scope">synced</span></div>
|
||||
<div class="set-group-body">
|
||||
<div class="set-row" data-search="claude voice dictation subscription no api key">
|
||||
<div class="set-row-text">
|
||||
<span class="set-row-label">Transcribe with this server's Claude login</span>
|
||||
<span class="set-row-desc">Dictation with no API key, through the same service Claude Code's own /voice mode uses. Microphone audio goes to Anthropic and is billed to this machine's Claude subscription, for everyone who can reach this UI.</span>
|
||||
</div>
|
||||
<label class="switch switch-sm"><input type="checkbox" id="appSettingsClaudeVoice"><span class="slider"></span></label>
|
||||
</div>
|
||||
<div class="set-row" data-search="claude voice status credentials">
|
||||
<div class="set-row-text"><span class="set-row-label">Server status</span></div>
|
||||
<span class="voice-provider-status" id="voiceClaudeStatus" data-i18n-skip>—</span>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="set-group">
|
||||
<div class="set-group-head"><h4>Deepgram Nova-3</h4><span class="set-scope">device</span></div>
|
||||
<div class="set-group-body">
|
||||
|
||||
Reference in New Issue
Block a user