feat(voice): dictate through the server's Claude Code login, no API key

The mic button previously needed a Deepgram API key, or fell back to the
browser's Web Speech engine. It can now transcribe through the same
speech-to-text service Claude Code's own /voice mode uses, so anyone signed
in to Claude Code on the server gets dictation with no third-party account.

Claude Code's voice mode cannot be driven directly: it opens the HOST's
microphone (sox/arecord), and the CLI runs in a headless tmux pane while the
human is in a browser somewhere else. So capture stays in the browser and only
the transcription backend is borrowed.

Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the
page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since
MediaRecorder cannot emit raw PCM) and receives text.

- GET /api/voice/status reports readiness and never the token
- GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade
  guard as the terminal socket, plus caps on concurrency, stream length and
  frame size
- credentials are read-only: Codeman never refreshes them, since a refresh
  rotates the refresh token and could sign the user out of their own CLI
- claudeVoiceEnabled (synced, default OFF) gates the whole server side
- voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers
  Claude, then a configured Deepgram key, then the browser

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-08-10 11:19:51 +02:00
parent 752374abc7
commit 4b51ba306e
19 changed files with 1904 additions and 22 deletions
+12
View File
@@ -166,6 +166,7 @@ import {
registerMeRoutes,
registerAdminRoutes,
registerWsRoutes,
registerVoiceRoutes,
registerWebviewRoutes,
tryWebviewRefererFallback,
} from './routes/index.js';
@@ -635,6 +636,7 @@ export class WebServer extends EventEmitter {
getClaudeModeConfig: this.getClaudeModeConfig.bind(this),
getTerminalHistoryConfig: this.getTerminalHistoryConfig.bind(this),
getAgentSkillEnabled: this.getAgentSkillEnabled.bind(this),
getClaudeVoiceEnabled: this.getClaudeVoiceEnabled.bind(this),
getDefaultClaudeMdPath: this.getDefaultClaudeMdPath.bind(this),
getLightState: this.getLightState.bind(this),
getLightSessionsState: this.getLightSessionsState.bind(this),
@@ -982,6 +984,7 @@ export class WebServer extends EventEmitter {
registerCronRoutes(this.app, { ...ctx, cron: this.cronService });
registerWsRoutes(this.app, ctx, () => this.getHostPolicy());
registerVoiceRoutes(this.app, ctx, () => this.getHostPolicy());
}
/**
@@ -1704,6 +1707,15 @@ export class WebServer extends EventEmitter {
return settings.agentSkillEnabled === true;
}
// Whether browser dictation may use this machine's Claude Code credentials
// (synced `claudeVoiceEnabled` setting, default OFF; docs/claude-voice-plan.md).
// OFF by default because turning it on spends the operator's Claude subscription
// on transcription for anyone who can reach the UI.
private async getClaudeVoiceEnabled(): Promise<boolean> {
const settings = await this.readSettings();
return settings.claudeVoiceEnabled === true;
}
/**
* Read My Mind predictor model (docs/readmymind-plan.md): `readMyMindModel`
* setting, defaulting to the AI-checker opus model. Prediction quality is