diff --git a/CLAUDE.md b/CLAUDE.md index 764ffd03..f2f8220b 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -35,7 +35,7 @@ When user says "COM": 1. Increment version in BOTH `package.json` AND `CLAUDE.md` (verify they match with `grep version package.json && grep Version CLAUDE.md`) 2. Run: `git add -A && git commit -m "chore: bump version to X.XXXX" && git push && npm run build && systemctl --user restart claudeman-web` -**Version**: 0.1634 (must match `package.json`) +**Version**: 0.1635 (must match `package.json`) ## Project Overview diff --git a/docs/voice-input-plan.md b/docs/voice-input-plan.md index ccbfc444..16f4d7a0 100644 --- a/docs/voice-input-plan.md +++ b/docs/voice-input-plan.md @@ -1,442 +1,155 @@ -# Voice Input Implementation Plan +# Voice Input V2 — Implementation Plan ## Executive Summary -Add a microphone button to Claudeman (desktop + mobile) that uses the **Web Speech API** for real-time speech-to-text. Users tap the mic, speak their prompt, see live interim transcription, review the text, and press Enter to send. Zero server cost, sub-200ms perceived latency, ~90% browser coverage (Chrome + Safari). +Fix and improve the existing VoiceInput implementation. The core class is solid but has **critical integration bugs** that prevent it from working on mobile, plus several UX improvements needed to make it feel fast and polished. --- -## Architecture Decision: Web Speech API (Primary, No Fallback for MVP) +## Current State: What Exists -### Why Web Speech API -- **Free** — no API keys, no server-side processing, no cost per minute -- **Fast** — interim results in ~150ms, feels real-time -- **Simple** — ~80 lines of JS, no server changes needed for MVP -- **Good coverage** — Chrome (desktop + Android) + Safari (desktop + iOS) = ~90% of users +The `VoiceInput` singleton (app.js:602-830) is already committed and uses the Web Speech API with: +- Toggle mode (tap start/stop), 5s silence auto-stop +- `interimResults: true` for streaming transcription preview +- iOS Safari `isFinal` workaround (750ms stability timer) +- Desktop button in `toolbar-right`, mobile button in `KeyboardAccessoryBar` +- `voice-pulse` CSS animation, `.voice-preview` overlay +- Cleanup on SSE reconnect, haptic feedback on mobile -### Why NOT a server-side fallback (for MVP) -- Firefox has <5% desktop share, even less on mobile -- Edge's Web Speech API is broken despite being Chromium (events don't fire) -- Adding Whisper/Deepgram requires API keys, server endpoints, audio streaming — significant complexity -- **Decision**: Hide the mic button on unsupported browsers. Add server fallback in Phase 2 if demand exists. +## Critical Bugs Found (Must Fix) -### Browser Support Matrix +### Bug 1: Mobile button NEVER shows (CRITICAL) +`KeyboardAccessoryBar.init()` runs at line 2239, BEFORE `VoiceInput.init()` at line 2240. The accessory bar template checks `VoiceInput.supported` at render time — but `init()` hasn't run yet, so `supported` is still `false`. The inline `style="${VoiceInput.supported ? '' : 'display:none'}"` always resolves to `display:none`. -| Browser | Support | Notes | -|---------|---------|-------| -| Chrome Desktop | YES | `webkitSpeechRecognition`, sends to Google servers | -| Chrome Android | YES | Same as desktop | -| Safari Desktop | YES | `webkitSpeechRecognition`, requires Siri enabled | -| Safari iOS | YES | Shows system permission modal | -| Firefox | NO | API exists behind flag but disabled by default | -| Edge | NO | Events don't fire despite API present | +**Fix:** Move `VoiceInput.init()` BEFORE `KeyboardAccessoryBar.init()`, OR remove the inline style check and have `VoiceInput.init()` show/hide the mobile button after the fact (like it does for desktop). -### Microphone Permissions -- HTTPS required, but **localhost is exempt** (Claudeman default) -- Claudeman with `--https` also works -- Chrome: persistent permission per domain (ask once) -- Safari: re-asks per page reload (less persistent, no workaround) -- Can pre-request via `getUserMedia()` on first interaction to avoid delay later +### Bug 2: `_showButtons()` ignores mobile button +`_showButtons()` only targets `#voiceInputBtn` (desktop). It never removes `display:none` from the mobile `[data-action="voice"]` button. ---- +**Fix:** Add mobile button selector to `_showButtons()`. -## UX Design +### Bug 3: Recognition instance leak on cleanup +`cleanup()` stops recording and removes the preview element, but doesn't null out `this.recognition`. After `cleanup()` + `init()` on SSE reconnect, the old `SpeechRecognition` instance with its handlers is orphaned. -### Interaction Model: Toggle Mode -- **Tap mic** → start recording (button turns red, pulses) -- **Tap again** → stop recording (text inserted into terminal) -- Auto-stop after **5 seconds of silence** as safety net -- Single-utterance mode (`continuous: false`) — perfect for prompt dictation +**Fix:** Add `this.recognition = null` in `cleanup()`. -Why toggle over push-to-talk: Terminal prompts are short, toggle is simpler, works one-handed on mobile, impossible to accidentally leave recording on with the auto-stop safety net. +## UX Improvements (Should Fix) -### Visual Feedback +### Improvement 1: Consider auto-sending after voice +Currently, voice text is inserted but the user must press Enter. This is safe but adds friction. Two options: +- **Option A (safe, current):** Insert text, user presses Enter — good for a terminal where wrong commands matter +- **Option B (fast):** Insert text + auto-send `\r` after a brief 500ms delay — feels more "voice assistant"-like +- **Recommendation:** Keep Option A as default, but add an optional setting for auto-send -**While recording:** -1. Mic button turns red with CSS pulsing animation (1.5s breathing cycle) -2. Interim text appears in a small overlay near the terminal bottom, styled in dim/italic to indicate "draft" -3. Text updates in real-time as user speaks (~150ms updates) +### Improvement 2: Shorter silence timeout for commands +5 seconds of silence before auto-stop feels slow for short terminal commands. Consider: +- 3 seconds for auto-stop (still generous for natural pauses) +- Or make it configurable via settings -**On completion:** -1. Final text inserted at terminal prompt via `sendInput()` -2. User reviews text, presses Enter when ready -3. Button returns to idle state (gray/outlined) +### Improvement 3: Better visual state on mobile +The blue-tinted voice button in the accessory bar is distinctive but subtle. When recording: +- The `.recording` class turns it red with pulse — good +- But the button is small among other buttons — easy to miss the state change +- Consider: also show a small red dot indicator in the header or terminal area during recording -### Text Flow +## Architecture Decision: Keep Web Speech API + +Confirmed by research: Web Speech API is the right choice. +- **Free, fast (150-300ms interim), trivial complexity** +- Chrome + Safari = ~70% of users, ~95% of Claudeman's target audience (devs on Chrome) +- Works on localhost without HTTPS +- Accuracy is adequate for English command dictation +- Deepgram streaming (Phase 2 optional) only if accuracy complaints arise +- Skip Whisper batch entirely (too slow for interactive voice input) + +## Implementation Plan + +### Phase 1: Fix Critical Bugs (Priority) + +**File: `src/web/public/app.js`** + +1. **Fix init order** — Move `VoiceInput.init()` BEFORE `KeyboardAccessoryBar.init()`: ``` -User speaks → SpeechRecognition interim result → Show in preview overlay - → Replace preview on each update - → SpeechRecognition final result → Insert text via sendInput() - → Clear preview overlay - → User presses Enter to submit +// Current (broken): +KeyboardAccessoryBar.init(); +VoiceInput.init(); + +// Fixed: +VoiceInput.init(); +KeyboardAccessoryBar.init(); ``` -**Critical: NO auto-submit.** User must press Enter. This is a terminal where wrong commands could be destructive. +2. **Fix `_showButtons()` to handle mobile** — Add mobile button selector: +```javascript +_showButtons() { + const desktopBtn = document.getElementById('voiceInputBtn'); + if (desktopBtn) desktopBtn.style.display = ''; + // Also show mobile button (may not exist yet if KeyboardAccessoryBar hasn't init'd) + const mobileBtn = document.querySelector('[data-action="voice"]'); + if (mobileBtn) mobileBtn.style.display = ''; +} +``` -### Error States +3. **Fix cleanup leak** — Null out recognition instance: +```javascript +cleanup() { + if (this.isRecording) this.stop(); + if (this.previewEl) { + this.previewEl.remove(); + this.previewEl = null; + } + this.recognition = null; // <-- add this + clearTimeout(this.silenceTimeout); + clearTimeout(this._stabilityTimer); + // ... rest +} +``` -| Error | Behavior | -|-------|----------| -| Browser unsupported | Hide mic button entirely (feature detection) | -| No mic permission | Show toast: "Microphone access needed" with retry link | -| Permission denied | Show toast: "Mic blocked. Check browser settings" | -| No speech detected (5s) | Auto-stop, show brief "No speech detected" toast | -| Network error | Show toast: "Voice input requires internet" | -| Accidental tap | Tap again immediately to cancel; ignore <0.5s recordings | +4. **Remove inline style from mobile button template** — Since `_showButtons()` will handle visibility, the template should always render the button visible and let `init()` hide it if unsupported: +``` +// Current (broken): +style="${VoiceInput.supported ? '' : 'display:none'}" ---- +// Fixed: remove the style attr entirely, let _showButtons/_hideButtons manage it +``` +Actually better: **always show the button** if we init VoiceInput before KeyboardAccessoryBar. The `VoiceInput.supported` will be set correctly by then. -## Implementation Details +### Phase 2: UX Polish -### Phase 1: MVP (Single PR) +5. **Reduce silence timeout** from 5s to 3s for snappier feel -#### Files to Modify +6. **Add recording indicator** — When recording, add a subtle pulsing red dot to the session header or status area so the recording state is visible even if the button is off-screen + +7. **Voice input setting** — Add a toggle in App Settings to enable/disable voice input (some users may not want the button). Default: enabled on supported browsers. + +### Phase 3: Future Enhancements (Not in this PR) + +- Language selector (currently hardcoded `en-US`) +- Auto-send option (insert text + `\r` automatically) +- Deepgram WebSocket fallback for Firefox/Edge +- Waveform visualization during recording +- Voice command recognition ("clear", "compact", "new session") + +## Files to Modify | File | Changes | |------|---------| -| `src/web/public/app.js` | `VoiceInput` class, keyboard shortcut, integration with `KeyboardAccessoryBar` | -| `src/web/public/index.html` | Desktop mic button in `toolbar-right` | -| `src/web/public/styles.css` | Desktop voice button styles, recording animation | -| `src/web/public/mobile.css` | Mobile voice button in accessory bar, recording animation | - -**No server-side changes needed.** Voice runs entirely in the browser. - -#### 1. VoiceInput Class (`app.js`) - -New singleton class (~80 lines) managing the Web Speech API lifecycle: - -```javascript -class VoiceInput { - constructor() { - this.recognition = null; - this.isRecording = false; - this.supported = !!(window.SpeechRecognition || window.webkitSpeechRecognition); - this.silenceTimeout = null; - this.previewEl = null; - - if (this.supported) { - const SR = window.SpeechRecognition || window.webkitSpeechRecognition; - this.recognition = new SR(); - this.recognition.continuous = false; - this.recognition.interimResults = true; - this.recognition.lang = 'en-US'; - this.recognition.maxAlternatives = 1; - - this.recognition.onresult = (e) => this._onResult(e); - this.recognition.onerror = (e) => this._onError(e); - this.recognition.onend = () => this._onEnd(); - } - } - - toggle() { this.isRecording ? this.stop() : this.start(); } - - start() { - if (!this.supported || this.isRecording) return; - this.isRecording = true; - this._updateButtons('recording'); - this._showPreview(''); - this.recognition.start(); - // Auto-stop after 5s silence - this._resetSilenceTimeout(); - } - - stop() { - if (!this.isRecording) return; - this.isRecording = false; - clearTimeout(this.silenceTimeout); - this._updateButtons('idle'); - this.recognition.stop(); - } - - _onResult(event) { - this._resetSilenceTimeout(); - let interim = '', final = ''; - for (let i = event.resultIndex; i < event.results.length; i++) { - const transcript = event.results[i][0].transcript; - if (event.results[i].isFinal) { - final += transcript; - } else { - interim += transcript; - } - } - - if (final) { - this._hidePreview(); - this._insertText(final); - this.stop(); - } else { - this._showPreview(interim); - // iOS workaround: isFinal is always false - this._iosStabilityCheck(interim); - } - } - - _onError(event) { - this.stop(); - if (event.error === 'not-allowed') { - app.showToast('Microphone access denied. Check browser settings.', 'error'); - } else if (event.error === 'no-speech') { - // Silent fail — auto-stop is enough - } else if (event.error === 'network') { - app.showToast('Voice input requires internet connection.', 'error'); - } - } - - _onEnd() { - if (this.isRecording) this.stop(); // Cleanup if ended unexpectedly - } - - _insertText(text) { - if (!app.activeSessionId || !text.trim()) return; - app.sendInput(text.trim()); - // Don't add \r — let user review and press Enter - } - - _resetSilenceTimeout() { - clearTimeout(this.silenceTimeout); - this.silenceTimeout = setTimeout(() => this.stop(), 5000); - } - - // iOS Safari: isFinal is always false. Detect stability. - _lastTranscript = ''; - _stabilityTimer = null; - _iosStabilityCheck(transcript) { - if (transcript !== this._lastTranscript) { - this._lastTranscript = transcript; - clearTimeout(this._stabilityTimer); - this._stabilityTimer = setTimeout(() => { - this._hidePreview(); - this._insertText(transcript); - this.stop(); - }, 750); - } - } - - _showPreview(text) { /* Update DOM overlay with interim text */ } - _hidePreview() { /* Remove DOM overlay */ } - _updateButtons(state) { /* Toggle .recording class on mic buttons */ } -} -``` - -#### 2. Mobile Button (KeyboardAccessoryBar) - -**Insert in HTML template** (after `/compact` button, before `paste`): -```html - -``` - -**Add to handleAction switch:** -```javascript -case 'voice': - if (typeof voiceInput !== 'undefined') voiceInput.toggle(); - break; -``` - -**Conditionally show:** Only render the button if `VoiceInput.supported` is true. - -#### 3. Desktop Button (index.html) - -**Insert in `toolbar-right`:** -```html -
- -
-``` - -Show via JS on init: `if (voiceInput.supported) voiceInputBtn.style.display = '';` - -#### 4. Keyboard Shortcut - -**Ctrl+Shift+V** (in `setupEventListeners` keydown handler): -```javascript -if ((e.ctrlKey || e.metaKey) && e.shiftKey && e.key === 'V') { - e.preventDefault(); - if (typeof voiceInput !== 'undefined') voiceInput.toggle(); -} -``` - -Why Ctrl+Shift+V: Ctrl+V is paste (sacred), but Ctrl+Shift+V ("paste without formatting") has no meaning in a terminal context. V for Voice is memorable. - -#### 5. CSS Styles - -**styles.css (desktop):** -```css -.btn-voice.recording { - background: rgba(239, 68, 68, 0.15); - border-color: rgba(239, 68, 68, 0.4); - color: #ef4444; - animation: voice-pulse 1.5s ease-in-out infinite; -} - -@keyframes voice-pulse { - 0%, 100% { transform: scale(1); opacity: 1; } - 50% { transform: scale(1.08); opacity: 0.8; } -} - -.voice-preview { - position: fixed; - bottom: 48px; - left: 50%; - transform: translateX(-50%); - background: rgba(0, 0, 0, 0.85); - color: rgba(255, 255, 255, 0.6); - font-style: italic; - padding: 6px 16px; - border-radius: 8px; - font-size: 0.85rem; - max-width: 80%; - white-space: nowrap; - overflow: hidden; - text-overflow: ellipsis; - z-index: 100; - pointer-events: none; -} -``` - -**mobile.css:** -```css -.accessory-btn-voice.recording { - background: rgba(239, 68, 68, 0.2); - border-color: rgba(239, 68, 68, 0.4); - color: #ef4444; - animation: voice-pulse 1.5s ease-in-out infinite; -} -``` - -#### 6. Interim Text Preview Overlay - -A small floating `
` showing live transcription: -```javascript -_showPreview(text) { - if (!this.previewEl) { - this.previewEl = document.createElement('div'); - this.previewEl.className = 'voice-preview'; - document.body.appendChild(this.previewEl); - } - this.previewEl.textContent = text || 'Listening...'; - this.previewEl.style.display = ''; -} - -_hidePreview() { - if (this.previewEl) { - this.previewEl.style.display = 'none'; - } -} -``` - -#### 7. Cleanup (Memory Leak Prevention) - -Following Claudeman's cleanup patterns: -- Store `VoiceInput` instance as `window.voiceInput` singleton -- On SSE reconnect (`handleInit()`): stop any active recording, reset state -- Remove preview element from DOM in cleanup -- Clear all timeouts (silenceTimeout, stabilityTimer) - -#### 8. Accessibility - -```html -