Compare commits

...
Author SHA1 Message Date
Codeman maintainer 26416f98de chore: version packages 2026-08-10 13:23:45 +02:00
Codeman maintainer 084d7b7328 fix(run-menu): let recent-session rows use the width the menu was given
PR #274 lifted the Run menu's 250px cap to `calc(100vw - 24px)` so a
recent-session row would have room for its worktree pill and parent path.
The rows never took it: `.run-mode-history` is a block scroller, so its
<button> rows are shrink-to-fit and stayed at ~250px inside a 1376px menu,
leaving ~1100px of empty dropdown and no space for `.hist-dir`'s
`flex: 1` + `text-align: right` to expand into.

Rows now fill the menu, and the menu is capped at the 760px one full row
actually costs rather than the whole window.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 13:21:40 +02:00
Ark0N a4cdb352be Merge pull request #244 from Lint111/feat/mobile-terminal-taps
fix(mobile): route terminal taps without breaking keyboard focus
2026-08-10 13:09:29 +02:00
Ark0N d81454b6f9 Merge pull request #275 from Ark0N/feat/claude-voice-integration
feat(voice): dictate through the server's Claude Code login, no API key
2026-08-10 13:09:24 +02:00
Ark0N 00f1b9228a Merge pull request #274 from jordan8037310/fix/run-menu-recent-sessions
fix(run-menu): make Recent Sessions rows legible on macOS (home-prefix regex + width + worktree)
2026-08-10 13:09:18 +02:00
Codeman maintainer 13d069e1e5 Merge remote-tracking branch 'origin/master' into feat/claude-voice-integration
# Conflicts:
#	CLAUDE.md
2026-08-10 12:56:40 +02:00
Codeman maintainer fa4c36c2a5 Merge remote-tracking branch 'origin/master' into pr274-rebase
# Conflicts:
#	src/web/public/session-ui.js
2026-08-10 12:55:32 +02:00
Ark0N fe2c03b2cc Merge pull request #276 from Ark0N/fix/home-path-abbreviation
fix(paths): one home-prefix helper, so path labels abbreviate on Linux and macOS
2026-08-10 12:53:22 +02:00
Ark0N 4e3f7ac36b Merge pull request #277 from Ark0N/feat/readmymind-phase3-part2
feat(readmymind): rethink steer note (phase 3 part 2)
2026-08-10 12:53:19 +02:00
Ark0N 089283e0b3 Merge pull request #278 from Ark0N/appsettings-details
One settings surface: App Settings, Session Options and Add Case
2026-08-10 12:52:43 +02:00
liorandClaude Opus 5 3b85001fed fix(mobile): keep the keyboard reachable when the viewport is scrolled up
Addresses the review on #244.

BLOCKING (item 1). selectSession() ends with scrollToLastNonEmptyLine(), which
parks the viewport above the bottom for any session taller than the screen, so
after a tab switch every tap classified as 'history' — touchstart ran
preventDefault() + blur, and touchend's early return skipped focus. Both routes
to focus closed on one gesture, the same mechanism as #173.

Suppressing the mouse REPORT while scrolled up is right and is kept; suppressing
FOCUS is not. touchstart now only preventDefaults 'content' taps (a scrolled-up
viewport sends nothing, so there is no compatibility click worth cancelling), and
the 'history' branch focuses instead of blurring.

Verified against the maintainer's own test, which was already on master and red:
`keeps the terminal input focusable after a tab switch parks the viewport
off-bottom` fails without this change and passes with it.

Item 2: dropped both `terminal-action-pending` guards. The class exists nowhere
in the repo, so both branches were permanently false and the comment promised
coverage that did not exist.

Item 3: removed the `Working` literals. Live claude 2.1.226 prints
"Cooked for 2m 6s" with a different bullet and a randomised verb, so they were
dead code. The status row is matched by its affordance ("esc to interrupt")
instead, which is what makes it actionable. The affordance regex is also
tightened to require a key or gesture name, so prose like "click here to open
the file" no longer dismisses the keyboard.

Item 4: removed _shouldForwardTouchScrollToApp and its test. It was never called,
and wiring it as written would have restricted forwarding to claude only,
dropping gemini from the path #205 established — a behaviour change this PR has
no reason to make.

Smaller items: the touchstart classification is cached and reused for the
touchend of the same gesture (keyed on exact coordinates, so a moved finger
re-classifies), removing two of the three full-viewport scans per gesture; the
duplicated touchLastX assignment is gone; and the no-touch bail-out returns null
rather than claiming 'history'.

test/mobile/keyboard.test.ts: 51 tests, 5 failed | 46 passed — the same 5
pre-existing failures as master, unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 13:32:11 +03:00
liorandClaude Opus 5 623fedf5b7 fix(mobile): keep the keyboard reachable on inert transcript taps
A mid-terminal tap on a claude-mode session left document.activeElement on
<body>, so the on-screen keyboard could not be raised and there was no way to
type — the blocker reduced upstream in #173.

_classifyMobileTerminalTap returns 'content' for any non-prompt row, and
_handleMobileTerminalTap blurred on every 'content' tap while touchstart's
preventDefault had already cancelled the compatibility click that would
otherwise focus xterm. Both routes to focus were closed on the same gesture.

Blur now applies only to rows that are actually TUI-owned. The distinguishing
signal is the affordance a CLI prints on or beside the row ("ctrl+r to expand",
"tap to collapse", "esc to interrupt"), not the row's title text — a readback's
title row carries no hint of its own, so the adjacent row is consulted too.
Keying on titles would recognise only the exact strings a fixture happens to
use and would let a real readback keep the keyboard open.

Measured with a real touchstart/touchend gesture, iPhone-class viewport,
claude-mode session, tapping mid-transcript:

  before  document.activeElement = body
  after   document.activeElement = xterm-helper-textarea

Note: upstream master already passes this assertion, so the added test is a
regression guard for this branch, not a test that fails on master.

test/mobile/keyboard.test.ts: 40 tests, 5 failed | 35 passed — the same 5
pre-existing failures as master (stale layout/accessory-bar expectations and a
CJK timeout), unchanged by this commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 13:22:59 +03:00
lior 1410362e5b fix(mobile): keep promptless terminal input focusable 2026-08-10 13:22:33 +03:00
lior 92ae46246c fix(mobile): route Claude terminal gestures 2026-08-10 13:21:15 +03:00
lior 6831d79127 fix(mobile): route terminal content taps to the CLI 2026-08-10 13:21:15 +03:00
lior b01ed611c4 fix(mobile): keep keyboard focus taps non-activating 2026-08-10 13:21:15 +03:00
Codeman maintainer ecc6f30e24 fix(voice): move Language and Domain keywords into the Provider group
Both are read by every engine (the Claude path sends the language as its base
tag and the keyterms as a recognition hint), but they sat under the "Deepgram
Nova-3" heading, which read as if they only applied to Deepgram. That group now
holds just the API key.

Ids are unchanged, so the getElementById load/save contract in settings-ui.js is
untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:17:54 +02:00
Codeman maintainer 4b51ba306e feat(voice): dictate through the server's Claude Code login, no API key
The mic button previously needed a Deepgram API key, or fell back to the
browser's Web Speech engine. It can now transcribe through the same
speech-to-text service Claude Code's own /voice mode uses, so anyone signed
in to Claude Code on the server gets dictation with no third-party account.

Claude Code's voice mode cannot be driven directly: it opens the HOST's
microphone (sox/arecord), and the CLI runs in a headless tmux pane while the
human is in a browser somewhere else. So capture stays in the browser and only
the transcription backend is borrowed.

Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the
page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since
MediaRecorder cannot emit raw PCM) and receives text.

- GET /api/voice/status reports readiness and never the token
- GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade
  guard as the terminal socket, plus caps on concurrency, stream length and
  frame size
- credentials are read-only: Codeman never refreshes them, since a refresh
  rotates the refresh token and could sign the user out of their own CLI
- claudeVoiceEnabled (synced, default OFF) gates the whole server side
- voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers
  Claude, then a configured Deepgram key, then the browser

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 11:19:51 +02:00
Codeman maintainer aaad031510 fix(paths): one home-prefix helper, so labels abbreviate on both platforms
The rule "show ~/project rather than /home/<user>/project" had three
implementations in the frontend, two of them platform-specific in opposite
directions, so each looked correct to whoever wrote it.

- The Run menu's Recent Sessions rows matched /home/<user>/ only. On macOS
  nothing was stripped, so every row spent its first ~19 characters on an
  identical /Users/<user>/ prefix and the left-to-right ellipsis removed the
  tail that identifies the row. That is #273, reported by @jordan8037310, who
  also traced why the menu's 250px cap made it worse: the width was chosen on
  the assumption the abbreviation had run.
- The case-manage list matched /Users/<user> only, the mirror image, so on a
  Linux host no case path was ever abbreviated there. Unreported.

Both now call _shortenHomePath(), which was already correct for both layouts
and already used by the Resume list, Cmd+K, the desktop home rail and the phone
overview. Its regex collapses to one alternation with a lookahead, so a path
that is exactly $HOME renders "~" instead of being left raw, matching what the
case-manage list used to do on macOS.

test/home-path-abbreviation.test.ts pins the helper on both layouts and the
rendered case-manage label, and fails if a fourth copy of the pattern appears in
src/web/public. The Run-menu guard counts helper calls rather than pinning a
source line, so it survives the row restructure in #274.

test/run-mode-ui.test.ts gains a _shortenHomePath stub: its harness loads
session-ui.js without terminal-ui.js, which the real app never does.

Verified against an isolated instance with 27 real cases and 50 history rows:
27 of 27 case paths and 17 of 20 Run menu rows abbreviate, the other 3 are
/tmp paths that correctly stay raw, tooltips keep the full path, no page errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 11:16:58 +02:00
Codeman maintainer 29efd0e970 fix(readmymind): style the modal footer, point the empty-result copy at the steer note
The footer buttons shipped with class="btn btn-secondary/primary", but no
.btn or .btn-secondary rule exists in this codebase, so all four rendered
as unstyled UA buttons. Moved them to the btn-toolbar convention every
other modal footer uses, with a scoped flex-row footer rule (btn-toolbar
is display:flex, block-level) mirroring the runSummaryModal footer.

Send's accent needs a (0,4,0) re-assert: the skin block's bare
.btn-toolbar rule is (0,2,1) under html:not([data-skin="og"]) and beats
.btn-toolbar.btn-primary (0,2,0), the same specificity trap CLAUDE.md
documents for mobile.css. Scoped to this modal; the repo-wide greying of
btn-primary on non-OG skins is pre-existing and left as a design call.

The empty-result copy now points at the steer note sitting right below
it ("Add a steer note and Rethink to try again"), zh-CN updated.

Verified with the steer E2E (still green) plus desktop, phone (390px),
and error-phase screenshots; static guards extended to pin the footer
convention and the accent re-assert.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 11:15:47 +02:00
Codeman maintainer 831af88579 feat(readmymind): rethink steer note (phase 3 part 2)
Adds the optional free-text steer note to the Read My Mind modal: a
dashed input under the suggestions ("no, I meant the mobile bug") that
rides along as `steer` on every Rethink. The API already accepted it;
this wires the frontend end of the contract.

- Shown whenever Rethink is live (ready AND empty-result phases),
  hidden only while a prediction runs; typed text survives re-runs.
- Enter in the field triggers Rethink, mirroring the prompt field's
  Enter-to-send; a fresh open clears it with the rethink memory.
- Trimmed and capped to the schema's 2000 chars on the way out; a
  plain open still sends an empty body (neither steer nor rejected).
- zh-CN strings for the placeholder and aria-label, phone-sized
  touch target in mobile.css, static guards in the phase-3 test.

Verified with a browser E2E against a live dev server (stubbed predict
endpoint): payload contents, phase visibility, Enter wiring, and
reset-on-reopen all asserted with real keystrokes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 10:41:39 +02:00
Jordan RyanandClaude Opus 5 3a106bd048 fix(run-menu): make Recent Sessions rows legible on macOS
Closes #273. Every row in the Run dropdown's Recent Sessions list rendered as
`/Users/<user>/co…`, indistinguishable from every other row.

The width was the symptom. The cause is that the home-prefix abbreviation
matched `/home/<user>/` only:

    s.workingDir.replace(/^\/home\/[^/]+\//, '~/')

On macOS the prefix is `/Users/<user>/`, so nothing was stripped and every row
spent its first ~19 characters on an identical prefix, with left-to-right
ellipsis cutting the only part that identifies it. The 250px menu cap was
chosen, per its own comment, as "the width at which the common `~/<dir>/<repo>`
+ timestamp recent-session row still fits whole" — sizing that assumes the
abbreviation ran. On Linux it does. On macOS the menu was permanently too
narrow for content it was never actually shortening, which is why this reads
as fine on one platform and broken on the other.

Changes:

- the regex matches `/home/` and `/Users/`
- the row leads with the identifying folder in semibold, with the parent path
  trailing, dimmed and right-aligned, so truncation removes context instead of
  identity
- the menu goes full width above 769px and the history list grows 200px -> 320px.
  Phones keep the compact popover deliberately: mobile.css positions this menu
  itself and a viewport-wide drawer there would cover the composer
- a worktree pill renders from the fields /api/history/sessions already returns
  unprojected (#266/#269), since a worktree's directory basename is often just
  the worktree name and rows stayed ambiguous without it
- a trailing `/.claude/worktrees` is trimmed from the displayed parent path once
  the pill states it, so the repo name stays visible

Verified in a browser at 1440px against a real 38-session history: menu 1416px,
0 of 34 rows clip their project name (was: all of them), 9 worktree pills
render, no page errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uTqt8ttmsBLXbm5JFHis3
2026-08-10 01:44:14 -04:00
35 changed files with 3049 additions and 109 deletions
+18
View File
@@ -1,5 +1,23 @@
# aicodeman
## 1.16.4
### Patch Changes
- **Voice dictation through your Claude Code login (no API key).** The mic button can now transcribe using this machine's existing Claude Code subscription, via the same speech-to-text service the CLI's own `/voice` mode uses. Off by default (`claudeVoiceEnabled`, synced): turning it on spends the server owner's Claude subscription on transcription for anyone who can reach the UI. The OAuth token never leaves the server process, credentials are read-only (Codeman never refreshes them, which would rotate the refresh token out from under the CLI), streams are capped at 5 minutes and 4 concurrent, and the WebSocket carries the same allowed-Host + same-site Origin guard as the terminal socket. A new Speech engine picker (Auto / Claude / Deepgram / Browser) sits alongside the existing Deepgram and Web Speech paths, which are untouched.
**One settings surface.** Session Options and Add Case now use the same `set-*` chrome as App Settings instead of the old modal-tab chrome, with a left rail, grouped rows, per-group device/synced scope badges and a search box. App Settings leads with version + update; the Session Options rail stays a real switcher (one section at a time) because Summary and Respawn are each long enough to bury the other. Collapsed Add Case blocks gained a disclosure chevron.
**Read My Mind: rethink steer note (phase 3 part 2).** Rethink now carries an optional free-text note ("no, I meant the mobile bug") sent as `steer`, the highest-authority signal the predictor gets. It stays in the field across re-runs, clears on each open, and the empty-result copy points at it. The modal footer moved to the styled `btn-toolbar` convention; the bare `btn btn-*` classes it shipped with match no CSS in this codebase and rendered as unstyled browser buttons.
**Mobile terminal taps no longer fight the keyboard.** Taps on TUI-owned rows (expandable readbacks, tool results, decision menus, the working/status row) now act on the CLI without popping the keyboard, while a tap on inert transcript text keeps the keyboard reachable. Rows are told apart by the affordance the CLI prints (`ctrl+r to expand`, `tap to collapse`, `esc to interrupt`) rather than by row titles, which vary per CLI and per version. A tap with the viewport scrolled up sends no mouse report at all but still restores focus, so the keyboard is reachable after every tab switch. Thanks to @Lint111.
**Path labels abbreviate `$HOME` on both platforms.** The "show `~/project`" rule had three implementations and two were platform-specific in opposite directions: the Run menu's matched `/home/<user>/` only, so on macOS every Recent Sessions row spent its first ~19 characters on an identical `/Users/<user>/` prefix and ellipsized away the tail that identifies it (#273); the case-manage list's matched `/Users/<user>` only, so no Linux case path was ever abbreviated. Both now route through one helper, with a static guard against a fourth copy appearing.
**Run menu Recent Sessions rows are legible.** Rows now read as folder, worktree pill, dimmed parent path, timestamp, with only the parent path allowed to shrink, so truncation can never hide which project (or which worktree) a row refers to. `<repo>/.claude/worktrees` is dropped from the parent path as noise. Thanks to @jordan8037310. Follow-up fix: the widened menu was not actually usable by its rows, since `.run-mode-history` is a block scroller and its `<button>` rows stayed shrink-to-fit at ~250px inside a full-window-width menu; rows now fill the menu and it is capped at the 760px one full row costs.
**Desktop home screen** no longer clips, and shows full tab names.
## 1.16.3
### Patch Changes
+6 -4
View File
@@ -74,7 +74,7 @@ When user says "COM":
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
**Version**: 1.16.3 (must match `package.json`)
**Version**: 1.16.4 (must match `package.json`)
## Project Overview
@@ -159,7 +159,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| **Search** | `src/search-service.ts` | Pure in-memory core for `GET /api/search` |
| **Attachments** | `src/attachment-registry.ts`, `attachment-magic`, `generated-artifact-attachments`, `session-attachment-history`, `document-preview-cache`, `document-thumbnailer`, `document-conversion-limiter`, `config/attachment-guard` | See Key Patterns |
| **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/` (`claude-md.ts` + `case-template.md`) | `templates/` holds the CLAUDE.md scaffold generated into new cases |
| **Web** | `src/web/server.ts` ★, `sse-events.ts`, `routes/*.ts` (23 modules + barrel; `session-routes.ts` ★), `route-helpers.ts`, `ports/*.ts`, `middleware/auth.ts`, `schemas.ts`, `self-update.ts`, `plan-usage-latest.ts`, `ws-connection-registry.ts`, `heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` | |
| **Web** | `src/web/server.ts` ★, `sse-events.ts`, `routes/*.ts` (24 modules + barrel; `session-routes.ts` ★), `route-helpers.ts`, `ports/*.ts`, `middleware/auth.ts`, `schemas.ts`, `self-update.ts`, `plan-usage-latest.ts`, `ws-connection-registry.ts`, `heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` | |
| **Frontend** | `src/web/public/app.js` (~5K lines, core) + 29 modules + `sw.js` | See Frontend section for the load order, which is authoritative |
| **Types** | `src/types/index.ts` (barrel) → 22 domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts |
@@ -210,7 +210,9 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Approvals Inbox** (cross-session queue of prompts waiting on a human; `approvalsInboxEnabled`, SYNCED, default OFF: every surface is opt-in; only the store and answer endpoints run regardless, so flipping it ON shows anything already pending): `web/approval-inbox.ts` is a `sessionWaits`-style singleton fed by `/api/hook-event`, holding at most ONE item per session (a new prompt supersedes), claude-mode only, in-memory. Cards are answered via `POST /api/approvals/:id/answer`, which sends a digit / Esc / idle-prompt text through `writeViaMux` (menu answers never carry `\r`). ⚠️ `option` digits are accepted ONLY when they match options parsed from the captured pane frame, and the answer path RE-CAPTURES the pane first (a dialog that no longer parses on screen means the keystroke would land in the composer, so refuse with 409). ⚠️ Resolution on the heuristic `working` signal is restricted to `idle` items; permission/question items clear only on definitive signals (`stop`, `elicitation_complete`/`elicitation_response`, exit/delete, answer, supersede, 12h TTL). The frontend seeds from `GET /api/approvals` in `handleInit` (which is what makes tab alerts survive reloads), but only with the setting ON; push Approve/Deny buttons are also gated on it (`sendPushNotifications` strips `actions`/`approvalId` when OFF) and are answered from `sw.js` directly so they work with no tab open. Surfaces (all gated on the setting): header bell (marker-hidden until count > 0, phones never show it) + drawer (`approvals-ui.js`), phone overview NEEDS YOU answer strips (`mobile-overview.js`). Design: `docs/approvals-inbox-plan.md`.
**Read My Mind intent profiles** (phase 1 of `docs/readmymind-plan.md`; `readMyMindEnabled`, SYNCED, default OFF): per-CASE profiles (user-stated `goals` + the user's recent real prompts), keyed by owner + realpath(workingDir) so they survive `/clear`/respawns and multi-user scoping is structural. Capture rides the transcript (`transcript:user_prompt` from `transcript-watcher.ts`), NOT the input paths: `POST /input` sees only programmatic prompts and the WS channel is raw keystrokes. The listener lives inside `startTranscriptWatcher()`'s `if (!watcher)` block (outside it would duplicate per hook event) and is claude-only + gated on the setting per event. Store: `src/intent-store.ts` singleton, `intents.json` written 0600 tmp+rename (prompts can contain secrets; never fed to `/api/search`). Endpoints: GET/PUT/DELETE `/api/sessions/:id/intent` + POST `/api/sessions/:id/readmymind` (`readmymind-routes.ts`, ownership via `findSessionOrFail` WITH `req`; registrations stay the bare `app.<method>('path')` shape, the endpoints.md drift scanner cannot see generics). **Phase 2 (predictor + 🧠 button)**: `readmymind-context.ts` is the PURE budgeted assembler (9 ranked sources, drop order siblings→away→workspace→tools, sections 1-4 truncate only); IO lives in `readmymind-collectors.ts` (transcript TAIL read — the live watcher keeps only a 500-char snippet — + git signals, skipped for remote-SSH cases) and the route; `readmymind-predictor.ts` reuses the AiCheckerBase spawn mechanics standalone (verdict-shaped base vs freeform JSON) as a mutable singleton routes call and tests stub. Claude-mode only (400), one in flight per session (409 CONFLICT), model = `readMyMindModel` setting defaulting to `AI_CHECK_MODEL` (opus, decided). Frontend `readmymind-ui.js`: header 🧠 marker-hidden (`btn-readmymind--hidden`) until the setting is ON; phones hide it in mobile.css and get a keyboard-accessory 🧠 key instead (ships in BOTH bar templates, revealed by the `rmm-enabled` class on the BAR element — setMode() rebuilds button innerHTML, so per-key state would be wiped; synced at init + every `applyHeaderVisibilitySettings()`). Alternate suggestions render as tappable rows that swap into the editable field without losing edits; Rethink rejects the whole shown set. Suggestions render via value/`textContent` ONLY and Send/Insert go through `POST /input` (server-side, so the sendEnterKey/local-echo trap does not apply) — nothing auto-sends, ever. User guide: `docs/readmymind.md`.
**Read My Mind intent profiles** (phase 1 of `docs/readmymind-plan.md`; `readMyMindEnabled`, SYNCED, default OFF): per-CASE profiles (user-stated `goals` + the user's recent real prompts), keyed by owner + realpath(workingDir) so they survive `/clear`/respawns and multi-user scoping is structural. Capture rides the transcript (`transcript:user_prompt` from `transcript-watcher.ts`), NOT the input paths: `POST /input` sees only programmatic prompts and the WS channel is raw keystrokes. The listener lives inside `startTranscriptWatcher()`'s `if (!watcher)` block (outside it would duplicate per hook event) and is claude-only + gated on the setting per event. Store: `src/intent-store.ts` singleton, `intents.json` written 0600 tmp+rename (prompts can contain secrets; never fed to `/api/search`). Endpoints: GET/PUT/DELETE `/api/sessions/:id/intent` + POST `/api/sessions/:id/readmymind` (`readmymind-routes.ts`, ownership via `findSessionOrFail` WITH `req`; registrations stay the bare `app.<method>('path')` shape, the endpoints.md drift scanner cannot see generics). **Phase 2 (predictor + 🧠 button)**: `readmymind-context.ts` is the PURE budgeted assembler (9 ranked sources, drop order siblings→away→workspace→tools, sections 1-4 truncate only); IO lives in `readmymind-collectors.ts` (transcript TAIL read — the live watcher keeps only a 500-char snippet — + git signals, skipped for remote-SSH cases) and the route; `readmymind-predictor.ts` reuses the AiCheckerBase spawn mechanics standalone (verdict-shaped base vs freeform JSON) as a mutable singleton routes call and tests stub. Claude-mode only (400), one in flight per session (409 CONFLICT), model = `readMyMindModel` setting defaulting to `AI_CHECK_MODEL` (opus, decided). Frontend `readmymind-ui.js`: header 🧠 marker-hidden (`btn-readmymind--hidden`) until the setting is ON; phones hide it in mobile.css and get a keyboard-accessory 🧠 key instead (ships in BOTH bar templates, revealed by the `rmm-enabled` class on the BAR element — setMode() rebuilds button innerHTML, so per-key state would be wiped; synced at init + every `applyHeaderVisibilitySettings()`). Alternate suggestions render as tappable rows that swap into the editable field without losing edits; Rethink rejects the whole shown set and carries the optional steer note (`#readMyMindSteer`, sent as `steer`, shown in ready + empty-result phases, cleared on each open). Suggestions render via value/`textContent` ONLY and Send/Insert go through `POST /input` (server-side, so the sendEnterKey/local-echo trap does not apply) — nothing auto-sends, ever. User guide: `docs/readmymind.md`.
**Voice dictation via Claude** (`claudeVoiceEnabled`, SYNCED, default OFF): the mic button can transcribe through this machine's Claude Code login instead of a Deepgram key, using the same speech-to-text service the CLI's own `/voice` mode uses. ⚠️ **Claude Code's voice mode itself is unusable here**: it opens the HOST's microphone (`sox`/`arecord`), and the CLI runs in a headless tmux pane while the human is in a browser elsewhere. So Codeman captures in the browser and borrows only the backend. Audio goes browser → Codeman → Anthropic (`src/web/voice-stream.ts`): the OAuth token never reaches the page, and the browser only sends PCM and receives text. ⚠️ Credentials are **read-only** (`src/claude-credentials.ts`) and Codeman never refreshes them — a refresh rotates the refresh token and could sign the user out of their own CLI; an elapsed token reports `expired` instead. ⚠️ Capture MUST be linear16/16 kHz/mono, so it uses an **AudioWorklet**, not MediaRecorder (which cannot emit raw PCM); `voice-pcm-worklet.js` is fetched from JS, so it is invisible to `cacheBustAssets` and borrows voice-input.js's `?v=` token — **edit the two together**. ⚠️ Transcript frames carry the WHOLE running transcript, not deltas: the Claude path replaces where the Deepgram path appends. Provider choice is `voiceSettings.provider` (`auto` prefers Claude → Deepgram → Web Speech). → `docs/claude-voice-plan.md`
**Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `docs/agent-teams/`.
@@ -318,7 +320,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
### API Routes
~200 handlers across 23 route files in `src/web/routes/`: system (45), sessions (34), cases (29), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), approvals (3), readmymind (4), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
~200 handlers across 24 route files in `src/web/routes/`: system (45), sessions (34), cases (29), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), approvals (3), readmymind (4), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), voice (1 + the `/ws/voice/stream` relay), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
**HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`).
+22
View File
@@ -471,6 +471,28 @@ All four enforce session ownership in multi-user mode; a foreign session id
answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the
same directory are distinct by construction.
## Voice dictation
Browser dictation transcribed through this server's Claude Code login, i.e. the
same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synced
`claudeVoiceEnabled` setting (default OFF). Design:
[`claude-voice-plan.md`](claude-voice-plan.md).
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
signed in to Claude Code on the server), `expired` (the access token elapsed;
running any Claude session refreshes it) or `malformed`. The OAuth token
itself is never returned by this or any other endpoint.
- `GET /ws/voice/stream?language=&keyterms=` (WebSocket, not under `/api`)
relays one dictation. Client sends binary frames of signed 16-bit
little-endian PCM, 16 kHz mono (<= 64 KB per frame), plus JSON control frames
`{"t":"finalize"}` (ask for the final transcript) and `{"t":"stop"}`. Server
sends `{"t":"ready"}`, `{"t":"transcript","text","final"}` (each frame is the
WHOLE running transcript, not a delta), `{"t":"error","message"}` and
`{"t":"closed"}`. Close codes: `4003` disallowed Host/Origin, `4004`
unavailable (reason in the close reason), `4008` too many concurrent streams.
Streams are capped in count and length (`src/config/voice.ts`).
## Authentication
Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque
+121
View File
@@ -0,0 +1,121 @@
# Claude voice dictation in Codeman
Wire Codeman's existing mic button to the same speech-to-text service Claude Code's own
`/voice` mode uses, so dictation works with **no third-party API key** for anyone already
signed in to Claude Code on the server.
## Why the CLI's own voice mode cannot be reused directly
Claude Code 2.1.x ships voice input: `/voice hold|tap|off` arms it, the CLI opens the
**host's** microphone (native `audio-capture-napi`, falling back to `sox`/`arecord` on Linux
after probing `/proc/asound/cards`), streams PCM upstream and types the transcript into its
own composer.
Every part of that is on the wrong machine for Codeman. The CLI runs inside a tmux pane on
the server, which is typically headless and has no sound card at all, while the human is in
a browser on a phone somewhere else. Toggling `/voice` in the pane from Codeman would arm a
microphone nobody is sitting in front of. So Codeman keeps capturing audio in the browser,
where the user actually is, and only borrows the CLI's **transcription backend**.
## The backend, as the CLI uses it
Extracted from the 2.1.226 binary (`connectVoiceStream`):
| | |
| --- | --- |
| URL | `wss://api.anthropic.com/api/ws/speech_to_text/voice_stream` |
| Query | `encoding=linear16`, `sample_rate=16000`, `channels=1`, `endpointing_ms=300`, `utterance_end_ms=1000`, `language=<lang>`, `use_conversation_engine=true`, `stt_provider=deepgram-nova3` |
| Headers | `Authorization: Bearer <Claude Code OAuth access token>`, `User-Agent`, `x-app: cli`, `anthropic-client-platform`, optional `x-config-keyterms` |
| Audio | raw binary frames, PCM signed 16-bit little-endian, 16 kHz, mono |
| Keepalive | `{"type":"KeepAlive"}` on open, then every 8 s |
| Finalize | `{"type":"CloseStream"}`, then wait for the endpoint frame |
| Downstream | `{"type":"TranscriptText"\|"TranscriptInterim","data":"…"}` (running interim), `{"type":"TranscriptEndpoint"}` (promotes the pending interim to final), `{"type":"TranscriptError",…}`, `{"type":"error","message":…}` |
Deepgram Nova-3 runs server-side, so the Deepgram-quality result arrives without a Deepgram
account. Verified against the live endpoint before this design was written: connect, stream
PCM, receive interims and an endpoint frame.
## Architecture
The browser cannot call that endpoint itself: it would need the OAuth bearer token in page
JavaScript (and CORS would refuse anyway). So the audio goes browser → Codeman → Anthropic,
and Codeman is the only thing that ever touches the token.
```
mic → AudioWorklet (Float32 → PCM16 @16 kHz)
→ wss://<codeman>/ws/voice/stream [cookie/basic auth, Origin+Host guarded]
→ VoiceStreamRelay (reads ~/.claude/.credentials.json per connect)
→ wss://api.anthropic.com/api/ws/speech_to_text/voice_stream
← {"t":"transcript","text":…,"final":…} → existing _insertText() path
```
Nothing about the insert path changes: the transcript lands in the same preview overlay,
the same direct/compose insert modes, the same green Send button.
### Server pieces
- **`src/claude-credentials.ts`** — locate and parse the Claude Code OAuth credentials.
`parseClaudeCredentials()` is pure (JSON string + `now` → status) and unit-tested;
`readClaudeOAuthToken()` wraps it with IO: `$CLAUDE_CONFIG_DIR/.credentials.json` or
`~/.claude/.credentials.json`, and on macOS the login keychain
(`security find-generic-password -s "Claude Code-credentials"`).
**Read-only, always.** Codeman never writes credentials and never refreshes the token: a
refresh rotates the refresh token, and racing Claude Code's own refresh could sign the
user out of their CLI. An expired token surfaces as a plain "run a Claude session to
refresh" error instead.
The token is never logged, never returned by any endpoint, and never sent to the browser.
- **`src/web/voice-stream.ts`** — pure `buildVoiceStreamUrl()` / `buildVoiceStreamHeaders()` /
`sanitizeKeyterms()` (ASCII-only, deduped, 1024-char cap, mirroring the CLI), plus
`VoiceStreamRelay`, which owns one upstream socket: keepalive timer, audio passthrough,
transcript translation, finalize, and the caps below.
- **`src/web/routes/voice-routes.ts`**
- `GET /api/voice/status` → `{ available, reason, subscriptionType?, expiresAt? }`. Never
the token. `available:false` with a machine-readable `reason` (`disabled`, `no-credentials`,
`expired`) is what the settings row and the provider resolver read.
- `GET /ws/voice/stream?language=&keyterms=` → the relay. Same upgrade guard as
`/ws/sessions/:id/terminal`: allowed Host, same-site Origin, and the global auth hook has
already run on the handshake.
Caps, because an open mic is an open pipe: one stream per connection, `MAX_VOICE_STREAMS`
concurrent server-wide, a hard `MAX_STREAM_MS` per stream, and a per-frame size cap. A tab
left recording cannot bill an unbounded amount of upstream audio.
### Frontend pieces
- **`voice-pcm-worklet.js`** — an `AudioWorkletProcessor` converting Float32 blocks to PCM16
and posting ~256 ms frames back. `MediaRecorder` cannot produce raw PCM, which is why the
existing Deepgram path (container audio, auto-detected) cannot be reused as-is. Falls back
to `ScriptProcessorNode` where AudioWorklet is unavailable.
- **`ClaudeVoiceProvider`** in `voice-input.js` — mirrors `DeepgramProvider`'s shape
(`start({language, keyterms, onStream, onResult, onError, onEnd})`) so `VoiceInput` treats
the three providers uniformly.
- **Provider resolution** — new `voiceSettings.provider`: `auto` (default) | `claude` |
`deepgram` | `webspeech`. `auto` picks Claude when `/api/voice/status` reports it
available, else Deepgram when a key is set, else Web Speech. Pinning a provider always
wins, so an existing Deepgram user can keep exactly what they have.
### Settings
- `claudeVoiceEnabled` — synced, **default OFF**, gating the whole server side. Off is the
honest default: turning it on means this machine's Claude subscription starts paying for
transcription for whoever can reach the UI, and the audio goes to Anthropic rather than to
wherever it went before. One switch in Settings → Voice, and the mic works with no key.
- `voiceSettings.provider` — per the resolution table above; joins the existing synced
`voiceSettings` object.
## Things worth knowing
- **This uses an undocumented endpoint with subscription credentials.** It is the user's own
token, on the user's own machine, driving the user's own Claude Code install, but it is not
a published API and Anthropic can change or restrict it. Default-OFF is deliberate; the
Deepgram and Web Speech paths stay untouched as the supported fallbacks.
- **Multi-user mode**: every user's dictation would run on the server owner's Claude
credentials, exactly as every user's *sessions* already run on them. Consistent, but worth
stating out loud in the settings copy.
- **Token lifetime** is about 8 hours, refreshed by Claude Code itself whenever it runs. The
relay re-reads the file on every connect rather than caching, so a refresh is picked up on
the next press of the mic.
- **HTTPS or localhost**: `getUserMedia` needs a secure context. Prod is HTTPS behind
`tailscale serve`, so this is already satisfied; the existing error copy covers the rest.
+1 -1
View File
@@ -124,7 +124,7 @@ Agent use cases this unlocks: a lead session records intentions as the user stat
1. **Intent store + capture + intent endpoints + skill docs.** Immediately useful to agents even before any UI exists.
2. **Context assembler + predictor + predict endpoint + desktop button/modal.** The feature as pitched. The assembler ships with all collectors it can serve from day one (transcript, intent, git, run-summary, siblings); the approvals collector activates when PR #245 lands.
3. **Phone accessory key, rethink steering, alternates row.** Part 1 (shipped): the alternates row (tappable, swap into the field without losing edits; Rethink rejects the whole shown set), the phone 🧠 keyboard-accessory key (both bar templates, `rmm-enabled` marker class on the bar), and a phone-sized modal (small dialog, not full-screen). Part 2: rethink steering (the free-text steer note; the API already accepts `steer`).
3. **Phone accessory key, rethink steering, alternates row.** Part 1 (shipped): the alternates row (tappable, swap into the field without losing edits; Rethink rejects the whole shown set), the phone 🧠 keyboard-accessory key (both bar templates, `rmm-enabled` marker class on the bar), and a phone-sized modal (small dialog, not full-screen). Part 2 (shipped): rethink steering, the free-text steer note under the suggestions, sent as `steer`, visible whenever Rethink is live (ready and empty-result phases), cleared on each open; the empty-result copy points at the note, and the footer buttons moved to the styled `btn-toolbar` convention (the bare `btn btn-*` classes they shipped with match no CSS in this codebase and rendered as unstyled UA buttons).
4. Explicitly later: proactive predict-on-idle (ghost suggestion chip), auto-compaction of `recentPrompts` into `goals` via a cheap model, codex/gemini capture, cross-case "global" intent.
## Open questions
+2 -2
View File
@@ -27,7 +27,7 @@ On a Claude session, press the brain button in the header (desktop) or the 🧠
- **Send** submits it to the session (with Enter).
- **Insert** drops it on the CLI composer *without* Enter, so you can edit it in the terminal before sending.
- **Rethink** re-runs with everything shown (the field and the alternates) recorded as rejected.
- **Rethink** re-runs with everything shown (the field and the alternates) recorded as rejected. An optional steer note below the suggestions ("no, I meant the mobile bug") rides along as your own words, the highest-authority signal the predictor gets; it stays in the field across re-runs until you clear it or reopen the modal.
- **Dismiss** closes; nothing happens.
A prediction takes 5-90 seconds and costs real tokens; one runs per session at a time. If the session is sitting on a permission/question dialog, the suggestion is usually an answer to that dialog: that is intentional.
@@ -87,7 +87,7 @@ The `codeman` agent skill documents the same verbs (SKILL.md §3 plus `reference
## What comes next
A steer-note input on Rethink ("no, I meant the mobile bug"; the API already accepts `steer`). Explicitly later: proactive predict-on-idle, auto-compaction of the prompt history into goals, non-Claude capture. See the phases section of [`readmymind-plan.md`](readmymind-plan.md).
Explicitly later: proactive predict-on-idle, auto-compaction of the prompt history into goals, non-Claude capture. See the phases section of [`readmymind-plan.md`](readmymind-plan.md).
## Troubleshooting
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.16.3",
"version": "1.16.4",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.16.3",
"version": "1.16.4",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.16.3",
"version": "1.16.4",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
+126
View File
@@ -0,0 +1,126 @@
/**
* @fileoverview Read-only access to the Claude Code OAuth credentials.
*
* Claude Code stores its subscription OAuth tokens in
* `$CLAUDE_CONFIG_DIR/.credentials.json` (default `~/.claude/.credentials.json`,
* mode 0600) on Linux/Windows, and in the login keychain on macOS. Codeman reads
* the access token to authenticate the voice-dictation relay
* (`src/web/voice-stream.ts`) against the same speech-to-text service the CLI's
* own `/voice` mode uses.
*
* ⚠️ READ-ONLY, deliberately. Codeman never writes this file and never performs
* an OAuth refresh: a refresh ROTATES the refresh token, so racing Claude Code's
* own refresh could invalidate the user's CLI login. An expired access token is
* reported as `expired` and the caller tells the user to run a Claude session
* (which refreshes it) instead.
*
* ⚠️ The token is a bearer secret: it is never logged, never persisted, never
* included in any API response, and never sent to the browser.
*/
import { readFile } from 'fs/promises';
import { execFile } from 'child_process';
import { homedir, userInfo } from 'os';
import { join } from 'path';
/** Result of inspecting the credential store. The token is present only on 'ok'. */
export type ClaudeCredentialStatus = 'ok' | 'expired' | 'missing' | 'malformed';
export interface ClaudeOAuthCredentials {
status: ClaudeCredentialStatus;
/** Bearer token. Present only when status is 'ok'. Never log or serialize this. */
accessToken?: string;
/** Epoch ms the access token expires at, when the store reports one. */
expiresAt?: number;
/** e.g. 'max', 'pro'. Display-only, safe to surface. */
subscriptionType?: string;
}
/** Skew applied to the stored expiry so a token that dies mid-stream is refused up front. */
const EXPIRY_SKEW_MS = 60_000;
/** macOS keychain service holding the same JSON blob as `.credentials.json`. */
const KEYCHAIN_SERVICE = 'Claude Code-credentials';
/** Keychain lookups shell out; keep them short so a locked keychain cannot hang a request. */
const KEYCHAIN_TIMEOUT_MS = 3000;
/**
* Parse a `.credentials.json` payload. Pure: no IO, no clock read (pass `now`),
* so the expiry and shape handling are unit-testable.
*
* Returns 'malformed' for anything that is not the expected `claudeAiOauth`
* shape rather than throwing — a hand-edited or half-written file must degrade
* to "voice unavailable", never to a 500.
*/
export function parseClaudeCredentials(raw: string, now: number): ClaudeOAuthCredentials {
let parsed: unknown;
try {
parsed = JSON.parse(raw);
} catch {
return { status: 'malformed' };
}
if (!parsed || typeof parsed !== 'object') return { status: 'malformed' };
const oauth = (parsed as { claudeAiOauth?: unknown }).claudeAiOauth;
if (!oauth || typeof oauth !== 'object') return { status: 'malformed' };
const record = oauth as Record<string, unknown>;
const accessToken = typeof record.accessToken === 'string' ? record.accessToken.trim() : '';
if (!accessToken) return { status: 'malformed' };
const expiresAt = typeof record.expiresAt === 'number' ? record.expiresAt : undefined;
const subscriptionType = typeof record.subscriptionType === 'string' ? record.subscriptionType : undefined;
// An expired token is a real state (the CLI refreshes on its next run), not a
// malformed store: report it separately so the UI can say something useful.
if (expiresAt !== undefined && expiresAt - EXPIRY_SKEW_MS <= now) {
return { status: 'expired', expiresAt, subscriptionType };
}
return { status: 'ok', accessToken, expiresAt, subscriptionType };
}
/** Path of the credentials file, honoring CLAUDE_CONFIG_DIR like the CLI does. */
export function claudeCredentialsPath(env: NodeJS.ProcessEnv = process.env): string {
const configDir = typeof env.CLAUDE_CONFIG_DIR === 'string' && env.CLAUDE_CONFIG_DIR.trim();
return join(configDir || join(homedir(), '.claude'), '.credentials.json');
}
/** Read the macOS keychain entry. Resolves to null on any failure (locked, absent, non-mac). */
function readKeychainCredentials(): Promise<string | null> {
return new Promise((resolve) => {
execFile(
'security',
['find-generic-password', '-a', userInfo().username, '-w', '-s', KEYCHAIN_SERVICE],
{ encoding: 'utf-8', timeout: KEYCHAIN_TIMEOUT_MS },
(err, stdout) => resolve(err ? null : stdout.trim() || null)
);
});
}
/**
* Locate and parse the Claude Code OAuth credentials.
*
* File first (present on every platform once the CLI has run there), keychain
* second on macOS. Never caches: Claude Code rewrites the store roughly every
* 8 hours, and a cached token would go stale inside a long-lived server.
*/
export async function readClaudeOAuthCredentials(now: number = Date.now()): Promise<ClaudeOAuthCredentials> {
let fileResult: ClaudeOAuthCredentials | null = null;
try {
fileResult = parseClaudeCredentials(await readFile(claudeCredentialsPath(), 'utf-8'), now);
} catch {
fileResult = null;
}
if (fileResult && fileResult.status !== 'malformed') return fileResult;
if (process.platform === 'darwin') {
const raw = await readKeychainCredentials();
if (raw) {
const keychainResult = parseClaudeCredentials(raw, now);
if (keychainResult.status !== 'malformed') return keychainResult;
}
}
return fileResult ?? { status: 'missing' };
}
+56
View File
@@ -0,0 +1,56 @@
/**
* @fileoverview Bounds and endpoint config for Claude voice dictation.
*
* Backs the browser → Codeman → Anthropic dictation relay (`src/web/voice-stream.ts`,
* `src/web/routes/voice-routes.ts`; design in `docs/claude-voice-plan.md`).
*
* Why everything here is bounded: an open microphone is an open pipe. Each live
* stream holds a browser socket, an upstream socket and a keepalive timer, and
* every second of audio is billed against the server owner's Claude subscription.
* A tab left recording (phone in a pocket, forgotten laptop) must cost a bounded
* amount, so streams die on their own at `MAX_STREAM_MS` and the server refuses
* more than `MAX_CONCURRENT_STREAMS` at once.
*
* The audio frame cap is a memory guard on a socket that carries attacker-shaped
* binary data: PCM16 at 16 kHz mono is 32 KB/s, so a 256 ms frame is ~8 KB and
* anything near 64 KB is either a broken client or an attempt to make the relay
* buffer for someone else.
*/
/** Upstream speech-to-text service (the one Claude Code's own `/voice` mode uses). */
export const VOICE_STREAM_HOST = 'wss://api.anthropic.com';
/** Path of the streaming speech-to-text endpoint. */
export const VOICE_STREAM_PATH = '/api/ws/speech_to_text/voice_stream';
/**
* Base override, for tests (point the relay at a local mock) and for users on an
* Anthropic-compatible gateway. Must be a ws:// or wss:// origin.
*/
export function voiceStreamBase(env: NodeJS.ProcessEnv = process.env): string {
const override = typeof env.CODEMAN_VOICE_STREAM_BASE === 'string' ? env.CODEMAN_VOICE_STREAM_BASE.trim() : '';
if (override && /^wss?:\/\//.test(override)) return override.replace(/\/+$/, '');
return VOICE_STREAM_HOST;
}
/** Upstream drops an idle socket; the CLI pings at 8s and so do we. */
export const KEEPALIVE_INTERVAL_MS = 8000;
/** Hard ceiling on one dictation. Long enough for any real utterance, short enough to bound a forgotten mic. */
export const MAX_STREAM_MS = 5 * 60_000;
/** Concurrent relays server-wide. Dictation is a human-paced, one-at-a-time act. */
export const MAX_CONCURRENT_STREAMS = 4;
/** Largest single audio frame accepted from the browser (~2s of PCM16 @16 kHz mono). */
export const MAX_AUDIO_FRAME_BYTES = 64 * 1024;
/** How long to wait for the final transcript after the client asks to finalize. */
export const FINALIZE_TIMEOUT_MS = 3000;
/** Upstream caps the keyterms header; mirrors the CLI's own limit. */
export const MAX_KEYTERMS_HEADER_CHARS = 1024;
/** Audio format the endpoint is opened with. The browser worklet must match exactly. */
export const AUDIO_SAMPLE_RATE = 16000;
export const AUDIO_CHANNELS = 1;
+2
View File
@@ -19,6 +19,8 @@ export interface ConfigPort {
getTerminalHistoryConfig(): Promise<TerminalHistoryConfig>;
/** Synced `agentSkillEnabled` app setting (default OFF); gates per-case agent-skill injection. */
getAgentSkillEnabled(): Promise<boolean>;
/** Synced `claudeVoiceEnabled` app setting (default OFF); gates the Claude voice dictation relay. */
getClaudeVoiceEnabled(): Promise<boolean>;
getDefaultClaudeMdPath(): Promise<string | undefined>;
getLightState(identity?: { username: string; role: 'admin' | 'user' }): unknown;
getLightSessionsState(): unknown[];
+4 -1
View File
@@ -255,12 +255,15 @@
'Read My Mind: predict your next prompt': '读心术:预测您的下一条提示',
'Predict my next prompt': '预测我的下一条提示',
'Reading your mind…': '正在读取您的想法…',
'No suggestion this time. Rethink to try again.': '这次没有建议。点击「重想」再试一次。',
'No suggestion this time. Add a steer note and Rethink to try again.':
'这次没有建议。可添加引导备注后点击「重想」再试一次。',
Rethink: '重想',
Insert: '插入',
"Put the text on the session's composer without submitting it": '将文本放入会话输入框但不提交',
'Predicted prompt, editable': '预测的提示,可编辑',
'Use this suggestion instead': '改用此建议',
"Steer the rethink, e.g. 'no, I meant the mobile bug'": '引导重想,例如:"不,我是指移动端的问题"',
'Steer note for Rethink': '重想的引导备注',
'Select a session first': '请先选择一个会话',
'Read My Mind works on Claude sessions only': '读心术仅适用于 Claude 会话',
'Prompt sent': '提示已发送',
+66 -24
View File
@@ -2144,6 +2144,18 @@
<div class="set-row-text"><span class="set-row-label">Active provider</span></div>
<span class="voice-provider-status" id="voiceProviderStatus">&mdash;</span>
</div>
<div class="set-row has-field" data-search="voice provider claude deepgram web speech engine">
<div class="set-row-text">
<span class="set-row-label">Speech engine</span>
<span class="set-row-desc">Auto prefers Claude when this server can transcribe, then Deepgram, then the browser.</span>
</div>
<select id="voiceProvider" class="set-select">
<option value="auto">Auto</option>
<option value="claude">Claude</option>
<option value="deepgram">Deepgram</option>
<option value="webspeech">Browser (Web Speech)</option>
</select>
</div>
<div class="set-row has-field" data-search="voice insert mode compose">
<div class="set-row-text">
<span class="set-row-label">Insert mode</span>
@@ -2154,6 +2166,45 @@
<option value="compose">Compose dialog</option>
</select>
</div>
<div class="set-row has-field" data-search="voice language">
<div class="set-row-text">
<span class="set-row-label">Language</span>
<span class="set-row-desc">Applies to whichever engine is active.</span>
</div>
<select id="voiceLanguage" class="set-select">
<option value="en-US">English (US)</option>
<option value="en-GB">English (UK)</option>
<option value="es">Spanish</option>
<option value="fr">French</option>
<option value="de">German</option>
<option value="ja">Japanese</option>
<option value="multi">Auto-detect (multi)</option>
</select>
</div>
<div class="set-row has-field" data-search="voice keyterms domain keywords">
<div class="set-row-text">
<span class="set-row-label">Domain keywords</span>
<span class="set-row-desc">Comma-separated terms to boost recognition accuracy. Sent to Claude and Deepgram alike.</span>
</div>
<input type="text" id="voiceKeyterms" class="set-input" placeholder="codeman, tmux, respawn, subagent">
</div>
</div>
</div>
<div class="set-group">
<div class="set-group-head"><h4>Claude</h4><span class="set-scope">synced</span></div>
<div class="set-group-body">
<div class="set-row" data-search="claude voice dictation subscription no api key">
<div class="set-row-text">
<span class="set-row-label">Transcribe with this server's Claude login</span>
<span class="set-row-desc">Dictation with no API key, through the same service Claude Code's own /voice mode uses. Microphone audio goes to Anthropic and is billed to this machine's Claude subscription, for everyone who can reach this UI.</span>
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsClaudeVoice"><span class="slider"></span></label>
</div>
<div class="set-row" data-search="claude voice status credentials">
<div class="set-row-text"><span class="set-row-label">Server status</span></div>
<span class="voice-provider-status" id="voiceClaudeStatus" data-i18n-skip>&mdash;</span>
</div>
</div>
</div>
@@ -2170,25 +2221,6 @@
<button class="btn-toolbar btn-sm" onclick="app.toggleDeepgramKeyVisibility()" id="voiceKeyToggleBtn" type="button">Show</button>
</div>
</div>
<div class="set-row has-field" data-search="voice language">
<div class="set-row-text"><span class="set-row-label">Language</span></div>
<select id="voiceLanguage" class="set-select">
<option value="en-US">English (US)</option>
<option value="en-GB">English (UK)</option>
<option value="es">Spanish</option>
<option value="fr">French</option>
<option value="de">German</option>
<option value="ja">Japanese</option>
<option value="multi">Auto-detect (multi)</option>
</select>
</div>
<div class="set-row has-field" data-search="voice keyterms domain keywords">
<div class="set-row-text">
<span class="set-row-label">Domain keywords</span>
<span class="set-row-desc">Comma-separated terms to boost recognition accuracy.</span>
</div>
<input type="text" id="voiceKeyterms" class="set-input" placeholder="codeman, tmux, respawn, subagent">
</div>
</div>
</div>
</section>
@@ -3107,13 +3139,23 @@
<div class="readmymind-why" id="readMyMindWhy" data-i18n-skip></div>
<div class="readmymind-alternates" id="readMyMindAlternates" data-i18n-skip style="display:none"></div>
</div>
<div class="readmymind-error" style="display:none">No suggestion this time. Rethink to try again.</div>
<div class="readmymind-error" style="display:none">No suggestion this time. Add a steer note and Rethink to try again.</div>
<!-- Rethink steer note: the user's own words about what they actually
meant, sent as `steer` with the next Rethink. Shown whenever
Rethink is live (ready AND empty-result phases), hidden while a
prediction runs. -->
<div class="readmymind-steer-row" id="readMyMindSteerRow" style="display:none">
<input type="text" id="readMyMindSteer" class="readmymind-steer-input" maxlength="2000" placeholder="Steer the rethink, e.g. 'no, I meant the mobile bug'" aria-label="Steer note for Rethink" onkeydown="if(event.key==='Enter')app.rethinkReadMyMind()">
</div>
</div>
<!-- btn-toolbar, not bare "btn btn-*": no .btn/.btn-secondary rule exists
in this codebase, so the bare classes render as unstyled UA buttons.
btn-toolbar is the convention every other modal footer uses. -->
<div class="modal-footer">
<button class="btn btn-secondary" onclick="app.closeReadMyMind()">Dismiss</button>
<button class="btn btn-secondary" id="readMyMindRethink" onclick="app.rethinkReadMyMind()">Rethink</button>
<button class="btn btn-secondary" onclick="app.sendReadMyMind(false)" title="Put the text on the session's composer without submitting it">Insert</button>
<button class="btn btn-primary" onclick="app.sendReadMyMind(true)">Send</button>
<button class="btn-toolbar" onclick="app.closeReadMyMind()">Dismiss</button>
<button class="btn-toolbar" id="readMyMindRethink" onclick="app.rethinkReadMyMind()">Rethink</button>
<button class="btn-toolbar" onclick="app.sendReadMyMind(false)" title="Put the text on the session's composer without submitting it">Insert</button>
<button class="btn-toolbar btn-primary" onclick="app.sendReadMyMind(true)">Send</button>
</div>
</div>
</div>
+7 -5
View File
@@ -1318,20 +1318,22 @@ html.mobile-init .file-browser-panel {
width: calc(100% - 2rem);
}
/* Four footer buttons on a narrow phone: let them wrap instead of clipping,
and give buttons + alternate rows finger-sized targets. */
and give buttons + alternate rows finger-sized targets. The flex row
itself comes from the base rule in styles.css. */
.readmymind-modal .modal-footer {
display: flex;
flex-wrap: wrap;
justify-content: flex-end;
gap: 0.5rem;
}
.readmymind-modal .modal-footer .btn {
.readmymind-modal .modal-footer .btn-toolbar {
flex: 1 1 auto;
justify-content: center;
min-height: 38px;
}
.readmymind-alt {
min-height: 38px;
}
.readmymind-steer-input {
min-height: 38px;
}
/* Modal safe area padding - all sides for full-screen modals */
.ios-device .modal-content {
+20 -4
View File
@@ -10,7 +10,8 @@
* suggestions render as tappable alternate rows that swap into the field
* without losing edits. Buttons are Send (with Enter), Insert (drop on the CLI
* composer WITHOUT Enter, for editing), Rethink (re-run with the whole shown
* set, main + alternates, recorded as rejected), Dismiss.
* set, main + alternates, recorded as rejected, plus the optional free-text
* steer note, e.g. "no, I meant the mobile bug", sent as `steer`), Dismiss.
*
* Suggestions are NEVER auto-sent: the explicit click here is the security
* boundary for observed/injectable predictor inputs, so suggestion text is
@@ -47,8 +48,11 @@ Object.assign(CodemanApp.prototype, {
this.showToast('Read My Mind works on Claude sessions only', 'warning');
return;
}
// Rethink memory resets on each open (a fresh open is a fresh question).
// Rethink memory resets on each open (a fresh open is a fresh question),
// and the steer note resets with it.
this._rmm = { sessionId, suggestions: [], selected: 0, rejected: [], busy: false };
const steer = document.getElementById('readMyMindSteer');
if (steer) steer.value = '';
document.getElementById('readMyMindModal')?.classList.add('active');
this._readMyMindPredict();
},
@@ -65,7 +69,13 @@ Object.assign(CodemanApp.prototype, {
state.busy = true;
this._rmmSetPhase('loading');
const body = state.rejected.length > 0 ? { rejected: state.rejected.slice(-10) } : {};
const body = {};
if (state.rejected.length > 0) body.rejected = state.rejected.slice(-10);
// The steer note rides every re-run while it stays in the field: what the
// user sees in the box is what the predictor gets. Empty on first open
// (openReadMyMind clears it), so a plain predict sends neither key.
const steer = document.getElementById('readMyMindSteer')?.value.trim() ?? '';
if (steer) body.steer = steer.slice(0, 2000);
const data = await this._apiJson(`/api/sessions/${state.sessionId}/readmymind`, { method: 'POST', body });
// The modal may have been dismissed (or reopened for another session) while
@@ -176,7 +186,8 @@ Object.assign(CodemanApp.prototype, {
},
/** Re-run with the whole shown set (main + alternates) recorded as rejected:
* the user saw every row and asked for something else. */
* the user saw every row and asked for something else. The steer note (if
* any) is read from the field by _readMyMindPredict itself. */
rethinkReadMyMind() {
const state = this._rmm;
if (!state || state.busy) return;
@@ -193,6 +204,11 @@ Object.assign(CodemanApp.prototype, {
modal.querySelector('.readmymind-loading').style.display = phase === 'loading' ? '' : 'none';
modal.querySelector('.readmymind-result').style.display = phase === 'ready' ? '' : 'none';
modal.querySelector('.readmymind-error').style.display = phase === 'error' ? '' : 'none';
// The steer note belongs to Rethink, so it shows wherever Rethink is live:
// the ready phase AND the empty-result phase (typed text survives the
// loading round-trip, only the row's visibility toggles).
const steerRow = document.getElementById('readMyMindSteerRow');
if (steerRow) steerRow.style.display = phase === 'loading' ? 'none' : '';
const rethinkBtn = document.getElementById('readMyMindRethink');
if (rethinkBtn) rethinkBtn.disabled = phase === 'loading';
},
+42 -7
View File
@@ -489,23 +489,56 @@ Object.assign(CodemanApp.prototype, {
const date = new Date(s.lastModified);
const timeStr = date.toLocaleDateString('en', { month: 'short', day: 'numeric' })
+ ' ' + date.toLocaleTimeString('en', { hour: '2-digit', minute: '2-digit', hour12: false });
const shortDir = s.workingDir.replace(/^\/home\/[^/]+\//, '~/');
// Shared helper, not a local regex: the copy that used to live here
// matched `/home/<user>/` only, so on macOS (`/Users/<user>/`) nothing was
// stripped and every row spent its first ~19 characters on an identical
// prefix — with the tail ellipsized, all rows rendered as
// `/Users/jordanryan/co…` and became indistinguishable (#273).
const shortDir = this._shortenHomePath(s.workingDir);
// Lead with the folder that identifies the row; the parent path trails and
// is what gets truncated. Truncation must never eat the identity.
const lastSlash = shortDir.lastIndexOf('/');
const leafName = lastSlash === -1 ? shortDir : shortDir.slice(lastSlash + 1);
// `<repo>/.claude/worktrees` in the parent path is pure noise once the pill
// says which worktree it is — drop it so the repo stays visible instead.
const parentDir = (lastSlash === -1 ? '' : shortDir.slice(0, lastSlash)).replace(/\/\.claude\/worktrees$/, '');
const btn = document.createElement('button');
btn.className = 'run-mode-option';
btn.className = 'run-mode-option run-mode-hist-row';
btn.title = s.workingDir;
btn.dataset.sessionId = s.sessionId;
btn.dataset.workingDir = s.workingDir;
const dirSpan = document.createElement('span');
dirSpan.className = 'hist-dir';
dirSpan.textContent = shortDir;
const nameSpan = document.createElement('span');
nameSpan.className = 'hist-name';
nameSpan.textContent = leafName;
const parts = [nameSpan];
// Worktree pill, same data the session rows use (#266). A worktree's
// directory basename is often just the worktree name, so without this two
// worktrees of one repo still read alike.
const wt = this._worktreeLabel ? this._worktreeLabel(s) : '';
if (wt) {
const wtSpan = document.createElement('span');
wtSpan.className = 'hist-wt';
wtSpan.textContent = wt;
parts.push(wtSpan);
}
if (parentDir) {
const dirSpan = document.createElement('span');
dirSpan.className = 'hist-dir';
dirSpan.textContent = parentDir;
parts.push(dirSpan);
}
const metaSpan = document.createElement('span');
metaSpan.className = 'hist-meta';
metaSpan.textContent = timeStr;
parts.push(metaSpan);
btn.append(dirSpan, metaSpan);
btn.append(...parts);
btn.addEventListener('click', (e) => {
e.stopPropagation();
this.resumeHistorySession(s.sessionId, s.workingDir, s.name);
@@ -2705,7 +2738,9 @@ Object.assign(CodemanApp.prototype, {
cases.forEach((c, idx) => {
const isFirst = idx === 0;
const isLast = idx === cases.length - 1;
const pathDisplay = c.path ? c.path.replace(/^\/Users\/[^/]+/, '~') : '';
// Was `/Users/<user>` only, the mirror image of the Run menu's bug: every
// case path on a Linux host rendered in full, unabbreviated.
const pathDisplay = c.path ? this._shortenHomePath(c.path) : '';
html += `
<div class="case-manage-item" data-case="${escapeHtml(c.name)}">
<div class="case-manage-info">
+44 -6
View File
@@ -482,17 +482,18 @@ Object.assign(CodemanApp.prototype, {
const voiceCfg = VoiceInput._getDeepgramConfig();
document.getElementById('voiceDeepgramKey').value = voiceCfg.apiKey || '';
document.getElementById('voiceLanguage').value = voiceCfg.language || 'en-US';
document.getElementById('voiceKeyterms').value = voiceCfg.keyterms || 'refactor, endpoint, middleware, callback, async, regex, TypeScript, npm, API, deploy, config, linter, env, webhook, schema, CLI, JSON, CSS, DOM, SSE, backend, frontend, localhost, dependencies, repository, merge, rebase, diff, commit, com';
document.getElementById('voiceKeyterms').value = voiceCfg.keyterms || DEFAULT_VOICE_KEYTERMS;
document.getElementById('voiceInsertMode').value = voiceCfg.insertMode || 'direct';
document.getElementById('voiceProvider').value = voiceCfg.provider || 'auto';
document.getElementById('appSettingsClaudeVoice').checked = settings.claudeVoiceEnabled ?? false;
// Reset key visibility to hidden
const keyInput = document.getElementById('voiceDeepgramKey');
keyInput.type = 'password';
document.getElementById('voiceKeyToggleBtn').textContent = 'Show';
// Update provider status
const providerName = VoiceInput.getActiveProviderName();
const providerEl = document.getElementById('voiceProviderStatus');
providerEl.textContent = providerName;
providerEl.className = 'voice-provider-status' + (providerName.startsWith('Deepgram') ? ' active' : '');
// Update provider status. The Claude row needs a fresh server probe: the
// setting is synced, so another device may have flipped it since page load.
this._renderVoiceProviderStatus();
VoiceInput.refreshClaudeStatus().then(() => this._renderVoiceProviderStatus());
// Updates section — show current version, reset transient result/progress UI.
this._initUpdatesSection();
@@ -1915,6 +1916,37 @@ Object.assign(CodemanApp.prototype, {
}
},
/**
* Paint both Voice status rows: which provider a mic press would use, and what
* the server reports about its Claude login. Called on open and again once the
* /api/voice/status probe resolves.
*/
_renderVoiceProviderStatus() {
const providerEl = document.getElementById('voiceProviderStatus');
if (providerEl) {
const providerName = VoiceInput.getActiveProviderName();
providerEl.textContent = providerName;
const live = providerName.startsWith('Deepgram Nova') || providerName.startsWith('Claude (this');
providerEl.className = 'voice-provider-status' + (live ? ' active' : '');
}
const claudeEl = document.getElementById('voiceClaudeStatus');
if (!claudeEl) return;
const status = VoiceInput._claudeStatus;
const text = !status
? 'Checking...'
: status.available
? `Ready${status.subscriptionType ? ` (${status.subscriptionType})` : ''}`
: status.reason === 'expired'
? 'Login expired - run a Claude session to refresh'
: status.reason === 'no-credentials'
? 'No Claude Code login on the server'
: status.reason === 'malformed'
? 'Claude credentials unreadable'
: 'Off - enable it above';
claudeEl.textContent = text;
claudeEl.className = 'voice-provider-status' + (status?.available ? ' active' : '');
},
async saveAppSettings() {
// Gesture overlay is injected at page render (server-side), so a change to it
// only takes effect on reload — remember the prior value to decide below.
@@ -1977,6 +2009,7 @@ Object.assign(CodemanApp.prototype, {
// Claude Permissions settings
agentTeamsEnabled: document.getElementById('appSettingsAgentTeams').checked,
agentSkillEnabled: document.getElementById('appSettingsAgentSkill').checked,
claudeVoiceEnabled: document.getElementById('appSettingsClaudeVoice').checked,
claudeModel: document.getElementById('appSettingsClaudeModel').value,
opusContext1mEnabled: document.getElementById('appSettingsOpusContext1m').checked,
remoteAutoReconnect: document.getElementById('appSettingsRemoteAutoReconnect').checked,
@@ -2016,6 +2049,7 @@ Object.assign(CodemanApp.prototype, {
// Save voice settings to localStorage + include in server payload for cross-device sync
const voiceSettings = {
provider: document.getElementById('voiceProvider').value,
apiKey: document.getElementById('voiceDeepgramKey').value.trim(),
language: document.getElementById('voiceLanguage').value,
keyterms: document.getElementById('voiceKeyterms').value.trim(),
@@ -2195,6 +2229,10 @@ Object.assign(CodemanApp.prototype, {
this.closeAppSettings();
// Voice availability is a server-side answer, so re-probe after a save:
// otherwise the mic keeps using the pre-save provider until the next reload.
VoiceInput.refreshClaudeStatus();
// The gesture overlay is injected at page render (server reads
// gestureControlEnabled from settings.json), so a change only takes effect on
// reload. Reload when it actually changed — the server PUT above already
+110
View File
@@ -4455,6 +4455,28 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
max-width: 250px;
box-shadow: 0 8px 32px rgba(0, 0, 0, 0.5), 0 2px 8px rgba(0, 0, 0, 0.3);
}
/* Above phone width the menu becomes a full-width drawer across the bottom of the
window. The 250px cap above was sized for a `~/<dir>/<repo>` row, which assumed
the home prefix had been abbreviated — it never was on macOS, so real rows blew
straight through it and ellipsized into `/Users/<user>/co…`, identical on every
line. Width is what buys room for the worktree + branch, so the cap is lifted
rather than nudged. Phones keep the compact popover: mobile.css positions this
menu itself and a viewport-wide drawer there would cover the composer. */
@media (min-width: 769px) {
.run-mode-menu {
/* Bounded, not viewport-wide. `100vw - 24px` spanned the whole window while
every row stayed shrink-to-fit at ~250px (see below), so the menu grew by
~1100px of dead space. 760px is what one full row costs: name + worktree
pill + parent path + timestamp. */
max-width: none;
width: min(760px, calc(100vw - 24px));
}
.run-mode-history {
/* More rows are worth showing once each one is legible. */
max-height: 320px;
}
}
.run-mode-menu.active {
display: flex;
flex-direction: column;
@@ -4524,6 +4546,37 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
simply pinned the menu at its max-width. */
.run-mode-history .run-mode-option {
max-width: 100%;
/* `.run-mode-history` is a block scroller, not a flex container, so these
<button> rows are shrink-to-fit and ignore the menu's width. Without this
the `flex: 1` + `text-align: right` on `.hist-dir` below have nothing to
expand into: the path never right-aligns and widening the menu only adds
empty space to its right. */
width: 100%;
}
/* Recent-session row: identity first, path last.
The row reads <folder> [⑂ worktree · branch] <parent path> <time>
with only the parent path allowed to shrink, so truncation can never hide
which project (or which worktree) a row refers to. */
.run-mode-hist-row {
align-items: baseline;
gap: 8px;
}
.run-mode-option .hist-name {
flex: 0 1 auto;
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
font-weight: 600;
}
.run-mode-option .hist-wt {
flex: 0 0 auto;
font-size: 0.78em;
padding: 0 5px;
border-radius: 3px;
background: color-mix(in srgb, var(--accent) 16%, transparent);
color: var(--accent);
white-space: nowrap;
}
.run-mode-option .hist-dir {
flex: 1;
@@ -4534,6 +4587,11 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
/* The parent path is context, not identity — dim it and let it be the part
that loses characters first. */
color: var(--text-muted);
font-size: 0.85em;
text-align: right;
}
.run-mode-option .hist-meta {
font-size: 0.85em;
@@ -10933,6 +10991,58 @@ kbd {
white-space: nowrap;
font-family: var(--mono-font, monospace);
}
/* Rethink steer note: the user's own words about what they actually meant,
sent as `steer` on the next re-run (readmymind-ui.js _readMyMindPredict).
Dashed border marks it as the optional side channel, distinct from the
primary suggestion field above; it solidifies on focus. */
.readmymind-steer-row {
margin-top: 10px;
}
.readmymind-steer-input {
width: 100%;
min-width: 0;
font-size: 12px;
padding: 7px 10px;
background: transparent;
color: var(--text);
border: 1px dashed var(--control-border);
border-radius: 8px;
}
.readmymind-steer-input:focus {
outline: none;
border-style: solid;
border-color: var(--accent);
background: var(--bg-dark);
}
.readmymind-steer-input::placeholder {
color: var(--text-dim);
}
/* Footer row: the buttons are btn-toolbar (display: flex, block-level), so
without this rule the four of them stack vertically. Mirrors the
runSummaryModal footer; the ≤430px block in mobile.css adds wrapping. */
.readmymind-modal .modal-footer {
display: flex;
justify-content: flex-end;
gap: 0.5rem;
padding: 0.75rem 1rem;
border-top: 1px solid rgba(255, 255, 255, 0.06);
}
/* Send keeps the primary accent: the skin block's bare .btn-toolbar rule is
(0,2,1) under html:not([data-skin="og"]) and outranks the base
.btn-toolbar.btn-primary (0,2,0) — the specificity trap CLAUDE.md documents
for mobile.css — so the accent is re-asserted here at (0,4,0). Scoped to
this modal on purpose; un-greying every btn-primary on the new skins is a
design call, not this feature's. */
.readmymind-modal .modal-footer .btn-toolbar.btn-primary {
background: var(--accent);
border-color: var(--accent);
color: #fff;
}
.readmymind-modal .modal-footer .btn-toolbar.btn-primary:hover {
background: var(--accent-hover);
border-color: var(--accent-hover);
color: #fff;
}
/* Keyboard-accessory 🧠 key: the phone surface for the same opt-in setting
(the header 🧠 button stays phone-hidden in mobile.css). The key ships in
+276 -38
View File
@@ -53,6 +53,7 @@
// Bound on page keys emitted from one gesture batch, mirroring the SGR tick
// cap: a fling must not build a backlog that keeps paging after it stops.
const PAGE_KEY_MAX_PER_BATCH = 3;
const TUI_PROMPT_DEFAULT_ROWS_FROM_BOTTOM = 4;
// Composer navigation keys as xterm.js encodes user keystrokes: plain and
// modified arrows (CSI A-D, CSI 1;mA-D, SS3 A-D), Home/End (CSI H/F, SS3
// H/F, CSI 1~/4~), Insert/Delete/PgUp/PgDn (CSI 2~/3~/5~/6~, optional
@@ -180,6 +181,7 @@
KEY_PAGE_DOWN,
PAGE_KEY_SCREEN_FRACTION,
PAGE_KEY_MAX_PER_BATCH,
TUI_PROMPT_DEFAULT_ROWS_FROM_BOTTOM,
};
global.CODEMAN_XTERM_THEMES = CODEMAN_XTERM_THEMES;
global.codemanCurrentXtermTheme = currentXtermTheme;
@@ -650,6 +652,8 @@ Object.assign(CodemanApp.prototype, {
let didScroll = false; // track whether touchmove fired (tap vs scroll)
let touchStartY = 0;
let tapStartedWithTerminalFocus = false;
let tapStartIntentCache = null;
const TAP_THRESHOLD = 8; // px — ignore micro-drift to distinguish tap from scroll
container.addEventListener(
'touchstart',
@@ -662,6 +666,28 @@ Object.assign(CodemanApp.prototype, {
pixelAccum = 0;
isTouching = true;
didScroll = false;
tapStartedWithTerminalFocus = this._isMobileTerminalInputFocused();
// Classifying scans the whole viewport with translateToString, and
// this runs at the start of EVERY gesture including scroll drags.
// Cache the result for the touchend of this same gesture rather than
// recomputing it; the cache is keyed on the exact start coordinates
// so a finger that moved re-classifies at its real position.
const touchStartIntent = this._classifyMobileTerminalTap(touchLastX, touchLastY);
tapStartIntentCache = { x: touchLastX, y: touchLastY, intent: touchStartIntent };
if (touchStartIntent === 'content') {
// Cancel xterm/browser focus before the compatibility click can
// open the OS keyboard. Content taps are re-emitted as SGR on
// touchend.
//
// 'history' is deliberately NOT included. A scrolled-up viewport
// sends nothing, so there is no compatibility click worth
// cancelling — and preventDefault() here, paired with touchend's
// early return, closes both routes to focus at once. Since
// selectSession() ends with scrollToLastNonEmptyLine(), that made
// the keyboard unreachable after every tab switch.
ev.preventDefault();
this._blurMobileTerminalInput();
}
lastTime = 0;
if (scrollFrame) {
cancelAnimationFrame(scrollFrame);
@@ -669,7 +695,7 @@ Object.assign(CodemanApp.prototype, {
}
}
},
{ passive: true }
{ passive: false }
);
container.addEventListener(
@@ -721,44 +747,19 @@ Object.assign(CodemanApp.prototype, {
scrollFrame = requestAnimationFrame(scrollLoop);
}
if (!didScroll && this.terminal) {
// ── Tap-to-position cursor ──────────────────────────────────
// Synthesize a click from the real touch point so the foreground app
// moves its cursor to the tapped cell (iOS doesn't reliably do this
// itself under touch-action:none). CRITICAL: only when mouse tracking
// is ON. xterm disables its local SelectionService while mouse events
// are active, so the synthetic click is forwarded to the PTY as an SGR
// report (cursor moves). But when tracking is OFF, that same click
// drives xterm's LOCAL selection (detail 1/2/3 → char/word/line) — a
// tap on CJK text would select & copy it instead of positioning. So
// gate strictly on the live mouse-tracking mode.
const touch = ev.changedTouches && ev.changedTouches[0];
const mouseMode = this.terminal.modes?.mouseTrackingMode;
const mouseTrackingOn = !!mouseMode && mouseMode !== 'none';
if (touch) {
this._suppressTrustedTapMouseEvents();
}
if (touch && mouseTrackingOn) {
this._dispatchSyntheticTerminalClick(touch.clientX, touch.clientY);
} else if (touch && this._sessionUsesServerMouseStrip()) {
// The server strips mouse-tracking DECSETs from claude/codex/gemini
// output (isAltScreenStripMode, session.ts) so the wheel keeps
// scrolling scrollback — which leaves THIS xterm permanently at
// mouseTrackingMode 'none' even though the TUI on the PTY side has
// tracking ON and still understands SGR reports. Encode the report
// ourselves and send it straight to the PTY: no DOM click is
// dispatched, so xterm's local selection can't trigger either.
this._sendSyntheticSgrTap(touch.clientX, touch.clientY);
}
this._syncMobileHelperTextareaToCursor();
// Route subsequent typing to the right place: keep the CJK input
// field focused when Chinese input is on, otherwise the terminal.
const cjkInput = document.getElementById('cjkInput');
if (cjkInput?.classList.contains('cjk-input-visible')) {
cjkInput.focus();
} else {
this.terminal.focus();
const cached =
tapStartIntentCache &&
tapStartIntentCache.x === touch.clientX &&
tapStartIntentCache.y === touch.clientY
? tapStartIntentCache.intent
: null;
this._handleMobileTerminalTap(touch, tapStartedWithTerminalFocus, cached);
}
}
tapStartedWithTerminalFocus = false;
},
{ passive: true }
);
@@ -769,6 +770,7 @@ Object.assign(CodemanApp.prototype, {
isTouching = false;
velocity = 0;
pixelAccum = 0;
tapStartedWithTerminalFocus = false;
},
{ passive: true }
);
@@ -1610,11 +1612,21 @@ Object.assign(CodemanApp.prototype, {
return workingDir.split('/').pop() || workingDir;
},
/** Normalize home prefixes to "~/" on both Linux and macOS */
/**
* Normalize a home prefix to "~" on both Linux (`/home/<user>`) and macOS
* (`/Users/<user>`). The lookahead lets the home directory ITSELF match, so a
* path that is exactly `$HOME` renders "~" instead of being left raw.
*
* This is the only place that pattern belongs. Two hand-rolled copies had
* drifted, each broken on the platform its author was not using: the Run
* menu's matched `/home/` only, so on macOS nothing was stripped and every
* Recent Sessions row spent its first ~19 characters on an identical
* `/Users/<user>/` prefix (#273); the case-manage list's matched `/Users/`
* only, so no Linux path was ever abbreviated there. Route new path labels
* through here rather than writing a third copy.
*/
_shortenHomePath(p) {
return (p || '')
.replace(/^\/home\/[^/]+\//, '~/')
.replace(/^\/Users\/[^/]+\//, '~/');
return (p || '').replace(/^\/(?:home|Users)\/[^/]+(?=\/|$)/, '~');
},
/**
@@ -3363,6 +3375,232 @@ Object.assign(CodemanApp.prototype, {
} catch {}
},
_isMobileTerminalInputFocused() {
const active = document.activeElement;
return (
active === this.terminal?.textarea ||
active?.classList?.contains('xterm-helper-textarea') ||
active?.id === 'cjkInput'
);
},
/**
* Separate terminal input from TUI-owned content on touch devices. A hidden
* keyboard must not consume taps on expandable readbacks, tool results, or
* decision rows; those taps belong to the foreground CLI. The visible prompt
* row remains the deliberate keyboard target.
*/
_classifyMobileTerminalTap(clientX, clientY) {
if (!this._terminalViewportAtBottom()) return 'history';
const pos = this._clientPointToCell(clientX, clientY);
if (!pos || !this.terminal) return 'input';
const mouseMode = this.terminal.modes?.mouseTrackingMode;
const mouseTrackingOn = !!mouseMode && mouseMode !== 'none';
if (!mouseTrackingOn && !this._sessionUsesServerMouseStrip()) return 'input';
const buffer = this.terminal.buffer?.active;
if (!buffer?.getLine) return 'input';
const rows = Math.max(1, this.terminal.rows || 1);
const lines = [];
const wrappedRows = [];
let hasVisibleContent = false;
for (let row = 0; row < rows; row++) {
const line = buffer.getLine(buffer.viewportY + row);
const text = line?.translateToString?.(true) || '';
lines.push(text);
wrappedRows.push(Boolean(line?.isWrapped));
if (text.trim()) hasVisibleContent = true;
}
if (!hasVisibleContent) return 'input';
const cursorRow = Math.max(0, Math.min(rows - 1, buffer.cursorY || 0));
const mode = this.sessions?.get(this.activeSessionId)?.mode || 'claude';
let promptRow = -1;
let menuSelectionVisible = false;
if (mode === 'opencode') {
if (lines[cursorRow]?.includes('\u2503')) promptRow = cursorRow;
} else {
for (let row = rows - 1; row >= 0; row--) {
const promptMatch = lines[row].match(/^\s*[❯›]/);
if (!promptMatch) continue;
const tail = lines[row].slice(promptMatch[0].length).trim();
// A highlighted numbered choice is a menu row, not an editable prompt.
const hasSiblingChoice = lines.some(
(line, choiceRow) => choiceRow !== row && /^\s+\d+[.)]\s/.test(line)
);
if (/^\d+[.)]\s/.test(tail) && hasSiblingChoice) {
menuSelectionVisible = true;
break;
}
promptRow = row;
break;
}
}
const tappedRow = pos.row - 1;
let logicalLineStart = tappedRow;
while (logicalLineStart > 0 && wrappedRows[logicalLineStart]) logicalLineStart--;
let logicalLineEnd = tappedRow;
while (logicalLineEnd + 1 < rows && wrappedRows[logicalLineEnd + 1]) logicalLineEnd++;
const tappedLine = lines.slice(logicalLineStart, logicalLineEnd + 1).join('');
// Claude's status row is TUI-owned: tapping it opens the teammate view, so it
// must not be treated as a keyboard target. Match the AFFORDANCE, not the
// wording — the bullet and verb are both unstable (claude 2.1.226 prints
// "✻ Cooked for 2m 6s", "✻ Baked for 9m 47s"; earlier builds printed
// "• Working …"), while "esc to interrupt" / "background" are what make the
// row actionable in the first place.
if (mode === 'claude' && /\b(?:esc to interrupt|background)\b/i.test(tappedLine)) {
return 'content';
}
if (menuSelectionVisible) return 'content';
if (promptRow >= 0) {
const inputEnd = cursorRow >= promptRow ? cursorRow : promptRow;
if (tappedRow >= promptRow && tappedRow <= inputEnd) return 'input';
} else if (
tappedRow === cursorRow ||
tappedRow >=
Math.max(
0,
rows -
window.CodemanTerminalInput
.TUI_PROMPT_DEFAULT_ROWS_FROM_BOTTOM
)
) {
// During redraws a CLI can temporarily omit its prompt marker or place
// the cursor above a status footer. Keep the live cursor and a stable
// lower-screen focus band usable without turning transcript rows above
// that band into keyboard targets.
return 'input';
}
return 'content';
},
_blurMobileTerminalInput() {
const active = document.activeElement;
if (
active === this.terminal?.textarea ||
active?.classList?.contains('xterm-helper-textarea') ||
active?.id === 'cjkInput'
) {
active.blur?.();
}
},
/**
* Which 'content' taps should DISMISS the mobile keyboard. Expandable
* readbacks, tool results and decision rows are TUI-owned: tapping them acts
* on the CLI, so popping the keyboard there is wrong. An inert transcript row
* still sends its mouse report, but must keep the keyboard reachable —
* touchstart's preventDefault cancels the compatibility click that would
* otherwise focus xterm, so focus has to be restored explicitly.
*/
_isActionableMobileTerminalTap(clientX, clientY) {
const pos = this._clientPointToCell(clientX, clientY);
const buffer = this.terminal?.buffer?.active;
if (!pos || !buffer?.getLine) return false;
const rows = Math.max(1, this.terminal.rows || 1);
const lines = [];
const wrappedRows = [];
for (let row = 0; row < rows; row++) {
const line = buffer.getLine(buffer.viewportY + row);
lines.push(line?.translateToString?.(true) || '');
wrappedRows.push(Boolean(line?.isWrapped));
}
const tappedRow = pos.row - 1;
let logicalLineStart = tappedRow;
while (logicalLineStart > 0 && wrappedRows[logicalLineStart]) logicalLineStart--;
let logicalLineEnd = tappedRow;
while (logicalLineEnd + 1 < rows && wrappedRows[logicalLineEnd + 1]) logicalLineEnd++;
const tappedLine = lines.slice(logicalLineStart, logicalLineEnd + 1).join('');
// Match the AFFORDANCE a CLI prints, not the row's title text: an
// expandable readback, tool result or status row advertises how to act on
// it ("ctrl+r to expand", "tap to collapse", "esc to interrupt"). Keying on
// titles instead would only recognise the exact strings a fixture happens
// to use, and would let a real readback keep the keyboard open.
//
// The hint sits on its own row, so a readback's TITLE row — the one a
// finger actually lands on — carries no affordance text itself. Look at the
// adjacent row too, which is how these blocks are laid out in practice.
// Keyed on the ACTION VERB, and deliberately not on prose verbs. A CLI hint
// names a key or a gesture ("ctrl+r to expand", "tap to collapse",
// "esc to interrupt"); "click here to open the file" is transcript content
// and must keep the keyboard, so `click` and bare `here` are excluded.
// The hint may sit mid-line — Claude's status row is
// "✻ Cooked for 2m 6s · esc to interrupt" — so this is not anchored.
const affordance =
/\b(?:ctrl\+\w+|shift\+\w+|esc|enter|tab|tap)\s+to\s+(?:expand|collapse|view|open|interrupt|see)\b/i;
const blockStart = Math.max(0, logicalLineStart - 1);
const blockEnd = Math.min(rows - 1, logicalLineEnd + 1);
for (let row = blockStart; row <= blockEnd; row++) {
if (affordance.test(lines[row])) return true;
}
// A Claude status row ("✻ Cooked for 2m 6s · esc to interrupt") is caught by
// the affordance above; there is deliberately no verb literal here, because
// the verb is randomised per build.
const hasMenuPrompt = lines.some((line) => /^\s*[❯›]\s+\d+[.)]\s/.test(line));
const hasMenuChoice = lines.some((line) => /^\s+\d+[.)]\s/.test(line));
return hasMenuPrompt && hasMenuChoice;
},
_focusMobileTerminalInput() {
this._syncMobileHelperTextareaToCursor();
const cjkInput = document.getElementById('cjkInput');
if (cjkInput?.classList.contains('cjk-input-visible')) {
cjkInput.focus();
} else {
this.terminal?.focus();
}
},
_handleMobileTerminalTap(touch, startedWithTerminalFocus, cachedIntent = null) {
// A guard bail-out, not a classification: there is nothing to classify. It is
// deliberately NOT 'history', which would claim the viewport was scrolled up.
if (!touch || !this.terminal) return null;
// touchstart already classified this exact point; reuse it rather than paying
// a second full-viewport scan for the same gesture.
const intent = cachedIntent ?? this._classifyMobileTerminalTap(touch.clientX, touch.clientY);
if (intent === 'history') {
// Scrolled up: send NO mouse report — a tap on old output must not be
// delivered to the CLI as a click on whatever row now occupies that cell.
// Focus is a separate question, and the answer is yes: the user tapped the
// terminal, so let them type. Blurring here stranded activeElement on
// <body> with no way back to the keyboard.
this._focusMobileTerminalInput();
return intent;
}
const mouseMode = this.terminal.modes?.mouseTrackingMode;
const mouseTrackingOn = !!mouseMode && mouseMode !== 'none';
const shouldActivate = intent === 'content' || startedWithTerminalFocus;
if (shouldActivate && mouseTrackingOn) {
// xterm's mouse encoder owns live DECSET modes. The synthetic DOM click
// follows the same path as a desktop click.
this._dispatchSyntheticTerminalClick(touch.clientX, touch.clientY);
} else if (shouldActivate && this._sessionUsesServerMouseStrip()) {
// Claude/Codex/Gemini DECSETs are stripped from the browser stream, so
// report directly to the PTY while retaining local touch scrollback.
this._sendSyntheticSgrTap(touch.clientX, touch.clientY);
}
if (intent === 'content' && this._isActionableMobileTerminalTap(touch.clientX, touch.clientY)) {
// A synthetic xterm click can focus its helper textarea. Blur after the
// report so collapsing a readback never opens or retains the keyboard.
this._blurMobileTerminalInput();
} else {
this._focusMobileTerminalInput();
}
return intent;
},
// ═══════════════════════════════════════════════════════════════
// Synthetic tap → mouse report
// ═══════════════════════════════════════════════════════════════
+421 -13
View File
@@ -1,7 +1,13 @@
/**
* @fileoverview Voice input with Deepgram Nova-3 (primary) and Web Speech API (fallback).
* @fileoverview Voice input with three providers: Claude (this server's Claude Code
* login), Deepgram Nova-3, and the Web Speech API.
*
* Defines two singleton objects:
* Defines three singleton objects:
*
* - ClaudeVoiceProvider — Dictation through Codeman's own `/ws/voice/stream`, which
* relays to the speech-to-text service Claude Code's `/voice` mode uses. No API key:
* the server holds the OAuth token, the browser only sends PCM16 @16 kHz (AudioWorklet,
* since MediaRecorder cannot emit raw PCM) and receives text. See docs/claude-voice-plan.md.
*
* - DeepgramProvider — Direct browser-to-Deepgram WebSocket connection for speech-to-text.
* Captures audio via MediaRecorder, streams chunks every 250ms, handles KeepAlive pings,
@@ -14,6 +20,7 @@
* Includes a temporary green Send button that replaces the settings gear icon after voice input.
* Web Speech API has auto-retry (up to 2x) for premature onend and iOS Safari stability check.
*
* @globals {object} ClaudeVoiceProvider
* @globals {object} DeepgramProvider
* @globals {object} VoiceInput
*
@@ -22,9 +29,13 @@
* @loadorder 3 of 15 — loaded after mobile-handlers.js, before notification-manager.js
*/
// Codeman — Voice input with Deepgram Nova-3 and Web Speech API fallback
// Codeman — Voice input with Claude, Deepgram Nova-3 and Web Speech API
// Loaded after mobile-handlers.js, before app.js
/** Dev vocabulary sent to the recognizer as a hint. Shared by every provider and the settings form. */
const DEFAULT_VOICE_KEYTERMS =
'refactor, endpoint, middleware, callback, async, regex, TypeScript, npm, API, deploy, config, linter, env, webhook, schema, CLI, JSON, CSS, DOM, SSE, backend, frontend, localhost, dependencies, repository, merge, rebase, diff, commit, com';
// ═══════════════════════════════════════════════════════════════
// Voice Input (Deepgram Nova-3 + Web Speech API fallback)
// ═══════════════════════════════════════════════════════════════
@@ -245,7 +256,282 @@ const DeepgramProvider = {
};
/**
* VoiceInput - Speech-to-text with Deepgram Nova-3 (primary) and Web Speech API (fallback).
* ClaudeVoiceProvider - Speech-to-text through this Codeman server's Claude Code
* login, i.e. the same service the CLI's own `/voice` mode uses. No API key.
*
* Audio goes browser -> Codeman -> Anthropic: the OAuth token never leaves the
* server, so the browser only ever sends PCM and receives text
* (docs/claude-voice-plan.md).
*
* ⚠️ The upstream endpoint is opened as linear16 / 16 kHz / mono, so capture MUST
* be raw PCM at that rate. MediaRecorder cannot emit raw PCM (container formats
* only), which is why this path uses an AudioWorklet rather than reusing
* DeepgramProvider's recorder. The AudioContext is constructed at 16000 Hz so the
* browser does the resampling.
*
* ⚠️ Transcript frames carry the WHOLE running transcript, not deltas. Callers
* must replace, never concatenate.
*/
const ClaudeVoiceProvider = {
_ws: null,
_stream: null,
_audioContext: null,
_workletNode: null,
_sourceNode: null,
_scriptNode: null,
_silenceTimeout: null,
_onResult: null,
_onError: null,
_onEnd: null,
_finalized: false,
/** How long without any transcript before the recording gives up on its own. */
SILENCE_MS: 6000,
/**
* Start streaming.
* @param {object} opts - { language, keyterms[], onResult(text, isFinal), onError(msg), onEnd(), onStream(stream) }
*/
async start(opts) {
this._onResult = opts.onResult;
this._onError = opts.onError;
this._onEnd = opts.onEnd;
this._finalized = false;
if (!navigator.mediaDevices?.getUserMedia) {
this._onError?.('Microphone requires a secure context (HTTPS). Use --https flag or access via localhost.');
this._cleanup();
return;
}
try {
this._stream = await navigator.mediaDevices.getUserMedia({
audio: { noiseSuppression: true, echoCancellation: true, autoGainControl: true }
});
} catch (err) {
const msg = err.name === 'NotAllowedError'
? 'Microphone access denied. Check browser settings.'
: 'Microphone error: ' + err.message;
this._onError?.(msg);
this._cleanup();
return;
}
opts.onStream?.(this._stream);
const params = new URLSearchParams();
if (opts.language) params.set('language', opts.language);
if (opts.keyterms?.length) params.set('keyterms', opts.keyterms.join(','));
const proto = location.protocol === 'https:' ? 'wss:' : 'ws:';
try {
this._ws = new WebSocket(`${proto}//${location.host}/ws/voice/stream?${params}`);
} catch (err) {
this._onError?.('Failed to open voice stream: ' + err.message);
this._cleanup();
return;
}
this._ws.binaryType = 'arraybuffer';
this._ws.onopen = () => {
// Capture starts only once the socket is up: PCM buffered before that would
// be the oldest audio, and dropping it keeps the transcript aligned with what
// the user hears themselves saying.
this._startCapture().catch((err) => {
this._onError?.('Microphone capture failed: ' + err.message);
this.stop();
});
this._resetSilenceTimeout();
};
this._ws.onmessage = (event) => {
let msg;
try {
msg = JSON.parse(event.data);
} catch (_e) {
return;
}
if (msg.t === 'transcript' && msg.text) {
this._resetSilenceTimeout();
this._onResult?.(msg.text, msg.final === true);
} else if (msg.t === 'error') {
this._onError?.(msg.message || 'Voice transcription failed');
}
};
this._ws.onerror = () => {
// onclose carries the actionable detail (close code); nothing useful here.
};
this._ws.onclose = (event) => {
if (event.code === 4004) {
this._onError?.(this._unavailableMessage(event.reason));
} else if (event.code === 4008) {
this._onError?.('Too many voice streams are already running on this server.');
} else if (event.code === 4003) {
this._onError?.('Voice stream refused (origin not allowed).');
} else if (event.code !== 1000 && !this._finalized) {
this._onError?.('Voice stream closed: ' + (event.reason || `code ${event.code}`));
}
this._stopCapture();
const onEnd = this._onEnd;
this._onEnd = null;
onEnd?.();
};
},
/** Map the server's close reason onto something a user can act on. */
_unavailableMessage(reason) {
if (reason === 'expired') return 'Claude login expired. Run a Claude session to refresh it, then try again.';
if (reason === 'disabled') return 'Claude voice is off. Enable it in Settings > Voice.';
return 'No Claude Code login found on the server. Sign in with `claude` there, or use Deepgram.';
},
/** Wire mic -> 16 kHz PCM16 frames -> WebSocket. */
async _startCapture() {
const Ctx = window.AudioContext || window.webkitAudioContext;
// Ask for 16 kHz directly so the browser resamples; Safari may hand back its
// own rate, which _pcmFromFloat32 then downsamples to match.
this._audioContext = new Ctx({ sampleRate: 16000 });
if (this._audioContext.state === 'suspended') await this._audioContext.resume();
this._sourceNode = this._audioContext.createMediaStreamSource(this._stream);
if (this._audioContext.audioWorklet) {
await this._audioContext.audioWorklet.addModule(this._workletUrl());
this._workletNode = new AudioWorkletNode(this._audioContext, 'pcm-frame-processor');
this._workletNode.port.onmessage = (event) => this._sendAudio(event.data);
this._sourceNode.connect(this._workletNode);
// A worklet with no destination is not pulled in some engines; a zero-gain
// sink keeps the graph running without echoing the mic to the speakers.
const sink = this._audioContext.createGain();
sink.gain.value = 0;
this._workletNode.connect(sink).connect(this._audioContext.destination);
return;
}
// Fallback for engines without AudioWorklet (older Safari): deprecated, but
// it is this or no dictation at all there.
this._scriptNode = this._audioContext.createScriptProcessor(4096, 1, 1);
this._scriptNode.onaudioprocess = (event) => {
this._sendAudio(this._pcmFromFloat32(event.inputBuffer.getChannelData(0), this._audioContext.sampleRate));
};
this._sourceNode.connect(this._scriptNode);
this._scriptNode.connect(this._audioContext.destination);
},
/**
* Worklet URL carrying this page's cache-bust token.
*
* ⚠️ Static assets are served `immutable` for a year, and `cacheBustAssets`
* only rewrites `.js` refs in `<script>`/`<link>` tags — a URL built here in JS
* is invisible to it. So the token is borrowed from voice-input.js's own script
* tag, which the server DID rewrite. Consequence: **edit the worklet and this
* file together**, or the browser keeps serving the old worklet.
*/
_workletUrl() {
const src = document.querySelector('script[src*="voice-input.js"]')?.getAttribute('src') || '';
const q = src.indexOf('?');
return 'voice-pcm-worklet.js' + (q === -1 ? '' : src.slice(q));
},
/** Float32 [-1,1] at any rate -> Int16 PCM at 16 kHz (nearest-neighbour decimation). */
_pcmFromFloat32(input, sampleRate) {
const ratio = sampleRate / 16000;
const outLength = Math.floor(input.length / ratio);
const out = new Int16Array(outLength);
for (let i = 0; i < outLength; i++) {
const sample = Math.max(-1, Math.min(1, input[Math.floor(i * ratio)]));
out[i] = sample < 0 ? sample * 0x8000 : sample * 0x7fff;
}
return out.buffer;
},
_sendAudio(arrayBuffer) {
if (this._finalized) return;
if (this._ws?.readyState !== WebSocket.OPEN) return;
try {
this._ws.send(arrayBuffer);
} catch (_e) {
/* socket died mid-frame */
}
},
_resetSilenceTimeout() {
clearTimeout(this._silenceTimeout);
this._silenceTimeout = setTimeout(() => this.stop(), this.SILENCE_MS);
},
/**
* Ask for the final transcript and let the server close the socket. Capture stops
* immediately, but the WebSocket stays open: the last (and usually best) transcript
* arrives AFTER the audio does, so closing here would throw away the utterance.
*/
stop() {
clearTimeout(this._silenceTimeout);
this._silenceTimeout = null;
if (this._finalized) return;
this._finalized = true;
this._stopCapture();
if (this._ws?.readyState === WebSocket.OPEN) {
try {
this._ws.send(JSON.stringify({ t: 'finalize' }));
} catch (_e) {
/* ignore */
}
} else {
const onEnd = this._onEnd;
this._onEnd = null;
onEnd?.();
}
},
/** Tear down the audio graph and release the mic. Idempotent. */
_stopCapture() {
if (this._workletNode) {
this._workletNode.port.onmessage = null;
try { this._workletNode.disconnect(); } catch (_e) { /* ignore */ }
this._workletNode = null;
}
if (this._scriptNode) {
this._scriptNode.onaudioprocess = null;
try { this._scriptNode.disconnect(); } catch (_e) { /* ignore */ }
this._scriptNode = null;
}
if (this._sourceNode) {
try { this._sourceNode.disconnect(); } catch (_e) { /* ignore */ }
this._sourceNode = null;
}
if (this._audioContext) {
try { this._audioContext.close(); } catch (_e) { /* ignore */ }
this._audioContext = null;
}
if (this._stream) {
this._stream.getTracks().forEach(t => t.stop());
this._stream = null;
}
},
/** Hard stop: drop the socket without waiting for a final transcript. */
_cleanup() {
this._finalized = true;
clearTimeout(this._silenceTimeout);
this._silenceTimeout = null;
this._stopCapture();
if (this._ws) {
this._ws.onclose = null;
this._ws.onmessage = null;
this._ws.onerror = null;
if (this._ws.readyState === WebSocket.OPEN) {
try { this._ws.close(1000); } catch (_e) { /* ignore */ }
}
this._ws = null;
}
this._onResult = null;
this._onError = null;
this._onEnd = null;
}
};
/**
* VoiceInput - Speech-to-text with Claude (this server's Claude Code login),
* Deepgram Nova-3, or the Web Speech API.
* Toggle mode: tap mic to start, tap again to stop. Auto-stops after silence.
* Shows interim transcription in a floating preview overlay.
* Inserts final text into the active session (user presses Enter to submit).
@@ -273,6 +559,29 @@ const VoiceInput = {
this._initRecognition();
// Always show buttons — if unsupported, toggle() shows a toast
this._showButtons();
// Probe the server's Claude voice availability in the background. `auto`
// resolution reads the cached answer, so the first mic press does not wait
// on a round trip; a miss just falls through to the next provider.
this.refreshClaudeStatus();
},
/** Last /api/voice/status answer, or null before the first probe resolves. */
_claudeStatus: null,
/**
* Re-probe whether this server can transcribe with its Claude Code login.
* Called at init and whenever App Settings opens (the setting is server-side,
* so another device could have flipped it).
*/
async refreshClaudeStatus() {
try {
const res = await fetch('/api/voice/status');
const json = await res.json();
this._claudeStatus = json?.success ? json.data : { available: false, reason: 'disabled' };
} catch (_e) {
this._claudeStatus = { available: false, reason: 'disabled' };
}
return this._claudeStatus;
},
// --- Deepgram config (localStorage only, never sent to server) ---
@@ -294,11 +603,37 @@ const VoiceInput = {
return !!(cfg.apiKey && cfg.apiKey.trim());
},
_claudeAvailable() {
return this._claudeStatus?.available === true;
},
/**
* Which provider a press of the mic would use.
*
* An explicit pick always wins, even when it cannot run — the resulting error
* ("Claude voice is off", "no Deepgram key") is more useful than silently
* transcribing somewhere the user did not choose. `auto` prefers Claude because
* it needs no key and no per-word billing, then the configured Deepgram key,
* then the browser's own engine.
*/
_resolveProvider() {
const pinned = this._getDeepgramConfig().provider;
if (pinned === 'claude' || pinned === 'deepgram' || pinned === 'webspeech') return pinned;
if (this._claudeAvailable()) return 'claude';
if (this._shouldUseDeepgram()) return 'deepgram';
return 'webspeech';
},
/** Get the active provider name for display */
getActiveProviderName() {
if (this._shouldUseDeepgram()) return 'Deepgram Nova-3';
if (this.supported) return 'Web Speech API';
return 'None';
switch (this._resolveProvider()) {
case 'claude':
return this._claudeAvailable() ? 'Claude (this server’s login)' : 'Claude (unavailable)';
case 'deepgram':
return this._shouldUseDeepgram() ? 'Deepgram Nova-3' : 'Deepgram (no API key)';
default:
return this.supported ? 'Web Speech API' : 'None';
}
},
/** Try to create a SpeechRecognition instance */
@@ -334,13 +669,81 @@ const VoiceInput = {
}
this._retryCount = 0;
if (this._shouldUseDeepgram()) {
const provider = this._resolveProvider();
if (provider === 'claude') {
this._startClaude();
} else if (provider === 'deepgram') {
this._startDeepgram();
} else {
this._startWebSpeech();
}
},
_startClaude() {
if (!this._claudeAvailable()) {
const reason = this._claudeStatus?.reason;
app.showToast(
reason === 'expired'
? 'Claude login expired on the server. Run a Claude session to refresh it.'
: reason === 'no-credentials'
? 'No Claude Code login found on the server. Sign in there with `claude`.'
: 'Claude voice is off. Enable it in Settings > Voice.',
'warning'
);
// Re-probe so a setting flipped on another device is picked up by the next press.
this.refreshClaudeStatus();
return;
}
const cfg = this._getDeepgramConfig();
this.isRecording = true;
this._activeProvider = 'claude';
this._accumulatedFinal = '';
this._lastTranscript = '';
this._hasReceivedResult = false;
this._recordingStartedAt = Date.now();
this._updateButtons('recording');
this._showPreview('Listening...', 'claude');
this._startDurationTimer();
const keyterms = (cfg.keyterms || DEFAULT_VOICE_KEYTERMS)
.split(',').map(t => t.trim()).filter(Boolean);
ClaudeVoiceProvider.start({
// The upstream endpoint wants a bare language tag; the Deepgram picker's
// 'en-US' style narrows to its base, and 'multi' means auto-detect.
language: (cfg.language || 'en-US').split('-')[0],
keyterms,
onStream: (stream) => this._startLevelMeter(stream),
onResult: (text, isFinal) => {
if (!this.isRecording) return;
this._hasReceivedResult = true;
// Each frame is the WHOLE running transcript, so replace rather than append.
this._accumulatedFinal = text;
if (isFinal) {
this._hidePreview();
this._insertText(text);
this.stop();
} else {
this._showPreview(text, 'claude');
}
},
onError: (msg) => {
const wasRecording = this.isRecording;
this.stop();
if (wasRecording) app.showToast(msg, 'error');
},
onEnd: () => {
if (this.isRecording) {
if (this._accumulatedFinal) this._insertText(this._accumulatedFinal);
this.stop();
}
}
});
if (navigator.vibrate) navigator.vibrate(50);
},
_startDeepgram() {
const cfg = this._getDeepgramConfig();
this.isRecording = true;
@@ -353,7 +756,7 @@ const VoiceInput = {
this._showPreview('Listening...', 'deepgram');
this._startDurationTimer();
const keyterms = (cfg.keyterms || 'refactor, endpoint, middleware, callback, async, regex, TypeScript, npm, API, deploy, config, linter, env, webhook, schema, CLI, JSON, CSS, DOM, SSE, backend, frontend, localhost, dependencies, repository, merge, rebase, diff, commit, com')
const keyterms = (cfg.keyterms || DEFAULT_VOICE_KEYTERMS)
.split(',').map(t => t.trim()).filter(Boolean);
DeepgramProvider.start({
@@ -452,7 +855,10 @@ const VoiceInput = {
this._updateButtons('idle');
this._hidePreview();
if (this._activeProvider === 'deepgram') {
if (this._activeProvider === 'claude') {
// Finalize, don't hang up: the last transcript arrives after the audio does.
ClaudeVoiceProvider.stop();
} else if (this._activeProvider === 'deepgram') {
DeepgramProvider.stop();
} else if (this._activeProvider === 'webspeech') {
try {
@@ -803,11 +1209,12 @@ const VoiceInput = {
timerEl.textContent = '0:00';
indicator.appendChild(timerEl);
this.previewEl.appendChild(indicator);
// Provider badge for Deepgram
if (provider === 'deepgram') {
// Provider badge (Web Speech gets none — it is the fallback, not a choice)
const badgeText = provider === 'deepgram' ? 'DG' : provider === 'claude' ? 'CLAUDE' : '';
if (badgeText) {
const badge = document.createElement('span');
badge.className = 'voice-preview-badge';
badge.textContent = 'DG';
badge.textContent = badgeText;
this.previewEl.appendChild(badge);
this.previewEl.appendChild(document.createTextNode(' '));
}
@@ -861,6 +1268,7 @@ const VoiceInput = {
if (this.isRecording) this.stop();
this._hideVoiceSendBtn();
DeepgramProvider._cleanup();
ClaudeVoiceProvider._cleanup();
this.recognition = null;
this._activeProvider = null;
this._stopDurationTimer();
+57
View File
@@ -0,0 +1,57 @@
/**
* @fileoverview AudioWorklet that turns microphone audio into the PCM frames the
* Claude voice endpoint expects.
*
* The endpoint is opened as `encoding=linear16, sample_rate=16000, channels=1`,
* i.e. raw signed 16-bit little-endian mono. MediaRecorder cannot produce that
* (it only emits container formats — webm/opus, mp4), which is why the Deepgram
* path's capture code cannot be reused here: Deepgram sniffs the container,
* Anthropic's endpoint does not.
*
* Sample rate is handled by the AudioContext, constructed at 16000 Hz so the
* browser resamples the mic for us. This processor only converts Float32 [-1,1]
* to Int16 and batches, because a raw 128-sample render quantum is a ~4 ms
* WebSocket frame — 250 frames a second of pure overhead.
*
* Loaded via `audioWorklet.addModule()` from voice-input.js. Runs on the audio
* thread: no DOM, no globals from the page.
*
* ⚠️ Edit this file and voice-input.js together. Static assets are served
* `immutable` for a year and this one is fetched from JS, so it inherits its
* cache-bust token from voice-input.js's script tag (see `_workletUrl()`); a
* change here alone would keep serving the old copy to every returning browser.
*/
/** ~256 ms at 16 kHz. Big enough to keep frame overhead down, small enough that interim transcripts stay live. */
const FRAME_SAMPLES = 4096;
class PcmFrameProcessor extends AudioWorkletProcessor {
constructor() {
super();
this._buffer = new Int16Array(FRAME_SAMPLES);
this._offset = 0;
}
process(inputs) {
const channel = inputs[0]?.[0];
// No input yet (mic still warming) — keep the processor alive.
if (!channel) return true;
for (let i = 0; i < channel.length; i++) {
// Clamp before scaling: values slightly outside [-1,1] are legal in Web Audio
// and would wrap around to the opposite sign as Int16, which sounds like a click.
const sample = Math.max(-1, Math.min(1, channel[i]));
this._buffer[this._offset++] = sample < 0 ? sample * 0x8000 : sample * 0x7fff;
if (this._offset === FRAME_SAMPLES) {
// Transfer a copy: the worklet keeps reusing its own buffer.
const frame = new Int16Array(this._buffer);
this.port.postMessage(frame.buffer, [frame.buffer]);
this._offset = 0;
}
}
return true;
}
}
registerProcessor('pcm-frame-processor', PcmFrameProcessor);
+1
View File
@@ -24,4 +24,5 @@ export { registerSearchRoutes } from './search-routes.js';
export { registerMeRoutes } from './me-routes.js';
export { registerAdminRoutes } from './admin-routes.js';
export { registerWsRoutes } from './ws-routes.js';
export { registerVoiceRoutes } from './voice-routes.js';
export { registerWebviewRoutes, tryWebviewRefererFallback } from './webview-routes.js';
+194
View File
@@ -0,0 +1,194 @@
/**
* @fileoverview Claude voice dictation routes.
*
* - `GET /api/voice/status` — can this server transcribe? (settings gate + credential state)
* - `GET /ws/voice/stream` — one dictation: PCM16 audio up, transcripts down
*
* Design and the upstream protocol: `docs/claude-voice-plan.md`. The relay itself
* lives in `../voice-stream.ts`; this file is the auth, gating and lifetime shell
* around it.
*
* ⚠️ `/api/voice/status` reports STATE, never the token: `{ available, reason,
* subscriptionType?, expiresAt? }`. The Claude OAuth access token stays inside the
* server process — the browser sends audio and receives text, nothing else.
*
* ⚠️ The WebSocket carries the same upgrade guard as `/ws/sessions/:id/terminal`
* (allowed Host + same-site Origin, on top of the global auth hook that already ran
* on the handshake). Without it a cross-site page could open a dictation stream on
* the user's credentials and bill their subscription.
*
* ⚠️ The feature is OFF unless `claudeVoiceEnabled` is set: turning it on spends the
* server owner's Claude subscription on transcription for anyone who can reach the
* UI, which is a decision for the operator rather than a default.
*/
import { createRequire } from 'module';
import { FastifyInstance } from 'fastify';
import type { WebSocket } from 'ws';
import { ApiErrorCode, createErrorResponse } from '../../types.js';
import { isAllowedRequestHost, isAllowedRequestOrigin, type HostPolicy } from '../network-auth-policy.js';
import { readClaudeOAuthCredentials } from '../../claude-credentials.js';
import { VoiceStreamRelay } from '../voice-stream.js';
import { MAX_AUDIO_FRAME_BYTES, MAX_CONCURRENT_STREAMS } from '../../config/voice.js';
import type { ConfigPort } from '../ports/index.js';
const require = createRequire(import.meta.url);
const { version: APP_VERSION } = require('../../../package.json') as { version: string };
/** Why voice is unavailable, in a form the frontend can branch on. */
export type VoiceUnavailableReason = 'disabled' | 'no-credentials' | 'expired' | 'malformed';
export interface VoiceStatus {
available: boolean;
reason?: VoiceUnavailableReason;
/** Display-only ('max', 'pro'); present when the credential store reported one. */
subscriptionType?: string;
expiresAt?: number;
}
/**
* Resolve the server's dictation readiness. Split out and exported so the status
* endpoint and the WebSocket upgrade cannot drift apart: the socket must never
* accept a stream the status endpoint calls unavailable.
*/
export async function resolveVoiceStatus(enabled: boolean): Promise<VoiceStatus> {
if (!enabled) return { available: false, reason: 'disabled' };
const creds = await readClaudeOAuthCredentials();
switch (creds.status) {
case 'ok':
return { available: true, subscriptionType: creds.subscriptionType, expiresAt: creds.expiresAt };
case 'expired':
return { available: false, reason: 'expired', expiresAt: creds.expiresAt };
case 'malformed':
return { available: false, reason: 'malformed' };
default:
return { available: false, reason: 'no-credentials' };
}
}
/** Live relays, server-wide. Dictation is human-paced, so the cap is small. */
let activeStreams = 0;
/** Test seam: the cap is process-wide state, so suites must be able to reset it. */
export function _resetVoiceStreamCountForTesting(): void {
activeStreams = 0;
}
/** Split a comma-separated keyterms query value into terms. */
function parseKeyterms(raw: unknown): string[] {
if (typeof raw !== 'string' || !raw) return [];
return raw
.split(',')
.map((t) => t.trim())
.filter(Boolean)
.slice(0, 100);
}
export function registerVoiceRoutes(app: FastifyInstance, ctx: ConfigPort, getHostPolicy: () => HostPolicy): void {
app.get('/api/voice/status', async (_req, reply) => {
try {
return { success: true, data: await resolveVoiceStatus(await ctx.getClaudeVoiceEnabled()) };
} catch {
reply.code(500);
return createErrorResponse(ApiErrorCode.INTERNAL_ERROR, 'Failed to read voice status');
}
});
app.get<{ Querystring: { language?: string; keyterms?: string } }>(
'/ws/voice/stream',
{ websocket: true },
async (socket: WebSocket, req) => {
// Cross-site upgrade guard first: this socket spends the operator's Claude
// subscription, so it must be reachable only from Codeman's own origin.
const policy = getHostPolicy();
if (!isAllowedRequestHost(req.headers.host, policy) || !isAllowedRequestOrigin(req.headers.origin, policy)) {
socket.close(4003, 'Forbidden');
return;
}
const status = await resolveVoiceStatus(await ctx.getClaudeVoiceEnabled());
if (!status.available) {
socket.close(4004, status.reason ?? 'unavailable');
return;
}
// Re-read rather than trusting resolveVoiceStatus's discarded token: the
// status helper deliberately never returns it.
const creds = await readClaudeOAuthCredentials();
if (creds.status !== 'ok' || !creds.accessToken) {
socket.close(4004, 'no-credentials');
return;
}
if (activeStreams >= MAX_CONCURRENT_STREAMS) {
socket.close(4008, 'Too many voice streams');
return;
}
activeStreams++;
let released = false;
const release = () => {
if (released) return;
released = true;
activeStreams--;
};
const send = (payload: Record<string, unknown>) => {
if (socket.readyState !== 1) return;
try {
socket.send(JSON.stringify(payload));
} catch {
/* client vanished mid-write */
}
};
const relay = new VoiceStreamRelay({
accessToken: creds.accessToken,
appVersion: APP_VERSION,
language: req.query.language,
keyterms: parseKeyterms(req.query.keyterms),
onReady: () => send({ t: 'ready' }),
onTranscript: (text, final) => send({ t: 'transcript', text, final }),
onError: (message) => send({ t: 'error', message }),
onClose: () => {
release();
send({ t: 'closed' });
if (socket.readyState === 1) {
try {
socket.close(1000, 'Voice stream ended');
} catch {
/* already closing */
}
}
},
});
// Handlers are attached synchronously before any further await
// (@fastify/websocket drops messages that arrive before they exist).
socket.on('message', (raw: Buffer, isBinary: boolean) => {
if (isBinary) {
if (raw.length === 0 || raw.length > MAX_AUDIO_FRAME_BYTES) return;
relay.sendAudio(raw);
return;
}
try {
const msg = JSON.parse(String(raw)) as { t?: string };
if (msg.t === 'finalize') relay.finalize();
else if (msg.t === 'stop') relay.close();
} catch {
/* non-JSON control frame — ignore */
}
});
socket.on('close', () => {
relay.close();
release();
});
socket.on('error', () => {
relay.close();
release();
});
relay.connect();
}
);
}
+12
View File
@@ -857,6 +857,16 @@ export const SettingsUpdateSchema = z
* add-only at create; a marker keeps user-authored copies untouched.
*/
agentSkillEnabled: z.boolean().optional(),
/**
* Let browser dictation transcribe through this machine's Claude Code login,
* the same speech-to-text service the CLI's own `/voice` mode uses
* (docs/claude-voice-plan.md). SYNCED, default OFF: enabling it spends the
* operator's Claude subscription on transcription for anyone who can reach
* the UI, and routes microphone audio to Anthropic rather than to whichever
* provider was configured before. The Deepgram and Web Speech paths are
* untouched by this flag.
*/
claudeVoiceEnabled: z.boolean().optional(),
/**
* Approvals Inbox (header bell + drawer, phone overview answer buttons,
* push Approve/Deny action buttons). SYNCED, default OFF (opt-in): even
@@ -970,6 +980,8 @@ export const SettingsUpdateSchema = z
// Voice settings (cross-device sync)
voiceSettings: z
.object({
/** 'auto' | 'claude' | 'deepgram' | 'webspeech'. Unknown values fall back to auto client-side. */
provider: z.string().max(20).optional(),
apiKey: z.string().max(200).optional(),
language: z.string().max(20).optional(),
keyterms: z.string().max(500).optional(),
+12
View File
@@ -166,6 +166,7 @@ import {
registerMeRoutes,
registerAdminRoutes,
registerWsRoutes,
registerVoiceRoutes,
registerWebviewRoutes,
tryWebviewRefererFallback,
} from './routes/index.js';
@@ -635,6 +636,7 @@ export class WebServer extends EventEmitter {
getClaudeModeConfig: this.getClaudeModeConfig.bind(this),
getTerminalHistoryConfig: this.getTerminalHistoryConfig.bind(this),
getAgentSkillEnabled: this.getAgentSkillEnabled.bind(this),
getClaudeVoiceEnabled: this.getClaudeVoiceEnabled.bind(this),
getDefaultClaudeMdPath: this.getDefaultClaudeMdPath.bind(this),
getLightState: this.getLightState.bind(this),
getLightSessionsState: this.getLightSessionsState.bind(this),
@@ -982,6 +984,7 @@ export class WebServer extends EventEmitter {
registerCronRoutes(this.app, { ...ctx, cron: this.cronService });
registerWsRoutes(this.app, ctx, () => this.getHostPolicy());
registerVoiceRoutes(this.app, ctx, () => this.getHostPolicy());
}
/**
@@ -1704,6 +1707,15 @@ export class WebServer extends EventEmitter {
return settings.agentSkillEnabled === true;
}
// Whether browser dictation may use this machine's Claude Code credentials
// (synced `claudeVoiceEnabled` setting, default OFF; docs/claude-voice-plan.md).
// OFF by default because turning it on spends the operator's Claude subscription
// on transcription for anyone who can reach the UI.
private async getClaudeVoiceEnabled(): Promise<boolean> {
const settings = await this.readSettings();
return settings.claudeVoiceEnabled === true;
}
/**
* Read My Mind predictor model (docs/readmymind-plan.md): `readMyMindModel`
* setting, defaulting to the AI-checker opus model. Prediction quality is
+300
View File
@@ -0,0 +1,300 @@
/**
* @fileoverview Upstream half of Claude voice dictation: one browser recording
* relayed to the speech-to-text service Claude Code's own `/voice` mode uses.
*
* The browser cannot talk to that service directly — it would need the Claude
* OAuth bearer token in page JavaScript, and the endpoint is not CORS-open — so
* Codeman sits in the middle and is the only thing that ever holds the token.
* See `docs/claude-voice-plan.md` for the protocol table this implements.
*
* Wire contract (mirrors the CLI's `connectVoiceStream`):
* - Query pins the audio format: linear16 PCM, 16 kHz, mono. The browser worklet
* produces exactly that; a mismatch transcribes as silence or noise, never an error.
* - `{"type":"KeepAlive"}` on open and every 8s, or upstream drops the socket
* between utterances.
* - Audio frames go up as raw binary.
* - Downstream, `TranscriptText`/`TranscriptInterim` carry the RUNNING transcript
* (each frame supersedes the previous one — they are not deltas to concatenate),
* and `TranscriptEndpoint` promotes the pending interim to final.
* - `{"type":"CloseStream"}` finalizes; the endpoint frame that follows is the
* last transcript, so `finalize()` waits briefly for it rather than closing.
*
* The pure builders at the top are unit-tested; `VoiceStreamRelay` owns the socket,
* the keepalive timer and the lifetime cap.
*/
import WebSocket from 'ws';
import {
AUDIO_CHANNELS,
AUDIO_SAMPLE_RATE,
FINALIZE_TIMEOUT_MS,
KEEPALIVE_INTERVAL_MS,
MAX_KEYTERMS_HEADER_CHARS,
MAX_STREAM_MS,
VOICE_STREAM_PATH,
voiceStreamBase,
} from '../config/voice.js';
const KEEPALIVE_FRAME = '{"type":"KeepAlive"}';
const CLOSE_STREAM_FRAME = '{"type":"CloseStream"}';
export interface VoiceStreamParams {
/** BCP-47-ish language hint. Anything unusable falls back to 'en'. */
language?: string;
/** Domain vocabulary sent as a recognition hint. */
keyterms?: string[];
}
/**
* Collapse keyterms into the single ASCII header value upstream accepts.
*
* Commas separate terms, so a comma INSIDE a term would silently split it; it is
* replaced with a space rather than dropped. Non-ASCII is stripped because the
* value travels as an HTTP header, where anything outside the visible ASCII range
* is not portable. Deduped and truncated on a term boundary so a long list degrades
* to a shorter list instead of a mangled final term.
*/
export function sanitizeKeyterms(terms: string[]): string {
const seen = new Set<string>();
const out: string[] = [];
let length = 0;
for (const term of terms) {
const cleaned = term
.replace(/,/g, ' ')
.replace(/[^\x20-\x7E]/g, '')
.replace(/\s+/g, ' ')
.trim();
if (!cleaned || seen.has(cleaned)) continue;
const cost = cleaned.length + (out.length > 0 ? 1 : 0);
if (length + cost > MAX_KEYTERMS_HEADER_CHARS) break;
seen.add(cleaned);
out.push(cleaned);
length += cost;
}
return out.join(',');
}
/** Normalize a language hint to what the endpoint expects, defaulting to English. */
export function normalizeVoiceLanguage(language: string | undefined): string {
const trimmed = (language ?? '').trim();
if (!trimmed || !/^[a-zA-Z]{2,3}(-[a-zA-Z0-9]{2,8})?$|^multi$/.test(trimmed)) return 'en';
return trimmed;
}
/** Full upstream URL with the audio format pinned. */
export function buildVoiceStreamUrl(params: VoiceStreamParams = {}, env: NodeJS.ProcessEnv = process.env): string {
const query = new URLSearchParams({
encoding: 'linear16',
sample_rate: String(AUDIO_SAMPLE_RATE),
channels: String(AUDIO_CHANNELS),
endpointing_ms: '300',
utterance_end_ms: '1000',
language: normalizeVoiceLanguage(params.language),
use_conversation_engine: 'true',
stt_provider: 'deepgram-nova3',
});
return `${voiceStreamBase(env)}${VOICE_STREAM_PATH}?${query.toString()}`;
}
/**
* Upstream headers. Codeman identifies itself honestly (it is not the CLI), which
* the endpoint accepts; the bearer token is the only thing that authenticates.
*/
export function buildVoiceStreamHeaders(
accessToken: string,
appVersion: string,
keyterms: string[] = []
): Record<string, string> {
const headers: Record<string, string> = {
Authorization: `Bearer ${accessToken}`,
'User-Agent': `codeman/${appVersion} (voice-bridge)`,
'x-app': 'codeman',
'anthropic-client-platform': 'codeman_web',
};
const sanitized = sanitizeKeyterms(keyterms);
if (sanitized) headers['x-config-keyterms'] = sanitized;
return headers;
}
export interface VoiceStreamRelayOptions extends VoiceStreamParams {
accessToken: string;
appVersion: string;
/** Called once the upstream socket is open and audio may flow. */
onReady: () => void;
/** Running transcript. `final` marks the utterance as complete. */
onTranscript: (text: string, final: boolean) => void;
/** Human-readable failure. The relay is dead (or dying) by the time this fires. */
onError: (message: string) => void;
/** Terminal: the relay released its socket and timers. Fires exactly once. */
onClose: () => void;
}
/**
* One dictation, upstream. Owns exactly one WebSocket and dies with it: every
* exit path (error, upstream close, lifetime cap, caller close) funnels through
* `_teardown()`, which fires `onClose` once and clears both timers.
*/
export class VoiceStreamRelay {
private ws: WebSocket | null = null;
private keepAlive: ReturnType<typeof setInterval> | null = null;
private lifetimeTimer: ReturnType<typeof setTimeout> | null = null;
private finalizeTimer: ReturnType<typeof setTimeout> | null = null;
private closed = false;
private finalizing = false;
/** Latest interim, held so a close/finalize can promote it to final. */
private pendingTranscript = '';
constructor(private readonly opts: VoiceStreamRelayOptions) {}
/** Open the upstream socket. Safe to call once; a second call is a no-op. */
connect(): void {
if (this.ws || this.closed) return;
const url = buildVoiceStreamUrl({ language: this.opts.language, keyterms: this.opts.keyterms });
const ws = new WebSocket(url, {
headers: buildVoiceStreamHeaders(this.opts.accessToken, this.opts.appVersion, this.opts.keyterms ?? []),
});
this.ws = ws;
ws.on('open', () => {
// Ping immediately: the gap between upgrade and the browser's first audio
// frame is long enough (mic permission, worklet boot) for upstream to drop us.
this.safeSend(KEEPALIVE_FRAME);
this.keepAlive = setInterval(() => this.safeSend(KEEPALIVE_FRAME), KEEPALIVE_INTERVAL_MS);
this.lifetimeTimer = setTimeout(() => {
this.opts.onError('Voice stream reached its maximum length');
this.close();
}, MAX_STREAM_MS);
this.opts.onReady();
});
ws.on('message', (raw) => this.handleMessage(String(raw)));
// An upgrade rejection never reaches 'open', so its status is the only signal
// that the token was refused rather than the network being down.
ws.on('unexpected-response', (_req, res) => {
const status = res.statusCode ?? 0;
res.resume();
this.opts.onError(
status === 401 || status === 403
? 'Claude rejected the voice credentials. Run a Claude session to refresh your login.'
: `Voice service refused the connection (HTTP ${status})`
);
this.teardown();
});
ws.on('error', (err: Error) => {
if (this.closed) return;
this.opts.onError(`Voice stream error: ${err.message}`);
});
ws.on('close', () => {
this.promotePending();
this.teardown();
});
}
/** Relay one raw PCM16 frame upstream. Dropped after finalize, as upstream ignores it. */
sendAudio(chunk: Buffer): void {
if (this.finalizing || this.closed) return;
if (this.ws?.readyState !== WebSocket.OPEN) return;
this.ws.send(chunk);
}
/**
* Ask upstream for the final transcript. The endpoint frame usually follows
* within a few hundred ms; the timer is the backstop so a silent upstream still
* yields whatever interim we already have instead of hanging the caller.
*/
finalize(): void {
if (this.finalizing || this.closed) return;
this.finalizing = true;
if (this.ws?.readyState !== WebSocket.OPEN) {
this.promotePending();
this.close();
return;
}
this.safeSend(CLOSE_STREAM_FRAME);
this.finalizeTimer = setTimeout(() => {
this.promotePending();
this.close();
}, FINALIZE_TIMEOUT_MS);
}
/** Terminal shutdown. Idempotent. */
close(): void {
if (this.closed) return;
const ws = this.ws;
this.teardown();
if (ws && (ws.readyState === WebSocket.OPEN || ws.readyState === WebSocket.CONNECTING)) {
try {
ws.close();
} catch {
/* already closing */
}
}
}
private handleMessage(raw: string): void {
let msg: { type?: string; data?: string; description?: string; error_code?: string; message?: string };
try {
msg = JSON.parse(raw);
} catch {
return;
}
switch (msg.type) {
case 'TranscriptText':
case 'TranscriptInterim': {
// Each frame is the whole running transcript, not a delta.
if (typeof msg.data === 'string' && msg.data) {
this.pendingTranscript = msg.data;
this.opts.onTranscript(msg.data, false);
}
break;
}
case 'TranscriptEndpoint': {
this.promotePending();
if (this.finalizing) this.close();
break;
}
case 'TranscriptError': {
this.opts.onError(msg.description || msg.error_code || 'Transcription failed');
break;
}
case 'error': {
this.opts.onError(msg.message || 'Voice service error');
break;
}
default:
break;
}
}
/** Emit the held interim as final, exactly once per utterance. */
private promotePending(): void {
if (!this.pendingTranscript) return;
const text = this.pendingTranscript;
this.pendingTranscript = '';
this.opts.onTranscript(text, true);
}
private safeSend(frame: string): void {
if (this.ws?.readyState !== WebSocket.OPEN) return;
try {
this.ws.send(frame);
} catch {
/* socket died between the check and the send */
}
}
private teardown(): void {
if (this.closed) return;
this.closed = true;
if (this.keepAlive) clearInterval(this.keepAlive);
if (this.lifetimeTimer) clearTimeout(this.lifetimeTimer);
if (this.finalizeTimer) clearTimeout(this.finalizeTimer);
this.keepAlive = null;
this.lifetimeTimer = null;
this.finalizeTimer = null;
this.opts.onClose();
}
}
+83
View File
@@ -0,0 +1,83 @@
/**
* Claude Code credential parsing.
*
* The voice relay authenticates with the token this parser returns, so every
* degraded store (absent, truncated, hand-edited, expired) must resolve to a
* status the caller can act on rather than a throw or a silently empty token.
*/
import { describe, it, expect } from 'vitest';
import { parseClaudeCredentials, claudeCredentialsPath } from '../src/claude-credentials.js';
const NOW = 1_800_000_000_000;
function store(overrides: Record<string, unknown> = {}): string {
return JSON.stringify({
claudeAiOauth: {
accessToken: 'sk-ant-oat01-test',
refreshToken: 'sk-ant-ort01-test',
expiresAt: NOW + 3_600_000,
subscriptionType: 'max',
...overrides,
},
});
}
describe('parseClaudeCredentials', () => {
it('returns the token and display metadata for a live store', () => {
const result = parseClaudeCredentials(store(), NOW);
expect(result.status).toBe('ok');
expect(result.accessToken).toBe('sk-ant-oat01-test');
expect(result.subscriptionType).toBe('max');
expect(result.expiresAt).toBe(NOW + 3_600_000);
});
it('reports an elapsed token as expired and withholds it', () => {
const result = parseClaudeCredentials(store({ expiresAt: NOW - 1000 }), NOW);
expect(result.status).toBe('expired');
expect(result.accessToken).toBeUndefined();
});
it('treats a token expiring within the skew as already expired', () => {
// A token with 30s left would die mid-dictation; refusing up front turns a
// confusing mid-utterance disconnect into a clear "refresh your login".
expect(parseClaudeCredentials(store({ expiresAt: NOW + 30_000 }), NOW).status).toBe('expired');
});
it('accepts a store with no expiry at all', () => {
const raw = JSON.stringify({ claudeAiOauth: { accessToken: 'sk-ant-oat01-test' } });
expect(parseClaudeCredentials(raw, NOW).status).toBe('ok');
});
it.each([
['not json at all', 'malformed'],
['{}', 'malformed'],
['null', 'malformed'],
['[]', 'malformed'],
['{"claudeAiOauth":null}', 'malformed'],
['{"claudeAiOauth":{}}', 'malformed'],
['{"claudeAiOauth":{"accessToken":""}}', 'malformed'],
['{"claudeAiOauth":{"accessToken":" "}}', 'malformed'],
['{"claudeAiOauth":{"accessToken":123}}', 'malformed'],
])('reports %s as malformed instead of throwing', (raw, expected) => {
expect(parseClaudeCredentials(raw, NOW).status).toBe(expected);
});
it('trims whitespace around a token written by hand', () => {
const raw = JSON.stringify({ claudeAiOauth: { accessToken: ' sk-ant-oat01-test\n' } });
expect(parseClaudeCredentials(raw, NOW).accessToken).toBe('sk-ant-oat01-test');
});
});
describe('claudeCredentialsPath', () => {
it('honors CLAUDE_CONFIG_DIR like the CLI does', () => {
expect(claudeCredentialsPath({ CLAUDE_CONFIG_DIR: '/tmp/alt-claude' })).toBe('/tmp/alt-claude/.credentials.json');
});
it('falls back to ~/.claude when the override is blank', () => {
expect(claudeCredentialsPath({ CLAUDE_CONFIG_DIR: ' ' })).toMatch(/\.claude\/\.credentials\.json$/);
});
it('falls back to ~/.claude when unset', () => {
expect(claudeCredentialsPath({})).toMatch(/\.claude\/\.credentials\.json$/);
});
});
+166
View File
@@ -0,0 +1,166 @@
/**
* @fileoverview Issue #273 and its mirror image: abbreviating `$HOME` in path labels.
*
* The rule ("show `~/project` rather than `/home/<user>/project`") had three
* implementations in the frontend, and two of them were platform-specific in
* opposite directions, so each looked correct to whoever wrote it:
*
* - the Run menu's Recent Sessions rows matched `/home/<user>/` only, so on
* macOS nothing was stripped, every row spent its first ~19 characters on an
* identical `/Users/<user>/` prefix, and the left-to-right ellipsis removed
* the tail that identifies the row (#273),
* - the case-manage list matched `/Users/<user>` only, so on a Linux host no
* case path was ever abbreviated at all.
*
* Both now call `_shortenHomePath()`, which is pinned here for both layouts, and
* a static guard fails if a fourth copy of the pattern appears.
*
* Loaded via `vm` against a stub CodemanApp with a fake DOM, same harness as
* history-list-controls.test.ts. Port: none (no browser, no server).
*/
import { readdirSync, readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it, vi } from 'vitest';
/* eslint-disable @typescript-eslint/no-explicit-any */
const PUBLIC = resolve(import.meta.dirname, '../src/web/public');
/**
* The container the vm's `document.getElementById` resolves for the case list.
* Swapped per test: the closure lives in THIS realm, so the shipping code inside
* the vm reads whatever the current test installed.
*/
let currentCaseList: { innerHTML: string } | null = null;
function loadTerminalUiPrototype(): Record<string, any> {
const source = readFileSync(resolve(PUBLIC, 'terminal-ui.js'), 'utf8');
const context = vm.createContext({
console,
CodemanApp: class CodemanApp {},
setInterval: vi.fn(),
clearInterval: vi.fn(),
setTimeout,
clearTimeout,
requestAnimationFrame: vi.fn(),
document: { addEventListener: vi.fn(), getElementById: () => null, createElement: () => ({}) },
window: { addEventListener: vi.fn(), removeEventListener: vi.fn() },
});
vm.runInContext(`${source}\nglobalThis.__proto = CodemanApp.prototype;`, context);
return (context as unknown as { __proto: Record<string, any> }).__proto;
}
function loadSessionUiPrototype(): Record<string, any> {
const source = readFileSync(resolve(PUBLIC, 'session-ui.js'), 'utf8');
const context = vm.createContext({
console,
CodemanApp: class CodemanApp {},
VoiceInput: {},
escapeHtml: (t: unknown) => String(t ?? ''),
setTimeout,
clearTimeout,
localStorage: { getItem: () => null, setItem: () => {} },
document: { getElementById: (id: string) => (id === 'caseManageList' ? currentCaseList : null) },
window: { addEventListener: vi.fn() },
});
vm.runInContext(`${source}\nglobalThis.__proto = CodemanApp.prototype;`, context);
return (context as unknown as { __proto: Record<string, any> }).__proto;
}
const terminalProto = loadTerminalUiPrototype();
const sessionProto = loadSessionUiPrototype();
const shorten = (p: unknown) => terminalProto._shortenHomePath.call(terminalProto, p);
describe('_shortenHomePath', () => {
it('abbreviates the Linux home prefix', () => {
expect(shorten('/home/arkon/default/claudeman')).toBe('~/default/claudeman');
});
it('abbreviates the macOS home prefix, which the Run menu never did (#273)', () => {
expect(shorten('/Users/jordanryan/code/facet/facet-agency-ops')).toBe('~/code/facet/facet-agency-ops');
});
it('abbreviates the home directory itself, not only paths below it', () => {
// The case-manage list's old regex had no trailing slash and did collapse
// this to "~"; keep that, or a case whose path IS $HOME would regress.
expect(shorten('/home/arkon')).toBe('~');
expect(shorten('/Users/jordanryan')).toBe('~');
});
it('leaves paths that only look like a home prefix alone', () => {
expect(shorten('/homer/bob/x')).toBe('/homer/bob/x');
expect(shorten('/Userspace/bob/x')).toBe('/Userspace/bob/x');
expect(shorten('/home')).toBe('/home');
expect(shorten('/mnt/d/work')).toBe('/mnt/d/work');
expect(shorten('/opt/codeman')).toBe('/opt/codeman');
});
it('replaces only the leading occurrence', () => {
expect(shorten('/home/arkon/home/bob/x')).toBe('~/home/bob/x');
});
it('tolerates empty and missing input', () => {
expect(shorten('')).toBe('');
expect(shorten(undefined)).toBe('');
expect(shorten(null)).toBe('');
});
});
describe('renderCaseManageList path labels', () => {
function render(cases: Array<{ name: string; path: string; location?: string }>): string {
currentCaseList = { innerHTML: '' };
const app: any = {
cases,
_shortenHomePath: terminalProto._shortenHomePath,
renderCaseManageList: sessionProto.renderCaseManageList,
};
app.renderCaseManageList();
const html = currentCaseList.innerHTML;
currentCaseList = null;
return html;
}
it('abbreviates a Linux case path (the mirror of #273)', () => {
const html = render([{ name: 'demo', path: '/home/arkon/codeman-cases/demo' }]);
expect(html).toContain('~/codeman-cases/demo');
expect(html).not.toContain('/home/arkon/codeman-cases/demo');
});
it('still abbreviates a macOS case path', () => {
const html = render([{ name: 'demo', path: '/Users/jordanryan/codeman-cases/demo' }]);
expect(html).toContain('~/codeman-cases/demo');
expect(html).not.toContain('/Users/jordanryan/codeman-cases/demo');
});
it('renders the row when a case has no path at all', () => {
const html = render([{ name: 'demo', path: '' }]);
expect(html).toContain('demo');
expect(html).toContain('class="case-manage-path"');
});
});
describe('single implementation of the home-prefix rule', () => {
/** Every top-level frontend module (vendor/ and subdirs are not ours). */
const sources = readdirSync(PUBLIC)
.filter((name) => name.endsWith('.js'))
.map((name) => ({ name, text: readFileSync(resolve(PUBLIC, name), 'utf8') }));
it('has exactly one home-prefix regex, in terminal-ui.js', () => {
// Any regex literal anchored at a home root. Three of these had drifted
// apart; a fourth would drift the same way.
const pattern = /\/\^\\\/(?:\(\?:home\|Users\)|home|Users)\\\//g;
const hits = sources.flatMap(({ name, text }) => (text.match(pattern) ?? []).map(() => name));
expect(hits).toEqual(['terminal-ui.js']);
});
it('routes both session-ui path labels through the helper', () => {
// Deliberately counts calls rather than pinning source lines: the Run menu
// row is being restructured in #274, and this guard should survive that as
// long as the label still goes through the helper.
const sessionUi = sources.find((s) => s.name === 'session-ui.js')!.text;
const calls = sessionUi.match(/this\._shortenHomePath\(/g) ?? [];
expect(calls.length).toBeGreaterThanOrEqual(2);
});
});
+277
View File
@@ -845,6 +845,283 @@ describe('Virtual Keyboard', () => {
expect(state.sentInputs).toEqual([]);
});
it('collapses a terminal readback without focusing the hidden textarea', async () => {
const point = await page.evaluate(async () => {
window.__sentInputs = [];
app.activeSessionId = 'mobile-readback-tap-test';
app.sessions.set('mobile-readback-tap-test', {
id: 'mobile-readback-tap-test',
mode: 'codex',
status: 'running',
});
app._sendInputAsync = (_sessionId: string, input: string) => {
window.__sentInputs.push(input);
};
app.hideWelcome();
const settings = app.loadAppSettingsFromStorage();
settings.cjkInputEnabled = false;
app.saveAppSettingsToStorage(settings);
app._updateCjkInputState();
app.terminal.reset();
await new Promise<void>((resolve) =>
app.terminal.write('Agent readback\r\n tap to collapse\r\n\r\n› ask', resolve)
);
app.terminal.focus();
const screen = app.terminal.element?.querySelector('.xterm-screen');
const cell = app.terminal._core?._renderService?.dimensions?.css?.cell;
const rect = screen?.getBoundingClientRect();
if (!rect || !cell?.width || !cell?.height) return null;
return {
x: rect.left + cell.width * 2,
y: rect.top + cell.height / 2,
};
});
expect(point).not.toBeNull();
await page.touchscreen.tap(point!.x, point!.y);
const state = await page.evaluate(() => ({
activeClass: document.activeElement?.className,
sentInputs: window.__sentInputs,
}));
expect(state.activeClass).not.toContain('xterm-helper-textarea');
expect(state.sentInputs).toHaveLength(1);
expect(state.sentInputs[0]).toMatch(/^\x1b\[<0;\d+;1M\x1b\[<0;\d+;1m$/);
});
it('keeps the hidden keyboard input focused after an inert Claude transcript tap', async () => {
const point = await page.evaluate(async () => {
window.__sentInputs = [];
app.activeSessionId = 'mobile-claude-transcript-tap-test';
app.sessions.set('mobile-claude-transcript-tap-test', {
id: 'mobile-claude-transcript-tap-test',
mode: 'claude',
cliVersion: '2.1.220',
status: 'working',
});
app._sendInputAsync = (_sessionId: string, input: string) => {
window.__sentInputs.push(input);
};
app.hideWelcome();
const settings = app.loadAppSettingsFromStorage();
settings.cjkInputEnabled = false;
app.saveAppSettingsToStorage(settings);
app._updateCjkInputState();
app.terminal.reset();
await new Promise<void>((resolve) =>
app.terminal.write(
'Transcript row one\r\nTranscript row two\r\nTranscript row three\r\nTranscript row four\r\nTranscript row five\r\nTranscript row six\r\nTranscript row seven\r\nTranscript row eight\r\nTranscript row nine\r\nTranscript row ten\r\n\r\n❯ ',
resolve
)
);
app.terminal.focus();
const screen = app.terminal.element?.querySelector('.xterm-screen');
const cell = app.terminal._core?._renderService?.dimensions?.css?.cell;
const rect = screen?.getBoundingClientRect();
if (!screen || !rect || !cell?.width || !cell?.height) return null;
const cursorRow = app.terminal.buffer.active.cursorY;
const transcriptRow = Math.max(1, Math.floor(cursorRow / 2));
const x = rect.left + cell.width * 2;
const y = rect.top + cell.height * (transcriptRow + 0.5);
return {
x,
y,
intent: app._classifyMobileTerminalTap(x, y),
activeClass: document.activeElement?.className,
};
});
expect(point).toEqual(
expect.objectContaining({
intent: 'content',
activeClass: expect.stringContaining('xterm-helper-textarea'),
})
);
await page.touchscreen.tap(point!.x, point!.y);
const activeClass = await page.evaluate(() => document.activeElement?.className);
expect(activeClass).toContain('xterm-helper-textarea');
});
it('prevents Claude subagent status taps from opening the hidden keyboard input', async () => {
const point = await page.evaluate(async () => {
window.__sentInputs = [];
app.activeSessionId = 'mobile-claude-subagent-tap-test';
app.sessions.set('mobile-claude-subagent-tap-test', {
id: 'mobile-claude-subagent-tap-test',
mode: 'claude',
cliVersion: '2.1.220',
status: 'working',
});
app._sendInputAsync = (_sessionId: string, input: string) => {
window.__sentInputs.push(input);
};
app.hideWelcome();
app.terminal.reset();
const statusRow = Math.max(0, app.terminal.rows - 2);
await new Promise<void>((resolve) =>
app.terminal.write(
`${'\r\n'.repeat(statusRow)}• Working (1m 50s • esc to interrupt) · 1 background teammate`,
resolve
)
);
app.terminal.focus();
const screen = app.terminal.element?.querySelector('.xterm-screen');
const cell = app.terminal._core?._renderService?.dimensions?.css?.cell;
const rect = screen?.getBoundingClientRect();
if (!screen || !rect || !cell?.width || !cell?.height) return null;
const cursorRow = app.terminal.buffer.active.cursorY;
const x = rect.left + cell.width * 2;
const y = rect.top + cell.height * (cursorRow + 0.5);
return {
x,
y,
intent: app._classifyMobileTerminalTap(x, y),
cursorRow,
screenBottom: rect.bottom,
};
});
expect(point).toEqual(
expect.objectContaining({
intent: 'content',
})
);
const dispatch = await page.evaluate(({ x, y }) => {
const target = document.querySelector('#terminalContainer .xterm-screen');
if (!(target instanceof Element)) {
return { prevented: false, insideTerminal: false, targetClass: null };
}
const touch = new Touch({
identifier: 3,
target,
clientX: x,
clientY: y,
pageX: x,
pageY: y,
});
const allowed = target.dispatchEvent(
new TouchEvent('touchstart', {
touches: [touch],
changedTouches: [touch],
bubbles: true,
cancelable: true,
})
);
target.dispatchEvent(
new TouchEvent('touchend', {
touches: [],
changedTouches: [touch],
bubbles: true,
cancelable: true,
})
);
return {
prevented: !allowed,
insideTerminal: Boolean(target.closest('#terminalContainer')),
targetClass: target.className,
};
}, point!);
const state = await page.evaluate(() => ({
activeClass: document.activeElement?.className,
sentInputs: window.__sentInputs,
}));
expect(dispatch).toEqual(
expect.objectContaining({
prevented: true,
insideTerminal: true,
})
);
expect(state.activeClass).not.toContain('xterm-helper-textarea');
expect(state.sentInputs).toHaveLength(1);
});
it('focuses the terminal helper textarea when the visible prompt is tapped', async () => {
const point = await page.evaluate(async () => {
window.__sentInputs = [];
app.activeSessionId = 'mobile-focus-visible-input-test';
app.sessions.set('mobile-focus-visible-input-test', {
id: 'mobile-focus-visible-input-test',
mode: 'codex',
status: 'running',
});
app._sendInputAsync = (_sessionId: string, input: string) => {
window.__sentInputs.push(input);
};
app.hideWelcome();
const settings = app.loadAppSettingsFromStorage();
settings.cjkInputEnabled = false;
app.saveAppSettingsToStorage(settings);
app._updateCjkInputState();
app.terminal.reset();
await new Promise<void>((resolve) =>
app.terminal.write('Agent readback\r\n tap to collapse\r\n\r\n› ask', resolve)
);
(document.activeElement as HTMLElement | null)?.blur?.();
const screen = app.terminal.element?.querySelector('.xterm-screen');
const cell = app.terminal._core?._renderService?.dimensions?.css?.cell;
const rect = screen?.getBoundingClientRect();
if (!rect || !cell?.width || !cell?.height) return null;
return {
x: rect.left + cell.width * 2,
y: rect.top + cell.height * (app.terminal.buffer.active.cursorY + 0.5),
};
});
expect(point).not.toBeNull();
await page.touchscreen.tap(point!.x, point!.y);
const state = await page.evaluate(() => ({
activeClass: document.activeElement?.className,
sentInputs: window.__sentInputs,
}));
expect(state.activeClass).toContain('xterm-helper-textarea');
expect(state.sentInputs).toEqual([]);
});
it('focuses the live Claude cursor when a redraw omits the prompt glyph', async () => {
const point = await page.evaluate(async () => {
window.__sentInputs = [];
app.activeSessionId = 'mobile-focus-promptless-claude-test';
app.sessions.set('mobile-focus-promptless-claude-test', {
id: 'mobile-focus-promptless-claude-test',
mode: 'claude',
status: 'running',
});
app._sendInputAsync = (_sessionId: string, input: string) => {
window.__sentInputs.push(input);
};
app.hideWelcome();
app.terminal.reset();
await new Promise<void>((resolve) => app.terminal.write('Claude response\r\nready for input', resolve));
(document.activeElement as HTMLElement | null)?.blur?.();
const screen = app.terminal.element?.querySelector('.xterm-screen');
const cell = app.terminal._core?._renderService?.dimensions?.css?.cell;
const rect = screen?.getBoundingClientRect();
if (!rect || !cell?.width || !cell?.height) return null;
return {
x: rect.left + cell.width * 2,
y: rect.top + cell.height * (app.terminal.buffer.active.cursorY + 0.5),
};
});
expect(point).not.toBeNull();
await page.touchscreen.tap(point!.x, point!.y);
const state = await page.evaluate(() => ({
activeClass: document.activeElement?.className,
sentInputs: window.__sentInputs,
}));
expect(state.activeClass).toContain('xterm-helper-textarea');
expect(state.sentInputs).toEqual([]);
});
it('keeps terminal touch drag available for scrollback with the visible textarea enabled', async () => {
const calls = await page.evaluate(async () => {
app.activeSessionId = 'mobile-touch-scroll-test';
+7 -1
View File
@@ -15,7 +15,11 @@ import { resolveTerminalHistoryConfig } from '../../src/config/terminal-history.
* Creates a mock context that satisfies all port interfaces.
* Pre-populated with one session for convenience.
*/
export function createMockRouteContext(options?: { sessionId?: string; agentSkillEnabled?: boolean }) {
export function createMockRouteContext(options?: {
sessionId?: string;
agentSkillEnabled?: boolean;
claudeVoiceEnabled?: boolean;
}) {
const sessionId = options?.sessionId ?? 'test-session-1';
const session = createMockSession(sessionId);
const sessions = new Map<string, MockSession>();
@@ -90,6 +94,8 @@ export function createMockRouteContext(options?: { sessionId?: string; agentSkil
// case's .claude/skills. Overridable per test because the create-time
// injection call sites are otherwise unreachable from a route test.
getAgentSkillEnabled: vi.fn(async () => options?.agentSkillEnabled ?? false),
// Default OFF mirrors the shipped setting: no test opens a voice relay by accident.
getClaudeVoiceEnabled: vi.fn(async () => options?.claudeVoiceEnabled ?? false),
getDefaultClaudeMdPath: vi.fn(async () => undefined),
getLightState: vi.fn(() => ({ sessions: [], status: 'ok' })),
getLightSessionsState: vi.fn(() => {
+32
View File
@@ -73,4 +73,36 @@ describe('read my mind phone key + alternates (static guards)', () => {
expect(ui).toContain('readMyMindAlternates');
expect(ui).toMatch(/\.textContent = suggestion\.prompt/);
});
// Phase 3 part 2: the Rethink steer note (docs/readmymind-plan.md phase 3).
it('wires the rethink steer note end to end: field, payload, phase visibility, reset', () => {
// The field lives in the modal, capped to the schema's 2000-char limit,
// and Enter in it triggers a rethink (mirroring the prompt field's
// Enter-to-send).
expect(html).toMatch(/id="readMyMindSteer"[^>]*maxlength="2000"/);
expect(html).toMatch(/id="readMyMindSteer"[^>]*onkeydown="[^"]*rethinkReadMyMind\(\)"/);
// Predict sends the trimmed note as `steer`, bounded to the schema cap.
expect(ui).toMatch(/body\.steer = steer\.slice\(0, 2000\)/);
// The row hides ONLY during loading: Rethink is live in both the ready
// and the empty-result phases, so the note must be reachable in both.
expect(ui).toMatch(/steerRow\.style\.display = phase === 'loading' \? 'none' : ''/);
// A fresh open resets the note along with the rethink memory.
expect(ui).toMatch(/steer\.value = ''/);
});
it('styles the footer with btn-toolbar (bare "btn btn-*" matches no CSS in this codebase)', () => {
const modal = html.slice(html.indexOf('id="readMyMindModal"'), html.indexOf('id="approvalsDrawer"'));
// The unstyled classes the footer originally shipped with must not return.
expect(modal).not.toMatch(/class="btn /);
expect(modal.match(/class="btn-toolbar/g)?.length).toBe(4);
expect(modal).toMatch(/class="btn-toolbar btn-primary"[^>]*sendReadMyMind\(true\)/);
// btn-toolbar is display:flex (block-level): without the desktop footer
// row rule the four buttons would stack vertically.
expect(styles).toMatch(/\.readmymind-modal \.modal-footer \{[^}]*display: flex/);
// The skin block's bare .btn-toolbar (0,2,1) greys out .btn-primary
// (0,2,0), so Send's accent must be re-asserted at higher specificity.
expect(styles).toMatch(/\.readmymind-modal \.modal-footer \.btn-toolbar\.btn-primary \{[^}]*var\(--accent\)/);
// The phone block sizes the same class for finger targets.
expect(phoneBlock).toMatch(/\.readmymind-modal \.modal-footer \.btn-toolbar/);
});
});
+302
View File
@@ -0,0 +1,302 @@
/**
* @fileoverview Claude voice dictation routes.
*
* Covers the status endpoint's gating and the full relay round trip against a
* mock upstream (a local `ws` server speaking the Anthropic voice-stream
* protocol, selected via CODEMAN_VOICE_STREAM_BASE). WebSocket routes need a
* real listening server — app.inject() cannot do upgrades.
*
* What these pin, beyond "it works":
* - the OAuth token never appears in an API response,
* - the socket refuses exactly what the status endpoint calls unavailable,
* - a cross-site upgrade cannot open a stream on the operator's subscription.
*
* Port: 3230 (routes), 3231 (mock upstream)
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import Fastify, { type FastifyInstance } from 'fastify';
import fastifyWebsocket from '@fastify/websocket';
import WebSocket, { WebSocketServer } from 'ws';
import { mkdirSync, writeFileSync, rmSync } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
import { registerVoiceRoutes, _resetVoiceStreamCountForTesting } from '../../src/web/routes/voice-routes.js';
import { MAX_CONCURRENT_STREAMS } from '../../src/config/voice.js';
const PORT = 3230;
const UPSTREAM_PORT = 3231;
const TOKEN = 'sk-ant-oat01-voice-route-test';
/** State captured by the mock upstream, so tests can assert what Codeman sent. */
interface UpstreamCapture {
headers: Record<string, string | string[] | undefined>;
url: string;
binaryFrames: Buffer[];
textFrames: string[];
socket: WebSocket | null;
}
function writeCredentials(expiresAt: number | undefined): void {
const dir = join(homedir(), '.claude');
mkdirSync(dir, { recursive: true });
writeFileSync(
join(dir, '.credentials.json'),
JSON.stringify({ claudeAiOauth: { accessToken: TOKEN, expiresAt, subscriptionType: 'max' } })
);
}
function removeCredentials(): void {
rmSync(join(homedir(), '.claude', '.credentials.json'), { force: true });
}
function waitForClose(ws: WebSocket, timeoutMs = 3000): Promise<{ code: number; reason: string }> {
return new Promise((resolve, reject) => {
const timer = setTimeout(() => reject(new Error('WS close timeout')), timeoutMs);
ws.on('close', (code, reason) => {
clearTimeout(timer);
resolve({ code, reason: reason.toString() });
});
});
}
/** Wait for the first message satisfying `match`, ignoring earlier frames. */
function waitForMessage(
ws: WebSocket,
match: (msg: Record<string, unknown>) => boolean,
timeoutMs = 3000
): Promise<Record<string, unknown>> {
return new Promise((resolve, reject) => {
const timer = setTimeout(() => reject(new Error('WS message timeout')), timeoutMs);
const onMessage = (raw: WebSocket.RawData) => {
let msg: Record<string, unknown>;
try {
msg = JSON.parse(String(raw));
} catch {
return;
}
if (!match(msg)) return;
clearTimeout(timer);
ws.off('message', onMessage);
resolve(msg);
};
ws.on('message', onMessage);
});
}
function waitUntil(predicate: () => boolean, timeoutMs = 3000): Promise<void> {
return new Promise((resolve, reject) => {
const deadline = Date.now() + timeoutMs;
const tick = () => {
if (predicate()) return resolve();
if (Date.now() > deadline) return reject(new Error('condition not met in time'));
setTimeout(tick, 10);
};
tick();
});
}
describe('voice-routes', () => {
let app: FastifyInstance;
let ctx: MockRouteContext;
let upstream: WebSocketServer;
let capture: UpstreamCapture;
let voiceEnabled: boolean;
beforeEach(async () => {
_resetVoiceStreamCountForTesting();
voiceEnabled = true;
capture = { headers: {}, url: '', binaryFrames: [], textFrames: [], socket: null };
upstream = new WebSocketServer({ port: UPSTREAM_PORT, host: '127.0.0.1' });
upstream.on('connection', (socket, req) => {
capture.headers = req.headers;
capture.url = req.url ?? '';
capture.socket = socket;
socket.on('message', (raw, isBinary) => {
if (isBinary) capture.binaryFrames.push(Buffer.from(raw as Buffer));
else capture.textFrames.push(String(raw));
});
});
await new Promise<void>((resolve) => upstream.once('listening', resolve));
process.env.CODEMAN_VOICE_STREAM_BASE = `ws://127.0.0.1:${UPSTREAM_PORT}`;
writeCredentials(Date.now() + 3_600_000);
app = Fastify({ logger: false });
await app.register(fastifyWebsocket);
ctx = createMockRouteContext();
ctx.getClaudeVoiceEnabled = (async () => voiceEnabled) as typeof ctx.getClaudeVoiceEnabled;
registerVoiceRoutes(app, ctx as never, () => ({ bindHost: '127.0.0.1', allowedHosts: [], tunnelHost: null }));
await app.listen({ port: PORT, host: '127.0.0.1' });
});
afterEach(async () => {
delete process.env.CODEMAN_VOICE_STREAM_BASE;
removeCredentials();
await app.close();
await new Promise<void>((resolve) => upstream.close(() => resolve()));
});
describe('GET /api/voice/status', () => {
it('reports available with display metadata when enabled and signed in', async () => {
const res = await app.inject({ method: 'GET', url: '/api/voice/status' });
expect(res.statusCode).toBe(200);
expect(res.json().data).toMatchObject({ available: true, subscriptionType: 'max' });
});
it('never returns the access token', async () => {
const res = await app.inject({ method: 'GET', url: '/api/voice/status' });
expect(res.body).not.toContain(TOKEN);
expect(res.body).not.toContain('sk-ant');
});
it('reports disabled when the setting is off, without touching credentials', async () => {
voiceEnabled = false;
const res = await app.inject({ method: 'GET', url: '/api/voice/status' });
expect(res.json().data).toEqual({ available: false, reason: 'disabled' });
});
it('reports no-credentials when nothing is signed in', async () => {
removeCredentials();
const res = await app.inject({ method: 'GET', url: '/api/voice/status' });
expect(res.json().data).toEqual({ available: false, reason: 'no-credentials' });
});
it('reports expired separately, so the UI can say how to fix it', async () => {
writeCredentials(Date.now() - 1000);
const res = await app.inject({ method: 'GET', url: '/api/voice/status' });
expect(res.json().data.available).toBe(false);
expect(res.json().data.reason).toBe('expired');
});
});
describe('GET /ws/voice/stream', () => {
it('relays audio up and transcripts down, finalizing on request', async () => {
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream?language=en&keyterms=tmux,respawn`);
const ready = waitForMessage(ws, (m) => m.t === 'ready');
await new Promise((resolve) => ws.once('open', resolve));
await ready;
ws.send(Buffer.alloc(3200));
await waitUntil(() => capture.binaryFrames.length > 0);
expect(capture.binaryFrames[0].length).toBe(3200);
const interim = waitForMessage(ws, (m) => m.t === 'transcript' && m.final === false);
capture.socket!.send(JSON.stringify({ type: 'TranscriptText', data: 'run the type check' }));
expect((await interim).text).toBe('run the type check');
ws.send(JSON.stringify({ t: 'finalize' }));
await waitUntil(() => capture.textFrames.some((f) => f.includes('CloseStream')));
const final = waitForMessage(ws, (m) => m.t === 'transcript' && m.final === true);
capture.socket!.send(JSON.stringify({ type: 'TranscriptEndpoint' }));
expect((await final).text).toBe('run the type check');
const { code } = await waitForClose(ws);
expect(code).toBe(1000);
});
it('authenticates upstream with the bearer token and forwards keyterms', async () => {
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream?keyterms=tmux,respawn`);
await waitForMessage(ws, (m) => m.t === 'ready');
expect(capture.headers.authorization).toBe(`Bearer ${TOKEN}`);
expect(capture.headers['x-config-keyterms']).toBe('tmux,respawn');
expect(capture.url).toContain('encoding=linear16');
expect(capture.url).toContain('sample_rate=16000');
ws.close();
});
it('pings upstream immediately so the idle gap before first audio cannot drop it', async () => {
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`);
await waitForMessage(ws, (m) => m.t === 'ready');
await waitUntil(() => capture.textFrames.some((f) => f.includes('KeepAlive')));
ws.close();
});
it('surfaces an upstream transcription error to the browser', async () => {
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`);
await waitForMessage(ws, (m) => m.t === 'ready');
const err = waitForMessage(ws, (m) => m.t === 'error');
capture.socket!.send(JSON.stringify({ type: 'TranscriptError', description: 'no audio' }));
expect((await err).message).toBe('no audio');
ws.close();
});
it('drops an oversized audio frame instead of relaying it', async () => {
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`);
await waitForMessage(ws, (m) => m.t === 'ready');
ws.send(Buffer.alloc(200_000));
ws.send(Buffer.alloc(1600));
await waitUntil(() => capture.binaryFrames.length > 0);
// The legal frame arrived; the oversized one was never forwarded.
expect(capture.binaryFrames.every((f) => f.length === 1600)).toBe(true);
ws.close();
});
it('closes 4004 with the reason when the setting is off', async () => {
voiceEnabled = false;
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`);
const { code, reason } = await waitForClose(ws);
expect(code).toBe(4004);
expect(reason).toBe('disabled');
});
it('closes 4004 when no Claude login exists on the server', async () => {
removeCredentials();
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`);
const { code, reason } = await waitForClose(ws);
expect(code).toBe(4004);
expect(reason).toBe('no-credentials');
});
it('closes 4004 rather than streaming on an expired login', async () => {
writeCredentials(Date.now() - 1000);
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`);
const { code, reason } = await waitForClose(ws);
expect(code).toBe(4004);
expect(reason).toBe('expired');
});
it('refuses a cross-site upgrade (a foreign page must not spend the subscription)', async () => {
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`, {
headers: { origin: 'https://evil.example' },
});
const { code } = await waitForClose(ws);
expect(code).toBe(4003);
expect(capture.socket).toBeNull();
});
it('caps concurrent streams', async () => {
const open: WebSocket[] = [];
for (let i = 0; i < MAX_CONCURRENT_STREAMS; i++) {
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`);
await waitForMessage(ws, (m) => m.t === 'ready');
open.push(ws);
}
const extra = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`);
const { code, reason } = await waitForClose(extra);
expect(code).toBe(4008);
expect(reason).toBe('Too many voice streams');
for (const ws of open) ws.close();
});
it('frees a stream slot when the browser hangs up', async () => {
const first = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`);
await waitForMessage(first, (m) => m.t === 'ready');
first.close();
await waitForClose(first);
// The slot is reusable: MAX_CONCURRENT more streams must still be admitted.
const reopened: WebSocket[] = [];
for (let i = 0; i < MAX_CONCURRENT_STREAMS; i++) {
const ws = new WebSocket(`ws://127.0.0.1:${PORT}/ws/voice/stream`);
await waitForMessage(ws, (m) => m.t === 'ready');
reopened.push(ws);
}
for (const ws of reopened) ws.close();
});
});
});
+5
View File
@@ -760,6 +760,11 @@ describe('case selector refresh', () => {
];
app.showToast = vi.fn();
// deleteCase re-renders the case-manage list, whose path label goes through
// _shortenHomePath. That method lives in terminal-ui.js, which this harness
// does not load (the real app always has it: load order 7 before 12).
app._shortenHomePath = (p: string) => p;
await app.deleteCase('deleted-case');
expect(quickStartCase.blur).toHaveBeenCalled();
+143
View File
@@ -6,8 +6,17 @@ import { describe, expect, it, vi } from 'vitest';
function loadTerminalUiHarness() {
const CodemanApp = function CodemanApp(this: any) {};
let now = 1_000;
let keyboardVisible = false;
let activeElement: unknown = null;
const context = vm.createContext({
window: {},
document: {
body: { classList: { contains: () => false } },
get activeElement() {
return activeElement;
},
getElementById: () => null,
},
CodemanApp,
console: { warn: vi.fn(), log: vi.fn() },
_crashDiag: { log: vi.fn() },
@@ -25,6 +34,11 @@ function loadTerminalUiHarness() {
MobileDetection: {
isTouchDevice: () => true,
},
KeyboardHandler: {
get keyboardVisible() {
return keyboardVisible;
},
},
DEC_SYNC_STRIP_RE: /\x1b\[\?2026[hl]/g,
TERMINAL_CHUNK_SIZE: 32 * 1024,
});
@@ -38,6 +52,12 @@ function loadTerminalUiHarness() {
setNow: (value: number) => {
now = value;
},
setKeyboardVisible: (visible: boolean) => {
keyboardVisible = visible;
},
setActiveElement: (element: unknown) => {
activeElement = element;
},
};
}
@@ -55,7 +75,130 @@ function createElementHarness() {
};
}
function createTerminalGrid(lines: string[], cursorY: number, wrappedRows = new Set<number>()) {
const textarea = {
classList: { contains: (name: string) => name === 'xterm-helper-textarea' },
blur: vi.fn(),
};
return {
cols: 80,
rows: lines.length,
modes: { mouseTrackingMode: 'none' },
buffer: {
active: {
viewportY: 0,
baseY: 0,
cursorY,
getLine: (row: number) =>
row >= 0 && row < lines.length
? { isWrapped: wrappedRows.has(row), translateToString: () => lines[row] }
: undefined,
},
},
element: {
querySelector: (selector: string) =>
selector === '.xterm-screen' ? { getBoundingClientRect: () => ({ left: 0, top: 0 }) } : null,
},
_core: { _renderService: { dimensions: { css: { cell: { width: 8, height: 16 } } } } },
textarea,
focus: vi.fn(),
};
}
describe('terminal touch tap mouse guard', () => {
it('recognizes focus only when a terminal input owns the active element', () => {
const { app, setActiveElement } = loadTerminalUiHarness();
const textarea = { classList: { contains: () => true } };
app.terminal = { textarea };
setActiveElement(null);
expect(app._isMobileTerminalInputFocused()).toBe(false);
setActiveElement(textarea);
expect(app._isMobileTerminalInputFocused()).toBe(true);
});
it('routes a readback row to the TUI while keeping the prompt row as keyboard input', () => {
const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'codex' }]]);
app.terminal = createTerminalGrid(
['Agent readback mentions › inline', ' tap to collapse', '', '', '› ask', 'gpt-5 · Context 80% left'],
4
);
expect(app._classifyMobileTerminalTap(9, 1)).toBe('content'); // inline marker is not a prompt
expect(app._classifyMobileTerminalTap(9, 17)).toBe('content'); // row 2: readback
expect(app._classifyMobileTerminalTap(9, 65)).toBe('input'); // row 5: prompt
expect(app._classifyMobileTerminalTap(9, 81)).toBe('content'); // row 6: status
});
it('classifies Claude background-agent status as content rather than keyboard input', () => {
const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.220' }]]);
app.terminal = createTerminalGrid(
['', '', '', '• Working (1m 50s • esc to ', 'interrupt) · 1 background teammate', ''],
4,
new Set([4])
);
expect(app._classifyMobileTerminalTap(9, 65)).toBe('content');
});
it('keeps the live cursor focusable when Claude temporarily omits its prompt glyph', () => {
const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'claude' }]]);
app.terminal = createTerminalGrid(['Prior response', '', 'ready for input', '', 'status footer', ''], 2);
expect(app._classifyMobileTerminalTap(9, 33)).toBe('input');
expect(app._classifyMobileTerminalTap(9, 1)).toBe('content');
});
it('treats a highlighted numbered choice as TUI content, not an input prompt', () => {
const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'claude' }]]);
app.terminal = createTerminalGrid(['Would you like to proceed?', '', '❯ 1. Yes', ' 2. No', '', ''], 2);
expect(app._classifyMobileTerminalTap(9, 33)).toBe('content');
expect(app._classifyMobileTerminalTap(9, 49)).toBe('content');
});
it('collapses TUI readback content without opening or retaining the keyboard', () => {
const { app, setActiveElement } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'codex' }]]);
app.terminal = createTerminalGrid(
['Agent readback', ' tap to collapse', '', '', '› ask', 'gpt-5 · Context 80% left'],
4
);
app._sendInputAsync = vi.fn();
setActiveElement(app.terminal.textarea);
expect(app._handleMobileTerminalTap({ clientX: 9, clientY: 17 }, true)).toBe('content');
expect(app._sendInputAsync).toHaveBeenCalledWith('sess-1', '\x1b[<0;2;2M\x1b[<0;2;2m');
expect(app.terminal.textarea.blur).toHaveBeenCalledOnce();
expect(app.terminal.focus).not.toHaveBeenCalled();
});
it('keeps the first prompt tap focus-only so it cannot activate a CLI row', () => {
const { app, setActiveElement } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'codex' }]]);
app.terminal = createTerminalGrid(
['Agent readback', ' tap to collapse', '', '', '› ask', 'gpt-5 · Context 80% left'],
4
);
app._sendInputAsync = vi.fn();
setActiveElement(null);
expect(app._handleMobileTerminalTap({ clientX: 9, clientY: 65 }, false)).toBe('input');
expect(app._sendInputAsync).not.toHaveBeenCalled();
expect(app.terminal.focus).toHaveBeenCalledOnce();
});
it('suppresses browser trusted compatibility mouse events during the tap window', () => {
const { app } = loadTerminalUiHarness();
const { element, dispatch } = createElementHarness();
+111
View File
@@ -0,0 +1,111 @@
/**
* Voice stream request building.
*
* The audio format lives in the query string, so a drift between these params
* and the browser worklet does not fail loudly — it transcribes as noise. These
* tests pin the contract, and pin that the bearer token never leaks into a URL
* (which would land it in proxy logs).
*/
import { describe, it, expect } from 'vitest';
import {
buildVoiceStreamHeaders,
buildVoiceStreamUrl,
normalizeVoiceLanguage,
sanitizeKeyterms,
} from '../src/web/voice-stream.js';
import { MAX_KEYTERMS_HEADER_CHARS } from '../src/config/voice.js';
describe('buildVoiceStreamUrl', () => {
it('pins linear16 / 16 kHz / mono, matching the browser worklet', () => {
const url = new URL(buildVoiceStreamUrl({}, {}));
expect(url.protocol).toBe('wss:');
expect(url.host).toBe('api.anthropic.com');
expect(url.pathname).toBe('/api/ws/speech_to_text/voice_stream');
expect(url.searchParams.get('encoding')).toBe('linear16');
expect(url.searchParams.get('sample_rate')).toBe('16000');
expect(url.searchParams.get('channels')).toBe('1');
expect(url.searchParams.get('stt_provider')).toBe('deepgram-nova3');
});
it('carries the language hint', () => {
expect(new URL(buildVoiceStreamUrl({ language: 'de' }, {})).searchParams.get('language')).toBe('de');
});
it('accepts a ws:// or wss:// base override for tests and gateways', () => {
const url = buildVoiceStreamUrl({}, { CODEMAN_VOICE_STREAM_BASE: 'ws://127.0.0.1:3199/' });
expect(url.startsWith('ws://127.0.0.1:3199/api/ws/speech_to_text/voice_stream?')).toBe(true);
});
it('ignores a non-websocket override rather than building a broken URL', () => {
expect(buildVoiceStreamUrl({}, { CODEMAN_VOICE_STREAM_BASE: 'https://evil.example' })).toContain(
'wss://api.anthropic.com'
);
});
it('never puts credentials in the URL', () => {
expect(buildVoiceStreamUrl({ language: 'en' }, {})).not.toMatch(/token|Bearer|sk-ant/i);
});
});
describe('normalizeVoiceLanguage', () => {
it.each([
['en', 'en'],
['en-US', 'en-US'],
['multi', 'multi'],
['', 'en'],
[undefined, 'en'],
['not a language', 'en'],
['../../etc/passwd', 'en'],
])('normalizes %s to %s', (input, expected) => {
expect(normalizeVoiceLanguage(input as string | undefined)).toBe(expected);
});
});
describe('sanitizeKeyterms', () => {
it('joins terms with commas', () => {
expect(sanitizeKeyterms(['tmux', 'respawn'])).toBe('tmux,respawn');
});
it('replaces an inner comma with a space so one term cannot become two', () => {
expect(sanitizeKeyterms(['hello, world'])).toBe('hello world');
});
it('drops non-ASCII, which is not portable in a header value', () => {
expect(sanitizeKeyterms(['café', 'naïve'])).toBe('caf,nave');
});
it('strips CR/LF so a term cannot inject a header', () => {
const result = sanitizeKeyterms(['ok\r\nX-Evil: 1']);
expect(result).not.toContain('\r');
expect(result).not.toContain('\n');
expect(result).toBe('okX-Evil: 1');
});
it('dedupes and skips empties', () => {
expect(sanitizeKeyterms(['a', 'a', '', ' ', 'b'])).toBe('a,b');
});
it('truncates on a term boundary rather than mangling the last term', () => {
const terms = Array.from({ length: 500 }, (_, i) => `term${i}`);
const result = sanitizeKeyterms(terms);
expect(result.length).toBeLessThanOrEqual(MAX_KEYTERMS_HEADER_CHARS);
for (const term of result.split(',')) expect(term).toMatch(/^term\d+$/);
});
});
describe('buildVoiceStreamHeaders', () => {
it('sends the bearer token and identifies Codeman honestly', () => {
const headers = buildVoiceStreamHeaders('sk-ant-oat01-test', '1.2.3');
expect(headers.Authorization).toBe('Bearer sk-ant-oat01-test');
expect(headers['User-Agent']).toBe('codeman/1.2.3 (voice-bridge)');
expect(headers['x-app']).toBe('codeman');
});
it('omits the keyterms header when nothing survives sanitizing', () => {
expect(buildVoiceStreamHeaders('t', '1.0.0', ['', ' '])).not.toHaveProperty('x-config-keyterms');
});
it('includes sanitized keyterms when present', () => {
expect(buildVoiceStreamHeaders('t', '1.0.0', ['tmux', 'respawn'])['x-config-keyterms']).toBe('tmux,respawn');
});
});