# Codeman Gesture Control — Prototype Build Plan > Canonical spec for this project. A Jarvis-style hand-tracking input layer for > the Codeman dashboard. Webcam sees your hands; you pinch-drag session tabs > between Screen columns and fire discrete gesture commands. Runs entirely > in-browser, camera feed never leaves the machine. Build it as a **standalone prototype first** (its own folder, fake tabs) so the input feel can be validated on real hardware before any integration with Codeman's existing drag/command code. --- ## Goal & success criteria Build a `GestureController` module + a self-contained demo page that: 1. Opens the webcam and runs MediaPipe Gesture Recognizer at ~30fps. 2. Emits a smoothed cursor position and a `pinch` state (grab/release) from hand landmarks. 3. Lets you **drag fake tabs between columns by pinching**, dropping on release. 4. Fires **discrete gesture commands** (open palm, thumbs up, victory) onto an event bus. 5. Feels responsive — drag lag is not perceptible, jitter is filtered out. **Done when:** you can sit at your desk, pinch a tab, move it to another column, release, and it lands — reliably, without visible jitter, with the camera mounted at your chosen angle. --- ## Tech stack - **MediaPipe Tasks Vision** (`@mediapipe/tasks-vision`) — `GestureRecognizer` in `VIDEO` running mode, `numHands: 1` for v1 (add 2 later). Loads the prebuilt `gesture_recognizer.task` model + WASM from CDN. - **Vanilla TS + Vite** for the prototype (no framework needed; keep it portable so the module drops into Codeman regardless of its stack). If Codeman is React, the module stays framework-agnostic and you wrap it in a hook at integration time. - **One-Euro filter** for cursor smoothing — implement it directly, it's ~40 lines and is the correct tool for noisy interactive landmark streams (low lag at speed, heavy smoothing when still). - **Web Worker** for inference is a **Phase 4** optimization — do NOT start there. Get it working on the main thread first; only move to a worker if the dashboard UI stutters. --- ## File structure ``` gesture-proto/ ├── index.html # demo page: video preview + columns of fake tabs ├── package.json ├── vite.config.ts ├── src/ │ ├── main.ts # wires GestureController -> demo UI │ ├── gesture/ │ │ ├── GestureController.ts # core: camera + recognizer + state machine + events │ │ ├── OneEuroFilter.ts # cursor smoothing │ │ ├── pinch.ts # pinch detection w/ hysteresis │ │ ├── landmarks.ts # landmark index constants + helpers │ │ └── types.ts # event payload types, config │ └── demo/ │ ├── tabs.ts # fake tab/column model + render │ └── overlay.ts # draws hand skeleton + cursor dot over video (debug) └── README.md ``` --- ## Core algorithms (the parts that decide whether it feels good) ### 1. Cursor from landmarks The drag cursor is the **midpoint of thumb tip (landmark 4) and index tip (landmark 8)**, in normalized [0,1] coords from MediaPipe. - **Mirror X** (`x = 1 - x`) — the webcam image is flipped relative to the user. - Map normalized → screen pixels against the dashboard's bounding rect. - Run the resulting (x, y) through **two independent One-Euro filters** (one per axis) before using it. Raw landmarks jitter by several pixels even when the hand is still; this is the single most important quality step. ### 2. Pinch detection with hysteresis Compute euclidean distance between landmark 4 and landmark 8. **Normalize by hand size** (e.g. distance wrist→middle-finger-MCP, landmarks 0→9) so the threshold is robust to how close the hand is to the camera. - Use **two thresholds, not one** (hysteresis): enter pinch below `PINCH_ON` (e.g. 0.35 of hand size), exit only above `PINCH_OFF` (e.g. 0.5). This stops flickering between grab/release at the boundary — critical for not "dropping" a tab mid-drag. - Require the pinch state to persist N frames (e.g. 2–3) before firing, to reject single-frame noise. ### 3. State machine ``` IDLE ──hand detected──> HOVER ──pinch on──> GRABBED ──pinch off──> (drop) ──> HOVER ^ | | └────hand lost───────────┴────────────────hand lost────────────────────────┘ ``` - `HOVER`: cursor moves, highlights the tab/column under it (hit-test). - `GRABBED`: the grabbed tab follows the cursor; emit `drag` events. - On `pinch off` in GRABBED: hit-test cursor against drop columns, emit `drop {tabId, targetColumnId}` or `dropCancelled` if outside any column. ### 4. Discrete gestures → command bus From `result.gestures[0].categoryName`, debounced (fire once per gesture entry, not every frame while held): - `Open_Palm` held ~1s → `command: "halt-all"` (dead-man's-switch — pauses every session; genuinely useful for autonomous loops). - `Thumb_Up` → `command: "approve"`. - `Victory` → `command: "new-session"`. - Map these to the SAME command names your voice layer already dispatches, so both input sources converge on one dispatcher. --- ## GestureController public API (target shape) ```ts const gc = new GestureController({ video: videoEl, surface: dashboardEl, // coords mapped against this element's rect numHands: 1, pinchOn: 0.35, pinchOff: 0.5, palmHoldMs: 1000, }); gc.on("hover", ({ x, y, targetId }) => {...}); gc.on("grab", ({ x, y }) => {...}); gc.on("drag", ({ x, y }) => {...}); // throttled to frame rate gc.on("drop", ({ targetColumnId }) => {...}); gc.on("command",({ name }) => {...}); // halt-all | approve | new-session gc.on("status", ({ fps, handPresent, pinchDist }) => {...}); // debug HUD await gc.start(); // requests camera, loads model gc.stop(); ``` Keep it **transport-agnostic**: it emits semantic events, it does NOT know about Codeman's DOM. Integration is just subscribing to these events and calling Codeman's existing tab-move / command functions. --- ## Phased build (each phase is independently testable — stop and feel it before moving on) **Phase 0 — Scaffold & camera (½ day)** Vite + TS project. `index.html` with a mirrored `