fix(remote): a proxied host is reachability-unknown; scope remote: SSE per session

Review round 2 on #439.

1. The bare TCP probe connects to host:port, which a host behind a jump host
   or SOCKS proxy does not answer even while ssh works. Acting on that
   verdict drew a permanent banner over a healthy session, replaced a real
   "needs tmux" error with "not reachable" in quick-start, and - with a wake
   target - buffered every HTTP input for the life of the session, since the
   readiness poll could never succeed. `WakeableRemote` now carries
   `jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such
   a host into reachability-UNKNOWN: input is delivered, `checkReachable` /
   `checkHostReachable` answer `null` (never `false`), `ensureHostAwake`
   returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate
   fires on `=== false` only, and `GET …/reachability` reports
   `reachable: null, probeable: false` so the banner has nothing to key on.
   A wake target can still be fired for it, blind: no readiness poll, no
   reattach, no toast - the response says only whether the packet went out.

2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake
   has no session yet, so the registry names the requesting user
   (`ensureHostAwake({ requestedBy })` -> `username` in the payload) and
   `deriveSseHint` routes on it; with neither it fails closed to admins.
   Single-user mode is unaffected.

Smaller, from the same review:

- A flush write that fails now drops the remaining buffer (logged) instead
  of retaining it: the wake still resolved and marked the host reachable, so
  the retained chunk waited for the NEXT wake and was replayed hours later,
  after everything typed since. Same policy as the oversized paste.
- The banner polls on tab activation (a user action) and on its 30 s timer
  only for a host with a wake target; a timer connecting to a host Codeman
  cannot wake is the traffic invariant #2 rejects keepalives for. A proxied
  host is never polled.
- `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP
  socket refuse under VITEST, as remote-files.ts does. The guard caught a
  leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the
  probe but still polled readiness with the real one, so the shutdown test
  had been connecting to a production address. The poll now uses the
  injected probe.
- docs/remote-sessions.md is additions only again (the reformatting is
  gone); the architecture-invariants overlap resolved itself in the merge.

Live, against a throwaway instance with a non-routable ghost host: proxied
-> no probe, no wake, the genuine ssh error after 10 s; direct (control) ->
probe, magic packet, "did not come back" after the 40 s budget.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
This commit is contained in:
Randalix
2026-09-18 22:46:11 +02:00
co-authored by Claude Opus 5
parent e271a65e79
commit 1040f6c489
12 changed files with 609 additions and 83 deletions
+1 -1
View File
@@ -217,7 +217,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Remote sessions + remote SSH cases**: a case can point at a remote host. The agent runs inside a durable remote `tmux -L codeman-remote` (session name `codeman-ssh-<id>`, deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so an instance on the target host never adopts it), fronted by a LOCAL tmux pane running `ssh`. Attached (`owned:false`) sessions **detach, never kill** on tab close; owned ones propagate `kill-session`. A bounded-backoff watcher auto-reconnects dropped sessions (`remoteAutoReconnect`, default ON). ⚠️ **It revives ONLY when the durable remote tmux session is verifiably still alive** (`remoteTmuxSessionAlive()`, a `has-session` probe over ssh, #355): a clean agent exit (Ctrl-C, Ctrl-D, `exit`) tears that session down, and `isPaneDead()` cannot tell it from a transport drop, so the watcher used to relaunch a FRESH agent after every clean exit (claude only looked fine because its `|| --resume` fallback masked it). An unreachable host answers `undefined`, which also means do not revive. ⚠️ `has-session` prints NOTHING on success, so the probe is classified by EXIT STATUS (`classifyRemoteAliveExit`: 0 alive, ssh's 255 or a timeout unknown, anything else gone); reading stdout classified every live session as gone and silently disabled transport-drop reconnects. The answer is cached per session and forgotten whenever the pane is seen alive again, or a stale `true` from one transport drop would revive the next clean exit. ⚠️ **File reads in a remote case are the second ssh surface** (#415, `src/remote-files.ts`): they go through `buildSshConnectionArgs()` as well, a browser-supplied path is only ever a `shellescape`d token, an unreachable host answers 502 (never 404), the size cap uses the REMOTE size, and no remote file is ever copied onto the server's disk — which is why writes, office previews and thumbnails are deliberately unsupported over ssh (the `PUT` guard sits BEFORE the local path validation, or a same-named local directory such as an sshfs mount takes the write). The probe's symlink resolution FAILS CLOSED (a path it cannot canonicalize is a 404, never its own unresolved string: the directory-only fallback let a `notes.txt -> ~/.ssh/id_rsa` link pass containment), and ssh children are BOUNDED by `src/remote-ssh-limiter.ts` plus one batched probe per attachment-history listing, because terminal output in a remote session is written on the remote host and a prompt-injected agent can print hundreds of `codeman://attach` links. The ATTACHMENT routes (a clicked path outside the case dir) go through the same layer, and which host a record is read from follows the SESSION, never the path string. ⚠️ **Command-injection surface: every ssh command line must flow through `buildSshConnectionArgs()`**, which `shellescape`s every user field. Never hand-build an ssh line elsewhere. ⚠️ Run flows must route remote cases through `POST /api/quick-start`, not `POST /api/sessions` (which stat-validates `workingDir` locally and has no `caseName`). → [architecture-invariants#remote-sessions-over-ssh](docs/architecture-invariants.md#remote-sessions-over-ssh), [#remote-ssh-cases](docs/architecture-invariants.md#remote-ssh-cases), `docs/remote-sessions.md`
**Wake-on-LAN (`remote-wake.ts`)**: an optional `RemoteHost.wakeMac` (Codeman builds the magic packet itself) or `RemoteHost.wakeCommand` (single executable path, run without a shell, takes precedence) lets the INPUT route, `POST /api/sessions/:id/wake`, and the user's own create/attach request (`POST /api/quick-start`, `POST /api/sessions` with `attachRemoteSession`, via `ensureHostAwake`) wake a sleeping host instead of writing into a stalled ssh pane. ⚠️ An explicit request — input, the wake button, or the user pressing Run/Attach — and NOTHING else may wake: the auto-reconnect watcher, `handleRemoteSessionDropped`, boot recovery and `cron-service.ts` have no access to the registry (a wake there would re-wake the host seconds after every suspend, and the create wake is wired in the route rather than the shared session service for exactly that reason), which `test/remote-wake.test.ts` asserts as two wiring guards — the second also pins that `server.ts` holds the registry for its LIFETIME only (`drop` on cleanup, `stop` on shutdown) and never calls a waking method. `GET /api/sessions/:id/reachability` merely probes and never wakes. Detection is a throttled bare TCP probe — deliberately no `ServerAliveInterval`, because keepalives move bytes into an idle connection every interval and that is what a byte-threshold idle detector must not read as activity. Input arriving during a wake is buffered (a chunk over 4 KB is dropped whole, never delivered as a fragment) and flushed in order after `reattachRemote()`; send-and-wait blocks instead. ⚠️ Browser keystrokes travel over the WebSocket, which deliberately does NOT pass through the registry (that is the hot path), so only the HTTP input path ever queues anything — the banner must not promise queued input for the Wake button. A request that waits on the wake (create/attach, and the button) uses the 40 s `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS`, not the 90 s session default, because the dashboard's reverse proxy cuts a request at its own 60 s `proxy_read_timeout`. The wake fields are re-read from `remote-hosts.json` on recovery and, throttled+cached via `RemoteWakeDeps.resolveRemote`, for a LIVE session, since the persisted `remote` snapshot never sees a field added later. UI: the amber `#hostWakeBanner` (`host-wake-ui.js`) with Wake / "Configure WoL" → `#wakeConfigModal`.
**Wake-on-LAN (`remote-wake.ts`)**: an optional `RemoteHost.wakeMac` (Codeman builds the magic packet itself) or `RemoteHost.wakeCommand` (single executable path, run without a shell, takes precedence) lets the INPUT route, `POST /api/sessions/:id/wake`, and the user's own create/attach request (`POST /api/quick-start`, `POST /api/sessions` with `attachRemoteSession`, via `ensureHostAwake`) wake a sleeping host instead of writing into a stalled ssh pane. ⚠️ An explicit request — input, the wake button, or the user pressing Run/Attach — and NOTHING else may wake: the auto-reconnect watcher, `handleRemoteSessionDropped`, boot recovery and `cron-service.ts` have no access to the registry (a wake there would re-wake the host seconds after every suspend, and the create wake is wired in the route rather than the shared session service for exactly that reason), which `test/remote-wake.test.ts` asserts as two wiring guards — the second also pins that `server.ts` holds the registry for its LIFETIME only (`drop` on cleanup, `stop` on shutdown) and never calls a waking method. `GET /api/sessions/:id/reachability` merely probes and never wakes. Detection is a throttled bare TCP probe — deliberately no `ServerAliveInterval`, because keepalives move bytes into an idle connection every interval and that is what a byte-threshold idle detector must not read as activity. ⚠️ A host behind `jumpHost`/`socksProxy`/a `ProxyCommand` option is reachability-UNKNOWN (`isProbeable()`): the probe connects to `host:port`, which such a host does not answer even while ssh works, so the registry never buffers for it, never gates create/attach on it (`'unprobeable'`), and `/reachability` answers `reachable: null, probeable: false` — the banner keys on a PROVEN `false`, and the banner's 30 s poller runs only for a host with a wake target (a timer connecting to a host Codeman cannot wake is the same timer-driven traffic the keepalive rule forbids). Input arriving during a wake is buffered (a chunk over 4 KB is dropped whole, never delivered as a fragment) and flushed in order after `reattachRemote()` — a flush write that fails drops the rest (logged) rather than retaining it for a wake hours later; send-and-wait blocks instead. ⚠️ Browser keystrokes travel over the WebSocket, which deliberately does NOT pass through the registry (that is the hot path), so only the HTTP input path ever queues anything — the banner must not promise queued input for the Wake button. A request that waits on the wake (create/attach, and the button) uses the 40 s `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS`, not the 90 s session default, because the dashboard's reverse proxy cuts a request at its own 60 s `proxy_read_timeout`. The wake fields are re-read from `remote-hosts.json` on recovery and, throttled+cached via `RemoteWakeDeps.resolveRemote`, for a LIVE session, since the persisted `remote` snapshot never sees a field added later. UI: the amber `#hostWakeBanner` (`host-wake-ui.js`) with Wake / "Configure WoL" → `#wakeConfigModal`. The `remote:` SSE family is session-scoped in multi-user mode; a create/attach wake names its requester (`username`) since it has no session yet. `remote-wake.ts` refuses real IO under `VITEST` like `remote-files.ts`.
**Docker cases**: a case can point at a **container**, with any of the CLI run modes running inside it. Like remote-SSH this is a **LOCATION OVERLAY on cases, never a `SessionMode` of its own**. Exactly one long-lived container **per case**, shared by all its sessions, so killing a session kills only that session's in-container tmux and **never** `docker stop` while siblings remain. The workspace is a real host dir bind-mounted at the **same absolute path**, which is what keeps file-routes/watchers on real host bytes and makes the in-container transcript projHash match the host. Credentials are **seeded** (RO mount, copied into the container once) rather than shared RW, so in-container CLIs never write refreshed tokens back to the host, and bind mounts are excluded from `docker commit` so exports stay secret-free. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** Config drift is detected via a label hash and a drifted launch is REFUSED rather than silently launched with stale config. ⚠️ A case may instead **ADOPT** a container the user already runs (`DockerCase.owned === false`, mirror of remote-SSH's `owned:false`): Codeman only `exec`s into it and never creates, starts, stops, restarts or removes it, so a missing or stopped container FAILS CLOSED with an actionable message instead of being fixed. Absent = owned, so existing cases are byte-identical. ⚠️ An ADOPTED container may back SEVERAL cases at different in-container directories (`classifyAdoptContainerConflict` in `docker-hosts.ts`: an exact twin on the same container AND directory is refused, an owned container still backs exactly one case, and a container another user adopted is refused), which is what the Add Case panel's "copy an existing case" picker relies on; the wire carries `CaseInfo.docker.owned` ONLY when false, so the picker tests `=== false`, never truthiness. The guarantee is enforced at four independent layers because it cannot be observed by using the feature: `buildDockerStopCommand`/`buildDockerRemoveCommand` throw during pure STRING CONSTRUCTION, `removeDockerContainer` refuses again, drift reports "none" (an adopted container carries no `codeman.confighash` label, so a real comparison would 409 the launch forever), and the boot reaper skips it. ⚠️ Two lifecycle touches the original design missed and that are easy to re-introduce: the full-image export `docker commit`s the container (refused for an adopted case) and the workspace export `docker pause`s it first (skipped — it freezes the owner's processes for the length of the tar). ⚠️ `owned` is applied AFTER `dockerConfigHash`, which takes an explicit field list, or every pre-existing case would trip the drift gate at once. ⚠️ Run modes for a container case come from the CONTAINER (`availableModes`, live-probed): gating the run menu on HOST CLIs (#201) is right for local sessions and wrong here, since a host with no `claude` may run a container that ships one. ⚠️ **A failed probe means opposite things per ownership** — for an ADOPTED case it is a fault worth reporting, for an OWNED one it is the NORMAL state before the first session (the launch chain creates the container), so treating it as a fault hid every agent mode on every freshly linked Docker case behind "start it yourself first". That is why `CaseInfo.docker.owned` is on the wire. ⚠️ Claude is launched WITHOUT `--dangerously-skip-permissions` when the container's exec user is root (Claude Code refuses the flag as root and the refusal is visible only inside the container); which flag to drop is a per-CLI fact, so it is the registry's `overlays.docker.rootCommand`, never a branch. ⚠️ Adoption is **admin-only in multi-user mode**, unlike `docker-link`: linking creates OUR container, whose one bind mount `isWorkingDirAllowed` has already confined, while an adopted container's mounts belong to its owner and one mounting `/` hands the adopter the host. The same reasoning admin-gates the container listing and the in-container directory browser; the preflight instead admits a non-admin for a container already linked to a case they own, because the run menu probes it for every docker case. ⚠️ On the loopback-only prod bind a container cannot reach 127.0.0.1, so in-container hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; otherwise idle detection falls back to output-based. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` (user guide), `docs/docker-cases-plan.md` (design)
+1 -1
View File
@@ -54,7 +54,7 @@ Model is NOT a session field: it is a composition entry in the profile's config
### Remote SSH cases
**Remote host wake-on-LAN from user input**: an optional `RemoteHost.wakeMac` (magic packet built and broadcast by Codeman) or `RemoteHost.wakeCommand` (a single executable path, run WITHOUT a shell, and the explicit override) lets the input route — and an explicit `POST /api/sessions/:id/wake` — wake a SLEEPING host instead of writing into a stalled ssh pane; `tmux send-keys` succeeds against a stalled pane, so the bytes used to vanish silently. The wake flow lives in `src/remote-wake.ts` and is reachable **only** from an EXPLICIT user request: `POST /api/sessions/:id/input`, that explicit wake route, and the create/attach path (`POST /api/quick-start` for a remote case, `POST /api/sessions` with `attachRemoteSession`, via `ensureHostAwake`), because "the user pressed Run on a sleeping host" is the same kind of request and the tmux probe would otherwise fail with a misleading "needs tmux installed". Everything TIMER-driven must never wake a host: the COD-108 auto-reconnect watcher, `Server.handleRemoteSessionDropped` and boot recovery have no access to the registry, or a host would be re-woken seconds after each suspend and could never stay asleep (asserted by wiring guards in `test/remote-wake.test.ts`, not just documented — including that `ensureHostAwake` is called from the HTTP route only, since `cron-service.ts` builds sessions through the shared service with nobody waiting on the answer). `GET /api/sessions/:id/reachability` only ASKS — it never wakes — and feeds the amber "host unreachable" banner (`host-wake-ui.js`) whose action is either Wake or, with no target configured, "Configure WoL" → `#wakeConfigModal` (saved via `PUT /api/remote-hosts/:id`). Detection is a throttled bare TCP probe (no ssh, no `ServerAliveInterval` — keepalives would move bytes into an idle connection every interval), input is buffered and flushed in order after `reattachRemote()` (the send-and-wait path blocks instead, as does the create path, with a shorter request budget), and the wake fields are re-read from `remote-hosts.json` on recovery AND (throttled, cached) live for a running session, because the persisted `remote` snapshot would never see a field added later (`rehydrateRemoteHostFields` + `RemoteWakeDeps.resolveRemote`). Design + invariants: `docs/remote-sessions.md` §Wake-on-LAN from user input.
**Remote host wake-on-LAN from user input**: an optional `RemoteHost.wakeMac` (magic packet built and broadcast by Codeman) or `RemoteHost.wakeCommand` (a single executable path, run WITHOUT a shell, and the explicit override) lets the input route — and an explicit `POST /api/sessions/:id/wake` — wake a SLEEPING host instead of writing into a stalled ssh pane; `tmux send-keys` succeeds against a stalled pane, so the bytes used to vanish silently. The wake flow lives in `src/remote-wake.ts` and is reachable **only** from an EXPLICIT user request: `POST /api/sessions/:id/input`, that explicit wake route, and the create/attach path (`POST /api/quick-start` for a remote case, `POST /api/sessions` with `attachRemoteSession`, via `ensureHostAwake`), because "the user pressed Run on a sleeping host" is the same kind of request and the tmux probe would otherwise fail with a misleading "needs tmux installed". Everything TIMER-driven must never wake a host: the COD-108 auto-reconnect watcher, `Server.handleRemoteSessionDropped` and boot recovery have no access to the registry, or a host would be re-woken seconds after each suspend and could never stay asleep (asserted by wiring guards in `test/remote-wake.test.ts`, not just documented — including that `ensureHostAwake` is called from the HTTP route only, since `cron-service.ts` builds sessions through the shared service with nobody waiting on the answer). `GET /api/sessions/:id/reachability` only ASKS — it never wakes — and feeds the amber "host unreachable" banner (`host-wake-ui.js`) whose action is either Wake or, with no target configured, "Configure WoL" → `#wakeConfigModal` (saved via `PUT /api/remote-hosts/:id`). Detection is a throttled bare TCP probe (no ssh, no `ServerAliveInterval` — keepalives would move bytes into an idle connection every interval; and a host behind a jump host/SOCKS proxy is reachability-UNKNOWN, never "asleep": `isProbeable()` keeps the registry from buffering, gating or bannering on a probe that cannot reach it), input is buffered and flushed in order after `reattachRemote()` (the send-and-wait path blocks instead, as does the create path, with a shorter request budget), and the wake fields are re-read from `remote-hosts.json` on recovery AND (throttled, cached) live for a running session, because the persisted `remote` snapshot would never see a field added later (`rehydrateRemoteHostFields` + `RemoteWakeDeps.resolveRemote`). Design + invariants: `docs/remote-sessions.md` §Wake-on-LAN from user input.
**Remote SSH cases** (COD-94/#145): cases can point at a **remote host** (`~/.codeman/remote-hosts.json` + `remote-cases.json` via `src/remote-hosts.ts`; CRUD under `/api/cases` — cases route file). A remote session launches a LOCAL tmux pane running `ssh <host>` that creates a durable REMOTE tmux session on a **dedicated socket** `-L codeman-remote` with name `codeman-ssh-<id>` — deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so a Codeman instance on the target host never adopts it; no `-g` global tmux options are set remotely. `remotePath`/`identityFile` are schema-guarded against shell injection (backticks/`$` rejected — same approach as `extraSshOptions`); remote tmux availability is probed via `checkRemoteTmuxAvailable()` in quick-start (ssh args carry `-o ConnectTimeout=10`). Remote claude defaults to an idempotent `claude --session-id <id> || claude --resume <id>` pair under a login shell, so a respawn or reattach continues the SAME conversation rather than starting a fresh one (remote omp gets the same treatment via `--continue`; ⚠️ because the claude arm is an `a || b` pair under `-c`, that pane's PID is the login shell, not the agent); per-host `commands.*` override. Session kill best-effort kills the remote tmux too. `SessionState.remote`/`MuxSession.remote` round-trip through recovery (`restoreMuxSessions` passes `remote` back into the Session constructor). ⚠️ Run flows must route remote cases through `POST /api/quick-start` (which resolves the remote case and skips LOCAL CLI availability gates) — `POST /api/sessions` stat-validates `workingDir` locally and has no `caseName`. `envOverrides`/`effort`/`modelOverride`/`codexConfig`/`geminiConfig` are rejected for remote quick-starts (not silently dropped). UI: Create Case modal → Remote tab. Tests: `test/remote-hosts.test.ts`, `test/remote-ssh-options.test.ts`. ⚠️ **Reading a file in a remote case goes over ssh too** (#415): `src/remote-files.ts` is the single remote-READ layer (`buildRemoteFileCommand` = `buildSshConnectionArgs` + one shellescaped remote command; `remoteProbePaths` returns remote realpath + stat; `remoteCreateReadStream` streams a `Range` via `tail -c +N | head -c L` and its `close()` must be wired to the response's `close` or the ssh child outlives an aborted download). The guard order matches the local path exactly (`validateSessionFilePathLexical` → remote realpath of BOTH file and workspace root → containment → sensitive-path → size cap on the REMOTE size), a request path arrives from the browser and is only ever interpolated as a `shellescape`d token, and an unreachable host answers **502**, never a 404. ⚠️ The probe's symlink resolution FAILS CLOSED: `readlink -f` where it exists, otherwise a `cd -P`/`pwd -P` directory walk plus a bounded plain-`readlink` loop over the last component, and anything it cannot fully resolve is reported unresolvable (404), never as the unresolved string — the first version resolved the directory chain only, so on a host without `readlink -f` a `ws/notes.txt -> ~/.ssh/id_rsa` link passed containment under its own path while `cat` served the key. Records are NUL-separated and index-keyed so a newline in a filename cannot shift the mapping. ⚠️ ssh children are BOUNDED: probes and buffered reads go through `src/remote-ssh-limiter.ts` (a `document-conversion-limiter`-shaped semaphore, default 4), the attachment-history list probes its whole history in ONE batched call (`probeRemoteAttachmentHistory`, threaded into `registerExternalAttachment({remoteProbes})`), and probes chunk at 40 paths — a prompt-injected agent printing `codeman://attach` links in a remote session used to fork one `ssh` per link. `describeExecError` never returns Node's `Command failed: <ssh line>` message (identity path + probe script in a 502 body). The `PUT /file-content` guard sits AHEAD of `validateSessionFilePath`, which resolves LOCALLY, or a same-named local directory (an sshfs mount) takes the write. Under `VITEST` the three IO functions refuse rather than connect. This covers the ATTACHMENT routes too, which is the half a clicked path needs when the file is OUTSIDE the case directory (`_isExternalPreviewPath` sends it to `POST …/attachments`): registration, by-id `raw`, metadata and the history list all resolve over ssh (`registerExternalAttachment({remote})`, `resolveServableRemoteAttachment`), and what decides the host is the SESSION, never the path string — the same absolute path means a different file on each host. Deliberately NOT supported over ssh: writes (`edit=1`/`PUT` answer 400, `editable` is always false), office previews/thumbnails, the file tree/picker, `tail-file`. Tests: `test/remote-files.test.ts`, `test/routes/file-routes-remote.test.ts`.
+60 -19
View File
@@ -25,7 +25,7 @@ custom port, identity file, `-J` jump host, `-o ProxyCommand`).
Types live in `src/types/session.ts`; persistence in `src/remote-hosts.ts`.
| Type | Role |
| -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|------|------|
| `RemoteSshOptions` | The **HOW-to-reach** fields, shared by host + session: `identityFile`, `socksProxy` (`host:port`), `jumpHost` (`[user@]host[:port]`), `extraSshOptions` (`KEY=VALUE[]`). Every field optional — all-absent reproduces port-22, default-identity, directly-SSH-able behavior. |
| `RemoteHost` (extends `RemoteSshOptions`) | A saved host: `id`, `label`, `host`, `username`, `port?`, `commands?` (per-mode launch command override). |
| `RemoteCase` | A working directory on a host: `name`, `type: 'remote'`, `hostId`, `remotePath`. |
@@ -73,7 +73,7 @@ Rules that keep this safe — **do not bypass them by hand-building an ssh line
single-quote `shellescape`d (`'…'` with embedded `'\''`). The helper mirrors
the one in `tmux-manager.ts`.
- **`~`/`$HOME` in `identityFile` is expanded at build time** (`expandIdentityPath`),
_before_ escaping — ssh does not expand `~` inside `-i`, and the escaped value
*before* escaping — ssh does not expand `~` inside `-i`, and the escaped value
never reaches a shell that would.
- **The ProxyCommand is one shellescaped `-o KEY=VALUE` token**, so its spaces and
the `%h`/`%p` placeholders reach ssh as a single argument. `%h %p` survive
@@ -112,7 +112,7 @@ Key points:
asymmetry: **discovery/attach (COD-105) target the canonical `-L codeman`
socket** — they join sessions the remote's own Codeman manages, while owned
durable launches live on `-L codeman-remote`.
- **`exec <cli>`** replaces the pane shell with the agent, so the pane PID _is_
- **`exec <cli>`** replaces the pane shell with the agent, so the pane PID *is*
the agent. The per-mode command comes from `remote.commands?.[mode]` or
`defaultRemoteCommandForMode(mode)` (`exec claude` / `exec opencode` /
`exec codex` / `exec gemini` / `exec agy` / `exec bash -l`).
@@ -133,9 +133,9 @@ Because durable remote sessions require tmux on the remote host,
`checkRemoteTmuxAvailable(host)` runs `command -v tmux` over SSH **before**
creating a remote case/session and returns a structured, never-throwing result:
- empty stdout / non-zero exit → _"remote host `<host>` needs tmux installed for
durable remote sessions"_
- stderr present → _"could not verify tmux on remote host `<host>`: `<stderr>`"_
- empty stdout / non-zero exit → *"remote host `<host>` needs tmux installed for
durable remote sessions"*
- stderr present → *"could not verify tmux on remote host `<host>`: `<stderr>`"*
(a real connection failure, surfaced to the operator)
- success → `{ ok: true, tmuxPath }`
@@ -152,7 +152,7 @@ skipped; command construction is still asserted by unit tests.
## Ownership: launched vs. discovered-and-attached (COD-105)
COD-104 (above) was Phase 1 — Codeman _launches_ a remote session and owns it.
COD-104 (above) was Phase 1 — Codeman *launches* a remote session and owns it.
COD-105 is Phase 2 — Codeman can also **discover** `codeman-*` tmux sessions
already running on a remote host (created by the remote's own Codeman or another
instance) and **attach** to one it didn't launch. Ownership decides what happens
@@ -193,7 +193,7 @@ remote command line by ownership:
- **`owned === false`** → `buildRemoteAttachCommand(remote, name)` — emits
`ssh … -t … 'tmux -L codeman attach -t <remoteSessionName>'`. It uses **`attach`,
NOT `new-session -A`**, so it only _joins_ an existing session and never creates
NOT `new-session -A`**, so it only *joins* an existing session and never creates
one.
- **owned (default)** → `buildRemoteLaunchCommand` (the COD-104 path above).
@@ -201,7 +201,7 @@ remote command line by ownership:
`TmuxManager.killSession()` has an **early return for non-owned remote sessions**:
it tears down **only the LOCAL pane** holding the ssh client (`tmux -L codeman
kill-session` on _this_ host's socket). Killing the local ssh sends SIGHUP to the
kill-session` on *this* host's socket). Killing the local ssh sends SIGHUP to the
remote `tmux attach`, which **detaches** — the durable remote session survives.
The early return is a structural guarantee that **no code path can ever issue a
remote `kill-session` for a session we don't own** — the only `kill-session` run is
@@ -262,7 +262,7 @@ and it follows the same rule as the launch path: every ssh command line comes fr
`buildSshConnectionArgs()` — **never** a hand-built ssh line.
| Request | What happens |
| -------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|---------|--------------|
| `GET /api/sessions/:id/file-raw` | Streamed over `ssh` (`cat`, or `tail -c +N \| head -c L` for a `Range`); the same 200/206/416 contract as a local file, so `<video>`/`<audio>` seeking works |
| `GET /api/sessions/:id/file-content` | `cat` into memory, capped by the existing text limit; `edit=1` answers `400` (see below) and `editable` is always `false` |
| `PUT /api/sessions/:id/file-content` | `400` before any path is looked at: the guard sits AHEAD of the local path validation, because with a same-named directory on the Codeman host (an `sshfs` mount) the write would otherwise land on the local twin |
@@ -397,8 +397,13 @@ machine is asleep. With a wake target the action is **Wake** (`POST /api/session
with none it is **Configure WoL** and opens `#wakeConfigModal`, a small form for that host's
`wakeMac`/`wakeCommand` that saves with `PUT /api/remote-hosts/:id` (in multi-user mode that
GET is admin-only, so a non-admin is told the setting is admin-only instead of "host not
found"). Reachability for the banner comes from `GET /api/sessions/:id/reachability`, polled
for the active remote session (30 s, visible tab only). ⚠️ The button is pressed from the SAME
found"). Reachability for the banner comes from `GET /api/sessions/:id/reachability`: once
when the remote tab is activated (a user action), and every 30 s while the tab is visible
**only for a host with a wake target** — each poll is a TCP connect to the host, and a timer
that connects to a host Codeman could not wake anyway is exactly the timer-driven traffic
the keepalive rule below rejects (it cannot wake a host, but it can keep an activity-based
suspend timer from firing). A host the probe cannot reach (see the next section) is never
polled. ⚠️ The button is pressed from the SAME
dashboard as Run/Attach, so it holds its request open under the same proxy and uses the same
40 s budget — and it **queues nothing**: browser keystrokes travel over the WebSocket, which
deliberately does not pass through the registry (that is the hot path this feature keeps its
@@ -406,6 +411,23 @@ hands off), so the banner says "waiting for the host to come back" for the butto
claims "input is queued" when the HTTP input path actually buffered bytes
(`queuedInput` on the two SSE events).
**Hosts behind a jump host or SOCKS proxy are reachability-UNKNOWN.** The probe is a bare
TCP connect to `host:port`, and a host reached through `jumpHost`, `socksProxy` or a
`ProxyCommand`/`ProxyJump` in `extraSshOptions` does not answer that even while ssh works —
the direct address may not route at all (the cloudflared case). Acting on the resulting
"unreachable" verdict was wrong three times over: a permanent banner over a healthy session,
a create-path error that replaced a genuine "needs tmux" with "not reachable", and — with a
wake target configured — every HTTP input buffered for the life of the session, because the
readiness poll could never succeed. `isProbeable()` (`remote-wake.ts`) decides from the
proxy fields, which travel on `WakeableRemote`; for such a host the registry delivers input
unchanged, `GET …/reachability` answers `reachable: null, probeable: false` (unknown is not
`false`, and only a proven `false` raises the banner), the create/attach path is not gated
(`ensureHostAwake` → `'unprobeable'`, handled like `'no-target'`), and the quick-start
"not reachable" message is reserved for a **proven** unreachable host (`=== false`). A wake
target can still be fired for it through `POST /api/sessions/:id/wake`, blind: the packet or
command goes out and the response says only whether it did — no readiness poll, no reattach
(the COD-108 watcher owns the pane once ssh works again), no "waking" toast.
The invariants worth keeping:
- **Only an EXPLICIT request may wake a host:** user input on an established session, the wake
@@ -438,7 +460,11 @@ The invariants worth keeping:
is deliberately NOT wake-aware**, so typing into a sleeping host sends nothing and queues
nothing (the banner's Wake button is the recovery for that case, which is why it must not
promise queued input). The **send-and-wait** path blocks on the wake instead — its response
is open anyway, and buffering would break the wait contract.
is open anyway, and buffering would break the wait contract. ⚠️ A flush write that FAILS
drops the whole remaining buffer (logged) rather than retaining it: the wake still resolves
and marks the host reachable, so the next input takes the deliver path while a retained
chunk would wait for the NEXT wake — replayed hours later, after everything typed since,
possibly ending in a carriage return. Same policy as the oversized paste.
- **The command runs without a shell** (`spawn(path, [], { stdio: 'ignore' })` — `shell`
defaults to `false`), the schema
requires a single executable path (no arguments, no `$`/backtick), and `wakeMac` is a
@@ -458,20 +484,35 @@ The invariants worth keeping:
its handlers are the ONLY definitions, since a second one in another mixin would be
silently shadowed by script order. Both carry `queuedInput`, which is true only when the
server actually holds bytes for that session — the wording keys off that, not off "a wake
is running", so the button path never claims input is queued.
is running", so the button path never claims input is queued. In multi-user mode the
whole `remote:` family is **session-scoped** (`deriveSseHint`, `server.ts`): an event with
a `sessionId` reaches that session's owner, and the create/attach wake — which has no
session yet — carries the requesting `username` instead (`ensureHostAwake({ requestedBy })`),
since its payload names a `hostId`/`label` that `GET /api/remote-hosts` withholds from
non-admins. With neither, it reaches admins only.
- **No real IO under vitest.** `probeRemoteHostReachable`, `runRemoteWakeCommand` and the
default UDP socket of `sendWakePackets` throw under `VITEST` (as `remote-files.ts` does),
so a test that reaches the defaults fails loudly instead of connecting, spawning or
broadcasting from CI. Every consumer injects its IO (`RemoteWakeDeps`, the socket
factory); `createDefaultRemoteWakeDeps({ probe })` also polls readiness with THAT probe,
which is the leak the guard found.
Tests: `test/remote-wake.test.ts` (decision/throttle table, single-flight registry,
buffering + flush order, MAC parsing/magic packet, live host-config resolution, and the wiring
guard) and `test/routes/session-remote-wake.test.ts` (the input route buffers instead of writing
into a sleeping host, the reachability route never wakes, and the wake route reports the
no-target case the UI turns into "configure WoL").
buffering + flush order, MAC parsing/magic packet, live host-config resolution, the proxied
host, SSE payload routing, the vitest IO guard, and the wiring guard),
`test/routes/session-remote-wake.test.ts` (the input route buffers instead of writing into a
sleeping host — and writes straight into a proxied one —, the reachability route never wakes
and reports a proxied host as unknown, and the wake route reports the no-target case the UI
turns into "configure WoL"), `test/sse-routing-remote.test.ts` (multi-user routing of the
`remote:` family) and `test/host-wake-banner.test.ts` (banner visibility and when the poller
may connect).
## API
Routes are registered in `src/web/routes/case-routes.ts`:
| Method | Path | Purpose |
| -------- | ------------------------------------ | ---------------------------------------------------------------------------------------------- |
|--------|------|---------|
| `GET` | `/api/remote-hosts` | List saved hosts |
| `POST` | `/api/remote-hosts` | Create a host |
| `PUT` | `/api/remote-hosts/:id` | Update a host |
+138 -24
View File
@@ -147,6 +147,29 @@ export interface WakeableRemote {
label: string;
host: string;
port?: number;
/** SSH jump host (`-J`): the host is reached THROUGH it, never directly. */
jumpHost?: string;
/** SOCKS5 proxy (`ProxyCommand=nc -X 5 …`): same, the direct address may not even route. */
socksProxy?: string;
/** Extra `-o KEY=VALUE` options; a `ProxyCommand`/`ProxyJump` in here proxies the host too. */
extraSshOptions?: string[];
}
/**
* Whether the bare TCP probe can answer for this host at all. Pure.
*
* The probe connects straight to `host:port`. A host behind a jump host or a SOCKS
* proxy (the cloudflared case) is reachable ONLY through that proxy, so the direct
* connect fails while ssh works — and every consumer of the verdict would then act on
* a "sleeping" host that is fine: a permanent banner, a create-path gate that hides the
* real ssh error, and (with a wake target) input buffered for the life of the session
* because the readiness poll can never succeed. Such a host is reachability-UNKNOWN:
* the registry never buffers for it, never gates on it, and reports `null` rather than
* `false`. A wake target can still be fired for it, blind.
*/
export function isProbeable(remote: WakeableRemote): boolean {
if (remote.jumpHost || remote.socksProxy) return false;
return !(remote.extraSshOptions ?? []).some((option) => /^\s*proxy(command|jump)\s*=/i.test(option));
}
/**
@@ -270,12 +293,14 @@ export function wakeConfigured(remote: WakeableRemote | undefined): WakeConfigur
/**
* Outcome of waking a host for a caller that has NO session yet (the create/attach
* routes). A union rather than a boolean because the three cases need different
* handling: `'no-target'` must leave the caller's behavior byte-identical (no probe,
* no extra latency for a host without WoL), and only `'failed'` is an error that
* deserves its own message instead of the caller's usual one.
* routes). A union rather than a boolean because the cases need different handling:
* `'no-target'` must leave the caller's behavior byte-identical (no probe, no extra
* latency for a host without WoL), `'unprobeable'` likewise (a proxied host, see
* {@link isProbeable} — the probe cannot tell asleep from awake, so nothing is gated on
* it), and only `'failed'` is an error that deserves its own message instead of the
* caller's usual one.
*/
export type HostWakeOutcome = 'no-target' | 'ready' | 'failed';
export type HostWakeOutcome = 'no-target' | 'unprobeable' | 'ready' | 'failed';
/**
* State key for a host-scoped wake. Prefixed so it can never collide with a session
@@ -368,15 +393,21 @@ export class RemoteWakeRegistry {
}
/**
* Reachability for the UI: probe unless a recent result is still fresh.
* Reachability for the UI: probe unless a recent result is still fresh. `null` for a
* host the probe cannot reach (see {@link isProbeable}): unknown is not unreachable.
*
* Shares the per-session probe state with the input path on purpose — a fresh
* answer is exactly what the input ladder wants, and an `unreachable` verdict here
* makes the next keystroke buffer + wake instead of vanishing into a stalled pane.
*/
async checkReachable(session: WakeableSession, opts: { force?: boolean; ttlMs?: number } = {}): Promise<boolean> {
async checkReachable(
session: WakeableSession,
opts: { force?: boolean; ttlMs?: number } = {}
): Promise<boolean | null> {
const remote = await this._effectiveRemote(session);
if (!remote) return true;
// `null`, never `false`: the UI keys the banner on a PROVEN unreachable host.
if (!isProbeable(remote)) return null;
const state = this._state(session.id);
const ttl = opts.force ? 0 : (opts.ttlMs ?? REMOTE_WAKE_REACHABILITY_TTL_MS);
if (Date.now() - state.probedAt >= ttl) {
@@ -396,6 +427,10 @@ export class RemoteWakeRegistry {
*/
async handleInput(session: WakeableSession, data: string): Promise<RemoteInputOutcome> {
const remote = await this._effectiveRemote(session);
// A proxied host can never pass the readiness poll, so buffering for it would hold
// the bytes for the life of the session (reproduced upstream: three inputs, nothing
// written, no reattach). Deliver, as if the feature were off.
if (remote && !isProbeable(remote)) return 'deliver';
const state = this._state(session.id);
const target = resolveWakeTarget(remote);
const action = decideRemoteInputAction({
@@ -433,7 +468,11 @@ export class RemoteWakeRegistry {
async ensureAwake(session: WakeableSession, opts: { force?: boolean; timeoutMs?: number } = {}): Promise<boolean> {
if (this.stopped) return false;
const remote = await this._effectiveRemote(session);
if (!remote || !resolveWakeTarget(remote)) return true;
const target = resolveWakeTarget(remote);
if (!remote || !target) return true;
// A proxied host: the send-and-wait path has nothing to gate on (unknown is not
// asleep), so it delivers; the manual button still wakes, blind (see `wake`).
if (!isProbeable(remote)) return opts.force ? this.wake(session, opts) : true;
const state = this._state(session.id);
// `force` is the manual path (a user pressed "wake"): a cached "reachable" from
// seconds ago must not talk the button out of waking a host that just slept.
@@ -452,9 +491,16 @@ export class RemoteWakeRegistry {
* Host-scoped reachability, for a caller that has no session yet (create/attach).
* Shares the per-HOST probe state with {@link ensureHostAwake}, so the probe the
* wake flow just paid for also answers "was that ssh failure really a sleeping
* machine?". Never wakes anything — it is a question, not an action.
* machine?". Never wakes anything — it is a question, not an action. `null` when the
* question cannot be answered (see {@link isProbeable}).
*/
async checkHostReachable(remote: WakeableRemote, opts: { force?: boolean; ttlMs?: number } = {}): Promise<boolean> {
async checkHostReachable(
remote: WakeableRemote,
opts: { force?: boolean; ttlMs?: number } = {}
): Promise<boolean | null> {
// `null` for a proxied host: callers gate on `=== false` (proven unreachable), so an
// unknown verdict leaves their ordinary error path — "needs tmux" — intact.
if (!isProbeable(remote)) return null;
const state = this._state(hostWakeKey(remote.hostId));
const ttl = opts.force ? 0 : (opts.ttlMs ?? REMOTE_WAKE_REACHABILITY_TTL_MS);
if (Date.now() - state.probedAt >= ttl) {
@@ -472,8 +518,14 @@ export class RemoteWakeRegistry {
* nothing and behaves exactly as before. Single-flight per host, so a double click
* (or two cases on the same host) sends one packet and shares one readiness poll.
*/
async ensureHostAwake(remote: WakeableRemote, opts: { timeoutMs?: number } = {}): Promise<HostWakeOutcome> {
async ensureHostAwake(
remote: WakeableRemote,
opts: { timeoutMs?: number; requestedBy?: string } = {}
): Promise<HostWakeOutcome> {
if (!resolveWakeTarget(remote)) return 'no-target';
// The probe cannot tell a proxied host asleep from awake, and a wake that cannot
// verify readiness would only delay the request by its whole budget. Not gated.
if (!isProbeable(remote)) return 'unprobeable';
if (this.stopped) return 'failed';
const state = this._state(hostWakeKey(remote.hostId));
if (state.waking) return (await state.waking) ? 'ready' : 'failed';
@@ -493,12 +545,16 @@ export class RemoteWakeRegistry {
* user-initiated and the host only wakes once), and keying them together would mean a
* create request joining an unrelated session's wake and inheriting its budget.
*/
private async wakeHost(remote: WakeableRemote, opts: { timeoutMs?: number }): Promise<boolean> {
private async wakeHost(remote: WakeableRemote, opts: { timeoutMs?: number; requestedBy?: string }): Promise<boolean> {
const state = this._state(hostWakeKey(remote.hostId));
if (state.waking) return state.waking;
state.waking = (async (): Promise<boolean> => {
try {
return await this._wakeAndWait(remote, state, { timeoutMs: opts.timeoutMs, forNewSession: true });
return await this._wakeAndWait(remote, state, {
timeoutMs: opts.timeoutMs,
forNewSession: true,
requestedBy: opts.requestedBy,
});
} catch (err) {
// Injected IO is documented not to throw, but a rejected promise here would
// surface as an unhandled rejection AND take the route down with it (the
@@ -522,6 +578,7 @@ export class RemoteWakeRegistry {
const remote = await this._effectiveRemote(session);
const target = resolveWakeTarget(remote);
if (!remote || !target) return true;
if (!isProbeable(remote)) return this._wakeBlind(remote, target);
const state = this._state(session.id);
if (state.waking) return state.waking;
@@ -556,6 +613,24 @@ export class RemoteWakeRegistry {
return state.waking;
}
/**
* Fire the wake target for a host whose readiness cannot be verified (see
* {@link isProbeable}): no readiness poll (it could never succeed), no reattach (the
* COD-108 watcher owns the pane once ssh works again), no `hostWaking` broadcast (its
* toast promises a wait that does not happen). The caller learns only whether the
* packet/command went out — and a wake IO that throws is a failed wake, never a
* rejected route.
*/
private async _wakeBlind(remote: WakeableRemote, target: NonNullable<WakeTarget>): Promise<boolean> {
this.deps.log?.(`[RemoteWake] waking ${remote.label} (${remote.host}) via ${target.kind}, blind: proxied host`);
try {
return await this.deps.wake(target);
} catch (err) {
this.deps.log?.(`[RemoteWake] unexpected failure: ${err instanceof Error ? err.message : String(err)}`);
return false;
}
}
/**
* Broadcast + run the wake target + wait for SSH. Shared by the session flow (which
* then reattaches and flushes the buffer) and the create/attach flow (which has no
@@ -565,11 +640,18 @@ export class RemoteWakeRegistry {
private async _wakeAndWait(
remote: WakeableRemote,
state: WakeState,
opts: { sessionId?: string; timeoutMs?: number; forNewSession?: boolean } = {}
opts: { sessionId?: string; timeoutMs?: number; forNewSession?: boolean; requestedBy?: string } = {}
): Promise<boolean> {
const target = resolveWakeTarget(remote);
if (!target) return true;
const forWhat = opts.sessionId ? `for session ${opts.sessionId}` : 'for a new session';
// Routing for multi-user mode (server.ts `deriveSseHint`): a session-scoped event
// reaches its owner, and the create/attach wake has no session yet — so it names
// the requesting user instead, or it would reach admins only. The payload carries
// `hostId`/`label`, which non-admins are not shown elsewhere, so it must not go global.
const scope = opts.sessionId
? { sessionId: opts.sessionId }
: { forNewSession: true, ...(opts.requestedBy ? { username: opts.requestedBy } : {}) };
// No `sessionId` for a create-path wake: the toast handler is then the only one
// that acts (a banner for a session that does not exist yet would have no target),
// which is exactly the `forNewSession` distinction the UI renders.
@@ -580,7 +662,7 @@ export class RemoteWakeRegistry {
// something the user can disprove by typing (browser keystrokes go over the
// WebSocket, which never passes through this registry).
this.deps.broadcast?.('remote:hostWaking', {
...(opts.sessionId ? { sessionId: opts.sessionId } : { forNewSession: true }),
...scope,
hostId: remote.hostId,
label: remote.label,
queuedInput: state.pending.length > 0,
@@ -603,7 +685,7 @@ export class RemoteWakeRegistry {
`[RemoteWake] ${remote.label} did not come back — ${opts.forNewSession ? 'the session was not started' : 'input stays buffered'}`
);
this.deps.broadcast?.('remote:hostWakeFailed', {
...(opts.sessionId ? { sessionId: opts.sessionId } : { forNewSession: true }),
...scope,
hostId: remote.hostId,
label: remote.label,
queuedInput: state.pending.length > 0,
@@ -699,10 +781,15 @@ export class RemoteWakeRegistry {
state.pending = state.pending.slice(1);
const ok = await session.writeViaMux(chunk).catch(() => false);
if (!ok) {
// Retain it, IN ORDER: a failed write must not reorder the queue behind it.
state.pending = [chunk, ...state.pending];
// Drop the rest, and say so. Retaining it looked safer but was worse: the wake
// still resolves and marks the host reachable, so the NEXT input takes the
// deliver path while the old chunks sit here — to be replayed by the next wake,
// possibly hours later, after everything typed since, and maybe ending in a
// carriage return. Same policy as the oversized paste: gone, with a log line.
const dropped = state.pending.length + 1;
state.pending = [];
this.deps.log?.(
`[RemoteWake] flush failed for session ${session.id} — ${state.pending.length} chunk(s) retained`
`[RemoteWake] flush failed for session ${session.id} — ${dropped} buffered chunk(s) dropped rather than replayed on a later wake`
);
return;
}
@@ -713,7 +800,21 @@ export class RemoteWakeRegistry {
// ========== Default IO ==========
/**
* Cheap reachability probe: a bare TCP connect to the SSH port.
* Under vitest none of this may do real IO (a TCP connect, a child process, a UDP
* broadcast) — mirrors `remote-files.ts`. Every consumer injects its deps
* (`RemoteWakeDeps`, the socket factory); this is what makes that seam non-optional
* instead of a convention the next test can forget.
*/
function assertNotUnderTest(what: string): void {
if (process.env.VITEST) {
throw new Error(`remote-wake: ${what} is disabled under test — inject a fake (RemoteWakeDeps / WakeSocketFactory)`);
}
}
/**
* Cheap reachability probe: a bare TCP connect to the SSH port. Only meaningful for a
* host the registry deems probeable (see {@link isProbeable}); the registry never asks
* it about a proxied host.
*
* Deliberately NOT an `ssh … true` probe: that opens a full session (auth,
* remote log, process) every throttle window for a question a SYN already
@@ -725,6 +826,7 @@ export function probeRemoteHostReachable(
remote: WakeableRemote,
timeoutMs = REMOTE_WAKE_PROBE_TIMEOUT_MS
): Promise<boolean> {
assertNotUnderTest('the TCP probe');
const port = remote.port ?? DEFAULT_SSH_PORT;
return new Promise((resolve) => {
let settled = false;
@@ -748,6 +850,7 @@ export function probeRemoteHostReachable(
* than throwing: a broken wake command must not break the input route.
*/
export function runRemoteWakeCommand(command: string, timeoutMs = REMOTE_WAKE_COMMAND_TIMEOUT_MS): Promise<boolean> {
assertNotUnderTest('the wake command');
return new Promise((resolve) => {
let settled = false;
const finish = (value: boolean) => {
@@ -790,7 +893,10 @@ export function runRemoteWakeCommand(command: string, timeoutMs = REMOTE_WAKE_CO
export function sendWakePackets(
addresses: number[][],
port = 9,
createSocket: WakeSocketFactory = () => dgram.createSocket('udp4')
createSocket: WakeSocketFactory = () => {
assertNotUnderTest('the UDP broadcast');
return dgram.createSocket('udp4');
}
): Promise<boolean> {
if (addresses.length === 0) return Promise.resolve(false);
return new Promise((resolve) => {
@@ -884,12 +990,20 @@ function delayOrAbort(ms: number, signal?: AbortSignal): Promise<void> {
const delay = (ms: number): Promise<void> => new Promise((resolve) => setTimeout(resolve, ms));
/** Production wiring: all IO defaults, overridable for tests. */
/**
* Production wiring: all IO defaults, overridable for tests.
*
* The readiness poll uses the SAME probe as the rest of the deps, overridden or not.
* Wiring it to the module default instead let a caller that injected `probe` still
* poll the real host during the wait — under vitest, a TCP connect to a production
* address on every shutdown test (which the vitest guard is what finally caught).
*/
export function createDefaultRemoteWakeDeps(overrides: Partial<RemoteWakeDeps> = {}): RemoteWakeDeps {
const probe = overrides.probe ?? probeRemoteHostReachable;
return {
probe: probeRemoteHostReachable,
probe,
wake: (target) => (target.kind === 'command' ? runRemoteWakeCommand(target.command) : sendWakePackets(target.macs)),
waitUntilReady: (remote, opts) => waitUntilRemoteReady(remote, opts),
waitUntilReady: (remote, opts) => waitUntilRemoteReady(remote, { ...opts, probe }),
delay,
...overrides,
};
+35 -7
View File
@@ -9,10 +9,17 @@
* nothing happening" into one click.
*
* Behavior:
* - Polls `GET /api/sessions/:id/reachability` for the ACTIVE remote session only
* (on tab activation and every `POLL_MS` while the tab is visible). The endpoint
* shares the server's probe cache with the input path, so opening the tab also
* primes the wake path.
* - Asks `GET /api/sessions/:id/reachability` for the ACTIVE remote session only:
* once when the tab is activated (a user action), and every `POLL_MS` while the tab
* is visible ONLY for a host with a wake target. The timer is the one thing here that
* is not user-driven, and each poll is a TCP connect to the host — the same
* timer-driven traffic invariant #2 rejects keepalives for: it cannot wake a host,
* but it can keep an activity-based suspend timer from firing. So a host Codeman
* could not wake anyway is never polled on a timer. A host behind a jump host or
* SOCKS proxy (`probeable: false`) is never polled at all: the probe cannot reach
* it, so its answer would only ever be a false "asleep". The endpoint shares the
* server's probe cache with the input path, so opening the tab also primes the
* wake path.
* - Unreachable + a configured wake target → "Wake" button → `POST /api/sessions/:id/wake`
* (which wakes, waits, reattaches the pane and flushes buffered input).
* - Unreachable + NO wake target → "Configure WoL" → `#wakeConfigModal`, a small form
@@ -46,6 +53,11 @@ Object.assign(CodemanApp.prototype, {
wakeConfigured: 'none',
host: '',
label: '',
/**
* False for a host the server's probe cannot reach (behind a jump host or SOCKS
* proxy): its reachability is unknown, so there is no banner and no polling.
*/
probeable: true,
/** True between clicking Wake and the answer coming back. */
waking: false,
/**
@@ -79,17 +91,22 @@ Object.assign(CodemanApp.prototype, {
/** Create the page-wide poller once (interval + a visibility wake-up). */
_ensureHostWakePoller() {
if (this._hostWakeTimer) return;
this._hostWakeTimer = setInterval(() => this._hostWakeTick(), HOST_WAKE_POLL_MS);
this._hostWakeTimer = setInterval(() => this._hostWakeTick({ periodic: true }), HOST_WAKE_POLL_MS);
document.addEventListener('visibilitychange', () => {
if (document.visibilityState === 'visible') this._hostWakeTick();
if (document.visibilityState === 'visible') this._hostWakeTick({ periodic: true });
});
},
/**
* One poller tick: resolve the ACTIVE session, reset the banner when it changed, and
* ask the server. No-op while the page is hidden (a background tab must not poll).
*
* `periodic` marks the timer (and the visibility wake-up) as opposed to a tab
* activation: a periodic tick polls only a host with a wake target, see the module
* comment. The activation poll is what still offers "Configure WoL" for a sleeping
* host that has none — one connect, on a user action.
*/
_hostWakeTick() {
_hostWakeTick({ periodic = false } = {}) {
if (typeof document !== 'undefined' && document.visibilityState === 'hidden') return;
const sessionId = this.activeSessionId;
const session = sessionId && this.sessions ? this.sessions.get(sessionId) : null;
@@ -103,7 +120,9 @@ Object.assign(CodemanApp.prototype, {
return;
}
let state = this._hostWake;
let fresh = false;
if (!state || state.sessionId !== sessionId) {
fresh = true;
state = this._hostWake = this._hostWakeState();
state.sessionId = sessionId;
state.host = session.remote.host || '';
@@ -114,8 +133,14 @@ Object.assign(CodemanApp.prototype, {
// path is configured, so a command-only host is not mislabelled 'mac' until the
// first poll lands.
state.wakeConfigured = session.remote.wakeMac ? 'mac' : session.remote.wakeCommand ? 'command' : 'none';
// Known from the payload already: a proxied host is not probeable (the server
// says so too, on every answer), so not even the activation poll is worth a
// round trip whose verdict could only be a wrong "asleep".
state.probeable = !(session.remote.jumpHost || session.remote.socksProxy);
this._renderHostWakeBanner();
}
if (!state.probeable) return;
if (periodic && !fresh && state.wakeConfigured === 'none') return;
this._pollHostReachability();
},
@@ -130,7 +155,10 @@ Object.assign(CodemanApp.prototype, {
if (!data.success) return;
// The tab may have changed while this was in flight.
if (this._hostWake !== state || state.sessionId !== sessionId) return;
// `reachable` is `null` (unknown, not unreachable) for a host the probe cannot
// reach — only a PROVEN `false` may raise the banner.
state.reachable = data.data.reachable !== false;
if (data.data.probeable === false) state.probeable = false;
state.wakeConfigured = data.data.wakeConfigured || 'none';
if (data.data.host) state.host = data.data.host;
if (data.data.label) state.label = data.data.label;
+20 -3
View File
@@ -72,6 +72,7 @@ import {
RemoteWakeRegistry,
REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
createDefaultRemoteWakeDeps,
isProbeable,
type WakeableRemote,
} from '../../remote-wake.js';
import { clampWaitMs, MAX_BUFFER_SCAN_BYTES } from '../../config/agent-wait.js';
@@ -760,7 +761,8 @@ export function resolveOmpConfigForCreate(
/**
* `RemoteHost` → the wake registry's host shape. They differ in one field name only
* (`id` in host config vs `hostId` on a session's `remote`), but the rename is load-
* bearing: the registry keys its per-host wake state on `hostId`.
* bearing: the registry keys its per-host wake state on `hostId`. The proxy fields
* travel too: they are what tells the registry its probe cannot reach this host.
*/
function wakeableHost(host: RemoteHost): WakeableRemote {
return {
@@ -770,6 +772,9 @@ function wakeableHost(host: RemoteHost): WakeableRemote {
port: host.port,
wakeMac: host.wakeMac,
wakeCommand: host.wakeCommand,
jumpHost: host.jumpHost,
socksProxy: host.socksProxy,
extraSshOptions: host.extraSshOptions,
};
}
@@ -892,6 +897,8 @@ export function registerSessionRoutes(
// answers with an ssh failure that blames anything but the machine being asleep.
const hostWake = await remoteWake.ensureHostAwake(wakeableHost(host), {
timeoutMs: REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
// No session yet, so the wake events name their requester (multi-user routing).
requestedBy: ownerFor(req),
});
if (hostWake === 'failed') {
return createErrorResponse(
@@ -1535,13 +1542,18 @@ export function registerSessionRoutes(
const { id } = req.params as { id: string };
const session = findSessionOrFail(ctx, id, req);
const remote = session.remote;
if (!remote) return { success: true, data: { reachable: true, wakeConfigured: 'none' as const } };
if (!remote) {
return { success: true, data: { reachable: true, probeable: true, wakeConfigured: 'none' as const } };
}
const force = (req.query as { force?: string })?.force === '1';
// `reachable: null` + `probeable: false` for a host behind a jump host / SOCKS proxy:
// the probe cannot reach it, so the UI shows no banner and stops polling.
const reachable = await remoteWake.checkReachable(session, { force });
return {
success: true,
data: {
reachable,
probeable: isProbeable(remote),
wakeConfigured: await remoteWake.wakeConfigured(session),
host: remote.host,
label: remote.label,
@@ -3264,6 +3276,8 @@ export function registerSessionRoutes(
// host on every schedule (the failure invariant #1 exists to prevent).
const hostWake = await remoteWake.ensureHostAwake(wakeableHost(host), {
timeoutMs: REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
// No session yet, so the wake events name their requester (multi-user routing).
requestedBy: ownerFor(req),
});
if (hostWake === 'failed') {
return createErrorResponse(
@@ -3280,7 +3294,10 @@ export function registerSessionRoutes(
// An unreachable host and a host without tmux fail the same way over ssh, so the
// probe's own message would send the user hunting for a tmux install. Ask the
// registry (which just probed, when it woke the host) which of the two it is.
if (!(await remoteWake.checkHostReachable(wakeableHost(host)))) {
// `=== false` on purpose: a proxied host answers `null` (the probe cannot reach
// it), and an unknown verdict must not replace the real ssh error with
// "not reachable" over a host that is fine.
if ((await remoteWake.checkHostReachable(wakeableHost(host))) === false) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
hostWake === 'no-target'
+9 -1
View File
@@ -2349,11 +2349,19 @@ export class WebServer extends EventEmitter {
'scheduled:',
'team:',
'case:',
'remote:',
];
if (SESSION_PREFIXES.some((p) => event.startsWith(p))) {
const d = (data ?? {}) as { sessionId?: string; id?: string; session?: { id?: string } };
const d = (data ?? {}) as { sessionId?: string; id?: string; session?: { id?: string }; username?: string };
const sessionId = d.sessionId ?? d.id ?? d.session?.id;
const owner = sessionId ? this.sessions.get(sessionId)?.owner : undefined;
// `remote:hostWaking` / `remote:hostWakeFailed` for a create/attach wake have no
// session yet (nothing exists until the host is up), so the registry names the
// requesting user instead; the payload carries `hostId`/`label`, which non-admins
// are not shown elsewhere. No session and no requester: admins only (fail closed).
if (!sessionId && event.startsWith('remote:') && d.username) {
return { username: d.username, sessionScoped: true };
}
return { owner, sessionScoped: true };
}
// #20/#38: clipboard:write writes into the receiver's OS clipboard — route it to
+4
View File
@@ -177,6 +177,8 @@ export const MuxDied = 'mux:died' as const;
export const MuxStatsUpdated = 'mux:statsUpdated' as const;
// ─── Remote auto-reconnect (COD-108) + wake-on-LAN ───────────────────────────
// Session-scoped in multi-user mode (`deriveSseHint`): routed to the session's owner,
// or — for a wake with no session yet — to the requesting `username` in the payload.
/** A remote session's local ssh pane died; an auto-reconnect attempt is starting. */
export const RemoteSessionDropped = 'remote:sessionDropped' as const;
@@ -187,6 +189,8 @@ export const RemoteReconnectExhausted = 'remote:reconnectExhausted' as const;
/**
* User input arrived for a session whose host is unreachable, so a Wake-on-LAN
* command was started (see `remote-wake.ts`). Input sent meanwhile is buffered.
* Payload: `sessionId` (session wake) or `forNewSession: true` + `username`
* (create/attach wake), `hostId`, `label`, `queuedInput`.
*/
export const RemoteHostWaking = 'remote:hostWaking' as const;
/** The host did not come back within the wake timeout — buffered input is still held. */
+64 -2
View File
@@ -26,8 +26,12 @@ function fakeElement(): El {
return { hidden: false, textContent: '', disabled: false, classList: { add() {}, remove() {} } };
}
const PROXIED_ID = 'remote-session-proxied';
const NOWOL_ID = 'remote-session-nowol';
/** Load `host-wake-ui.js` with the minimal DOM it touches, and return a wired app. */
function loadWakeApp() {
const fetches: string[] = [];
const elements = new Map<string, El>([
['hostWakeBanner', fakeElement()],
['hostWakeBannerText', fakeElement()],
@@ -40,7 +44,10 @@ function loadWakeApp() {
console,
setInterval: () => 1,
clearInterval: () => {},
fetch: () => Promise.resolve({ json: () => Promise.resolve({ success: false }) }),
fetch: (url: string) => {
fetches.push(url);
return Promise.resolve({ json: () => Promise.resolve({ success: false }) });
},
document: {
visibilityState: 'visible',
getElementById: (id: string) => elements.get(id) ?? null,
@@ -58,9 +65,27 @@ function loadWakeApp() {
REMOTE_ID,
{ remote: { hostId: 'hufflepuff', host: '192.168.50.137', label: 'Hufflepuff', wakeMac: '04:d9:f5:80:c6:58' } },
],
[
PROXIED_ID,
{
remote: {
hostId: 'bastioned',
host: '10.20.0.5',
label: 'Behind bastion',
jumpHost: 'bastion',
wakeMac: '04:d9:f5:80:c6:58',
},
},
],
[NOWOL_ID, { remote: { hostId: 'plain', host: '10.0.0.9', label: 'Plain' } }],
[LOCAL_ID, {}],
]);
return { app, banner: elements.get('hostWakeBanner') as El, text: elements.get('hostWakeBannerText') as El };
return {
app,
fetches,
banner: elements.get('hostWakeBanner') as El,
text: elements.get('hostWakeBannerText') as El,
};
}
describe('host wake banner visibility', () => {
@@ -120,3 +145,40 @@ describe('host wake banner visibility', () => {
expect(banner.hidden).toBe(true);
});
});
describe('host wake banner polling', () => {
// Each poll is a TCP connect to the host from the server. The timer is the one
// trigger that is not a user action, so it must not fire for a host Codeman could
// not wake anyway (it cannot wake it, but it can keep an activity-based suspend timer
// from firing), and a proxied host is never polled: the probe cannot reach it.
const tick = (app: Record<string, unknown>, periodic: boolean) =>
(app._hostWakeTick as (o: { periodic: boolean }) => void)({ periodic });
it('polls a wake-configured host on activation and on the timer', () => {
const { app, fetches } = loadWakeApp();
app.activeSessionId = REMOTE_ID;
tick(app, false);
tick(app, true);
tick(app, true);
expect(fetches).toHaveLength(3);
expect(fetches[0]).toContain(`/api/sessions/${REMOTE_ID}/reachability`);
});
it('polls a host without a wake target once on activation, never on the timer', () => {
const { app, fetches } = loadWakeApp();
app.activeSessionId = NOWOL_ID;
tick(app, false);
tick(app, true);
tick(app, true);
expect(fetches).toHaveLength(1);
});
it('never polls a host behind a jump host or SOCKS proxy', () => {
const { app, fetches } = loadWakeApp();
app.activeSessionId = PROXIED_ID;
tick(app, false);
tick(app, true);
expect(fetches).toHaveLength(0);
expect((app._hostWake as { probeable: boolean }).probeable).toBe(false);
});
});
+175 -3
View File
@@ -22,8 +22,11 @@ import {
buildMagicPacket,
createDefaultRemoteWakeDeps,
decideRemoteInputAction,
isProbeable,
parseMacList,
probeRemoteHostReachable,
resolveWakeTarget,
runRemoteWakeCommand,
sendWakePackets,
waitUntilRemoteReady,
wakeConfigured,
@@ -217,6 +220,7 @@ interface Harness {
writeViaMux: ReturnType<typeof vi.fn>;
noteReconnected: ReturnType<typeof vi.fn>;
events: string[];
payloads: Array<{ event: string; payload: Record<string, unknown> }>;
}
function harness(
@@ -229,6 +233,7 @@ function harness(
const writeViaMux = vi.fn(async () => !opts.writesFail);
const noteReconnected = vi.fn();
const events: string[] = [];
const payloads: Array<{ event: string; payload: Record<string, unknown> }> = [];
const deps: RemoteWakeDeps = {
probe,
@@ -236,7 +241,10 @@ function harness(
waitUntilReady,
delay: async () => {},
noteReconnected,
broadcast: (event) => events.push(event),
broadcast: (event, payload) => {
events.push(event);
payloads.push({ event, payload });
},
log: () => {},
...(opts.resolveRemote ? { resolveRemote: opts.resolveRemote } : {}),
};
@@ -258,6 +266,7 @@ function harness(
writeViaMux,
noteReconnected,
events,
payloads,
};
}
@@ -361,15 +370,24 @@ describe('RemoteWakeRegistry', () => {
expect(h.events).not.toContain('remote:sessionReconnected');
});
it('retains input that could not be written and reports nothing lost', async () => {
it('drops the buffer when a flush write fails, so nothing is replayed by a later wake', async () => {
// Retaining the chunk was the earlier behaviour, and it was worse: the wake still
// resolves and marks the host reachable, so the next input takes the deliver path
// while the retained chunk waits for the NEXT wake — replayed hours later, after
// everything typed since. Same policy as the oversized paste: dropped, logged.
const h = harness({ writesFail: true });
h.probe.mockResolvedValue(false);
await h.registry.handleInput(h.session, 'abc');
await h.registry.handleInput(h.session, 'def');
await h.registry.wake(h.session);
expect(h.writeViaMux).toHaveBeenCalledTimes(1);
expect(h.registry.pendingBytes('sess-1')).toBe(3);
expect(h.registry.pendingBytes('sess-1')).toBe(0);
// And the recovered host takes the deliver path from here, with nothing behind it.
h.probe.mockClear();
await expect(h.registry.handleInput(h.session, 'g')).resolves.toBe('deliver');
expect(h.registry.pendingBytes('sess-1')).toBe(0);
});
it('flushes the chunk it is writing out of the buffer first, so a concurrent enqueue cannot drop a different one', async () => {
@@ -528,6 +546,160 @@ describe('RemoteWakeRegistry', () => {
// ========== Host-scoped wake (session create/attach) ==========
describe('isProbeable', () => {
const base: WakeableRemote = { hostId: 'h', label: 'H', host: '10.0.0.9', wakeMac: '04:d9:f5:80:c6:58' };
it('is true for a host reached directly', () => {
expect(isProbeable(base)).toBe(true);
expect(isProbeable({ ...base, extraSshOptions: ['ServerAliveCountMax=3', 'StrictHostKeyChecking=no'] })).toBe(true);
});
it('is false behind a jump host, a SOCKS proxy, or a ProxyCommand/ProxyJump option', () => {
expect(isProbeable({ ...base, jumpHost: 'bastion.example' })).toBe(false);
expect(isProbeable({ ...base, socksProxy: '127.0.0.1:1080' })).toBe(false);
expect(isProbeable({ ...base, extraSshOptions: ['ProxyCommand=cloudflared access ssh --hostname %h'] })).toBe(
false
);
expect(isProbeable({ ...base, extraSshOptions: ['proxyjump=bastion'] })).toBe(false);
});
});
describe('RemoteWakeRegistry — a proxied host is reachability-unknown', () => {
// The bare TCP probe connects to `host:port`, which a jump-host/SOCKS host does not
// answer even while ssh works. Acting on that verdict buffered input for the life of
// the session (the readiness poll could never succeed), showed a permanent banner and
// hid the real ssh error behind "not reachable". Unknown is not asleep.
const proxied: WakeableRemote = {
hostId: 'behind-bastion',
label: 'Behind bastion',
host: '10.20.0.5',
jumpHost: 'bastion.example',
wakeCommand: '/usr/local/bin/wake-behind-bastion',
};
it('delivers every input without probing, buffering or waking', async () => {
const h = harness({ remote: proxied });
await expect(h.registry.handleInput(h.session, 'ls\r')).resolves.toBe('deliver');
await expect(h.registry.handleInput(h.session, 'pwd\r')).resolves.toBe('deliver');
expect(h.probe).not.toHaveBeenCalled();
expect(h.wake).not.toHaveBeenCalled();
expect(h.registry.pendingBytes('sess-1')).toBe(0);
});
it('answers null (unknown), never false, so the UI has no banner to raise', async () => {
const h = harness({ remote: proxied });
await expect(h.registry.checkReachable(h.session, { force: true })).resolves.toBeNull();
await expect(h.registry.checkHostReachable(proxied, { force: true })).resolves.toBeNull();
expect(h.probe).not.toHaveBeenCalled();
});
it('does not gate a create/attach request on it (unprobeable, like no-target)', async () => {
const h = harness({ remote: proxied });
await expect(h.registry.ensureHostAwake(proxied)).resolves.toBe('unprobeable');
expect(h.probe).not.toHaveBeenCalled();
expect(h.wake).not.toHaveBeenCalled();
});
it('lets the send-and-wait path through, and fires the manual wake blind', async () => {
const h = harness({ remote: proxied });
await expect(h.registry.ensureAwake(h.session)).resolves.toBe(true);
expect(h.wake).not.toHaveBeenCalled();
// The button: the user asked, so the target goes out — but nothing can verify the
// host came back, so there is no readiness poll, no reattach and no "waking" toast
// promising a wait that does not happen.
await expect(h.registry.ensureAwake(h.session, { force: true })).resolves.toBe(true);
expect(h.wake).toHaveBeenCalledTimes(1);
expect(h.waitUntilReady).not.toHaveBeenCalled();
expect(h.reattachRemote).not.toHaveBeenCalled();
expect(h.events).toEqual([]);
h.wake.mockResolvedValueOnce(false);
await expect(h.registry.ensureAwake(h.session, { force: true })).resolves.toBe(false);
// A wake IO that throws is a failed wake, not a rejected route — and the public
// `wake()` takes the same blind path, so nobody can poll readiness through a proxy.
h.wake.mockRejectedValueOnce(new Error('udp socket exploded'));
await expect(h.registry.wake(h.session)).resolves.toBe(false);
expect(h.waitUntilReady).not.toHaveBeenCalled();
});
});
describe('RemoteWakeRegistry — SSE payload routing', () => {
const hostRemote: WakeableRemote = {
hostId: 'hufflepuff',
label: 'Hufflepuff',
host: '192.168.50.137',
wakeMac: '04:d9:f5:80:c6:58',
};
it('a session wake names its session, so the server routes it to the owner', async () => {
const h = harness({ remote: hostRemote });
h.probe.mockResolvedValue(false);
await h.registry.handleInput(h.session, 'x');
await h.registry.wake(h.session);
const waking = h.payloads.find((p) => p.event === 'remote:hostWaking')!;
expect(waking.payload).toMatchObject({ sessionId: 'sess-1', hostId: 'hufflepuff', label: 'Hufflepuff' });
expect(waking.payload).not.toHaveProperty('username');
});
it('a create/attach wake has no session, so it names the requesting user instead', async () => {
// Without it the server can only fail closed (admins only) — the requester would
// never see their own wake. The payload carries `hostId`/`label`, which non-admins
// are not shown elsewhere, so it must not go global either.
const h = harness({ remote: hostRemote });
h.probe.mockResolvedValue(false);
h.waitUntilReady.mockResolvedValue(false);
await expect(h.registry.ensureHostAwake(hostRemote, { requestedBy: 'alice' })).resolves.toBe('failed');
const [waking, failed] = ['remote:hostWaking', 'remote:hostWakeFailed'].map(
(event) => h.payloads.find((p) => p.event === event)!.payload
);
expect(waking).toMatchObject({ forNewSession: true, username: 'alice' });
expect(failed).toMatchObject({ forNewSession: true, username: 'alice' });
expect(waking).not.toHaveProperty('sessionId');
});
it('omits the requester when the route did not name one (single-user mode)', async () => {
const h = harness({ remote: hostRemote });
h.probe.mockResolvedValue(false);
await h.registry.ensureHostAwake(hostRemote);
expect(h.payloads.find((p) => p.event === 'remote:hostWaking')!.payload).not.toHaveProperty('username');
});
});
describe('real IO is refused under vitest', () => {
// Every consumer injects its IO (RemoteWakeDeps, the socket factory). The guard is
// what makes that seam mandatory: a test that reaches the defaults fails loudly here
// instead of opening a TCP connection, spawning a process or broadcasting UDP from CI.
const target: WakeableRemote = { hostId: 'h', label: 'H', host: '127.0.0.1', port: 1 };
it('the TCP probe', () => {
expect(() => probeRemoteHostReachable(target)).toThrow(/disabled under test/);
});
it('the wake command', () => {
expect(() => runRemoteWakeCommand('/bin/true')).toThrow(/disabled under test/);
});
it('the UDP broadcast — only with the DEFAULT socket, an injected one still works', async () => {
await expect(sendWakePackets([[1, 2, 3, 4, 5, 6]])).rejects.toThrow(/disabled under test/);
});
it('the readiness poll, which probes by default', async () => {
await expect(waitUntilRemoteReady(target, { timeoutMs: 10, intervalMs: 1 })).rejects.toThrow(/disabled under test/);
});
it('the default deps poll readiness with the INJECTED probe, never the real one', async () => {
// `createDefaultRemoteWakeDeps({ probe })` used to override `probe` alone while
// `waitUntilReady` kept the module default — so a shutdown test polled a production
// address until the guard above made it fail instead of connecting.
const probe = vi.fn(async () => true);
const deps = createDefaultRemoteWakeDeps({ probe });
await expect(deps.waitUntilReady(target, { timeoutMs: 10 })).resolves.toBe(true);
expect(probe).toHaveBeenCalledWith(target);
});
});
describe('RemoteWakeRegistry — host-scoped wake for a request that waits on it', () => {
const hostRemote: WakeableRemote = {
hostId: 'hufflepuff',
+24
View File
@@ -151,6 +151,19 @@ describe('POST /api/sessions/:id/input — wake-on-LAN', () => {
expect(h.wake).not.toHaveBeenCalled();
});
it('writes straight into a proxied host with a wake target: no probe, no buffer, no wake', async () => {
// With a target configured, the old verdict buffered EVERY input for the life of
// the session: the readiness poll can never succeed through a proxy, so nothing was
// ever flushed (three inputs, nothing written, buffer non-empty — reproduced upstream).
const h = await harness({ remote: { ...remoteSession, socksProxy: '127.0.0.1:1080' }, hostUp: false });
const session = h.ctx.sessions.get(SESSION_ID)!;
for (const input of ['a', 'b', 'c']) expect((await send(h.app, { input, useMux: true })).statusCode).toBe(200);
expect(session.writeBuffer).toEqual(['a', 'b', 'c']);
expect(h.probe).not.toHaveBeenCalled();
expect(h.wake).not.toHaveBeenCalled();
expect(h.registry.pendingBytes(SESSION_ID)).toBe(0);
});
it('wakes before writing on the send-and-wait path (no buffering, the response waits anyway)', async () => {
const h = await harness({ hostUp: false });
const session = h.ctx.sessions.get(SESSION_ID)!;
@@ -188,6 +201,17 @@ describe('GET /api/sessions/:id/reachability', () => {
expect(body.data.reachable).toBe(false);
expect(body.data.wakeConfigured).toBe('none');
});
it('reports a proxied host as unknown, not unreachable, and never probes it', async () => {
// A jump-host / SOCKS host does not answer the bare TCP probe even while ssh works;
// `reachable:false` here drew a permanent banner over a healthy session.
const h = await harness({ remote: { ...remoteSession, jumpHost: 'bastion.example' }, hostUp: false });
const body = (await get(h.app, `/api/sessions/${SESSION_ID}/reachability`)).json();
expect(body.data.reachable).toBeNull();
expect(body.data.probeable).toBe(false);
expect(body.data.wakeConfigured).toBe('command');
expect(h.probe).not.toHaveBeenCalled();
});
});
describe('POST /api/sessions/:id/wake', () => {
+56
View File
@@ -0,0 +1,56 @@
/**
* @fileoverview Multi-user routing of the `remote:*` SSE family (server.ts `deriveSseHint`).
*
* The wake events carry `hostId`/`label`, which `GET /api/remote-hosts` withholds from
* non-admins, and their toast fires before any session check on the client — so an
* event that falls through to the global branch shows every logged-in user "Waking
* <label>" for a session they do not own. Constructs the server without starting it:
* the hint is a pure function of the event, the payload and the sessions map.
*/
import { describe, expect, it } from 'vitest';
import { WebServer } from '../src/web/server.js';
type Hint = { owner?: string; username?: string; adminOnly?: boolean; sessionScoped?: boolean } | undefined;
function hintFor(event: string, payload: Record<string, unknown>, owners: Record<string, string> = {}): Hint {
const server = new WebServer(3999, false, true) as unknown as {
sessions: Map<string, { owner?: string }>;
deriveSseHint(event: string, data: unknown): Hint;
};
for (const [id, owner] of Object.entries(owners)) server.sessions.set(id, { owner });
return server.deriveSseHint(event, payload);
}
describe('deriveSseHint — remote: events are session-scoped', () => {
it('routes a session wake to that session’s owner', () => {
expect(hintFor('remote:hostWaking', { sessionId: 's1', hostId: 'h', label: 'H' }, { s1: 'alice' })).toEqual({
owner: 'alice',
sessionScoped: true,
});
expect(hintFor('remote:sessionReconnected', { sessionId: 's1' }, { s1: 'alice' })).toEqual({
owner: 'alice',
sessionScoped: true,
});
});
it('routes a create/attach wake (no session yet) to the user who asked for it', () => {
expect(hintFor('remote:hostWaking', { forNewSession: true, username: 'bob', hostId: 'h', label: 'H' })).toEqual({
username: 'bob',
sessionScoped: true,
});
expect(hintFor('remote:hostWakeFailed', { forNewSession: true, username: 'bob', hostId: 'h' })).toEqual({
username: 'bob',
sessionScoped: true,
});
});
it('fails closed (admins only) when it names neither a session nor a requester', () => {
const hint = hintFor('remote:hostWaking', { forNewSession: true, hostId: 'h', label: 'H' });
expect(hint).toEqual({ owner: undefined, sessionScoped: true });
});
it('never lets a wake event reach the global branch', () => {
expect(hintFor('remote:hostWaking', {})).not.toBeUndefined();
expect(hintFor('remote:reconnectExhausted', {})).not.toBeUndefined();
});
});