fix(remote): a proxied host is reachability-unknown; scope remote: SSE per session

Review round 2 on #439.

1. The bare TCP probe connects to host:port, which a host behind a jump host
   or SOCKS proxy does not answer even while ssh works. Acting on that
   verdict drew a permanent banner over a healthy session, replaced a real
   "needs tmux" error with "not reachable" in quick-start, and - with a wake
   target - buffered every HTTP input for the life of the session, since the
   readiness poll could never succeed. `WakeableRemote` now carries
   `jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such
   a host into reachability-UNKNOWN: input is delivered, `checkReachable` /
   `checkHostReachable` answer `null` (never `false`), `ensureHostAwake`
   returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate
   fires on `=== false` only, and `GET …/reachability` reports
   `reachable: null, probeable: false` so the banner has nothing to key on.
   A wake target can still be fired for it, blind: no readiness poll, no
   reattach, no toast - the response says only whether the packet went out.

2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake
   has no session yet, so the registry names the requesting user
   (`ensureHostAwake({ requestedBy })` -> `username` in the payload) and
   `deriveSseHint` routes on it; with neither it fails closed to admins.
   Single-user mode is unaffected.

Smaller, from the same review:

- A flush write that fails now drops the remaining buffer (logged) instead
  of retaining it: the wake still resolved and marked the host reachable, so
  the retained chunk waited for the NEXT wake and was replayed hours later,
  after everything typed since. Same policy as the oversized paste.
- The banner polls on tab activation (a user action) and on its 30 s timer
  only for a host with a wake target; a timer connecting to a host Codeman
  cannot wake is the traffic invariant #2 rejects keepalives for. A proxied
  host is never polled.
- `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP
  socket refuse under VITEST, as remote-files.ts does. The guard caught a
  leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the
  probe but still polled readiness with the real one, so the shutdown test
  had been connecting to a production address. The poll now uses the
  injected probe.
- docs/remote-sessions.md is additions only again (the reformatting is
  gone); the architecture-invariants overlap resolved itself in the merge.

Live, against a throwaway instance with a non-routable ghost host: proxied
-> no probe, no wake, the genuine ssh error after 10 s; direct (control) ->
probe, magic packet, "did not come back" after the 40 s budget.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
This commit is contained in:
Randalix
2026-09-18 22:46:11 +02:00
co-authored by Claude Opus 5
parent e271a65e79
commit 1040f6c489
12 changed files with 609 additions and 83 deletions
+20 -3
View File
@@ -72,6 +72,7 @@ import {
RemoteWakeRegistry,
REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
createDefaultRemoteWakeDeps,
isProbeable,
type WakeableRemote,
} from '../../remote-wake.js';
import { clampWaitMs, MAX_BUFFER_SCAN_BYTES } from '../../config/agent-wait.js';
@@ -760,7 +761,8 @@ export function resolveOmpConfigForCreate(
/**
* `RemoteHost` → the wake registry's host shape. They differ in one field name only
* (`id` in host config vs `hostId` on a session's `remote`), but the rename is load-
* bearing: the registry keys its per-host wake state on `hostId`.
* bearing: the registry keys its per-host wake state on `hostId`. The proxy fields
* travel too: they are what tells the registry its probe cannot reach this host.
*/
function wakeableHost(host: RemoteHost): WakeableRemote {
return {
@@ -770,6 +772,9 @@ function wakeableHost(host: RemoteHost): WakeableRemote {
port: host.port,
wakeMac: host.wakeMac,
wakeCommand: host.wakeCommand,
jumpHost: host.jumpHost,
socksProxy: host.socksProxy,
extraSshOptions: host.extraSshOptions,
};
}
@@ -892,6 +897,8 @@ export function registerSessionRoutes(
// answers with an ssh failure that blames anything but the machine being asleep.
const hostWake = await remoteWake.ensureHostAwake(wakeableHost(host), {
timeoutMs: REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
// No session yet, so the wake events name their requester (multi-user routing).
requestedBy: ownerFor(req),
});
if (hostWake === 'failed') {
return createErrorResponse(
@@ -1535,13 +1542,18 @@ export function registerSessionRoutes(
const { id } = req.params as { id: string };
const session = findSessionOrFail(ctx, id, req);
const remote = session.remote;
if (!remote) return { success: true, data: { reachable: true, wakeConfigured: 'none' as const } };
if (!remote) {
return { success: true, data: { reachable: true, probeable: true, wakeConfigured: 'none' as const } };
}
const force = (req.query as { force?: string })?.force === '1';
// `reachable: null` + `probeable: false` for a host behind a jump host / SOCKS proxy:
// the probe cannot reach it, so the UI shows no banner and stops polling.
const reachable = await remoteWake.checkReachable(session, { force });
return {
success: true,
data: {
reachable,
probeable: isProbeable(remote),
wakeConfigured: await remoteWake.wakeConfigured(session),
host: remote.host,
label: remote.label,
@@ -3264,6 +3276,8 @@ export function registerSessionRoutes(
// host on every schedule (the failure invariant #1 exists to prevent).
const hostWake = await remoteWake.ensureHostAwake(wakeableHost(host), {
timeoutMs: REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
// No session yet, so the wake events name their requester (multi-user routing).
requestedBy: ownerFor(req),
});
if (hostWake === 'failed') {
return createErrorResponse(
@@ -3280,7 +3294,10 @@ export function registerSessionRoutes(
// An unreachable host and a host without tmux fail the same way over ssh, so the
// probe's own message would send the user hunting for a tmux install. Ask the
// registry (which just probed, when it woke the host) which of the two it is.
if (!(await remoteWake.checkHostReachable(wakeableHost(host)))) {
// `=== false` on purpose: a proxied host answers `null` (the probe cannot reach
// it), and an unknown verdict must not replace the real ssh error with
// "not reachable" over a host that is fine.
if ((await remoteWake.checkHostReachable(wakeableHost(host))) === false) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
hostWake === 'no-target'