Files
Codeman/test/mobile-overview.test.ts
DevvynandClaude Opus 5.5 0a52a99ca9 feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis

Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI

Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).

Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.

Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).

Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.

Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.

Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.

27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else

window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.

Fixed in two places:

- server.ts: after building `available`, intersect the nine real
  SessionMode ids against `enabledClis()`. git/cloudflared (utility
  binaries, not CLI registry entries) and deepseekBinary (a secondary
  installed-only flag for the "add a profile" affordance) are deliberately
  left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
  `window.__codemanCliAvailable` in place and refreshes the welcome screen,
  the mobile overview and an already-open Run menu, mirroring the existing
  `installDeepSeekProfile()` pattern for the same "injected once, needs an
  explicit patch" reason — without this half, the server-side fix alone
  still left every surface stale until the next reload.

New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.

Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap

Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:

- A new "Status as of 2026-09-22" section up top: what's actually
  implemented (verified by grepping the routes/schema/UI, not just trusting
  the commit messages), the availability-flag staleness bug found and fixed
  in this session (commit 0c77dd0a) with its devbox verification record, and
  one real outstanding gap.

- The outstanding gap: a custom CLI created via Phase 5's write API has no
  way to actually be launched. The Run menu is static per-mode markup with
  no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
  custom entry, confirm it can be launched" verify step was never actually
  exercised against this. Documented with two candidate fixes, neither
  started.

- Each phase's checklist flipped to [x] where confirmed present in the tree,
  Status lines updated from TODO to DONE, and the two originally-open
  questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
  marked resolved against what actually shipped.

No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* docs: add the CLI-registry deployment plan and the parked Copilot plan

Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:

- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
  series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
  is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
  name in docs/cli-enable-disable-plan.md's own header as a sibling plan
  tracked separately — kept for continuity, not active on this branch.

The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* fix(cli-registry): render enabled CLIs in launch surfaces

* test(cli-registry): update frontend branch guard

* fix(test): isolate suite from deployment environment

* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent

shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.

Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): hide shell's toggle entirely instead of greying it out

A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical

renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* style: prettier fixes from the master merge

* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels

Four gaps found verifying #476 against the #343 review trail:

- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
  (the nine per-CLI resolvers and the generic registry one) caches in its own
  closure, with a negative-cache backoff of up to 5 minutes, and nothing
  cleared them. invalidateCliExecutableResolvers(binaries) now drops those
  caches per binary; install (success or failure), create, edit and delete
  call it plus invalidateCliResolverCache(id). Before this, a CLI installed
  from Settings could fail to launch for minutes, and an edited custom entry
  kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
  ignoring the entry's searchDirs and the login-shell lookup that spawn and
  the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
  sit behind an explicit confirm; the confirm now names the exact command,
  which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
  of the word ("Claude", "Codex"). It uses the registry label again, which is
  identical to the old static table for every stock CLI (now pinned).

14 new tests; 9 of them fail against the previous head and pass here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs

Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
  does not parse or has group/world permission bits instead of overwriting it
  (isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
  existence/duplicate checks inside the serialized step, plus a unique tmp
  name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
  section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md

Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
  body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
  map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
  label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
  "Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* test(cli-registry): pin the #343/#347 findings #476 makes reachable

A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 01:48:26 +02:00

594 lines
25 KiB
TypeScript

// Port: none (pure model + static markup assertions — no browser, no server).
//
// The phone home screen (src/web/public/mobile-overview.js) replaces the welcome
// overlay under 600px. Its grouping logic is the part that can silently go wrong:
// a session blocked on a permission prompt landing in "idle" is exactly the bug
// this surface exists to prevent. buildMobileOverviewModel() is pure for that
// reason, so it can be exercised here against plain objects.
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it } from 'vitest';
import { STOCK_CLIS } from '../src/config/cli-registry/stock.js';
const PUBLIC = resolve(import.meta.dirname, '../src/web/public');
/** Minimal fake DOM node — enough surface for mobile-overview.js's programmatic builders. */
function fakeElement(): any {
const el: any = {
className: '',
type: '',
dataset: {},
style: {},
attrs: {} as Record<string, string>,
children: [] as any[],
setAttribute(name: string, value: string) {
el.attrs[name] = value;
},
appendChild(child: any) {
el.children.push(child);
return child;
},
};
return el;
}
function loadOverviewApp(overrides: Record<string, any> = {}) {
const CodemanApp = function CodemanApp(this: any) {};
const context = vm.createContext({
CodemanApp,
console,
window: {},
document: {
getElementById: () => null,
createElement: () => fakeElement(),
createElementNS: () => fakeElement(),
},
MobileDetection: { getDeviceType: () => 'mobile' },
});
// constants.js first: it installs the row comparator (window.CodemanSessionOrder)
// that buildMobileOverviewModel() sorts every section with, shared with the
// desktop rail so the two home screens cannot order the same list differently.
for (const file of ['constants.js', 'mobile-overview.js']) {
vm.runInContext(readFileSync(resolve(PUBLIC, file), 'utf8'), context, { filename: file });
}
const app = new (CodemanApp as any)();
app.getSessionName = (session: any) => session.name || session.workingDir?.split('/').pop() || session.id.slice(0, 8);
app._shortenHomePath = (p: string) => (p || '').replace(/^\/home\/[^/]+\//, '~/');
app.loadAppSettingsFromStorage = () => ({});
Object.assign(app, overrides);
return app;
}
const CASES = [
{ name: 'claudeman', path: '/home/arkon/default/claudeman', location: 'local' },
{ name: 'beta', path: '/home/arkon/codeman-cases/beta', location: 'local' },
{ name: 'boxed', path: '/srv/boxed', location: 'docker' },
];
function session(over: Record<string, any>) {
return { id: 'x', status: 'idle', mode: 'claude', workingDir: '/home/arkon/default/claudeman', ...over };
}
describe('mobile overview model', () => {
it('routes a session with a pending permission prompt into NEEDS YOU, not idle', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [session({ id: 'a', status: 'idle' })],
cases: CASES,
pendingHooks: new Map([['a', new Set(['permission_prompt'])]]),
});
expect(model.needsYou.map((r: any) => r.id)).toEqual(['a']);
expect(model.current).toHaveLength(0);
expect(model.needsYou[0].state).toBe('needs');
expect(model.needsYou[0].pill).toBe('needs you');
});
it('ranks an action hook above an idle hook above a stale busy status', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
// An idle_prompt hook on a session the server still calls 'busy': the hook
// is the newer signal, so it must win.
sessions: [
session({ id: 'busy-with-idle-hook', status: 'busy' }),
session({ id: 'elicit', status: 'busy' }),
session({ id: 'plain-busy', status: 'busy' }),
],
cases: CASES,
pendingHooks: new Map([
['busy-with-idle-hook', new Set(['idle_prompt'])],
['elicit', new Set(['elicitation_dialog'])],
]),
});
expect(model.needsYou.map((r: any) => r.id)).toEqual(['elicit', 'busy-with-idle-hook']);
expect(model.current.map((r: any) => r.id)).toEqual(['plain-busy']);
});
it('buckets busy / idle / stopped / error and labels each pill', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [
session({ id: 'w', status: 'busy' }),
session({ id: 'i', status: 'idle' }),
session({ id: 'd', status: 'stopped' }),
session({ id: 'e', status: 'error' }),
],
cases: CASES,
});
// Everything that is not blocked on you shares one "current" section,
// most demanding first.
expect(model.current.map((r: any) => [r.id, r.pill])).toEqual([
['w', 'working'],
['i', 'idle'],
['d', 'done'],
]);
expect(model.needsYou.map((r: any) => r.pill)).toEqual(['error']);
expect(model.sessionCount).toBe(4);
});
it('keeps the user tab order as the tiebreak when nothing is stamped', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [session({ id: 'first' }), session({ id: 'second' }), session({ id: 'third' })],
cases: CASES,
sessionOrder: ['third', 'first', 'second'],
});
expect(model.current.map((r: any) => r.id)).toEqual(['third', 'first', 'second']);
});
it('sorts running sessions longest-turn-first and quiet ones most-recent-first', () => {
// A working pane repaints about once a second, so its last-activity stamp
// is always "now": the running group has to key off the pane's last Enter
// instead, or every turn ranks as freshly started.
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [
session({ id: 'quiet-old', status: 'idle', lastActivityAt: 2_000 }),
session({ id: 'turn-young', status: 'busy', lastSubmitAt: 9_000, lastActivityAt: 10_000 }),
session({ id: 'quiet-new', status: 'idle', lastActivityAt: 8_000 }),
session({ id: 'turn-old', status: 'busy', lastSubmitAt: 1_000, lastActivityAt: 10_000 }),
],
cases: CASES,
sessionOrder: ['quiet-old', 'turn-young', 'quiet-new', 'turn-old'],
});
expect(model.current.map((r: any) => r.id)).toEqual(['turn-old', 'turn-young', 'quiet-new', 'quiet-old']);
});
it('puts the longest-blocked session at the top of NEEDS YOU', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [
session({ id: 'just-asked', lastActivityAt: 9_000 }),
session({ id: 'starving', lastActivityAt: 1_000 }),
],
cases: CASES,
pendingHooks: new Map([
['just-asked', new Set(['permission_prompt'])],
['starving', new Set(['permission_prompt'])],
]),
});
expect(model.needsYou.map((r: any) => r.id)).toEqual(['starving', 'just-asked']);
});
it('matches a session started in a subdirectory to its case (longest prefix)', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [
session({ id: 'sub', workingDir: '/home/arkon/default/claudeman/src/web' }),
session({ id: 'outside', workingDir: '/tmp/scratch' }),
],
cases: [...CASES, { name: 'claudeman-web', path: '/home/arkon/default/claudeman/src/web' }],
});
const rows = Object.fromEntries(model.current.map((r: any) => [r.id, r.caseName]));
expect(rows.sub).toBe('claudeman-web');
expect(rows.outside).toBe('');
});
it('lists past conversations newest first and never repeats a live session', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [session({ id: 'live-1' })],
cases: CASES,
history: [
// Same id as the running session: the unified list includes live rows,
// and showing one in both sections would be a duplicate.
{ sessionId: 'live-1', workingDir: '/home/arkon/default/claudeman', lastActivityAt: 500 },
{
sessionId: 'old-a',
workingDir: '/home/arkon/codeman-cases/beta',
firstPrompt: 'fix the mobile header',
claudeSessionId: 'claude-uuid-a',
lastActivityAt: 100,
},
{
sessionId: 'old-b',
workingDir: '/home/arkon/default/claudeman',
name: 'w4-claudeman',
lastActivityAt: 400,
},
],
});
expect(model.past.map((r: any) => r.id)).toEqual(['old-b', 'old-a']);
expect(model.past[1]).toMatchObject({
title: 'fix the mobile header',
caseName: 'beta',
claudeSessionId: 'claude-uuid-a',
workingDir: '/home/arkon/codeman-cases/beta',
});
// A row with no prompt falls back to its name, so it is never a bare UUID.
expect(model.past[0].title).toBe('w4-claudeman');
});
it('carries resumeId through to the past row so a Codex tap resumes its thread', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [],
cases: CASES,
history: [
{
sessionId: 'rollout-1',
workingDir: '/home/arkon/codeman-cases/beta',
firstPrompt: 'port the parser',
mode: 'codex',
resumeId: 'codex-thread-id',
lastActivityAt: 300,
},
// A claude row carries none, and must not grow one.
{
sessionId: 'claude-1',
workingDir: '/home/arkon/default/claudeman',
claudeSessionId: 'claude-uuid-a',
lastActivityAt: 200,
},
],
});
// resumeMobileOverviewSession() reads row.resumeId off exactly this projection and
// hands it to resumeHistorySession(); an undefined here is a fresh codex session on
// a thread that already exists, which is the phone-only half of the resume feature.
expect(model.past[0]).toMatchObject({ mode: 'codex', resumeId: 'codex-thread-id' });
expect(model.past[1].resumeId).toBeUndefined();
});
it('does not title a past row with the transcript reader placeholder', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [],
cases: CASES,
history: [
{ sessionId: 'blank', workingDir: '/home/arkon/default/claudeman', firstPrompt: '(no content)' },
{ sessionId: 'spaces', workingDir: '/home/arkon/codeman-cases/beta', firstPrompt: ' ' },
],
});
expect(model.past.map((r: any) => r.title)).toEqual(['claudeman', 'beta']);
});
it('accepts the live Map as-is and survives an empty state', () => {
const app = loadOverviewApp();
const fromMap = app.buildMobileOverviewModel({
sessions: new Map([['a', session({ id: 'a' })]]),
cases: CASES,
});
expect(fromMap.current.map((r: any) => r.id)).toEqual(['a']);
const empty = app.buildMobileOverviewModel({});
expect(empty).toMatchObject({ needsYou: [], current: [], past: [], sessionCount: 0 });
});
it('anchors the "how long" stamp on last activity, and on the last Enter while working', () => {
const app = loadOverviewApp();
const now = Date.now();
const model = app.buildMobileOverviewModel({
sessions: [
// A working pane repaints about once a second, so lastActivityAt is
// always "now" and would report every running turn as 0m. The turn's
// own start is the last Enter.
session({
id: 'w',
status: 'busy',
createdAt: now - 7200_000,
lastActivityAt: now,
lastSubmitAt: now - 300_000,
}),
// A quiet pane prints nothing, so its last byte IS when it went idle.
session({ id: 'i', status: 'idle', createdAt: now - 7200_000, lastActivityAt: now - 900_000 }),
],
cases: CASES,
});
const rows = Object.fromEntries(model.current.map((r: any) => [r.id, r]));
expect(rows.w.since).toEqual({ key: 'working', at: now - 300_000 });
expect(rows.i.since).toEqual({ key: 'idle', at: now - 900_000 });
expect(rows.i.createdAt).toBe(now - 7200_000);
});
it('falls back to the sort anchor for a working row with no submit stamp', () => {
// A session that has never submitted has no turn start to measure from, but
// `sessionActivityAnchor` still RANKS it by lastActivityAt. The stamp must
// show that same number rather than nothing: a row sorted by a value it
// does not display reads as randomly placed.
const now = Date.now();
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [session({ id: 'w', status: 'busy', lastActivityAt: now })],
cases: CASES,
});
expect(model.current[0].since).toEqual({ key: 'working', at: now });
expect(model.current[0].createdAt).toBe(0);
});
it('still leaves the stamp off when there is no anchor at all', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [session({ id: 'w', status: 'busy' })],
cases: CASES,
});
expect(model.current[0].since).toBeNull();
});
it('formats a moment as "ago" and a span as a bare duration', () => {
const app = loadOverviewApp();
app.formatRelativeTime = () => '3d ago';
const now = Date.now();
expect(app._mobileOverviewStampText(now - 86_400_000, 'ago')).toBe('3d ago');
expect(app._mobileOverviewStampText(now - 20_000, 'for')).toBe('<1m');
expect(app._mobileOverviewStampText(now - 12 * 60_000, 'for')).toBe('12m');
expect(app._mobileOverviewStampText(now - 125 * 60_000, 'for')).toBe('2h 5m');
expect(app._mobileOverviewStampText(now - 3 * 3600_000, 'for')).toBe('3h');
expect(app._mobileOverviewStampText(now - 50 * 3600_000, 'for')).toBe('2d 2h');
// No anchor renders as a dash, never as "56 years ago" off epoch 0.
expect(app._mobileOverviewStampText(0, 'for')).toBe('—');
expect(app._mobileOverviewStampText(0, 'ago')).toBe('—');
});
it('no longer builds a spaces section', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({ sessions: [session({ id: 'a' })], cases: CASES });
expect(model.spaces).toBeUndefined();
});
});
describe('mobile overview gate', () => {
it('is phone-width only, off in solo windows, and off when explicitly disabled', () => {
expect(loadOverviewApp().shouldUseMobileOverview()).toBe(true);
expect(loadOverviewApp({ isSoloWindow: true }).shouldUseMobileOverview()).toBe(false);
expect(
loadOverviewApp({
loadAppSettingsFromStorage: () => ({ mobileOverviewEnabled: false }),
}).shouldUseMobileOverview()
).toBe(false);
// An unset value must read as ON: phones that already have saved settings
// from before this feature existed have no key for it.
expect(loadOverviewApp({ loadAppSettingsFromStorage: () => ({ skin: 'og' }) }).shouldUseMobileOverview()).toBe(
true
);
});
});
describe('mobile overview wiring', () => {
const html = readFileSync(resolve(PUBLIC, 'index.html'), 'utf8');
const mobileCss = readFileSync(resolve(PUBLIC, 'mobile.css'), 'utf8');
const moduleSrc = readFileSync(resolve(PUBLIC, 'mobile-overview.js'), 'utf8');
it('speaks the same status language as the session tabs', () => {
// A session that is fine reads green on the tabs; anything else here would
// mean two meanings for one color on the same screen.
expect(mobileCss).toMatch(/\.mobile-overview-dot--idle\s*\{\s*background:\s*var\(--green\)/);
expect(mobileCss).toMatch(/\.mobile-overview-dot--working\s*\{[^}]*var\(--green\)[^}]*animation:\s*pulse/);
// Waiting-for-input blinks yellow, asked-a-question blinks red, same as
// tab-alert-idle / tab-alert-action.
expect(mobileCss).toMatch(/\.mobile-overview-row--waiting\s*\{[^}]*animation:\s*mobile-overview-blink-yellow/);
expect(mobileCss).toMatch(/\.mobile-overview-row--needs\s*\{[^}]*animation:\s*mobile-overview-blink-red/);
expect(mobileCss).toContain('@keyframes mobile-overview-blink-red');
expect(mobileCss).toContain('@keyframes mobile-overview-blink-yellow');
// The alert must survive reduced-motion as a held color, not vanish.
expect(mobileCss).toMatch(/prefers-reduced-motion[^}]*\}[\s\S]*?\.mobile-overview-row--needs/);
});
it('reuses the toolbar Run button classes instead of its own palette', () => {
// The per-backend gradient lives in styles.css keyed on
// `.btn-toolbar.btn-run.mode-<backend>` (and light skins override exactly
// those); carrying the same classes keeps both Run buttons identical.
expect(moduleSrc).toContain('btn-toolbar btn-run mode-');
expect(moduleSrc).toContain('btn-toolbar btn-run-gear mode-');
// The two button rules (not the dropdown below them) must set no color at
// all, or they would win over the mode gradient.
const buttonRules = mobileCss.match(/\.mobile-overview-run(-caret)?\s*\{[^}]*\}/g) || [];
expect(buttonRules.length).toBe(2);
for (const rule of buttonRules) {
expect(rule).not.toMatch(/\b(background|color)\s*:/);
}
});
it('ships the container hidden and loads the module', () => {
expect(html).toMatch(/<div class="mobile-overview" id="mobileOverview" hidden><\/div>/);
expect(html).toContain('<script defer src="mobile-overview.js"></script>');
});
it('never gives .mobile-overview a bare display rule', () => {
// Desktop does not load mobile.css at all, so the [hidden] attribute is the
// only thing keeping the overview off desktop. A bare
// `.mobile-overview { display: … }` rule would beat the UA [hidden] rule.
const bareDisplay = /\.mobile-overview\s*\{[^}]*display\s*:/;
expect(bareDisplay.test(mobileCss)).toBe(false);
expect(mobileCss).toContain('.mobile-overview.visible {');
});
it('styles the overview from skin tokens rather than hardcoded colors', () => {
// Skins re-point the :root tokens, so a hex literal here is a rule that
// silently stays dark on the four light skins.
const rules = mobileCss.match(/\.mobile-overview[^{}]*\{[^}]*\}/g) || [];
expect(rules.length).toBeGreaterThan(10);
const hardcoded = rules.flatMap((rule) => rule.match(/:\s*#[0-9a-f]{3,8}\b/gi) || []);
expect(hardcoded).toEqual([]);
});
});
describe('mobile overview run picker (CLI availability gating)', () => {
function modeButtons(menu: any): string[] {
return menu.children.filter((c: any) => c.dataset.moAction === 'run-mode').map((c: any) => c.dataset.moMode);
}
// #201 gated the toolbar's #runModeMenu on isCliAvailable(); this phone-only
// picker (MOBILE_OVERVIEW_RUN_MODES / _buildMobileOverviewRunMenu) is a
// separate, hardcoded duplicate of that menu rather than a shared render, so
// it silently offered every backend regardless of what the server reported.
it('hides run modes the server reports as unavailable, keeps shell always', () => {
const app = loadOverviewApp({
runMode: 'claude',
isCliAvailable: (tool: string) => tool === 'claude',
});
const menu = app._buildMobileOverviewRunMenu();
expect(modeButtons(menu)).toEqual(['claude', 'shell']);
});
it('shows every mode when every CLI is available', () => {
const app = loadOverviewApp({
runMode: 'claude',
isCliAvailable: () => true,
});
const menu = app._buildMobileOverviewRunMenu();
expect(modeButtons(menu)).toEqual([
'claude',
'opencode',
'codex',
'gemini',
'antigravity',
'pi',
'grok',
'deepseek',
'omp',
'shell',
]);
});
it('gates every mode the picker actually offers', () => {
// Catches a new backend being added to MOBILE_OVERVIEW_RUN_MODES without
// being gated — the same class of bug that let this list drift from the
// toolbar menu's gating in the first place.
const src = readFileSync(resolve(PUBLIC, 'mobile-overview.js'), 'utf8');
const modesBlock = src.slice(
src.indexOf('const MOBILE_OVERVIEW_RUN_MODES'),
src.indexOf('];', src.indexOf('const MOBILE_OVERVIEW_RUN_MODES')) + 2
);
const offered = [...modesBlock.matchAll(/mode: '([^']+)'/g)].map((m) => m[1]);
expect(offered).toContain('omp');
const fn = src.slice(src.indexOf('_buildMobileOverviewRunMenu() {'));
const gate = fn.slice(0, fn.indexOf('const header'));
expect(gate).toContain('isCliAvailable');
});
});
describe('mobile overview watching badge', () => {
it('carries what the pane says is running in the background', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [session({ id: 'a', status: 'idle', watching: '1 monitor' }), session({ id: 'b', status: 'idle' })],
cases: CASES,
});
const rows = Object.fromEntries(model.current.map((r: any) => [r.id, r.watching]));
expect(rows).toEqual({ a: '1 monitor', b: '' });
});
it('leaves a session that is ALSO blocked on a dialog in NEEDS YOU', () => {
// An agent can arm a monitor and ask the user a question in the same breath, so the
// badge adds a fact to the row and never moves it out of the group that says a human
// is needed. Only the row's own pill decides that.
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [session({ id: 'a', status: 'idle', watching: '2 shells' })],
cases: CASES,
pendingHooks: new Map([['a', new Set(['permission_prompt'])]]),
});
expect(model.needsYou.map((r: any) => r.id)).toEqual(['a']);
expect(model.needsYou[0].pill).toBe('needs you');
expect(model.needsYou[0].watching).toBe('2 shells');
});
it('says one word and puts the detail where every surface can reach it', () => {
const app = loadOverviewApp();
const badge = app._buildWatchingBadge('1 monitor', 'mobile-overview-pill');
expect(badge.className).toBe('mobile-overview-pill mobile-overview-pill--watching');
expect(badge.textContent).toBe('watching');
expect(badge.title).toBe('Still running in the background: 1 monitor');
// A phone has no hover target and a screen reader reads neither the class nor the
// tooltip, so the label has to be here too or this surface says only "watching".
expect(badge.attrs['aria-label']).toBe('Still running in the background: 1 monitor');
});
it('takes the pill class of whichever surface asks for it', () => {
// The phone's pill styles live inside a media query the desktop rail never enters,
// so the rail passes its own base class and gets the same badge in its own clothes.
const app = loadOverviewApp();
expect(app._buildWatchingBadge('1 shell', 'home-sessions-pill').className).toBe(
'home-sessions-pill home-sessions-pill--watching'
);
});
});
describe('mobile overview Run picker, driven by the registry catalogue', () => {
function loadRunModes(catalog: unknown[] | undefined) {
const context = vm.createContext({
CodemanApp: function CodemanApp() {},
console,
window: catalog ? { __codemanCliCatalog: catalog } : {},
document: {
getElementById: () => null,
createElement: () => fakeElement(),
createElementNS: () => fakeElement(),
},
MobileDetection: { getDeviceType: () => 'mobile' },
});
for (const file of ['constants.js', 'mobile-overview.js']) {
vm.runInContext(readFileSync(resolve(PUBLIC, file), 'utf8'), context, { filename: file });
}
return {
modes: JSON.parse(JSON.stringify(vm.runInContext('mobileOverviewRunModes()', context))),
staticTable: JSON.parse(JSON.stringify(vm.runInContext('MOBILE_OVERVIEW_RUN_MODES', context))),
};
}
it('keeps the Run button on its word label, never the two-letter tab badge', () => {
const { modes } = loadRunModes([
{ id: 'claude', label: 'Claude Code', shortBadge: 'CC', kind: 'agent', enabled: true },
{ id: 'codex', label: 'Codex', shortBadge: 'CX', kind: 'agent', enabled: true },
{ id: 'grok', label: 'Grok', shortBadge: 'GK', kind: 'agent', enabled: false },
{ id: 'shell', label: 'Shell', shortBadge: 'SH', kind: 'shell', enabled: true },
]);
expect(modes).toEqual([
{ mode: 'claude', label: 'Claude Code', short: 'Claude Code' },
{ mode: 'codex', label: 'Codex', short: 'Codex' },
{ mode: 'shell', label: 'Terminal / Shell', short: 'Shell' },
]);
});
it('renders every stock CLI exactly as the static fallback table did', () => {
// The catalogue-driven path must be byte-identical to the pre-registry table for the
// shipped CLIs; only a CLI the table never knew (a custom entry) may differ.
const catalog = STOCK_CLIS.map((e) => ({
id: e.id,
label: e.label,
shortBadge: e.shortBadge,
kind: e.kind,
enabled: true,
}));
const { modes, staticTable } = loadRunModes(catalog);
const byMode = (list: Array<{ mode: string }>) => [...list].sort((a, b) => a.mode.localeCompare(b.mode));
expect(byMode(modes)).toEqual(byMode(staticTable));
});
});