RFC: Operator's Manual Generated from E2E Flows

Date: 2026-09-21

Status: Draft — unbuilt. The Maestro tooling it builds on merged in #4104; nothing in this RFC exists yet, including the manual itself.

Related: gui/apps/mobile/.maestro/README.md (Maestro conventions and gotchas), Onboarding QA (the nemo@architect.co identity), Local stack (dev_local), #4213 (Expo SDK 57 — required to run the mobile flows on macOS 27 / Xcode 27; implementation branches from it until it merges).

Author: (with Claude)

We have no operator's manual, and we have one smoke flow per platform. This RFC proposes building both as one artifact: a Maestro flow that walks a user journey is, with takeScreenshot steps and a caption per step, also the storyboard for that journey. The manual page is generated from the flow, so a reworded button breaks the test, and fixing the test regenerates the page. E2E coverage and documentation cannot drift apart because they are the same file.

Three things fall out of that one primitive: a versioned operator's manual, GUI-change screenshots on PRs (the git diff of the manual's images), and a realized storyboard of onboarding walked by the resettable QA user.

This is a design document. The Maestro behaviour it relies on (takeScreenshot, startRecording, runScript, addMedia, Maestro web on Chromium) is from its documentation and has not been exercised in this repo beyond the two login smoke flows. M0 exists to find out early.


Decision-first summary

D1 — The flow is the source of truth; the manual page is derived

A storyboard flow is an ordinary Maestro flow plus two conventions: every documented step carries a label: (the caption) and is followed by takeScreenshot: NN-slug. A generator reads the YAML and emits an MDX page — title and intro from the flow's header comment, one captioned image per labelled step. Nobody hand-edits the generated page.

Rejected: hand-written manual pages that embed generated screenshots. It halves the work saved and reintroduces drift — the prose describes a button the screenshot no longer shows.

D2 — The manual is indexed against dev_local, at the checkout's SHA

The manual's environment is the local docker stack: mock EP3, seeded fixtures, Clerk dev instance with the scriptable +clerk_test / 424242 login. Screenshots become a function of the git SHA. Demo and prod are rejected as the base: their data drifts, their deployed version is not your checkout, and real customer data would land in committed images.

This is a precondition for D3, not just a preference. Partial regeneration is only cheap if an unchanged screen re-renders the same.

D3 — "What changed" is decided on the output side, by perceptual diff

Predicting which flows a source diff affects means walking the React import graph from changed files to screens to flows; any shared component touches everything. We do not attempt it. Instead: run the flows, compare each PNG to the committed one with a perceptual diff, and overwrite only past a threshold. The comparison costs milliseconds; the cost is running the flows (est. 30–90 s per mobile flow, serial — one Maestro driver per simulator).

A flow may declare a coarse watches: glob list as a hint for which flows to run first on a feature branch. It is an optimization, never a correctness mechanism — the full regeneration (D5) is the backstop.

D4 — PR screenshots are the git diff of the images directory

Because unchanged screens do not churn (D2 + D3), the PNGs a PR changes are its GUI-change screenshots. GitHub renders image diffs in the files view, which sidesteps the absence of any API for attaching images to a PR. No separate PR-screenshot tooling is built. /ship and /cook gain one step: if the diff touches gui/, run just update-operators-manual for the relevant flows before committing.

This is local-first. Every workflow today runs on ubuntu-latest; iOS in CI needs macOS runners and an Expo build per PR and is out of scope. Web-only CI is plausible later (M5).

D5 — Partial regeneration per feature; full regeneration per version

Feature branches regenerate the flows they touch. Each advance-version is followed by a full regeneration that stamps VERSION and the SHA into the manual. The full pass normalizes differences between authors' machines, resolves binary merge conflicts from concurrent PRs, and catches screens a partial run missed.

It is a separate recipe run after the bump, landing as its own commit — not a step inside scripts/advance-version.py. A release must not depend on a booted simulator, a healthy docker stack, and twenty minutes of flows.

The manual tracks HEAD. We do not keep multiple versioned copies in Mintlify (each would duplicate every image); the manual for version N is the checkout of N's commit.

D6 — Onboarding gets two captures with different jobs

Onboarding is the one journey whose truth lives partly outside our code (Clerk, Plaid, Box, and on AIEX, Bitnomial). It gets two captures:

Mocked capture Real capture
Environment dev_local, externals mocked prod (or demo — see O1)
Question answered what does our code render? does the integrated system still work?
When pre-merge, every relevant PR per version, human-triggered
Role regression gate; canonical for steps we own post-deploy synthetic check; canonical for third-party-owned steps

The real capture is ex-post-facto by nature, and that is acceptable: the class of change only it can catch — a third party changing its UI, config drift, a broken linkage — is not caused by our PRs, so no pre-merge check could catch it. PR-caused changes are gated by the mocked capture.

D7 — Canonical source is chosen per step; convergence is a publish-time check

The published manual shows one storyboard for onboarding, not two parallel tracks. Steps we own come from the mocked capture. Steps a third party owns (the Plaid flow, the Clerk verification email) come from the real capture, because a mock of someone else's UI documents nothing.

Convergence is a check, not content: an internal report page puts the mocked and real captures of our-owned steps side by side, and --publish warns or refuses when they diverge. A divergence means either the mock has drifted from reality or prod is not running the code being documented — both worth knowing before a release. Dates and ids differ across environments, so this is a loose perceptual threshold plus a human looking at the report, not a pixel gate.

D8 — Mock at the narrowest seam; never fake a third party's screen

Frontend: the usePlaid hook. Backend: the plaid.rs and box_client.rs clients in onboarding-gateway. The mocked storyboard shows "hands off to Plaid" and resumes afterwards.

D9 — The real capture is human-triggered, not unattended

The QA-user reset is gated by Clerk MFA step-up, and the verification code goes to a real mailbox. Fully automating the run would mean adding a service-credential path to a production endpoint that deletes a user. We do not do that. A person runs the capture once per version: they reset and enter the code; Maestro drives and captures everything else. On AIEX the real capture stops before submit, per the existing warning in the Onboarding QA doc about Bitnomial's review queue. Real captures are web-only (Maestro web pointed at the environment's URL); mobile would need a prod build on the simulator.


Design

1. Storyboard flow conventions

# Title: Signing in
# Intro: How a returning trader reaches the markets screen.
# watches: gui/packages/app/screens/auth/**
appId: ${APP_ID}
---
- launchApp: { clearState: true }
- tapOn:
    text: '(?i).*email.*'
    label: Enter the email address you registered with
- takeScreenshot: 01-email

2. Determinism

Source of churn Mitigation
Status bar clock, battery, signal xcrun simctl status_bar override in the recipe
Simulator model / OS, web viewport pinned in the recipe, not left to the author's machine
Mock market data, timestamps seeded fixtures; screenshot only settled states; mask regions that cannot be fixed
Animations, caret blink waitForAnimationToEnd before each takeScreenshot
Residual sub-pixel noise the perceptual threshold in D3

3. Recipes

just update-operators-manual [flows…]   # named flows, or all; overwrite PNGs past threshold; regenerate MDX
just update-operators-manual --publish  # full regeneration + convergence check + VERSION/SHA stamp
just capture-nemo <env>                 # the human-triggered real onboarding capture (D9)

Output lands in docs/internal/operations/manual/ (MDX) and docs/internal/images/manual/<flow>/NN-slug.png, following the precedent of just record-tapes → docs/internal/images/tapes/. A walk-through video is the same flow under maestro record --local and is produced only on --publish.

4. The onboarding storyboard

Reset → sign up → questionnaire → documents → submit → admin review and approval → reset. The trader half is a trader-app flow; the approval half is a Maestro web flow against the admin GUI on the same stack. The full onboarding UI lives in apps/web; mobile has only onboarding/index.tsx and assessment.tsx, so the web flow is primary.

Known obstacles:


Milestones

Open questions

Non-goals