By Ryan Richardson · Published 8 October 2026
Generating page variants and testing them is step 36 of the matrix (Part VI, Build the machine). The CSV marks it standard practice at moderate difficulty, automated tooling, which is unusual for this part of the book: most of the surrounding steps are marked often skipped.
The mistake isn't running a test. It's reading an underpowered one as a verdict. Sample size, not creativity, decides whether a page test means anything.
On a real rebuild, the version that read better pushed 48% of visitors past the hero, against 31.6% on the control. By reading alone, the rebuild won. On settled sales it lost: one purchase against three, on 75 and 79 sessions, a sample too thin to carry either number.
A clearer page is not the same claim as a page that sells. The two can point opposite directions in the same test, and a reader-quality metric (scroll depth, time on page) will never catch that a sale was lost further down the funnel.
At a 2% purchase rate, detecting a 20% relative lift needs roughly 19,600 visitors per arm, about 39,200 total. Detecting a 50% relative lift needs about 3,136 per arm, 6,272 total. Halving the effect size you want to see quadruples the sample required to see it.
That arithmetic is also why the smallest-sounding test (one button, one headline) is quietly the largest ask: small effects need the biggest samples. Most book-funnel tests never clear it, which is why a before-and-after read of settled cash, not a significance test, decided the real result above.
Peeking at results daily and stopping the moment a threshold is crossed lifts the real false-positive rate from the stated 5% to about 26%. Write the stopping rule, and the sample size it implies, before the first visitor arrives, not after the numbers start looking good.
Calling an underpowered split a result. Reading engagement (scroll, time) as a proxy for sales when the two have already been shown to disagree. Running the test during a week something else also changed, so neither question gets a clean answer.
| Claim | Value | Source |
|---|---|---|
| Visitors past the hero, rebuild vs control | 48% vs 31.6% | The Sixty Steps manuscript |
| Sales, rebuild vs control | 1 sale vs 3 sales, on 75 and 79 sessions | The Sixty Steps manuscript |
| Visitors per arm to detect a 20% relative lift at 2% baseline | ~19,600 per arm, ~39,200 total | Measured in Real Money, Field Manual |
| Visitors per arm to detect a 50% relative lift at 2% baseline | ~3,136 per arm, ~6,272 total | Measured in Real Money, Field Manual |
| False positive rate from daily peeking | 5% stated rises to about 26% | Measured in Real Money, Field Manual |
| Matrix status for this step | Moderate difficulty, standard practice, automated tooling | The Sixty Steps matrix |