By Ryan Richardson · Published 8 October 2026
A small page test, a button colour, a headline tweak, runs for a week or two on ordinary traffic and never reaches a clear result either way.
Required sample size moves on an inverse square relationship with the effect you're trying to detect. A small, subtle change (a 20% relative lift) needs roughly four times the sample a larger, more obvious change (a 50% relative lift) would need to detect at the same confidence. At a 2% baseline conversion rate, detecting a 20% relative lift needs about 19,600 visitors per arm, or about 39,200 total. Detecting a 50% relative lift at the same baseline needs about 3,136 per arm, or about 6,272 total. The test that sounds smallest, a tweak to one button, is often the largest possible ask of your traffic, precisely because small changes produce small, hard-to-detect effects.
When a test produces zero purchases in a given number of sessions, the honest ninety-five percent upper bound on the true conversion rate is roughly three divided by that session count. That's usually enough on its own to kill a bad idea without running the test any further.
Calculate the required sample size before launching a test, not after reading a disappointing result. If your traffic can't reach the sample size the arithmetic calls for inside a reasonable window, either test a bigger change, one expected to produce a bigger effect, or treat the test as directional only rather than a formal read.
Before trusting any A/B test result, check the actual session count per arm against what the inverse-square formula says was needed for the effect size you were hoping to detect. A test that ran with a fraction of the required sample was never capable of returning a clear answer, whatever the result looked like.
| Claim | Value | Source |
|---|---|---|
| Visitors per arm formula | 16 x (1 - p) / (p x d x d), where p is the baseline conversion rate and d is the relative lift to detect | Measured in Real Money, Field Manual |
| Visitors per arm to detect a 20% relative lift at a 2% baseline purchase rate | about 19,600 | Measured in Real Money, Field Manual |
| Visitors per arm to detect a 50% relative lift at a 2% baseline purchase rate | about 3,136 | Measured in Real Money, Field Manual |
| Both arms combined, to detect a 20% relative lift at a 2% baseline purchase rate | about 39,200 | Measured in Real Money, Field Manual |
| Both arms combined, to detect a 50% relative lift at a 2% baseline purchase rate | about 6,272 | Measured in Real Money, Field Manual |
| Effect of halving the effect size you want to detect | quadruples the required sample | Measured in Real Money, Field Manual |
| Ninety-five percent upper bound on the true rate when a test shows zero purchases in n sessions | about 3 divided by n | Measured in Real Money, Field Manual |