sixtysteps.co

Why your smallest-sounding test needs the most traffic

By Ryan Richardson · Published 8 October 2026

Visitors per test arm scale roughly as 16 times (1 minus your conversion rate), divided by your conversion rate times the square of the relative lift you want to detect. At a 2% baseline purchase rate, detecting a 20% relative lift needs about 19,600 visitors per arm; detecting a 50% relative lift needs about 3,136 per arm. Halving the effect size you're trying to detect quadruples the sample required.

The symptom

A small page test, a button colour, a headline tweak, runs for a week or two on ordinary traffic and never reaches a clear result either way.

The cause

Required sample size moves on an inverse square relationship with the effect you're trying to detect. A small, subtle change (a 20% relative lift) needs roughly four times the sample a larger, more obvious change (a 50% relative lift) would need to detect at the same confidence. At a 2% baseline conversion rate, detecting a 20% relative lift needs about 19,600 visitors per arm, or about 39,200 total. Detecting a 50% relative lift at the same baseline needs about 3,136 per arm, or about 6,272 total. The test that sounds smallest, a tweak to one button, is often the largest possible ask of your traffic, precisely because small changes produce small, hard-to-detect effects.

When a test produces zero purchases in a given number of sessions, the honest ninety-five percent upper bound on the true conversion rate is roughly three divided by that session count. That's usually enough on its own to kill a bad idea without running the test any further.

The fix

Calculate the required sample size before launching a test, not after reading a disappointing result. If your traffic can't reach the sample size the arithmetic calls for inside a reasonable window, either test a bigger change, one expected to produce a bigger effect, or treat the test as directional only rather than a formal read.

How to spot it

Before trusting any A/B test result, check the actual session count per arm against what the inverse-square formula says was needed for the effect size you were hoping to detect. A test that ran with a fraction of the required sample was never capable of returning a clear answer, whatever the result looked like.

The numbers
ClaimValueSource
Visitors per arm formula16 x (1 - p) / (p x d x d), where p is the baseline conversion rate and d is the relative lift to detectMeasured in Real Money, Field Manual
Visitors per arm to detect a 20% relative lift at a 2% baseline purchase rateabout 19,600Measured in Real Money, Field Manual
Visitors per arm to detect a 50% relative lift at a 2% baseline purchase rateabout 3,136Measured in Real Money, Field Manual
Both arms combined, to detect a 20% relative lift at a 2% baseline purchase rateabout 39,200Measured in Real Money, Field Manual
Both arms combined, to detect a 50% relative lift at a 2% baseline purchase rateabout 6,272Measured in Real Money, Field Manual
Effect of halving the effect size you want to detectquadruples the required sampleMeasured in Real Money, Field Manual
Ninety-five percent upper bound on the true rate when a test shows zero purchases in n sessionsabout 3 divided by nMeasured in Real Money, Field Manual
Go deeper

This page covers one step. The full method is in the book.

Read the full part
sixtysteps.co
AI-powered growth for what's next. Sixty Steps is Onwards Analytics — a data and analytics firm.
Pages
Home Library The book — $7 About
Start
The Teardown — free [email protected]
© 2026 Onwards Analytics Every claim on this site carries a number, or is marked as reasoning.