By Ryan Richardson · Published 8 October 2026
A test is declared a winner the day it first crosses a significance threshold, often within the first few days of running. A later, more careful look at the same test shows the result wasn't as solid as the moment it was called.
A significance threshold's stated false-positive rate, typically 5%, assumes one single look at the data, decided on in advance. Checking the result every day and stopping the first time it clears the bar is a different procedure entirely, even though it uses the same threshold. Each daily check is another chance for random noise to cross the line at least once, and across repeated daily checks, the real false-positive rate climbs from the nominal 5% to roughly 26%.
Write the stopping rule down before the first visitor arrives: the sample size needed, and the single point at which the test will be read. Don't check the result with intent to stop early; checking to monitor is fine, acting on an early crossing is the mistake. If a test must be read before its planned sample size is reached, treat any early result as directional only, not as a decision.
If a test was called a winner before its planned sample size was reached, or without a sample size having been planned at all, the real false-positive rate on that call is higher than the threshold suggests. Check whether a stopping rule was written down before the test launched; if not, treat the result as provisional.
| Claim | Value | Source |
|---|---|---|
| Nominal false-positive rate at a standard significance threshold | 5% | Measured in Real Money, Field Manual |
| Real false-positive rate when peeking daily and stopping at the first threshold crossing | about 26% | Measured in Real Money, Field Manual |