By Ryan Richardson · Published 8 October 2026
A fixed-split A/B test keeps sending traffic to a losing variant for the full length of the test, because the split (50/50, or whatever's set) doesn't move. A bandit keeps the split moving: as soon as evidence favours one variant, more of the next batch of visitors sees it, which means less traffic is spent on the one that's losing.
The specific method is Thompson sampling. For each variant, the system tracks views and a weighted reward score, builds a Beta probability distribution from those counts, draws one random sample per visitor from each variant's distribution, and shows the visitor whichever variant drew the highest number. Variants with more evidence of winning get sampled from a narrower, higher distribution, so they win the draw more often, without ever being locked in completely.
Purchases are rare, so waiting for purchase data alone to tell variants apart takes a long time. The reward formula used here blends a small weight for clicks with a large weight for purchases: success-shape alpha = 0.1 per click plus 1.0 per purchase; failure-shape beta = views minus that alpha. A click moves the needle a little immediately; a purchase moves it a lot. That gives the bandit a usable signal from day one instead of waiting weeks for enough purchases to accumulate.
The bandit doesn't pool all visitors into one pot. It runs a separate instance per UTM source (what's called a stratum), because a visitor who clicked a specific ad angle is not behaviourally the same as a visitor who found the page through search. Pooling them would blur which hero wins for which audience. A never-seen source falls back to the default source's current standings until it has its own data, so a brand-new ad doesn't start from nothing.
It doesn't auto-declare a winner. Someone still has to read the numbers and decide whether to hardcode the leading variant or keep the bandit running. And it has a safety net: if the assignment call is slow, the visitor gets a deterministic fallback instead of an error, so a backend hiccup never breaks the page.
| Claim | Value | Source |
|---|---|---|
| Reward weighting: click weight | 0.1 | Sixty Steps bandit test logs |
| Reward weighting: purchase weight | 1.0 | Sixty Steps bandit test logs |
| Sticky assignment length before re-sampling | 7 days | Sixty Steps bandit test logs |
| Assignment fallback timeout | 800ms | Sixty Steps bandit test logs |