By Ryan Richardson · Published 8 October 2026
A standard bandit that only rewards a purchase has to wait for purchases to happen before it can tell any variant apart from any other. When purchases are rare, as they routinely are on cold traffic, every variant looks statistically identical for a long time, and the bandit has nothing to learn from until volume gets very high. A landing-page test run on this project hit exactly that wall: a 282-visit, 12 x 3 x 3 x 6 bandit test produced near-uniform probability-of-best scores across every dimension, because zero purchases meant zero learning signal.
The idea behind a depth-graduated design is local credit assignment: a hero section is judged only on whether the visitor scrolls past it, a call-to-action block is judged only on whether it gets clicked, and so on down the funnel. Each step earns its own reward from the step immediately after it, instead of every step competing for credit on one rare, downstream purchase event. This uses the abundant shallow signal (scrolls, clicks) first, and only pushes the reward deeper into the funnel once a variant has accumulated enough traffic to deserve that deeper test.
A test suite built to compare the two approaches, run against a simulated funnel with a known correct winner, found the depth-graduated bandit converged on identifying that true winning funnel using roughly 727 simulated visitors, versus roughly 3,872 simulated visitors for a bandit that only rewarded purchases: about 5.3 times fewer visitors to reach the same correct answer. This result comes from a software test suite simulating visitor behaviour against a known-correct funnel, not from a live ad campaign, and should be read as a proof of the underlying mechanism rather than a number to expect in a live test.
This mechanism was designed specifically to answer the failure mode seen in the real-world bandit test described in a companion page: 282 real visits with zero purchases gave a purchase-only bandit nothing to work with. A depth-graduated reward would have had scroll and click signal to learn from well before any purchase happened, which is the gap this approach is meant to close. It has not yet been run on live paid traffic.
| Claim | Value | Source |
|---|---|---|
| Simulated visitors needed for the depth-graduated bandit to converge on the true winning funnel | ~727 (simulated) | Velocity engine.md, 'v1 BUILT + PROVEN' paragraph |
| Simulated visitors needed for a purchase-only bandit to converge on the same answer | ~3,872 (simulated) | Velocity engine.md, 'v1 BUILT + PROVEN' paragraph |
| Improvement factor | ~5.3x fewer visitors (simulated) | Velocity engine.md, 'v1 BUILT + PROVEN' paragraph |
| Automated test coverage on the simulation engine | 27 of 27 tests passing | Velocity engine.md, 'v1 BUILT + PROVEN' paragraph |
| Real-world bandit test this simulation responds to: tracked visits across a 12x3x3x6 combination test | 282 visits, zero purchases | Sixty Steps bandit test logs |