sixtysteps.co

A funnel bandit that rewards scroll-depth, not just purchases

By Ryan Richardson · Published 8 October 2026

In simulation, judging each step of a funnel only on whether it earns the very next step, rather than on whether it eventually leads to a purchase, let a bandit find the true best-performing funnel using about 5.3 times fewer simulated visitors than a bandit that only counted purchases. This is a simulated result from a test suite, not a result from a live campaign.

The problem with purchase-only bandits

A standard bandit that only rewards a purchase has to wait for purchases to happen before it can tell any variant apart from any other. When purchases are rare, as they routinely are on cold traffic, every variant looks statistically identical for a long time, and the bandit has nothing to learn from until volume gets very high. A landing-page test run on this project hit exactly that wall: a 282-visit, 12 x 3 x 3 x 6 bandit test produced near-uniform probability-of-best scores across every dimension, because zero purchases meant zero learning signal.

The alternative: judge each step on the next step, not the sale

The idea behind a depth-graduated design is local credit assignment: a hero section is judged only on whether the visitor scrolls past it, a call-to-action block is judged only on whether it gets clicked, and so on down the funnel. Each step earns its own reward from the step immediately after it, instead of every step competing for credit on one rare, downstream purchase event. This uses the abundant shallow signal (scrolls, clicks) first, and only pushes the reward deeper into the funnel once a variant has accumulated enough traffic to deserve that deeper test.

What the simulation found

A test suite built to compare the two approaches, run against a simulated funnel with a known correct winner, found the depth-graduated bandit converged on identifying that true winning funnel using roughly 727 simulated visitors, versus roughly 3,872 simulated visitors for a bandit that only rewarded purchases: about 5.3 times fewer visitors to reach the same correct answer. This result comes from a software test suite simulating visitor behaviour against a known-correct funnel, not from a live ad campaign, and should be read as a proof of the underlying mechanism rather than a number to expect in a live test.

Where this sits relative to the page-pattern test

This mechanism was designed specifically to answer the failure mode seen in the real-world bandit test described in a companion page: 282 real visits with zero purchases gave a purchase-only bandit nothing to work with. A depth-graduated reward would have had scroll and click signal to learn from well before any purchase happened, which is the gap this approach is meant to close. It has not yet been run on live paid traffic.

The numbers
ClaimValueSource
Simulated visitors needed for the depth-graduated bandit to converge on the true winning funnel~727 (simulated)Velocity engine.md, 'v1 BUILT + PROVEN' paragraph
Simulated visitors needed for a purchase-only bandit to converge on the same answer~3,872 (simulated)Velocity engine.md, 'v1 BUILT + PROVEN' paragraph
Improvement factor~5.3x fewer visitors (simulated)Velocity engine.md, 'v1 BUILT + PROVEN' paragraph
Automated test coverage on the simulation engine27 of 27 tests passingVelocity engine.md, 'v1 BUILT + PROVEN' paragraph
Real-world bandit test this simulation responds to: tracked visits across a 12x3x3x6 combination test282 visits, zero purchasesSixty Steps bandit test logs
Go deeper

This page covers one step. The full method is in the book.

Read the full part
sixtysteps.co
AI-powered growth for what's next. Sixty Steps is Onwards Analytics — a data and analytics firm.
Pages
Home Library The book — $7 About
Start
The Teardown — free [email protected]
© 2026 Onwards Analytics Every claim on this site carries a number, or is marked as reasoning.