By Ryan Richardson · Published 8 October 2026
Set a quota before starting: sixty to a hundred and twenty review-length items, plus eight to fifteen real conversations. Take half the sample specifically from three- and four-star reviews. Five-star reviews describe the vendor. One-star reviews describe a delivery failure. The middle star ratings describe the actual trade-off a buyer made, which is the material worth collecting.
Harvest from six surfaces: review sites, retail reviews of comparable products, forum threads, comments under a competitor's posts, archives, and your own sales notes. Mine comments rather than posts wherever possible; the thread under a competitor's best post is a free, dated panel of people who have already declared the problem in public, in their own words.
Code in two passes, a day apart: buckets first, theme labels second. A delayed second pass gets a solo operator most of the benefit of a second coder.
Four buckets do most of the work: struggle situation (what was happening the week before they went looking, becomes the headline), desired outcome (what they said they wanted, becomes the promise), anxiety (what they feared about switching or paying, becomes the objection block and guarantee), and competitive alternative (what they tried instead, becomes the 'unlike' clause).
A fifth bucket, misfit, is the one that actually tells you something. If misfit climbs past fifteen per cent of the corpus, narrow the buyer definition and re-harvest, rather than coding your way through a sample that was never describing one buyer in the first place.
Write the anti-glossary alongside the corpus too: every word your business uses internally that the corpus never produced, now banned from cold-traffic copy. For an expert-led business, the gap between internal and buyer vocabulary is usually the entire conversion problem, and this is the cheapest way to see it in one sitting.
The common failure is paraphrasing while collecting, which turns the market's words back into your own without you noticing. Read ten rows aloud and listen for your own cadence creeping back in. A second, quieter failure is saturation theatre: stopping at sixty quotes because the themes repeated. That's fine for the theme, but it doesn't give you the causal story behind a single instance of it. The fix isn't thirty more quotes, it's three more actual conversations.
| Claim | Value | Source |
|---|---|---|
| Corpus quota | 60-120 review-length items plus 8-15 real conversations | The Sixty Steps manuscript |
| Misfit-bucket threshold that signals a re-harvest is needed | 15% | Measured in Real Money, Field Manual |
| Matrix difficulty/coverage for this step | Moderate / Often skipped | The Sixty Steps matrix |