sixtysteps.co

Two blind graders disagree: real problem, or a strict grader

By Ryan Richardson · Published 8 October 2026

Check which drafts fail an absolute floor first, before looking at why two graders disagree; only the drafts that pass the floor and fail on a softer tie-breaker are candidates for a calibration (grader-strictness) explanation, and even those need their written feedback read, not just their scores.

The setup

A set of 11 frozen email drafts was scored by two separate blind panels, each simulating 4 reader roles, on the same rubric, against byte-identical text. That's 220 reader-by-dimension score pairs. One panel passed 9 of the 11 drafts. The other passed 0 of 11. On the surface that reads as a broken grader. It wasn't quite that simple, and the way to find out which it actually is applies to any two-panel or two-reviewer grading setup.

Check the floor before the tie-breaker

The gate used two kinds of failure: a hard floor (every draft needs a total score of at least 80) and a softer tie-break rule (no more than a small number of individual 5-out-of-5 dimension scores can be missing). The first pass at this data concluded the entire gap was the soft tie-break rule, and recommended skipping edits on most of the drafts. That conclusion was wrong, and an audit caught it: 7 of the 11 drafts actually breached the hard floor outright, with totals as low as 64 points. Only 4 drafts were sitting exactly at the floor and failing purely on the soft tie-break. A 5-point rule can't explain a draft scoring 64 points instead of 80 points. The order matters: list floor failures first, tie-break failures second, and only apply a calibration explanation to the second group.

The diagnostic that separates a strict grader from a real gap

Across the 220 scored pairs, the two panels moved in the same direction far more than they moved apart: one direction of disagreement happened 122 times, scores matched exactly 95 times, and the other direction happened only 3 times. That lopsidedness is itself informative, but the dimension-level breakdown is what actually separates a strict grader from a real content gap: the gap was concentrated almost entirely in the dimensions that reward hard evidence (specificity, credibility, usability), and close to zero in dimensions that reward feel (insight, reading experience). A gap that's concentrated on evidence-based dimensions and near-zero on feel-based ones is a threshold difference in what counts as top marks, not a content defect.

Don't stop at the number

Even where the gap was explained by grader strictness, reading the written feedback (not just the scores) mattered: on one draft, all readers on the stricter panel independently dropped the same three dimensions and named the same missing element in their comments. Three independent readers naming the same gap is a real content issue a strict grader caught, even though its score pattern looked like pure calibration noise.

What not to do

Don't re-run the same panel on unchanged copy hoping for a friendlier result. And don't reason from "a grader-strictness gap exists somewhere in this data" straight to "this draft is fine" without checking its own floor score first. That shortcut is what produced the wrong initial recommendation here.

The numbers
ClaimValueSource
Number of frozen drafts in the test11Reference cross panel grader offset.md, header line
Total reader-dimension score pairs compared220Reference cross panel grader offset.md, header line
Drafts one panel passed9 of 11Reference cross panel grader offset.md, para 2
Drafts the other panel passed0 of 11Reference cross panel grader offset.md, para 2
Score floor for passingtotal >= 80Reference cross panel grader offset.md, para 3
Drafts that breached the floor outright7 of 11 (totals of 75, 77, 73, 67, 64, 67, 64)Reference cross panel grader offset.md, para 3
Drafts failing only on the soft tie-break rule4 (all at exactly 80)Reference cross panel grader offset.md, para 3
Score-transition distribution across 220 pairsone direction 122 times, identical 95 times, other direction 3 timesReference cross panel grader offset.md, para 'Claude lower in 122, identical in 95, higher in 3'
Mean score delta, specificity dimension-1.00Reference cross panel grader offset.md, dimension table
Mean score delta, credibility dimension-0.95Reference cross panel grader offset.md, dimension table
Mean score delta, insight dimension-0.23Reference cross panel grader offset.md, dimension table
External sources
Go deeper

This page covers one step. The full method is in the book.

Read the full part
sixtysteps.co
AI-powered growth for what's next. Sixty Steps is Onwards Analytics — a data and analytics firm.
Pages
Home Library The book — $7 About
Start
The Teardown — free [email protected]
© 2026 Onwards Analytics Every claim on this site carries a number, or is marked as reasoning.