By Ryan Richardson · Published 8 October 2026
A set of 11 frozen email drafts was scored by two separate blind panels, each simulating 4 reader roles, on the same rubric, against byte-identical text. That's 220 reader-by-dimension score pairs. One panel passed 9 of the 11 drafts. The other passed 0 of 11. On the surface that reads as a broken grader. It wasn't quite that simple, and the way to find out which it actually is applies to any two-panel or two-reviewer grading setup.
The gate used two kinds of failure: a hard floor (every draft needs a total score of at least 80) and a softer tie-break rule (no more than a small number of individual 5-out-of-5 dimension scores can be missing). The first pass at this data concluded the entire gap was the soft tie-break rule, and recommended skipping edits on most of the drafts. That conclusion was wrong, and an audit caught it: 7 of the 11 drafts actually breached the hard floor outright, with totals as low as 64 points. Only 4 drafts were sitting exactly at the floor and failing purely on the soft tie-break. A 5-point rule can't explain a draft scoring 64 points instead of 80 points. The order matters: list floor failures first, tie-break failures second, and only apply a calibration explanation to the second group.
Across the 220 scored pairs, the two panels moved in the same direction far more than they moved apart: one direction of disagreement happened 122 times, scores matched exactly 95 times, and the other direction happened only 3 times. That lopsidedness is itself informative, but the dimension-level breakdown is what actually separates a strict grader from a real content gap: the gap was concentrated almost entirely in the dimensions that reward hard evidence (specificity, credibility, usability), and close to zero in dimensions that reward feel (insight, reading experience). A gap that's concentrated on evidence-based dimensions and near-zero on feel-based ones is a threshold difference in what counts as top marks, not a content defect.
Even where the gap was explained by grader strictness, reading the written feedback (not just the scores) mattered: on one draft, all readers on the stricter panel independently dropped the same three dimensions and named the same missing element in their comments. Three independent readers naming the same gap is a real content issue a strict grader caught, even though its score pattern looked like pure calibration noise.
Don't re-run the same panel on unchanged copy hoping for a friendlier result. And don't reason from "a grader-strictness gap exists somewhere in this data" straight to "this draft is fine" without checking its own floor score first. That shortcut is what produced the wrong initial recommendation here.
| Claim | Value | Source |
|---|---|---|
| Number of frozen drafts in the test | 11 | Reference cross panel grader offset.md, header line |
| Total reader-dimension score pairs compared | 220 | Reference cross panel grader offset.md, header line |
| Drafts one panel passed | 9 of 11 | Reference cross panel grader offset.md, para 2 |
| Drafts the other panel passed | 0 of 11 | Reference cross panel grader offset.md, para 2 |
| Score floor for passing | total >= 80 | Reference cross panel grader offset.md, para 3 |
| Drafts that breached the floor outright | 7 of 11 (totals of 75, 77, 73, 67, 64, 67, 64) | Reference cross panel grader offset.md, para 3 |
| Drafts failing only on the soft tie-break rule | 4 (all at exactly 80) | Reference cross panel grader offset.md, para 3 |
| Score-transition distribution across 220 pairs | one direction 122 times, identical 95 times, other direction 3 times | Reference cross panel grader offset.md, para 'Claude lower in 122, identical in 95, higher in 3' |
| Mean score delta, specificity dimension | -1.00 | Reference cross panel grader offset.md, dimension table |
| Mean score delta, credibility dimension | -0.95 | Reference cross panel grader offset.md, dimension table |
| Mean score delta, insight dimension | -0.23 | Reference cross panel grader offset.md, dimension table |