Reviewer Agreement Depends on Failure Prevalence
By DX Research Group · · Trace evaluation
Report the agreement table alongside kappa so rare defects remain visible.
High reviewer agreement can coexist with weak evidence that reviewers recognize a rare defect. We would publish the two-rater label table, raw agreement and a chance-adjusted measure, then inspect the disagreements. A single agreement percentage hides whether both reviewers simply label almost everything acceptable.
Ninety-eight agreements, a modest kappa
In an illustrative one-hundred-trace audit, both reviewers mark one trace defective and ninety-seven acceptable. Each reviewer marks one additional trace defective that the other marks acceptable. They agree on ninety-eight traces, giving 98% raw agreement. Each reviewer labels two traces defective and ninety-eight acceptable.
Expected agreement from those marginal proportions is 0.02² + 0.98² = 0.9608. Cohen's kappa is (0.98 − 0.9608)/(1 − 0.9608), approximately 0.490. The calculation exposes how prevalence affects the chance-adjusted result. It supplies no universal threshold for deciding whether the audit is trustworthy.
The scikit-learn kappa documentation gives the statistic's definition. Our proposed audit keeps the full table because a scalar alone leaves the rare-label behavior obscure. Kappa also measures agreement rather than correctness against an independent reference.
Review the disagreements before changing the rubric
For the two disputed traces, ask whether the disagreement comes from missing evidence, an ambiguous policy clause or a different severity threshold. A reviewer who cannot see the order acknowledgement may be applying the same rubric to less information. Fixing access can resolve the label without rewriting the rule.
We would include deliberately constructed examples of the rare failure during reviewer training, then keep a separate natural-prevalence audit. The constructed set tests recognition under enriched prevalence. Its accuracy should stay separate from the rate of naturally occurring failures.
The operating-layer controls companion provides historical examples where trace labels matter. The continuous record companion illustrates evidence that depends on explicit measurement definitions. Neither supplies an agreement result for this proposed reviewer exercise.
An adjudicator should review original evidence while remaining blind to model identity when that identity is irrelevant. Preserve initial labels and the final resolution. Replacing the first labels with adjudicated agreement would erase the very disagreement being measured. Record whether the adjudicator found one correct interpretation or concluded that the evidence supports an unknown outcome.
For severe authorization failures, show positive agreement and the number of confirmed cases alongside global statistics. A rubric may perform well on routine rationale support and poorly on the rare exception that motivates release caution. The useful deliverable is therefore a compact agreement table, disagreement explanations and a revised fixture set. We would rerun only the changed rubric's affected cases before making broader claims about reviewer reliability.