Auditing Disagreement in Forecast Resolution

By DX Research Group · · Forecast evaluation

Separate price-source ambiguity from native LLM probability error.

When two valid-looking resolvers disagree, the target requires a resolution audit before the forecast earns a definitive score. Small price differences near a boundary can flip a binary outcome and dominate a comparison. Our evaluation should retain the competing observations, the declared canonical resolver, and a sensitivity result for ambiguous cases.

A boundary fixture with two receipts

An illustrative question asks whether price rises strictly above a reference of 100 at a fixed horizon. Resolver A records 100.01 and labels the outcome one. Resolver B records 99.99 and labels it zero. For a saved probability of 0.80, squared error is 0.04 under A and 0.64 under B, a difference of 0.60. A mere 0.02 difference in observed prices therefore changes the score substantially because the event definition is discontinuous.

Choose the resolver in advance

Specify venue, instrument, price type, timestamp matching, and permitted observation lag. A last trade, midpoint, mark price, and index price represent different measurements. Resolver selection after seeing model rankings invites a favorable target rewrite. Preserve raw receipts and resolve using the registered hierarchy. Cases without the required observation remain unresolved or excluded under a stated eligibility rule.

Report disagreement prevalence and where it occurs. Threshold-adjacent questions deserve a margin-to-boundary field, making it possible to show score sensitivity by ambiguity level. Human adjudication should be blinded to model identity and probability where feasible. A revised label version can be legitimate, but it needs a published change record and rescoring of all compared heads. Keep the original report available so the correction is traceable.

From fixture to a held-out result

The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.

Continue with freezing evaluation inputs for the related question of label resolution. Our benchmark card keeps this forecast-level comparison separate from action and execution results.

For ambiguous resolutions, publish the registered label and an alternate-resolver sensitivity on the same saved probabilities. If a ranking reverses, the benchmark should expose that dependence directly. A resolver correction is an evaluator repair. A model improvement requires a fresh comparison using the corrected, frozen target definition and the same eligibility rules for every candidate.

Sources

Related field notes