Label Noise and the Forecast Benchmark Ceiling

By DX Research Group · · Forecast evaluation

Quantify the score floor implied by an explicit observation-noise model.

Noisy labels can impose a score floor even on a forecast that knows the underlying state. A benchmark ceiling claim therefore needs an explicit noise model, independent evidence about its parameters, and a clear distinction between the latent event and the recorded target. We can calculate a sensitivity scenario without claiming that actual market labels follow it.

An illustrative symmetric-flip model

Let a binary latent state be known perfectly, while the recorded label independently flips with probability 0.10. The optimal probability for the observed label is 0.90 when the latent state is one and 0.10 when it is zero. Expected binary squared error is 0.90 × 0.01 + 0.10 × 0.81 = 0.09 in either state. Expected log loss is −0.90 ln(0.90) − 0.10 ln(0.10), approximately 0.325083. A deterministic latent prediction scores 0.10 in squared error against the noisy label.

Estimate noise through independent resolution

Duplicate resolution using the same upstream feed can reproduce the same error and reveal little. Compare independently sourced receipts and adjudicate disagreements under a registered rule. Symmetric independent flips are a simplifying assumption: timestamp errors may be directional, and ambiguous labels may cluster near thresholds. Publish sensitivity across plausible rates rather than promoting one invented number into a universal benchmark limit.

A model can sometimes predict features of the labeling process, including systematic feed delay. Its apparent skill against recorded labels may then differ from skill against the intended event. Preserve both questions in the report. When labels are corrected, rescore all models and retain versioned mappings. The noise floor is conditional on the assumed information set and error process; it is a diagnostic calculation, not proof that further forecasting improvement is impossible.

Where this enters the research workflow

The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.

Continue with the resolution-disagreement audit for the related question of label uncertainty. Our benchmark card keeps this forecast-level comparison separate from action and execution results.

Any benchmark ceiling estimate should specify which uncertainty is irreducible under the available information. A feed bug that can be corrected differs from genuine randomness in the target. Maintain that distinction in the label audit. Independent resolution evidence can tighten the assumed noise range; a purely illustrative flip calculation supplies a sensitivity mechanism rather than an empirical ceiling.

Sources

Related field notes