Paired Forecast Score Differences Versus Separate Means

By DX Research Group · · Forecast evaluation

Compare identical questions row by row before estimating uncertainty.

Paired score differences use the shared difficulty of each question to make a model comparison interpretable. Separate averages can hide missing rows and discard the covariance between models. We should join native LLM forecasts on the exact target identity, calculate each loss difference, and then analyze that difference over market blocks.

Four illustrative matched questions

Suppose model A has losses 0.10, 0.20, 0.30, and 0.40. Model B has 0.12, 0.22, 0.28, and 0.38. Both means are 0.25. Differences A minus B are −0.02, −0.02, 0.02, and 0.02, averaging zero. The sample standard deviation of differences is sqrt(0.0016/3), approximately 0.023094. Under an illustrative independent-row assumption, its standard error is 0.011547. The fixture demonstrates cancellation across cases, rather than evidence of equivalence.

The join is the experiment

Target identity should include instrument, reference time, horizon, label definition, and input snapshot. A join on symbol and day can pair different questions. Use explicit one-to-one validation and show unmatched counts for each model. If one head fails more frequently, publish complete scheduled coverage alongside the matched comparison. Shared successful rows answer a conditional question that can differ from full deployment reliability.

Estimate uncertainty at the level of independent information. Overlapping horizons and synchronized asset moves usually make the independent-row illustration too optimistic. Resample complete time blocks while keeping model pairs together, or use a declared dependence-aware method. Preserve sign conventions so a negative loss difference consistently favors A. Report the mean difference and its interval rather than interpreting overlapping separate confidence intervals as a formal comparison.

From fixture to a held-out result

The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.

Continue with shared outcome information for the related question of paired loss dependence. Our benchmark card keeps this forecast-level comparison separate from action and execution results.

When the paired mean is close to zero, inspect the distribution of differences and the calendar blocks producing each sign. A balanced mean can hide concentrated gains and losses. Preserve that conditional pattern as a hypothesis for a future frozen panel. It supplies a research direction while leaving general model-selection conclusions dependent on the reported interval.

Sources

Related field notes