Scoring Native LLM Forecasts by Volatility

By DX Research Group · · Forecast evaluation

Use lagged volatility strata to diagnose where probability error changes.

Volatility-stratified scoring can reveal a model’s dependence on market conditions that a pooled mean conceals. The strata must be constructed from information available at forecast time. We should preserve the volatility estimator and boundaries so a later researcher can distinguish difficult episodes from a change in the evaluation population.

An illustrative composition effect

Model A scores 0.10 in a low-volatility stratum and 0.40 in a high-volatility stratum. Model B scores 0.15 and 0.30 respectively. With 80 low and 20 high questions, A averages 0.16 and B averages 0.18. With 20 low and 80 high questions, A averages 0.34 and B averages 0.27. A change in the mixture reverses the ranking even though each model’s conditional scores remain identical.

Calculate volatility with the available past

Specify return interval, lookback window, missing-observation handling, and estimator. A realized volatility measure using future returns belongs to an ex-post diagnostic with a different interpretation. Training-period quantile boundaries can be frozen for holdout scoring; test-period boundaries describe relative conditions inside that test. Publish which choice was made and retain continuous volatility values so alternative cuts can be inspected.

Show positive-event prevalence inside each stratum. A fixed return threshold often becomes easier to cross when volatility rises, changing both target frequency and score difficulty. Compare against time-valid stratum references where support permits. Match model questions and report counts per asset and period. Thin extreme strata deserve uncertainty intervals and case inspection, particularly when several questions overlap the same volatile episode.

From fixture to a held-out result

The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.

Continue with the historical sizing finding for the related question of volatility strata. Our benchmark card keeps this forecast-level comparison separate from action and execution results.

The volatility report should show each stratum’s support and reference score beside the model score. A large loss in turbulent episodes can reflect a harder target, a worse relative forecast, or both. The reference comparison distinguishes these possibilities. Preserve the continuous estimator values so a reviewer can check whether the conclusion survives nearby boundaries and common weights.

Sources

Related field notes