Target Prevalence by Asset in LLM Forecast Evaluation
By DX Research Group · · Forecast evaluation
Measure event frequency separately for each asset before interpreting a pooled forecast score.
The same constant probability can look accurate on one instrument and poorly calibrated on another because the target occurs at different frequencies. Our evaluation should preserve those differences before it assigns credit to a native LLM forecast. An asset identifier is part of the target contract, alongside the price reference and resolution horizon.
An illustrative two-asset panel
Asset A has 20 positive labels among 100 questions; asset B has 80 among 100. A constant 0.20 forecast has binary squared-error mean 0.16 on A: (20 × 0.64 + 80 × 0.04)/100. On B it scores 0.52: (80 × 0.64 + 20 × 0.04)/100. The equally weighted pooled score is 0.34. A constant 0.50 scores 0.25 on both assets. The first forecast therefore loses overall despite matching A’s observed frequency. These are invented labels, used only to expose the arithmetic.
Keep the reference time valid
An asset-specific reference estimated from those evaluation labels would score 0.16 on each asset, but that is an in-sample diagnostic. A usable reference must be estimated from earlier resolved questions. Save the counts available at each forecast time and distinguish an expanding historical reference from a fixed training-period reference. Newly listed instruments need an explicit fallback, because a small denominator can create an extreme empirical frequency.
Inspect differences in event definitions before assigning them to asset behavior. A fixed percentage threshold produces a different prevalence profile from a volatility-scaled threshold. If the native head answers a direction question while the report resolves a threshold crossing, the score belongs to a different task. Include per-asset counts, positive counts, and forecast distributions. Combine assets only after stating whether weights represent questions, instruments, or elapsed market time.
The saved record that would make this reviewable
The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.
Continue with a time-valid base-rate reference for the related question of asset composition. Our benchmark card keeps this forecast-level comparison separate from action and execution results.
For an asset-prevalence report, the decisive receipt is a dated table of event counts and reference probabilities. Show whether a pooled advantage survives common asset weights. If it disappears, describe the result as sensitivity to market composition. That finding can guide the next question panel while leaving any proposed harness intervention awaiting its own evaluation.