How Forecast Scores Depend on the Resolution Horizon

By DX Research Group · · Forecast evaluation

Treat each horizon as a distinct target before comparing model quality.

Changing a forecast horizon changes the event, its prevalence, and the information that can predict it. A better score at one horizon therefore cannot establish that a native LLM is generally a better market forecaster. Our evaluation should bind every probability to its horizon and compare models within that target before summarizing across horizons.

A single forecast with two illustrative outcomes

At time zero, a hypothetical instrument price is 100. After one hour it is 101; after one day it is 99. A probability of 0.70 for positive movement has squared loss (0.70 − 1)² = 0.09 against the one-hour label, and (0.70 − 0)² = 0.49 against the one-day label. The 0.40 difference comes from changing the target. A valid one-day probability must be elicited as a one-day question rather than reusing the shorter-horizon answer.

Create a horizon-indexed question panel

Save origin time, terminal time, return convention, and the treatment of market closures or missing terminal observations. Give every horizon its own candidate definition and historical reference. Keep the origin snapshot identical across models while allowing horizon-specific prompts that are documented. If a label asks whether a threshold is crossed at any point before expiry, it differs from an endpoint-return label even at the same duration.

Longer horizons overlap more often at a fixed invocation rate. The number of scored rows can therefore grow faster than independent evidence. Use time blocks that respect the longest relevant label interval when assessing a cross-horizon comparison. Show per-horizon support, score differences, and references before calculating a weighted summary. Specify whether weights represent research priorities, question counts, or a fixed horizon portfolio. That choice determines what the overall ranking means.

Where this enters the research workflow

The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.

Continue with label-interval purging for the related question of resolution horizons. Our benchmark card keeps this forecast-level comparison separate from action and execution results.

A horizon comparison should include a small grid of origins and terminal timestamps before it displays a leaderboard. That grid catches reused probabilities and incorrectly joined outcomes. Publish one score difference per horizon with its support. Any overall summary then represents an explicit research preference, while each constituent task remains independently interpretable and reproducible.

Sources

Related field notes