Why Overlapping Forecast Labels Are Not Independent Evidence

By DX Research Group · · Forecast evaluation

A rolling-horizon fixture shows how many saved rows can represent far fewer distinct market intervals.

We count shared outcome information alongside rows when describing our evidence population.

Count the information shared by questions

An agent may forecast the next hour every five minutes. That creates many saved predictions, but neighboring labels share most of their price interval. A report with thousands of rows needs a separate assessment of independent market episodes.

The NIST autocorrelation reference describes correlation between a series and its lagged values and explains the role of randomness assumptions. Overlapping labels are a concrete reason to investigate dependence. Estimating autocorrelation requires examining actual data.

Calculate an overlap fixture

Imagine forecasts at 10:00, 10:05, and 10:10 UTC, each for the following 60 minutes. Their label intervals end at 11:00, 11:05, and 11:10. Adjacent intervals share 55 minutes, or 55 divided by 60, approximately 91.7% of their duration. The first and third share 50 minutes, approximately 83.3%.

Those percentages describe interval overlap rather than correlation coefficients. Price changes depend on more than overlap length, and nonlinear event labels can behave differently. The fixture establishes that three rows contain heavily shared outcome information. Estimating effective sample size requires additional evidence.

Now extend the schedule to 12 consecutive forecasts, five minutes apart. They cover 115 minutes from the first issuance to the last resolution. Counting 12 one-hour labels as 12 disjoint hours would describe 720 minutes, more than six times the actual union. This arithmetic helps expose the difference between question count and interval coverage without pretending to estimate independent evidence.

Preserve the rows and change the uncertainty design

Overlap is not necessarily a reason to discard useful forecasts. A deployed system may genuinely issue frequent decisions, and row-level scoring can accurately describe those decisions. The issue is uncertainty and generalization: resampling individual rows as though independent can understate variation when losses are correlated.

Inspect dependence in row-level losses and in paired loss differences between models. The latter is especially important because two systems may share market shocks that cancel in the comparison. Label overlap alone cannot tell you how much the difference series depends on its past.

Consider reporting scores by nonoverlapping market episodes alongside the full row-level score. A block-based resampling analysis can preserve local dependence, but the block unit and duration must match the target and observed dependence. A daily block is an assumption to document, not a universal correct choice.

Keep cross-sectional dependence visible

Two instruments can share a market event even when their own label intervals do not overlap. Grouping only by instrument may miss common shocks; grouping only by day may hide persistent relationships across days. Record the forecast cadence, horizons, instruments, and proposed grouping before fitting confidence claims.

A useful audit table contains total forecast rows, resolved rows, union of label intervals, nonoverlapping episode counts under a stated rule, and score by period. These quantities answer different questions and should remain separate.

In native LLM agent evaluation, frequent prompts measure repeated behavior in context. The independent market history depends on the underlying episodes. Any DXRG performance claim needs actual loss-series evidence and an uncertainty method appropriate to that population, while action economics and execution reliability require their own analysis.

Connect the metric to the research record

Our agent evaluation framework separates the forecast comparison from the action and execution layers. The continuous production record shows why reporting an unfavorable result is part of a useful research account.

Sources

Related field notes