Purging Label Intervals at an Agent Evaluation Boundary

By DX Research Group · · Forecast evaluation

A timestamp example explains why chronological rows alone do not prevent information leakage.

We audit our label intervals alongside row timestamps when checking evaluation boundaries.

Split the information interval

Chronological evaluation is essential for an agent whose questions resolve after its forecasts. Sorting rows by forecast time is only the first step. A training row can begin before the test boundary while its label depends on prices observed inside the test window. The row order is chronological, but the information used to fit the model crosses the boundary.

The mlfinlab cross-validation documentation describes purging training observations whose information overlaps testing observations. This note applies that documented idea to a small timestamp fixture. It does not prescribe a universal embargo length or claim that a particular DXRG experiment uses this implementation.

Audit a five-row boundary

Let the test window begin at 10:00 UTC. Assume training questions start at 09:20, 09:30, 09:40, 09:50, and 09:55. Their labels resolve at 09:40, 09:50, 10:00, 10:10, and 10:15 respectively. These times are illustrative.

If training requires every label to be fully known strictly before 10:00, only the first two rows qualify. The third row resolves exactly at the boundary and is excluded under this stated conservative convention. The last two resolve after it. A simple split by start time would retain all five and overlook the three label-availability conflicts.

That convention is a research choice whose rationale needs disclosure. A system with precisely defined simultaneous observation and ordering may allow a boundary-time label. State the choice explicitly, including clock precision and whether intervals are closed or half-open. Small timestamp ambiguities can matter when questions are issued frequently.

Separate purging from a generic gap

A fixed gap removes a set number of observations or a period around a boundary. Purging checks the actual information interval. They coincide only under particular assumptions about cadence and horizon.

For example, if every label resolves exactly 20 minutes after issuance and forecasts occur every 10 minutes, a carefully chosen time gap can enforce the desired boundary. With variable resolution horizons, a two-row gap might remove too little for one question and unnecessarily much for another. Save each row's start, last required observation, and resolution availability time.

The TimeSeriesSplit documentation offers a gap parameter measured in samples. That tool is useful, but its parameter alone does not verify variable label intervals, irregular sampling, or late source arrival. The audit must match the actual dataset.

Include the rest of the fitted pipeline

Apply the same training boundary to probability calibration, feature normalization, target thresholds, instrument selection, and prompt or policy selection. A purged model fit can still leak if its calibrator was trained on future labels. An externally updated LLM raises a separate provenance question: identify the model version and the historical claims the replay can support.

Record how many rows were removed and why. A large reduction changes the population and uncertainty, so publish both original and eligible counts. Do not replace a weak post-purge result with the pre-purge score.

For frontier agent evaluation, the useful artifact is an interval audit that a reviewer can reproduce from timestamps. It establishes the scope of information available to the fitted system. Forecast skill and action economics remain separate questions after that boundary is valid.

Connect the metric to the research record

Our agent evaluation framework separates the forecast comparison from the action and execution layers. The continuous production record shows why reporting an unfavorable result is part of a useful research account.

Sources

Related field notes