Censored Outcomes in an Agent Forecast Evaluation

By DX Research Group · · Forecast evaluation

Keep unresolved future outcomes visible in the denominator.

A forecast whose horizon extends beyond the dataset cutoff has a censored outcome. Treating it as a negative label invents an event; dropping it silently hides the forecast population. Our native LLM report should separate resolved outcomes, pending outcomes, and irrecoverable observation failures before computing a score.

A cutoff fixture

An illustrative run contains 100 saved questions. Eighty horizons have elapsed and reliable labels exist; 15 extend beyond the extraction cutoff; five lack their required price receipt. Suppose the 80 resolved rows have total squared loss 16. Their resolved score is 16/80 = 0.20. Temporal completion is 80/100 = 0.80. The report should state both figures. Dividing by 100 would produce 0.16 only by assigning zero loss to rows whose loss is unknown.

Distinguish maturity from missingness

Pending outcomes can resolve after sufficient calendar time. Missing price receipts need a recovery or exclusion policy. Save the reason and expected resolution time for each row. A model with more long-horizon questions can look underrepresented at a common cutoff even if its output pipeline is reliable. Compare mature cohorts at matched horizons and still publish the complete scheduled inventory.

If missingness depends on event severity or instrument outages, the resolved subset can be biased. Inverse-probability weighting requires an appropriate observation model and assumptions, which should be stated and checked; it is not a universal cure. A useful first sensitivity for bounded squared loss assigns unresolved rows hypothetical loss zero and one. Here the full-panel average could lie between 0.16 and 0.36, treating all 20 unknowns as unconstrained. These are bounds, not predictions of future scores.

The saved record that would make this reviewable

The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.

Continue with failed-turn denominators for the related question of outcome maturity. Our benchmark card keeps this forecast-level comparison separate from action and execution results.

For a censored panel, retain a scheduled follow-up date in the research protocol and a frozen cutoff in the current report. Recompute mature outcomes as a new version when labels become available. If attrition persists, inspect its reason distribution. A changing resolved denominator is part of the evidence and deserves the same visibility as the score itself.

Sources

Related field notes