Conformal Forecast Coverage Under Temporal Shift
By DX Research Group · · Forecast evaluation
Audit empirical coverage and set size when market distributions change.
A nominal conformal coverage level needs its assumptions and its realized time-series coverage reported together. Temporal market shift can break the exchangeability behind standard split-conformal guarantees. We should treat an observed coverage drop as a diagnostic for the forecast pipeline and calibration window, rather than turning a nominal percentage into a product assurance.
An illustrative quantile and shift
Suppose 19 calibration nonconformity scores are sorted and the target miscoverage is 0.10. The usual finite-sample rank is ceil((19 + 1) × 0.90) = 18. Select the eighteenth score, assuming the procedure’s required conventions and exchangeability conditions. Now imagine a later batch of 20 outcomes with 16 inside their prediction sets. Empirical coverage is 16/20 = 0.80, below nominal 0.90. That is an invented shift fixture, with a deliberately small denominator.
Retain the changing sets
Publish set size or interval width alongside coverage. Returning every class can achieve perfect classification-set coverage while offering little discrimination. Record calibration-window membership, score definition, quantile convention, and the model version. A native LLM head can feed a conformal procedure, but the wrapper’s assumptions remain separate from how the model produced probabilities.
Temporal diagnostics should use predeclared rolling windows and show both misses and support. Compare recent and earlier periods with identical outcome definitions. Adaptive methods have their own guarantees and update rules; they require direct inspection rather than inheriting the standard split-conformal statement. If the procedure updates after outcomes arrive, retain their availability timestamps so a replay cannot use a label that was still unresolved at the forecast time.
The saved record that would make this reviewable
The Adaptive Conformal Inference Under Distribution Shift provides the underlying methodological reference. The conformal tutorial details the split-conformal construction. The calculations and audit design here are illustrative extensions for an agent forecast record.
Continue with the distribution-shift protocol for the related question of temporal coverage assumptions. Our benchmark card keeps this forecast-level comparison separate from action and execution results.
The next conformal report should place empirical coverage, prediction-set size, and the date of each update side by side. A larger set can restore coverage while reducing usefulness. A model update can also change the score distribution beneath a fixed quantile. Version these changes so temporal failures can be attributed to the head, wrapper, or calibration data.