A Walk-Forward Protocol for Native LLM Forecast Comparisons

By DX Research Group · · Forecast evaluation

A three-window example separates model updating, probability calibration, and honest aggregate reporting.

We describe our updating procedure explicitly because its choices affect the measured comparison.

Specify what moves forward

Walk-forward evaluation advances through time, fitting or selecting a system on earlier data and evaluating it on later data. For native LLM forecasts, “fitting” can include more than weight updates: a prompt choice, probability map, baseline window, retrieval rule, and action threshold can all be selected using labels.

The scikit-learn time-series cross-validation guide describes successive temporal folds with growing training sets. The design below is a proposed evaluation protocol. Its main decision is whether a comparison concerns a frozen model or a repeatable updating procedure.

Work through three windows

Assume January through March is the initial development period. Freeze a candidate specification on March 31. Evaluate 40 April questions, then update the allowed components using only labels available by April 30. Evaluate 60 May questions, update again under the same rule, and evaluate 100 June questions.

Suppose illustrative Brier scores are 0.20, 0.24, and 0.22. The mean of these three window scores is 0.22. The score over all 200 questions is (40 times 0.20 plus 60 times 0.24 plus 100 times 0.22) divided by 200, or 0.222. Both are legitimate summaries of different weighting choices. One gives each month equal weight; the other gives each question equal weight.

Report the weighting choice rather than switching after seeing which number is smaller. Show each window's question count and eligible population. A changing count can reflect market coverage, label availability, or system failures, each of which affects interpretation.

Treat updates as part of the candidate

If one system is updated monthly and another stays frozen, the comparison includes the update procedure. That can be a valid research question, but it is not a clean test of model identity alone. To isolate model differences, match the available data, tuning budget, updating cadence, and permitted transformations.

Save the candidate version used for every forecast. A provider model identifier may be insufficient if the provider changes an endpoint behind the same public name. Keep whatever version evidence the route supplies and state the remaining reproducibility limit.

A forecast whose outcome resolves after the month-end boundary cannot become training data for that update. This matters when horizons vary. The update rule must use resolution availability, not merely question issuance date. Apply the boundary to calibration and reference forecasts as well as the principal model.

Avoid a retrospective winning path

Do not choose April's best prompt, May's best calibration map, and June's best threshold after seeing each month's test scores and then call their combined curve a single prospective strategy. That path has used evaluation labels for selection. If monthly selection is intended, implement it using earlier data and save the choice before the next window starts.

Keep a complete trial ledger with unsuccessful candidates and invalid runs. A final untouched period can test the selected updating procedure, although no finite holdout guarantees future generalization.

For agentic trading research, pair the probability report with a separately specified action policy and execution assumptions. A walk-forward probability improvement is evidence about forecasts in those periods. Realizable gains, product superiority, and safe deployment require additional evidence. The versioned procedure is what makes the conclusion inspectable.

Connect the metric to the research record

Our agent evaluation framework separates the forecast comparison from the action and execution layers. The continuous production record shows why reporting an unfavorable result is part of a useful research account.

Sources

Related field notes