Calculating Brier Score for an Agent Forecast Audit

By DX Research Group · · Forecast evaluation

A five-row worked example for checking probability scores, sample alignment, and comparison baselines.

We start our probability audit with arithmetic a reviewer can reproduce by hand.

Decide what the probability means

An agent can produce a convincing explanation and an unhelpful probability. To evaluate a native LLM forecast, first define the event independently of its prose. “Positive movement” needs an instrument, observation time, reference price, horizon, and rule for unchanged prices. The forecast must be saved before the outcome becomes observable. Otherwise a precise score describes a contaminated task.

This note uses the binary Brier convention: the mean of squared differences between probability p and outcome y, where y is zero or one. Lower is better. The scikit-learn Brier documentation documents the score and the distinction between binary and multiclass scaling. Record the convention because two implementations can report numbers differing by a factor of two.

Recalculate five rows

Consider five illustrative forecasts: 0.80, 0.60, 0.30, 0.20, and 0.70. Suppose the corresponding outcomes are 1, 0, 0, 1, and 1. These inputs are invented for the calculation.

The individual squared errors are 0.04, 0.36, 0.09, 0.64, and 0.09. Their sum is 1.22. Dividing by five gives a Brier score of 0.244. The fourth forecast contributes more than half of the total loss: it assigned only 20% to an event that occurred. That row deserves inspection, although one surprising event does not by itself prove miscalibration.

A constant forecast of 0.50 scores 0.25 on these same rows. The model's absolute improvement is only 0.006. Expressed as 1 minus model score divided by reference score, the improvement is 0.024, or 2.4%. Calling that a substantial advantage would require evidence about sampling uncertainty and repeated performance. Five rows supply neither.

Audit the comparison before interpreting it

Check that both forecasts use the same resolved questions. An agent that skips difficult cases cannot be compared fairly with a baseline scored on every case unless coverage is disclosed. Save an explicit failure outcome for missing or invalid probabilities rather than quietly removing them. Decide the treatment before examining the leaderboard.

Check label ordering as well. A probability for “down” scored against an “up” indicator can reverse the intended interpretation without causing a software error. Verify one favorable and one unfavorable example manually. Keep original probabilities rather than only rounded display values, and record whether the score uses sample weights.

If weights represent instruments, days, or mandates, show why those weights match the research question. Giving each forecast equal weight answers a different question from giving each market episode equal weight when one episode generates hundreds of repeated prompts.

Keep the score within its scope

A Brier score measures probability error on the chosen sample. It does not establish that the agent executes valid actions, uses affordable tools, or produces favorable realized economics. Those require separate artifacts. An agentic trading audit should preserve the forecast layer so a change in execution policy cannot masquerade as better prediction.

Use this calculation as a small arithmetic fixture in a larger evaluation: save the five inputs, compute each loss, verify the mean, compare the fixed reference, and state the eligible population. A transparent small fixture makes a large aggregate easier to trust, while leaving any DXRG performance conclusion dependent on actual held-out results.

Connect the metric to the research record

Our agent evaluation framework separates the forecast comparison from the action and execution layers. The continuous production record shows why reporting an unfavorable result is part of a useful research account.

Sources

Related field notes