Choosing Bins for an LLM Forecast Reliability Diagram

By DX Research Group · · Forecast evaluation

A worked example shows why bin counts and boundaries belong beside a calibration chart.

We choose binning rules that make sample support visible in our calibration charts.

Decide what the chart must reveal

A reliability diagram compares average forecast probabilities with observed event frequencies in groups of predictions. For native LLM forecasts, its useful question is whether probabilities retain their meaning across the region an agent actually uses. A smooth-looking curve is weak evidence if its points contain very few outcomes.

The scikit-learn calibration curve documentation distinguishes equal-width bins from quantile bins with similar sample counts. It also notes that empty bins are omitted. These implementation details matter when comparing two diagrams: missing points do not establish good calibration in unobserved regions.

Build a small diagram by hand

Take ten illustrative probabilities: 0.10, 0.15, 0.20, 0.25, 0.30, 0.35, 0.40, 0.45, 0.80, and 0.90. Assign outcomes 0, 0, 0, 0, 1, 0, 1, 1, 1, and 1. These numbers are constructed for explanation and contain no DXRG measurement.

With two equal-width bins divided at 0.50, the first group contains eight forecasts. Its mean probability is 2.20 divided by eight, or 0.275. Three events occurred, so its observed rate is 0.375. The second group contains two forecasts, with mean probability 0.85 and observed rate 1.00.

With two equal-count bins of five forecasts each, the first group has mean probability 0.20 and observed rate 0.20. The second group has mean probability 0.58 and observed rate 0.80. The underlying forecasts have not changed. The visual impression changes because grouping changes what differences remain visible.

Neither version proves that a model is calibrated. Ten observations are too few for a robust curve. The second equal-width point is particularly fragile: changing one outcome moves its observed rate by 0.50. Showing only a line would conceal that sensitivity.

Pair every point with its support

Display the number of resolved forecasts in each bin and a probability histogram or equivalent count table. Give the bin edges, treatment of exact boundary values, and rounding policy. For quantile bins, tied probabilities can complicate a promise of exactly equal counts, so inspect the actual grouping rather than trusting the requested count.

Keep one primary binning rule fixed before the comparison. A sensitivity view using another reasonable rule can expose structure, but selecting the diagram that looks most favorable is another form of result selection. Retain the primary chart even if the sensitivity view is more attractive.

For agentic trading, also identify the relevant population. A diagram over every saved forecast may differ from one restricted to forecasts selected by an entry policy. The latter is a conditional population and should be labeled with the policy version and count. It cannot silently replace the full population.

Use the chart to choose the next investigation

A repeated gap between probability and event rate can motivate a held-out recalibration test. Sparse bins motivate more measurement or a narrower claim. A gap confined to one market period motivates a regime check. Those are distinct decisions.

The diagram complements proper scores and coverage reporting. It does not establish execution quality or favorable outcomes from deployed actions. Its contribution is to show where probability meaning is supported by observations, and where the sample remains too thin to say much.

Connect the metric to the research record

Our agent evaluation framework separates the forecast comparison from the action and execution layers. The continuous production record shows why reporting an unfavorable result is part of a useful research account.

Sources

Related field notes