Separating Calibration from Discrimination in Agent Forecasts

By DX Research Group · · Forecast evaluation

Two constructed forecasters show why probability meaning and event ranking require separate checks.

We separate probability meaning from event ranking when planning our next evaluation.

Ask two different questions

When comparing native LLM forecasts, calibration asks whether probabilities agree with event frequencies. Discrimination asks whether higher forecasts tend to belong to events that occur. A system can perform well on one question and poorly on the other. Treating either as a complete model verdict obscures the reason a forecast might be useful.

The scikit-learn probability calibration guide explains reliability diagrams and notes that proper scoring rules reflect more than calibration alone. This is the narrow methodological point used here. The examples below are constructed arithmetic.

A calibrated constant can rank nothing

Imagine 100 questions, exactly 30 of which resolve positively. A constant forecaster assigns probability 0.30 to every question. Over this sample its single forecast group has average probability 0.30 and observed frequency 0.30. That group is calibrated in the sample.

Its binary Brier score is (30 times 0.70 squared plus 70 times 0.30 squared) divided by 100. That is (14.70 plus 6.30) divided by 100, or 0.21. However, the forecaster gives every question the same score. It provides no ranking among questions. Under the conventional treatment of ties, its ROC AUC is 0.50 when both classes are present.

This shows why discrimination remains a separate question even when a calibration curve lies close to the diagonal. A base-rate forecaster can be a strong and necessary reference without identifying which individual events will occur.

Excellent ranking can have weak probability meaning

Now construct four questions with outcomes 0, 0, 1, and 1. Assign probabilities 0.60, 0.70, 0.80, and 0.90 respectively. Every positive question has a higher probability than every negative question, so all four positive-negative pairs are correctly ranked and AUC is 1.00.

Its Brier score is (0.36 plus 0.49 plus 0.04 plus 0.01) divided by four, or 0.225. The negative questions received probabilities above 0.50 despite not occurring. Four observations cannot diagnose a population calibration curve, but the example demonstrates that perfect ranking does not force a perfect probability score.

A monotone probability transformation can preserve ranking while changing confidence. This makes a fitted calibration layer a plausible intervention when ranking is retained but probability meaning is weak. It remains an intervention to test, not an improvement to presume.

Keep the intervention identifiable

Save the raw model probabilities, the mapped probabilities, and the mapping version. Fit the mapping on a calibration sample and compare both versions on untouched questions. If the underlying model, prompt, target, and mapping all change together, the resulting score cannot isolate the benefit of calibration.

For an agentic trading evaluation, retain the action policy independently as well. A threshold change can alter the chosen questions without altering either probability quality or ranking on the full eligible population. Report selected-subset metrics as conditional results rather than replacing the original comparison.

Choose the next step based on the actual failure. Weak discrimination calls for investigating information, target, or representation. Weak calibration with retained discrimination calls for a probability-mapping experiment. Failure of realized action economics calls for policy and execution analysis. Each diagnosis requires evidence from its own layer of the agent system.

Connect the metric to the research record

Our agent evaluation framework separates the forecast comparison from the action and execution layers. The continuous production record shows why reporting an unfavorable result is part of a useful research account.

Sources

Related field notes