Calibration Slope and Intercept for a Native LLM Forecast Head

By DX Research Group · · Forecast evaluation

Use a logistic diagnostic to separate shifted probabilities from excessive confidence.

A calibration intercept diagnoses a common shift in forecast odds; a calibration slope diagnoses how strongly forecast odds vary relative to outcomes. These two numbers give our native LLM evaluation a more specific repair hypothesis than a single reliability plot. Their meaning depends on a defined population and a time-separated diagnostic fit.

Transform an illustrative probability

Write the diagnostic as q = sigmoid(a + b × logit(p)). Perfect logistic calibration corresponds to a = 0 and b = 1. For an illustrative forecast p = 0.80, logit(p) = ln(4), approximately 1.386294. With a = 0 and b = 0.5, q becomes sigmoid(0.693147), exactly 2/3. The slope pulls an extreme probability toward the center. With a = ln(2) and b = 1, odds of 4 become odds of 8, so q = 8/9. The intercept changes odds by a common multiplier.

Fit the diagnostic without reusing the test

Fit the outcome regression on a designated diagnostic or calibration period. Keep its coefficients fixed on a later score period. Estimating a and b and evaluating their correction on the same labels gives an optimistic result. Report coefficient uncertainty, available positive and negative counts, and whether identical market episodes generate repeated observations. Complete separation can make an ordinary logistic fit unstable or unbounded.

A slope below one suggests probabilities are too dispersed under this particular logistic description. The mechanism could involve temperature, changing target prevalence, or mixed asset populations. Check stratified residuals before treating a temperature adjustment as a complete explanation. Native candidate-token probabilities require a fixed candidate mapping; verbal confidence requires its own extraction contract. A stable transformation cannot repair a forecast pointed at the wrong event.

From fixture to a held-out result

The scikit-learn probability calibration provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.

Continue with calibration diagrams with sample support for the related question of fitted probability mapping. Our benchmark card keeps this forecast-level comparison separate from action and execution results.

Before adopting a logistic correction, examine coefficient stability across adjacent chronological windows. A wildly changing intercept may point to target shift rather than a permanent head defect. Publish the uncorrected and corrected held-out scores together. The useful outcome is a localized calibration diagnosis with enough evidence to decide whether another fitted transformation is warranted.

Sources

Related field notes