When Log Loss Exposes an Overconfident LLM Forecast
By DX Research Group · · Forecast evaluation
An illustrative comparison shows how a confident miss can outweigh several confident hits.
We examine confidence separately from classification accuracy in our forecast comparisons.
Make confident mistakes visible
A native LLM probability head can rank events sensibly while assigning extreme probabilities too readily. Accuracy alone may hide that behavior. Log loss retains the probability assigned to the outcome that actually occurred, making a confident miss expensive in score space.
For a binary outcome, use natural logarithms and loss equal to negative y times log(p), minus (1 minus y) times log(1 minus p). The scikit-learn log loss documentation states the formula and implementation details. Reporting the logarithm base and clipping convention makes the result reproducible. Mathematical probabilities of zero assigned to observed events produce infinite loss; numerical implementations may clip them.
Compare two five-question forecasts
Suppose five resolved questions contain four positive outcomes and one negative outcome. Forecaster A assigns 0.99 to the positive event for every question. Forecaster B assigns 0.80 throughout. This is a deliberately simple illustrative calculation.
A's four positive cases contribute approximately 0.01005 each. Its negative case contributes 4.60517. The mean is (4 times 0.01005 plus 4.60517) divided by five, approximately 0.92907.
B's positive cases contribute approximately 0.22314 each and its negative case contributes 1.60944. Its mean is approximately 0.50040. Both forecasters would classify all five events as positive at a 0.50 threshold, so both have 80% classification accuracy. Their probability scores are very different.
The lesson concerns the cost of confidence in a scoring rule. It is not a claim that 0.80 is always preferable to 0.99. If events truly occur at 99%, the more confident probability can be appropriate. Whether confidence is justified requires a larger, properly held-out population with a stable target definition.
Inspect the tail of the loss distribution
Alongside the mean, retain row-level losses and inspect the largest contributors. For an agent forecasting market events, a large loss might come from an extreme probability, a stale state snapshot, a malformed event, or a mistaken resolution rule. Each calls for a different repair.
Do not remove the largest losses because they are inconvenient. If a row is invalid, document a rule that applies to all systems and show the comparison with and without that rule. If the forecast was valid but wrong, it belongs in the measured result.
Track how often probabilities fall near zero and one. That frequency is a behavior diagnostic, not a verdict. A model issuing many extreme forecasts needs enough resolved examples in those regions to evaluate them. A handful of extreme hits can create reassuring accuracy while leaving the risk of confident misses almost unobserved.
Recalibration is another fitted choice
A temperature, probability map, or clipping threshold chosen after observing evaluation labels consumes those labels as fitting data. Score the resulting choice on a separate untouched sample. Preserve the original forecasts so reviewers can distinguish a changed probability mapping from a changed model.
For a frontier agentic trading study, report probability quality separately from action selection and execution outcomes. Log loss can reveal a confidence problem even when the agent's prose and classification accuracy look stable. It cannot determine the economic value of an action without the action policy, costs, and realizable outcomes.
Connect the metric to the research record
Our agent evaluation framework separates the forecast comparison from the action and execution layers. The continuous production record shows why reporting an unfavorable result is part of a useful research account.