Probability Rounding and Reproducible Forecast Scores
By DX Research Group · · Forecast evaluation
Keep scoring precision separate from the probability shown to a reader.
Round probabilities for display after computing scores from the saved values. Otherwise a harmless-looking formatting change can alter a model ranking, especially near zero or one. Our forecast trace should preserve the numerical input and the scoring implementation so a reviewer can reproduce the published aggregate independently.
A rounding discrepancy you can calculate
For an illustrative positive outcome, p = 0.0049 has binary squared error (1 − 0.0049)² = 0.99022401. Rounding to two decimal places yields 0.00 and squared error 1. The change is 0.00977599. Log loss changes more sharply: −ln(0.0049) is approximately 5.318520, while literal zero gives infinite loss. Clipping the displayed zero to 0.000001 yields about 13.815511. That clipping rule has become part of the evaluator.
Publish the arithmetic contract
Save the unrounded decimal string or a specified floating-point representation. Record the logarithm base, clipping epsilon, aggregation order, sample weights, and binary versus multiclass convention. A serialized probability should survive a round trip without silently becoming an integer percentage. A checksum of the scored input file helps bind a report to its actual numbers, although it establishes identity rather than forecasting quality.
Create a small fixture containing values close to both endpoints, a middle probability, and a failed extraction. Recompute per-row losses with an independent implementation, using a declared numerical tolerance. Sort and aggregate deterministically when extremely large panels make floating-point summation order visible. A displayed score may be rounded, but the comparison should use the preserved precision and report a difference smaller than its uncertainty plainly.
The saved record that would make this reviewable
The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.
Continue with the log-loss confidence example for the related question of precision and clipping. Our benchmark card keeps this forecast-level comparison separate from action and execution results.
A reproducibility package should contain the original probability bytes and the exact score settings. Include the rounded display column only as a presentation artifact. If an independent score differs, locate the first discrepant row before comparing aggregates. That approach makes a small precision mismatch diagnosable and prevents a formatting decision from becoming an unexplained model advantage.