Normalizing Multiclass Native LLM Forecasts
By DX Research Group · · Forecast evaluation
Make candidate probabilities coherent while preserving the mass outside the candidate set.
A multiclass forecast needs a complete, ordered set of mutually exclusive outcomes whose probabilities sum to one. For a native LLM head, candidate-token normalization also needs a receipt for probability mass outside the permitted answers. We should retain both quantities so a coherent conditional forecast does not conceal weak adherence to the answer contract.
A three-class normalization fixture
Consider an illustrative head with raw probability mass 0.40 for up, 0.30 for flat, and 0.10 for down. Candidate mass is 0.80. Renormalization gives 0.50, 0.375, and 0.125. Their sum is exactly one. If the observed class is up, multiclass log loss is −ln(0.50), approximately 0.693147. Using raw 0.40 instead would yield 0.916291 and answer a different scoring question. Keep the excluded 0.20 mass in the trace as a separate validity diagnostic.
Define the partition before extraction
Specify whether flat includes zero return, a band around zero, or prices rounded to a venue tick. The definitions must cover the entire eligible outcome space without overlap. A candidate token for a natural-language phrase can be multi-token; the probability of its first token alone may represent several completions. Use a documented extraction procedure and save token IDs, logits or probabilities, and the tokenizer version.
Compare heads using identical partitions and identical normalization rules. An invalid response can remain an invalid response even if a postprocessor can invent a distribution. Declare the fallback before evaluation, show validity rates separately, and preserve the original values. Where a head has nearly zero candidate mass, conditional probabilities can swing sharply under small logit changes. A minimum-mass rule is a protocol choice requiring its own coverage report.
Where this enters the research workflow
The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.
Continue with native versus verbal confidence for the related question of candidate probability extraction. Our benchmark card keeps this forecast-level comparison separate from action and execution results.
For a multiclass head, one acceptance fixture should deliberately permute class order. The evaluator should reject the altered mapping or reproduce its known changed loss. Include a candidate-mass collapse case as well. Passing these fixtures establishes that the normalization contract is implemented consistently, with conditional forecast quality and answer validity still reported as separate measurements.