Precision and Recall for Rare Market Event Forecasts

By DX Research Group · · Forecast evaluation

Use event counts to expose false alarms hidden by overall accuracy.

Rare-event evaluation needs positive-event support and false alarms alongside accuracy. A native LLM head can classify almost every quiet interval correctly while missing the event that defines the task. We should examine precision and recall over thresholds, retaining the event prevalence and the exact prediction population.

One thousand illustrative intervals

Suppose 10 of 1,000 questions resolve positive. A screening threshold produces eight true positives, 12 false positives, two false negatives, and 978 true negatives. Precision is 8/(8 + 12) = 0.40. Recall is 8/(8 + 2) = 0.80. Accuracy is (8 + 978)/1,000 = 0.986. A constant negative classifier has accuracy 0.990 but recall zero. The higher accuracy therefore answers a different question from useful event retrieval.

Preserve threshold and ranking evidence

The native probabilities should be retained before applying a threshold. Evaluate the threshold sequence with a declared tie convention. Average precision and a trapezoidal area under a precision-recall curve can differ; name the implementation and interpolation rule. At any chosen operating point, publish counts so a reader can reconstruct precision and recall rather than interpreting a curve in isolation.

Changes in prevalence change precision even when conditional error rates stay similar. A balanced evaluation subset has a different precision interpretation from the natural market stream. If sampling enriched positives, restore the appropriate weighting or label the result as enriched-sample performance. Sparse positives also make uncertainty large: eight detected events carry limited evidence about transfer to another period. Report time blocks and asset support before proposing a stronger event detector.

Where this enters the research workflow

The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.

Continue with base-rate timing for the related question of rare-event prevalence. Our benchmark card keeps this forecast-level comparison separate from action and execution results.

Before claiming useful rare-event detection, inspect each missed event and a blinded sample of false alarms. This can expose a timing error or a systematically mismatched event definition. Preserve the complete confusion counts for each threshold. Precision and recall then become reproducible statements about one population, rather than portable claims about any future market event.

Sources

Related field notes