Risk and Coverage for Selective Agent Forecasts

By DX Research Group · · Forecast evaluation

Evaluate confidence-based forecast selection across retained fractions.

A selective forecast can reduce error by answering fewer questions. The risk-coverage curve makes that exchange visible: coverage is the fraction retained, and risk is the chosen average loss on those retained questions. Our native LLM benchmark should compare complete curves on the same eligible questions before judging a confidence filter.

An illustrative ranked sequence

Take five forecasts ordered by a saved confidence score. Their classification errors are 0, 0, 1, 0, and 1. Retaining the first two gives coverage 2/5 = 0.40 and risk 0/2 = 0. Retaining the first three gives coverage 0.60 and risk 1/3. Retaining four gives coverage 0.80 and risk 0.25. At full coverage risk is 2/5 = 0.40. Conditional risk can rise and fall across individual prefixes; smooth improvement is an empirical property.

Choose the selector before the outcomes

The selector may use native entropy, candidate margin, input freshness, or an independently trained uncertainty model. Save its value before resolution. Choosing which questions to retain after seeing their loss creates an oracle curve. If two models have different confidence scales, compare at matched coverage rather than applying the same numeric threshold. Resolve ties deterministically and report the resulting retained counts.

A forecast abstention and an execution no-trade decision belong to different layers. A head can forecast every question while a policy acts on few of them. Conversely, a missing forecast may leave the runtime without required information. Show these denominators separately. If a screening rule removes volatile episodes, the remaining risk describes an easier population; a curve plus the retained asset and time composition makes that limitation inspectable.

Where this enters the research workflow

The Selective Classification for Deep Neural Networks provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.

Continue with action abstention and coverage for the related question of forecast eligibility. Our benchmark card keeps this forecast-level comparison separate from action and execution results.

Compare selectors at several predeclared coverage levels, including full coverage. Retain the confidence ties and the composition of every prefix. A selector whose apparent improvement depends on one favorable episode needs a time-block sensitivity report. The resulting curve measures a forecasting tradeoff; a separate policy experiment determines whether that tradeoff improves a complete agent workflow.

Sources

Related field notes