Confidence Intervals When Agent Outcomes Are Sparse

By DX Research Group · · Forecast evaluation

A zero-event example shows why a small observed rate can still leave substantial uncertainty.

We put the denominator beside our outcome rate before interpreting a sparse sample.

Zero observations are not a zero probability

An agent evaluation can contain few resolved events, few selected actions, or few operational failures. If no event occurs in a small sample, reporting an observed rate of zero is correct. Treating that number as a tight estimate of the underlying probability is a different and often unsupported claim.

This note concerns a binomial event proportion under independent trials with a common probability. Profitability uses a different target; applying this interval to dependent market labels requires a separate justification. The statsmodels proportion interval documentation lists Wilson and exact interval options. Method choice and assumptions belong beside the number.

Calculate the zero-event bound

Suppose zero events occur in 20 independent trials. The observed rate is 0 divided by 20, or zero. To find a one-sided 95% upper bound by exact binomial inversion, solve (1 minus p) raised to the twentieth power equal to 0.05. The solution is p equal to 1 minus 0.05 raised to the power 1/20, approximately 0.1391, or 13.9%.

For zero events in 100 independent trials, the analogous upper bound is approximately 0.0295, or 2.95%. These are illustrative arithmetic results. Applying them to DXRG reliability requires actual outcomes and evidence supporting independence.

The widely used approximation of three divided by the sample size gives 15% for 20 trials and 3% for 100. It is close here, but the exact expression states what is being calculated. A one-sided 95% bound is also different from the upper endpoint of a two-sided 95% interval. Label the sidedness explicitly.

Define the denominator before celebrating a low rate

If the outcome is an invalid action, are trials attempted actions, completed tool calls, or agent sessions? One session may contain many correlated attempts. A model that rarely acts can show few invalid actions while providing limited coverage. Report both the denominator and the eligible opportunities.

If the outcome is a forecast event, unresolved questions cannot be treated as negatives. If fast-resolving questions enter the sample first, the current resolved population may be biased toward particular event types. Show unresolved counts and the observation cutoff.

An observed zero can also result from a filter. For example, excluding failed traces before counting failures changes the question being answered. Preserve exclusions and reasons so the rate can be reconstructed.

Choose a next action from uncertainty

A wide interval motivates a narrower claim or more relevant observations. It can motivate improving the measurement design before considering more model training. If the denominator is wrong or dependence is severe, adding rows from the same episode may not resolve the uncertainty that matters.

For dependent agent outcomes, use an uncertainty analysis matching the unit of evidence, such as session or episode, and state its limitations. An exact binomial calculation is only exact under its statistical model. The precision of the arithmetic cannot repair a mismatched model.

In agentic trading research, sparse selected outcomes should be reported alongside full forecast coverage and action-policy rules. A small observed failure rate or favorable count is a scoped observation. Interpret that observation with its eligible population, artifacts, and assumptions.

Connect the metric to the research record

Our agent evaluation framework separates the forecast comparison from the action and execution layers. The continuous production record shows why reporting an unfavorable result is part of a useful research account.

Sources

Related field notes