Timeout Censoring Changes What an Agent Evaluation Can Estimate

By DX Research Group · · Trace evaluation

Distinguish deadline failure from an unobserved completion time and preserve timed-out cases in quality reporting.

A timed-out agent run establishes that the evaluation deadline was missed. It may leave the eventual completion time unobserved. We would report both facts, because dropping timed-out cases makes the surviving latency and quality sample look better precisely when slow cases are hardest.

Ten requests and one hidden tail

Consider an illustrative ten-request batch with a ten-second deadline. Eight requests finish in two seconds each. Two time out. The completed-only mean is two seconds. If both timed-out requests would have completed after ten seconds, the full mean exceeds (8 × 2 + 2 × 10)/10, or 3.6 seconds. The eventual mean cannot be recovered from these observations alone.

If the runtime cancels work at the deadline, we have observed a deadline-defined terminal outcome rather than merely stopped watching a continuing request. If work continues elsewhere, its later receipt should preserve the original deadline outcome and the observed eventual completion. These are different measurement mechanisms.

The Google SRE monitoring chapter distinguishes request latency and errors. Our agent-specific proposal adds deadline admission because a late market decision may have lost its intended use even when the inference eventually succeeds.

Give quality a deadline-aware denominator

Suppose the eight completions contain six correct proposals. Their conditional accuracy is 6/8, or 75%. The fraction of all ten assigned cases delivering a correct proposal within the deadline is 6/10, or 60%. The unresolved cases should receive an explicit timeout label; assigning them an invented semantic answer would erase uncertainty.

We would run a proposed paired fixture at ten and twenty seconds, using the same cases and a documented cancellation policy. Compare deadline-qualified proposal quality, eventual completion and resource cost. A longer deadline may reveal capable but slow behavior. For a trading mandate that demands review before a price expires, that recovered answer may still arrive too late for the specified task.

The continuous record companion distinguishes captured replay behavior from market performance. The operating-layer controls companion centers the path to settlement. Timeout reporting should preserve the stage where observation ended: model generation, tool response or venue reconciliation.

Censoring can also depend on case difficulty. Long tool chains and complicated restrictions may time out more often, so completed cases cease to represent the assigned population. Stratify deadline misses by intended task class and tool-chain depth, with counts rather than speculative adjustment weights.

The useful release artifact contains the assigned case list, deadline rule, cancellation receipts and observed finish times. It can support a completion lower bound and an on-time quality estimate. This fixture remains a proposed evaluation method. Its central decision is concrete: keep timeout cases visible, and choose the latency estimand that matches the agent's actual deadline.

Sources

Related field notes