Counting Failed and Missing Turns in Agent Evaluations

By DX Research Group · · Trace evaluation

Why completed-turn scores need scheduled-turn denominators and explicit missing-status reporting.

A trading agent can look reliable if its evaluator sees only completed turns. Timeouts, dropped records and failed finalization disappear before the score is calculated. We would make the turn population visible first, because a decision metric inherits the omissions of the table it reads.

Follow the population through the runtime

Define the eligible schedule or trigger events before reading model outputs. Each eligible event should resolve to a started turn, an explicit skipped event or an unresolved status. Started turns should resolve to a finalized decision, a known failure or a still-open record. Keep these counts distinct from orders and fills, since one turn can observe without trading.

The continuous record companion reports 231,638 finalized turns in its historical fleet. That word matters: a finalized-turn population establishes the population for analyses using those records. Establishing availability across every scheduled opportunity requires the schedule ledger and failure records as additional evidence.

For a new report, include the reconciliation from eligible events to finalized turns and state the observation cutoff. A turn still running at the cutoff differs from a turn that timed out hours earlier. Record when status was last checked so pending work cannot remain indefinitely hidden inside a success denominator.

A small example changes the interpretation

Suppose an illustrative experiment schedules 100 eligible turns. Ninety start, eighty finalize and seventy-six of those finalized decisions satisfy the policy rubric. The completed-decision adherence rate is 76/80, or 95%. The fraction of eligible opportunities producing a finalized, adherent decision is 76/100, or 76%.

Both figures are useful. The first describes decisions among completed turns. The second describes delivery of an adherent decision across the intended runtime population. Ten events failed to start and ten started turns failed to finalize, so an engineer has two separate failure groups to investigate. Calling the agent 95% reliable without the population qualification hides those groups.

For comparative scoring, show the missing-status rate for each candidate. If one model finishes fewer difficult scenarios, complete-case quality can favor it through selection. A proposed conservative analysis can score unresolved turns under a declared failure rule, while also publishing the complete-case result. The choice of rule should match the product commitment being measured.

Retain the empty economic decision

Our operating-layer controls paper describes buy, sell and observe actions in a bounded historical runtime. An observe action is a valid decision. A dropped turn is an availability failure. Treating both as an empty trade row merges deliberate abstention with absence of service.

DXAP explicitly records no-trade decisions. That public commitment makes the distinction meaningful for readers evaluating persistent agents. The proposed report should expose eligible opportunities, deliberate observations and unresolved records side by side. Then a user can assess decision behavior and operational continuity separately, and the research team can repair the runtime stage responsible for missing service.

Sources

Related field notes