Evaluating Trading-Agent Abstention Without Rewarding Silence
By DX Research Group · · Trace evaluation
A coverage-aware method for comparing no-trade behavior with conditional decision quality.
An agent that never trades avoids some execution losses and also supplies little evidence of trading judgment. An agent that trades every turn supplies many decisions and can spend away a weak signal. We need to measure the choice to act alongside the quality of the actions selected, especially when comparing models with different activity levels.
Define the opportunity set first
Use eligible decision scenarios as the denominator. Record whether the agent proposes an entry, exit, modification or deliberate observation. For a focused entry study, define coverage as the fraction of eligible entry opportunities receiving an entry proposal. Keep exits and protective actions in separate classes, because an abstention from entry can be sensible while an omitted required exit can violate the mandate.
The continuous record finds no demonstrated directional edge in its historical fleets. That historical null is useful context for designing tests that resist rewarding activity by itself. It also prevents a later no-trade-heavy replay from being read as evidence that profitability has improved merely because fewer risky actions appear.
An illustrative pair has 100 shared scenarios. Model A enters 20 and has a positive result on 12, a 60% positive-result fraction among entries. Model B enters 80 and has a positive result on 44, a 55% fraction. These counts establish conditional behavior, rather than a winner. Position sizing, loss magnitude, execution costs and the outcomes of the skipped scenarios can change the economic comparison.
Inspect the decisions that disappear
For each candidate, report action coverage and conditional score on the same saved scenarios. Then evaluate the difference on scenarios where both act, where only one acts and where both observe. This partition reveals whether an apparent quality gain comes from selecting an easier subset or improving decisions in shared situations.
A proposed abstention rubric should recognize explicit reasons grounded in the mandate or captured state. Missing input, breached limits and insufficient signal are different reasons for observation. A timeout belongs to operational completion metrics. It should retain its failure status rather than becoming a no-trade decision.
Keep the cost of acting explicit. A simulated scorer should include the stated fee and execution assumptions. A scorer that values every correct direction equally can favor frequent small actions whose simulated gain disappears after realistic costs. Conversely, a scorer that penalizes all missed upward moves encourages hindsight trading. Declare the decision objective before observing outcomes.
A research question for persistent agents
Our operating-layer controls paper documents an observe action and meaningful behavior differences from activity settings. That supports evaluating abstention as an intentional part of the action space. It also reminds us that coverage can be controlled by the harness configuration rather than arising solely from model confidence.
DXAP publicly includes no-trade turns in its record. A useful next evaluation would compare coverage under identical mandates, with shared scenario inputs and a declared risk policy. Publish conditional decision scores beside coverage and completion. That combination helps distinguish disciplined selectivity, inactive configuration and runtime failure, while leaving profitability claims dependent on the measured economics.