Outcome Bias in Trading Agent Grading: A Blinded Test
By DX Research Group · · Learning theories
A paired review experiment checks whether a profitable result changes how reviewers judge the same trading-agent decision.
We want a trading agent to make a defensible decision with the information available at the time. Reviewing its reasoning after the price move makes that harder. A rally can make a weak thesis sound perceptive; a drawdown can make a disciplined decision look careless. Our proposed experiment measures that distortion directly.
The subject is the grader, before it is the trading model. This distinction matters whenever human reviews or automated judges supply evaluation labels. If the grader rewards hindsight, an apparently improving agent may simply become better at producing explanations that fit the revealed outcome.
Two Copies of One Decision
Build a set of saved decision packets containing the mandate, market snapshot, tool outputs, proposed action, and policy result. Create two review versions. One ends at the decision time. The other adds the realized outcome at a declared horizon. Keep every preceding byte identical and assign versions to separate reviewers.
For an illustrative case, the mandate permits a small momentum entry when liquidity exceeds a threshold. The snapshot satisfies that threshold, the action fits the size cap, and the subsequent price falls. The grader should evaluate whether the thesis used the visible evidence and whether the action respected the mandate. The later loss belongs in a separate outcome field.
A second case should deliberately contain a mandate violation followed by a gain. These contrasting cases reveal whether reviewers excuse a breach when the market rewards it. Neither case needs a spectacular return to be useful; ordinary outcomes expose the same attribution error.
Measure the Judgment That Moves
Specify rubric dimensions before collecting ratings. We would score evidence support, mandate compliance, uncertainty acknowledgment, and action consistency separately. The economic result receives its own field. Compare paired score changes for each dimension and report the number of decisions and reviewers involved.
A change in compliance ratings is especially informative because the governing instruction already existed at decision time. A change in thesis-quality ratings needs closer inspection: some rubric wording may accidentally invite reviewers to score prediction accuracy instead of decision support. That finding calls for a rubric repair before a model comparison.
Keep the model identity hidden where feasible, and randomize order. Reviewer familiarity with a favored model can influence ratings independently of price outcomes. Preserve ties and disagreement rather than forcing a winner from a small batch.
Where We Would Use the Result
Our benchmark-card method asks evaluators to disclose the complete subject and comparison. This experiment extends that discipline to the person or model assigning a grade. The trace feedback article explains how stage-linked records make a blinded packet possible.
DXAP publicly describes recorded decisions, including turns that produce no trade. Those records create a practical basis for reviewing the decision as it occurred. Recording alone leaves the grader's bias unresolved, so the paired test remains a proposed research step.
A successful experiment would produce a more stable grading rule and disclose its remaining disagreements. Improved reviewer consistency would support the evaluation method. Claims about agent improvement would still require a separate frozen comparison, and claims about returns would require economic measurements under the stated costs and market conditions.