Trading Agent Feedback Labels: Which Outcome Should Teach the Agent?
By DX Research Group · · Learning theories
A proposed label design separates mandate compliance, execution quality and market outcomes before trading traces enter an agent evaluation.
A profitable trade can carry a broken instruction. A losing trade can follow the mandate exactly. Before we use a trading trace as feedback, we need to decide which event the label describes. Otherwise the same word, success, quietly mixes a user instruction, a valid order, and a favorable price move.
Our proposed design gives each stage its own label. It is an evaluation method for future experiments, rather than a claim that DXAP currently trains models from these labels. The starting point is our trace-derived regression method, which joins an outcome to the mandate and state that produced it.
A Label Belongs to a Specific Event
Consider an illustrative mandate that caps exposure at $1,000. The agent proposes a $1,400 order. Policy rejects it, the account stays unchanged, and the market subsequently rallies. At the proposal stage, the size violates the instruction. At the policy stage, the rejection performs its job. At the economic stage, there is no executed position to grade. Calling the whole turn a missed winning trade would teach the wrong behavior.
We would store the proposal label beside the requested exposure, the policy label beside the rejection reason, and the execution label beside the authoritative account change. A return label would carry an explicit horizon and only attach to an actual position. The absence of a trade remains an observable action, with its own reason and timing.
Construct a Small Adjudication Set
Start with saved cases covering size violations, stale inputs, ambiguous execution acknowledgments, and valid losses. Two reviewers independently label each stage using the same rubric. Resolve disagreements by inspecting the source field that controls the judgment. A policy reviewer should be able to point to the active mandate version; an execution reviewer should identify the settlement record.
The useful measurement is disagreement by field, rather than one overall agreement score. Frequent disagreement over profit horizons suggests an incomplete economic label. Frequent disagreement over exposure suggests a mandate or snapshot problem. Those diagnoses determine whether we revise the rubric, repair the recorder, or investigate agent behavior.
Include an unknown state whenever the required record is absent. Filling an unknown with a guessed failure makes missing instrumentation look like poor reasoning. Keep the missingness rate visible across models and time periods, because a recorder change can otherwise masquerade as an improvement.
The First Test Is Label Stability
For a held-out set, show reviewers the same decision with and without its later price path. Compliance and execution labels should remain stable. Economic labels may change only when the specified horizon supplies new information. This is a concrete test for contamination between decision quality and outcome knowledge.
The operating-layer paper supplies the reason for this separation: published pre-launch interventions measure trace behavior under named changes, while the deployment record measures different events. Their scopes stay attached to the labels.
DXAP's public loop separates model proposals, outside-model policy checks, venue execution, and recorded turns. That architecture provides places to attach precise feedback. The research opportunity is to improve those judgments over successive harness versions, with reviewer consistency and downstream behavior measured before any learning claim.