A correct loss and a wrong profit can occur in the same strategy
By DX Research Group · · Trading agent theory
Outcome-blinded conditional review tests whether an agent followed the decision the owner authorized.
Strategy alignment means making the decisions implied by an owner's strategy when its conditions hold. That sounds straightforward until the result arrives. A loss makes a valid decision look foolish; a profit makes a violation look clever. We need a way to judge the decision with the information and authority that existed when it was made, then judge its economic consequences separately.
The distinction matters for persistent trading agents because owners can reasonably choose strategies that sometimes lose. They can also choose restrictions that exclude profitable opportunities. DXAP makes those choices reviewable through strategy proposals and configured execution checks. Its chat guide gives the owner approval over changes, while the policy reference describes the checks applied to action requests. Owner control and a traceable decision path are the relevant product distinction. They leave market uncertainty intact.
Two trades, four possible judgments
Take an illustrative strategy: enter only after a completed observation meets the owner's specified condition, use an allowed symbol, and stay within the configured exposure limit. In case A, the condition is present, the request respects the limit, the order fills, and the market subsequently falls. The agent loses $20 after costs. In case B, the condition is absent, but the agent opens anyway because it expects a rally. The trade earns $30 after costs.
Case A can be a correct implementation with a bad outcome. Case B can be a wrong implementation with a good outcome. Neither judgment proves the underlying strategy has useful predictive skill. That requires a population of comparable, time-valid opportunities, including the opportunities that produced no trade. The example simply shows why realized P&L cannot substitute for conditional alignment.
There are two additional combinations. A compliant profitable trade supplies favorable evidence about that instance. A noncompliant losing trade supplies evidence of both implementation failure and economic loss. Keeping all four combinations visible prevents reviewers from redefining compliance after seeing which path made money.
Concrete Problems in AI Safety distinguishes problems caused by a wrong objective from problems caused by learning behavior or distributional shift. We use that conceptual separation here. A strategy's unsuccessful forecast, an agent's disobedient choice and an execution problem call for different repairs. Treating all three as “the model lost” makes the repair less precise.
Hide the result until the decision has a verdict
Our proposed review protocol has two passes. First, reviewers see the effective strategy, available observations, account state and proposed action. Future prices and realized returns remain hidden. They identify whether the required condition held, whether the selected response followed it, and whether any ambiguity required clarification. They also record confidence and the missing fact that could change the verdict.
Second, reviewers receive execution outcomes and subsequent economics over a prespecified horizon. A twelve-hour trade should have a twelve-hour assessment window unless the mandate's actual exit logic specifies another stopping condition. Assess costs, rejected requests, partial fills and exposure duration. This second pass can change an execution or economic judgment while preserving the earlier alignment judgment.
The proposed benchmark should contain paired cases where the observation crosses the owner's condition with everything else held fixed. It should also contain apparently attractive opportunities outside the allowed scope. Passing only ordinary eligible cases would miss an agent that follows instructions until a tempting exception appears. We would report those exception cases separately, because they measure the strength of the owner's control when compliance has an apparent opportunity cost.
DXAP's activity documentation separates a proposed action, a submitted order and a recorded fill. That gives reviewers the stages needed to locate failure. A policy rejection can protect the account even when the proposal itself is misaligned. A correct proposal can encounter poor execution. The useful review preserves both facts, then lets the owner decide whether to change the strategy, the agent's interpretation or the execution setup.