Counterfactual Trading Replay: Check Action Support Before Comparing Policies

By DX Research Group · · Learning theories

A proposed replay audit identifies which alternative trading actions the saved data and simulator can actually evaluate.

A replay can show the market path after a saved decision. It can also tempt us to assign outcomes to actions that the record never supported. A different trade size may change the fill, a different exit may require unavailable order-book data, and a rejected proposal may never have reached the venue.

Before comparing policies, we would write down the supported action set for each case. Supported means the data and declared simulator can evaluate that action under specified assumptions. This is an engineering requirement of the proposed comparison, rather than a guarantee that a replay reproduces live trading.

Start With an Action That Changes the Data Requirement

Imagine an illustrative record with a best bid and ask but no depth. A small market order and a large market order may receive very different prices. The quote supplies a reference price; it leaves the depth-dependent fill unresolved. Assigning both actions that quote silently favors the larger trade.

We would classify the small-order outcome as conditional on an explicit fill assumption and the large-order outcome as unsupported unless additional depth or a defensible impact model is available. The report should count both categories. Excluding unsupported actions without disclosure can make a policy appear better by selectively removing its difficult decisions.

A no-trade action usually requires fewer execution assumptions, yet its opportunity cost still depends on a declared alternative and horizon. The mere fact that the price rose later does not determine whether an admissible trade was available at the decision time.

A Support Table for Each Replay Case

The table would bind action type, size range, input availability, execution assumptions, and confidence category. It should also identify whether the action was actually proposed, accepted by policy, or executed. Those are separate branches of the trace.

Our execution and settlement framework describes why an acknowledgment and an authoritative position change need separate treatment. Our trace feedback method keeps the resulting economic event attached to its original mandate and state.

For a counterfactual policy, record how often its chosen actions fall outside support. Compare policies on common supported cases, then disclose the excluded fraction for each policy. A highly selective common subset can answer a narrow question while leaving the broader policy comparison unresolved.

Preserve the Counterfactual Boundary

A saved tape cannot include the market response to every hypothetical order. In a setting where agents move the market, the assumption that all policies see the same future path becomes especially consequential. Test sensitivity to execution and impact assumptions where the data permits, and retain unknown outcomes when it does not.

We would hold the entry and exit policies fixed when comparing model forecasts. If the policy is the subject, keep the forecasting inputs fixed instead. Changing both during replay produces a system comparison with wider attribution limits.

DXAP publicly describes model proposals, outside-model policy checks, Hyperliquid execution, and recorded turns. That full path offers useful observables for replay design, while counterfactual fills remain a separate research problem.

The contribution of an action-support audit is a clear denominator: which proposed alternatives were evaluated, under what assumptions, and which remain unknown. Better-looking simulated returns become interpretable only after that denominator is visible.

Sources

Related field notes