How to test parity between paper and live trading fills

By DX Research Group · · Frontier research

A proposed execution study measures fill, funding and margin differences before using paper outcomes to assess an agent.

A paper agent and a live agent can issue the same instruction and experience different outcomes. Our continuous record states that historical paper fills used live marks with zero slippage and zero funding, while its maintenance-margin placeholder differed from real venue margining. We would test those differences directly before using paper returns to judge current agent quality. This is a PROPOSED execution experiment, with a passive shadow stage before any separately authorized live orders.

Pair the instruction before comparing the result

Each observation would bind a proposed order to its market snapshot, order type, quantity and decision timestamp. The passive stage estimates what the paper engine would record and what a venue-faithful execution model would predict from available quotes and depth. Where authorized live observations already exist, the record can pair the actual acknowledgement and fills with that same proposal. A pairing must preserve rejected orders and partial fills, which are easy to lose when selecting only completed trades.

We would measure latency from decision to submission and from submission to acknowledgement. Fill price alone cannot distinguish a paper-price convention from an execution delay. A moving market can create a gap even when the order fills at a competitive price after it arrives. The report should attribute that gap using the same reference timestamp across the two paths.

Decompose the economic difference

An illustrative position has a $10,000 entry notional. A 5-basis-point worse entry costs $5 before any exit difference. If the exit is also 5 basis points worse on the same notional, the two price differences total $10. Fees and funding add separately under the recorded position history. This arithmetic demonstrates the accounting convention and supplies no estimate of typical DXAP execution.

We would report spread crossing, depth-related price impact, explicit fees and funding as separate components. Liquidation differences deserve their own analysis because changing margin assumptions can change whether a position survives at all. A simple cost deduction from a paper P&L series cannot repair a position that the live venue would have liquidated earlier.

Use parity to choose the next question

The paired sample should span order sizes, liquidity groups and volatile intervals. We would disclose coverage, including observations without adequate depth or reliable timestamps. Day-clustered intervals are appropriate when execution gaps expand together during a market shock. A small average difference can hide a consequential tail precisely where a risk-sensitive strategy needs reliable behavior.

The operating-layer controls paper treats settlement and reconciled state as part of the trading-agent record. DXAP describes execution through a user-authorized Hyperliquid agent wallet and logged decisions. That gives an identifiable production path to compare with a simulator. The resulting study would determine where paper evidence transfers, where it needs correction and which failures require a richer execution model before an evolving harness can earn a performance claim.

Sources

Related field notes