How to validate an execution-impact model before relying on it
By DX Research Group · · Frontier research
A proposed holdout study tests impact estimates with matched order-size and market-state controls.
An execution-impact model should predict incremental cost on held-out executions under a declared reference price. We propose a validation study that separates the cost of crossing a spread, market movement during the order and the additional effect associated with order size. The proposed study has no results and supplies no return estimate.
The continuous record companion states that historical paper fills used zero slippage and zero funding. Those assumptions make the historical engine unsuitable as evidence that an impact model is accurate. The controls paper companion supplies the execution and reconciliation trace needed to define quantities and timestamps in a future validation dataset.
Define the residual before fitting
In an illustrative buy, the decision midpoint is $100, the contemporaneous best ask is $100.02 and the execution average is $100.08. The midpoint-to-fill difference is eight basis points, while crossing the displayed half-spread explains two basis points. The remaining six basis points includes depth consumption, intervening movement and other effects. Calling all six market impact requires additional identification.
Freeze an arrival midpoint and retain the complete execution schedule. Measure a matched market return over the same interval, using an unaffected reference where the design supports one. State how much of the residual remains unassigned. A model can predict implementation shortfall well without identifying the causal effect of its own order, and that distinction belongs in the report.
Partition by event time before training. Group fills from one parent intent in the same partition. Otherwise fragments of one order can appear on both sides of the holdout and make the task artificially easy. Include attempted orders that produced partial fills or no fills in a companion coverage report, because completed trades alone are a selected sample.
Test the quantities the policy will use
Evaluate signed error, absolute error and tail underestimation across size-to-depth ratios, volatility and instrument groups. Suppose predicted cost is six basis points for each of three held-out orders whose realized shortfalls are four, eight and eighteen. Errors are minus two, plus two and plus twelve basis points when defined as realized minus predicted. Mean error is four basis points; mean absolute error is sixteen divided by three, about 5.33 basis points. The large final underprediction matters even if average error appears manageable.
Compare against a simple spread-and-depth baseline using identical available information. Report the range of supported order sizes. Extrapolation beyond that range should yield an explicit uncertainty status, so a policy can reduce size or defer action.
We would validate the model's predictions before evaluating a policy that consumes them. The policy experiment would keep its own cost, fill and risk accounting. A good prediction score alone would establish a more accurate execution estimate in the studied sample; the decision value of that estimate requires the second comparison.