Separating Execution Errors from Prediction Errors in Agent Trading

By DX Research Group · · Trace evaluation

A stage-by-stage diagnosis of trading losses that distinguishes forecasts, sizing, policy and venue execution.

A losing trade tells us that the realized position lost money. It does not, by itself, identify whether the agent predicted the wrong direction, chose the wrong size or failed to execute its intended protection. We need the decision trace and an explicit outcome definition to decide what to repair.

Attach the error to a stage

Start with the proposed action and its captured inputs. Then inspect policy acceptance, submitted payload and venue receipt. Compare intended position state with reconciled position state. A directional thesis can be correct while an oversized position liquidates before the move completes. A suitable protective order can be proposed while submission leaves the position unprotected.

The continuous record companion describes a historical window in which a trigger-quota error left 24 of 35 successful opens without their protective trigger. That is a specific execution-path finding. The same record reports a broader historical directional-edge null. Those observations establish different problems and should lead to different investigations.

A proposed diagnosis table would retain forecast direction, chosen exposure and intended protection alongside actual fills and protection status. Outcome labels can then point to the failed stage. More than one stage can contribute to a loss, so keep the measured events visible rather than forcing every position into one exclusive cause.

Work through an intended stop

Consider an illustrative entry at 100 with an intended stop at 98. The entry fills, the stop request fails, and the position later exits at 95. Ignoring fees, the intended stop loss would be 2% if a valid fill at 98 were available. The realized price loss is 5%. The three-percentage-point difference is a hypothetical execution comparison under that fill assumption.

It becomes evidence of avoidable loss only after checking whether the proposed stop was valid, whether the venue could have executed it and what slippage or gap conditions applied. A market moving directly from 100 to 95 may make the assumed fill at 98 unrealistic. Preserve both the missing protection receipt and the execution assumptions in the analysis.

A separate prediction evaluation should ask whether the original thesis outperformed an appropriate baseline over its declared horizon. Using the intended stop to relabel direction would mix risk management with forecasting. Similarly, blaming all losses on execution can hide a model whose entries have no measured directional advantage.

Repair the observable failure first

Our operating-layer controls account connects model output, policy and chain outcome in a trace. That structure gives an engineer a practical route from a loss to the payload or runtime stage that failed.

DXAP publicly separates model proposals from venue execution. A proposed review should preserve that separation in its metrics: decision quality on captured state, adherence to the mandate, execution completion and realized economics. Repairing a quota failure can improve delivery of intended protection. Establishing better forecasts requires the separate matched evidence that answers the prediction question.

Sources

Related field notes