Keep Execution Outcomes at the Right Training-Label Boundary
By DX Research Group · · Learning theories
A failed venue submission should teach the stage responsible for the failure.
Execution success is an appropriate label for an execution component, but it can be a misleading label for the model's decision. An economically sensible, authorized proposal may fail because of a transient submission error. We would bind each learning target to the stage that could have changed the outcome.
Our trace-feedback framework links diagnosis to an intervention boundary. The operating-layer paper records the path from mandate through policy and settlement. The proposed label audit below keeps those stages distinct, while reward-hacking research motivates care when a proxy substitutes for the real objective.
One failed trade, several possible lessons
In an illustrative set of twenty failed submissions, eight have malformed model proposals, seven have valid proposals followed by network timeouts, and five have valid proposals rejected after market conditions changed. Calling all twenty bad decisions makes model failure appear 100%. At the proposal-validity boundary, the count is eight of twenty, or 40%.
That percentage describes this selected failed-submission set, not the failure rate across all decisions. A full evaluation needs the total number of eligible proposals and submissions. The seven timeout cases may have unresolved economic outcomes until reconciliation, because a lost response can coexist with a successful venue action.
For the five market-change cases, inspect whether the decision input was stale at generation or whether the change occurred after a reasonable decision. The former can motivate data or timing repairs; the latter may motivate an execution guard or an expiry rule. A negative label without this distinction can teach the model to avoid valid actions whenever infrastructure is unreliable.
Bind labels to controllable inputs
We would create separate proposal-validity, mandate-compliance, submission, fill, and reconciliation fields. Each label cites the evidence record and its observation cutoff. Unknown execution outcomes remain unknown until resolved; they receive neither fabricated success nor automatic failure labels.
A paired fixture then keeps the exact model proposal fixed while changing the execution response between a confirmed fill, a confirmed rejection, and a timeout. Proposal and mandate labels should remain constant. Execution labels should follow the evidence. Another pair changes the proposed quantity across the mandate limit while preserving the same simulated network outcome; compliance should change even if submission happens to succeed.
If an end-to-end objective intentionally penalizes all failed turns, retain that aggregate alongside the component labels. The system-level penalty can guide runtime optimization without pretending each component caused every failure. A learned decision policy and a deterministic execution adapter have different action spaces.
The final artifact is a stage-by-stage confusion table and an unresolved-outcome count. It identifies which examples can supervise the model, which belong to execution engineering, and which need additional receipts before they can teach either component.