On-Policy Trading Data: Track Which Agent Produced Each Example

By DX Research Group · · Learning theories

A provenance design for distinguishing current-policy trading trajectories from historical traces before evaluating a learning proposal.

A saved trading trajectory carries the decisions of the agent that generated it. If we change that agent's model, mandate compiler, tools, or policy checks, the old trajectory represents a different decision process. Calling every historical trace on-policy erases the information needed to assess a learning proposal.

Here we use on-policy in a narrow sense: the recorded actions were sampled from the declared policy whose behavior the experiment is studying. For a language-model trading agent, that policy includes the relevant harness components and sampling settings, rather than just a model name. This is a proposed data-design method, with no claim about an active DXAP training route.

The Minimum Provenance Row

A trajectory record should bind a decision to the model version, compiled mandate version, input snapshot, tool schema, and action rules. Store the sampling settings and acquisition time as well. A model endpoint name that can change behind the same identifier needs a dated resolution or an explicit unknown version.

For an illustrative example, batch A uses model M1 with compiler C1. Batch B uses M1 with C2, which moves a risk instruction earlier. Those records share a model and differ in the decision policy. Combining them under M1 would conceal the intervention we need to evaluate.

The distinction is grounded in our published controls study. Its fee-placement intervention changed trace behavior with the model fixed. The result shows why a compiler revision belongs in provenance even when the weights stay constant.

Drift Can Enter Through a Tool

Imagine that an order-book tool begins returning an additional liquidity field halfway through collection. The agent now observes information unavailable in earlier examples. A timestamp alone identifies the period but may leave the cause unclear. Bind the tool schema and the actual returned payload to the turn.

Similarly, an execution adapter change can alter which proposals become positions. Preserve proposed action, policy decision, submitted payload, and reconciled outcome separately. Rejected proposals reveal a different part of the policy than completed fills. Missing these branches biases a dataset toward successful execution paths.

We would partition an evaluation set by the full declared provenance tuple and count the trajectories in each partition. Before aggregating, inspect whether action rates or rejection reasons differ by version. Small partitions should retain their uncertainty instead of being treated as precise evidence of progress.

The Reuse Decision

Older examples can still support regression tests, fault diagnosis, or a declared offline study. Their value depends on the question. Our harness-transfer method describes how to test across a component boundary rather than assuming a past result transfers.

A proposed learning run would document which examples came from the current policy, which came from earlier versions, and what correction method was chosen for the latter. It would also freeze a separate evaluation population before tuning. Provenance makes those decisions auditable; it does not solve distribution mismatch by itself.

DXAP describes model selection within a harness and recorded turns. As those components evolve, version-linked records can distinguish product iteration from a learning claim. The practical research contribution is a dataset whose rows retain the identity of the decision system that produced them.

Sources

Related field notes