Distribution Shift in Trading Agent Learning: Which Change Broke Transfer?
By DX Research Group · · Learning theories
A proposed shift audit separates changing markets, user mandates and agent infrastructure before judging learning transfer.
A trading agent can face a different problem tomorrow even when its model stays fixed. Liquidity changes, user mandates change, and a tool may deliver a new view of the market. If performance declines, the first task is to locate the changed distribution before choosing a learning intervention.
Our proposal treats shift as several observable changes rather than one general label. Market inputs, instructions, action constraints, and execution conditions each get their own comparison. This gives a research team a smaller question to test and reduces the temptation to retrain a model for a recorder or tool problem.
Build a Shift Ledger From Available Inputs
For each evaluation period, summarize market volatility and liquidity using definitions fixed before analysis. Record mandate types and exposure limits separately. Preserve model, compiler, tool, and policy versions. Inspect missing-field frequency and input age, since degraded observability can imitate a harder market.
An illustrative baseline contains low-turnover mandates during a quiet period. The next batch contains aggressive mandates during a volatile period and a revised liquidity tool. An aggregate decline cannot tell us which change matters. Separate the mandate cohorts, inspect the tool revision, and compare like scenarios before interpreting the model's response.
The ledger does not require an elaborate predictive model to be useful. Counts, quantiles, and missingness often identify obvious differences. Keep rare conditions visible rather than dropping them because a chart needs smoother lines.
Test One Boundary With Saved Cases
Our harness-transfer method makes the changed component explicit. For a tool shift, replay identical saved inputs through the old and new schema. For a mandate shift, hold the market snapshot fixed while substituting carefully defined instructions. For a market shift, hold the decision system fixed and evaluate a later untouched time window.
These tests answer different questions. A synthetic mandate substitution measures instruction sensitivity. A temporal holdout measures behavior on later observations under the chosen execution model. Neither alone establishes live transfer to a new venue.
Inspect action rate and policy rejection rate beside task performance. An agent that becomes inactive may appear to reduce losses simply by abandoning the intended opportunity set. An agent that increases activity may improve coverage while generating more costs. The mandate determines whether either behavior is acceptable.
Historical Scope Should Survive the Comparison
Our operating-layer paper documents a bounded 21-day market and specific pre-launch controls. Those observations supply failure cases and engineering hypotheses. Their value survives without assuming that the same rates apply in an ordinary perpetuals market.
DXAP places the current agent on Hyperliquid and describes evolving tools and models within a recorded loop. That creates several meaningful transfer boundaries from earlier research. We would name each boundary and test the relevant part instead of treating the historical deployment as a current-product performance estimate.
A useful shift audit ends with a specific next action: repair a changed field, revise a mandate test, investigate a market cohort, or run a declared model comparison. Learning becomes one possible response to a diagnosed change, with its own held-out evidence needed before claiming improvement.