A return paper and a profitable agent require different receipts
By DX Research Group · · Market predictability
A four-question translation audit connects academic predictability to the full decision system without importing a paper’s performance claim.
Published return predictability can justify a concrete agent experiment. The strongest translation begins by preserving the paper's question, then measuring what changes as the forecast enters a real decision system. We use four questions to organize that translation: what was predicted, what could be chosen, what was executed and what remained after costs?
The distinction follows directly from the research literature. Gu, Kelly and Xiu test historical monthly U.S. equity forecasts and portfolios. Kelly and Xiu's survey treats trading costs as part of portfolio design. Mosaics of Predictability studies heterogeneity in U.S. equities. Each supplies useful evidence for its own design. Applying a similar idea to a persistent agent adds an account, an action policy and an execution environment.
Follow one forecast through the decision
Suppose an illustrative agent forecasts a positive return for an eligible instrument over its next holding period. The estimate arrives after a market-data delay. The account already has exposure, the owner has a size limit and the venue offers a spread. Every subsequent choice changes the economic meaning of the original forecast.
First ask whether the forecast predicts the intended label on time-valid inputs. Save the target definition, input availability timestamps and model output before observing the outcome. Compare the candidate with a defined forecast benchmark on the same future cases. This assesses predictive skill independently of how frequently the agent trades.
Next ask what action the forecast induces. Keep the entry, sizing and exit policy explicit, including no-trade decisions. A positive expectation can be too small to warrant another order. Compare alternative forecasts through a fixed policy when estimating the value of prediction. Compare alternative policies through fixed forecasts when estimating the value of decision design. Changing both at once leaves attribution unresolved.
The third question concerns execution. A proposed order can be rejected, partially filled, delayed or filled at a different price from the simulation. Record actual submissions and outcomes alongside the intended action. This stage explains whether the agent obtained the exposure the policy requested, and what it paid to obtain it.
Finally, reconcile the resulting account ledger. Include the relevant venue charges, execution friction and financing or funding flows, plus the agent's operating cost where the comparison includes it. Net results answer whether the whole decision system produced useful economics over the tested period and capital scale.
Keep a matched comparison at each question
Our proposed translation audit uses a common set of saved opportunities, including cases where both candidates abstain. At the forecast stage, compare paired errors. At the policy stage, replay both candidates against the same state and mandate. At the execution stage, compare intended exposure with obtained exposure. At the economic stage, report the realized or simulated ledger and label which it is.
A paper-style result can survive one stage and fail at another. Forecast improvement might remain intact while the policy churns too often. The policy might work in a simulation while venue minimums make its small orders infeasible. Execution might be reliable while net results remain indistinguishable from the baseline. Each outcome identifies a specific development decision.
Simulation has a useful role here. It can test candidate policies with the same saved market opportunities and allow controlled changes to cost assumptions. Its fill model and state transitions need validation against the intended venue. A replay lacking actual depth or realistic margin behavior gives evidence about that modeled environment. The next receipt should establish which assumptions survive observed execution.
Population and horizon transfer deserve an explicit checkpoint. Monthly stock ranking, hourly perpetual-futures positioning and event-contract probability forecasting involve different targets and institutional rules. Shared model terminology creates no measured bridge between them. An agent researcher can borrow the hypothesis while designing new labels and a new holdout for the intended market.
We would publish the audit as four connected results, with uncertainty and failure cases attached to each. A forecasting null can coexist with a useful runtime repair. Better policy behavior can coexist with unproven net profitability. Preserving those distinctions lets academic findings guide ambitious engineering without turning a historical paper portfolio into an agent performance claim.