The best forecast and the best portfolio solve different problems
By DX Research Group · · Market predictability
Kelly and Xiu’s survey explains why trading costs and persistent signals change the objective for financial machine learning.
Financial machine learning gives researchers a way to estimate complicated return relationships and a way to build investment policies. Choosing between those uses changes what success means. A model minimizing forecast error can behave differently from a model choosing a cost-aware portfolio, even when both see exactly the same market information. We think this distinction belongs near the start of an agent research program.
Kelly and Xiu's Financial Machine Learning, issued as NBER Working Paper 31502 in July 2023, is a survey rather than a single return experiment. It covers prediction, risk-return relationships and portfolio choice, including trading costs and reinforcement learning. Readers should attach any numerical finding to its underlying study, population and horizon rather than treat the survey as one pooled performance result.
Its section on trading costs describes a central difficulty: models searching for forecastable returns can discover patterns that survive precisely because exploiting them is costly. The authors discuss integrating economic restrictions and trading costs into portfolio learning. Their review also explains why a signal's persistence affects allocation: the cost of entering a position can be spread across the periods over which the signal remains useful.
A forecast can be right and an action can be wasteful
Consider two illustrative signals. The first predicts a brief movement requiring repeated rebalancing. The second predicts a slower change while allowing a position to remain open. Under a forecast-error objective, the first may look superior. Under an objective that accounts for repeated spreads and fees, the second may create more useful investment opportunities.
This example also explains why one universal “AI trading accuracy” measure conceals the research question. Accuracy about what, over which horizon, and under which feasible actions? A correct direction forecast can accompany a return too small to pay for execution. An imperfect forecast can still improve a diversified allocation relative to a fixed baseline.
For an agent, the available action set further depends on current positions and the owner's mandate. A forecast about a stock can remain informative while opening more exposure would violate a cap. A system that respects the cap and chooses no action may perform its task correctly. Scoring every such turn as a missed opportunity would teach the wrong behavior.
We would preserve a forecast assessment that ignores those account constraints, alongside a policy assessment that includes them. This lets a researcher locate the failure. If forecasts deteriorate, investigate the model or inputs. If useful forecasts repeatedly trigger expensive churn, investigate the mapping from forecasts to actions. If intended actions fail to execute, inspect runtime and venue outcomes.
Compare objectives at a common research budget
Our proposed experiment holds the data universe, availability timestamps and development budget fixed. Candidate A estimates future returns and passes them to a predefined allocation rule. Candidate B learns allocations using an explicitly specified cost-aware objective. Both face the same eligible assets, capital constraints and execution model on untouched future periods.
The evaluation should retain two results. One compares the forecasts where the candidates expose comparable forecasts. The other compares resulting portfolios under the same cost assumptions. A candidate trained directly for allocation may sacrifice prediction accuracy while improving the chosen economic objective. That is a potentially useful result if the tradeoff is visible and the objective matches the owner's intended use.
Cost assumptions need their own sensitivity analysis. Vary execution costs over plausible conditions derived from the intended market, and record where the ranking reverses. A policy that wins only under the most favorable setting has a different deployment case from one whose benefit survives measured execution friction. Such a sensitivity curve is more informative than selecting a single convenient fee number.
There is also a research cost. Larger models, more frequent inference and wider policy search consume resources. Keep those resources in the ledger instead of reporting only portfolio returns. If two candidates have comparable economic outcomes, the cheaper candidate can offer a better operating choice without a claim about superior market skill.
The survey supports an ambitious research direction: combine flexible statistical estimation with the economics of feasible decisions. Our next step would be to test which objective produces a better policy for a defined mandate, with the same information and a reproducible cost model.