Monthly stock returns contain signal, even when most movement stays unpredictable

By DX Research Group · · Market predictability

Gu, Kelly and Xiu show why nonlinear interactions improve expected-return measurement and why the horizon matters.

Monthly equity returns contain measurable predictive structure. Gu, Kelly and Xiu's Empirical Asset Pricing via Machine Learning gives a concrete answer to a useful question: can flexible models extract more of that structure than familiar linear benchmarks? Their evidence supports the answer for a historical U.S. stock panel. We read the result as a finding about conditional expected returns, with a precise population and forecast horizon.

The published study uses monthly CRSP returns for NYSE, AMEX and NASDAQ stocks from March 1957 through December 2016. Its initial chronological split trains on 1957–1974, validates on 1975–1986 and tests on 1987–2016. Models are refitted annually with an expanding training period and rolling validation period. Stock characteristics interact with aggregate variables, giving the methods a common information basis.

The gain comes from relationships between signals

Trees and neural networks improve forecasts through nonlinear interactions. Momentum, liquidity and volatility repeatedly appear among important predictors. The lesson is more specific than “more data helps”: expanding ordinary least squares to the large feature set produces negative out-of-sample predictive R², while regularization and suitable nonlinear models recover useful signal. More layers also fail to produce a monotonic improvement in this experiment.

That comparison matters to anyone building a forecast pipeline. A feature can have a weak average association while becoming informative under a particular state. Conversely, a richly parameterized model can spend its flexibility fitting noise. The mechanism to investigate is whether a model represents the right relationships and controls estimation error on future observations.

An intuitive hypothetical illustrates the distinction. Imagine two stocks with the same recent price trend. One trades frequently with a narrow spread; the other rarely trades and has a wide spread. A linear additive specification can assign each characteristic a separate adjustment. An interaction model can allow the relevance of the trend itself to change with liquidity. The paper's empirical finding motivates testing such relationships; this example supplies no estimated coefficient or trade recommendation.

Predicting an expectation leaves substantial surprise

A conditional mean forecast answers what return is expected given the available variables. The next realized return adds news and other unexplained movement. A model can improve that expectation without correctly calling most individual price moves. Forecast R² measures an error comparison against a specified benchmark; it says something different from direction accuracy, confidence calibration or the fraction of profitable trades.

For a reader evaluating an agent, we would require the target to be written in plain language before seeing its score: next month's excess return for an eligible U.S. equity, for example. Then specify whether the task estimates a return, ranks stocks, predicts a direction or chooses an action. These tasks can share inputs while needing different labels and evaluation measures.

A ranked stock portfolio can benefit from consistently separating relatively stronger and weaker expected returns across many names. An agent taking a concentrated position on one asset needs a forecast appropriate to that asset and its holding period. A monthly cross-sectional finding supplies a research hypothesis for that agent, rather than a measured answer about its next intraday decision.

Our proposed translation check

We propose testing transfer in two controlled changes. First, preserve the paper-style target and compare candidate methods with the same point-in-time inputs and chronological split. Second, alter exactly the population or horizon relevant to the proposed agent and rerun the comparison. A gain that disappears at the second step identifies the boundary of transfer.

Keep portfolio construction frozen during the forecast comparison. Otherwise, an improvement could come from a changed sizing rule rather than improved estimates. Report paired forecast errors by month so common market shocks remain visible, and separately examine liquid and less liquid stocks. Select those slices using development data before inspecting the final test outcomes.

The finding earns a strong, bounded claim: historical monthly equity expectations have structure that well-chosen machine learning methods can estimate better. The next useful experiment establishes how much of that structure survives the agent's actual target and eligible universe.

Sources

Related field notes