Missing Market Candles: A Data Contract for Trading Agents
By DX Research Group · · Market data
A gap-classification workflow that keeps no-trade periods separate from collection failures.
A trading agent can encounter a missing bar during an otherwise valid decision cycle. Our proposed data contract keeps that absence visible, so an evaluation records whether the harness supplied a measured value or an imputation.
A missing candle is an absence of a row that requires a diagnosis. It can represent no observed trading, a provider omission, a failed request, or a broken collection pipeline. Filling every gap with the previous close makes these different situations look alike.
Start with an explicit time grid
Consider an illustrative five-minute window with expected one-minute buckets at 12:00, 12:01, 12:02, 12:03, and 12:04. The export contains only 12:00, 12:01, and 12:04. First construct the expected grid and mark 12:02 and 12:03 absent. Do this before calculating returns or rolling windows, because adjacent rows are now three minutes apart.
Coinbase's candle documentation warns that historical rate data may be incomplete and that intervals without ticks have no published data. This is a source-specific reason to avoid interpreting every absent candle as an outage. Each particular gap still needs evidence about its cause.
Classify with independent evidence
For each gap, check request logs, response status, covered time ranges, and any retained trade feed. If retained trades exist in 12:02 while its candle is absent, classify it as a candle omission. If the request failed, label collection failure. If the request succeeded but no corroborating trade record exists, leave the cause unresolved.
Use separate fields such as observed, confirmed_no_trade, collection_failure, and unknown_gap. These labels are proposed schema values, defined by the collector. Keep the evidence used for classification so another analyst can revisit the result without guessing.
Filling is a modeling choice
For an explicitly confirmed no-trade minute, a researcher might create a synthetic carry-forward price with zero observed trade volume. Mark the entire row synthetic and retain the original absence. This convention creates a regular feature grid with explicitly synthetic prices.
For a collection failure, carrying a price forward may conceal market movement. Leave price missing unless the experiment has a documented imputation policy. If you compare imputation methods, use identical gap masks and show how many decisions each method excludes. Otherwise a change in the evaluation population can masquerade as a feature improvement.
A small acceptance fixture
Create one gap with retained trades, one with a failed request, and one with independently confirmed no trades. The classifier should assign distinct labels. A rolling five-minute statistic should report its coverage, including the number of real and synthetic observations. The expected output is a transparent accounting of remaining missingness.
State what remains unknown
Absence of a trade in your collector leaves venue activity unresolved. Collection coverage determines how strong that inference can be. Report unknown gaps alongside classified gaps, and carry their flags into evaluation artifacts. Missingness often clusters during difficult market or infrastructure conditions. Deleting it silently can make the surviving sample easier than the system's actual operating environment.
Place this check in the agent loop
Our state and memory framework explains how this input contract fits a persistent trading agent. Use the harness-transfer test design to distinguish a data-adapter change from a change in model behavior.