Candle Close Leakage in Agentic Trading Replays

By DX Research Group · · Market data

A minute-bar example for separating bucket labels from feature availability.

When an autonomous trading agent reasons from minute bars, its input snapshot needs an explicit completed-bar policy. We use a synthetic minute below to expose how a harmless-looking bucket label can leak later information.

A candle timestamp often identifies a bucket. Its label leaves the availability of completed values unresolved. Using a completed candle at its opening timestamp can place later trades inside an earlier model input.

The minute that has not finished

Assume a synthetic one-minute bucket covering 09:30:00 through just before 09:31:00. Its opening trade is 100, its highest trade is 104, its lowest is 99, and its final trade is 103. A decision at 09:30:15 cannot know the final close of 103 or whether a later trade will produce the high of 104. Those are completed-bucket values.

Suppose an export labels that row 09:30:00. A naive join that includes all rows with timestamps before 09:30:15 supplies the full bar. The query is temporally ordered but informationally invalid. The fix is to represent bucket start, bucket end, and availability separately.

Treat the source format as a contract

Read the endpoint's bucket definitions rather than inferring them from a chart. Coinbase's candle documentation specifies the returned candle fields and supported granularities. A consumer still needs to establish when completed data was received and whether the latest bucket is provisional.

For a completed-bar-only workflow, compute an availability boundary no earlier than the bucket end, then include actual retrieval or receipt delay where recorded. An illustrative 250-millisecond delay makes the 09:30 bar available at 09:31:00.250. The delay is an illustrative fixture assumption.

If you intentionally use partial bars, store snapshots as partial bars. A snapshot at 09:30:15 should contain only the trades observed by then. Keep the earlier partial snapshot immutable when the completed minute arrives.

A boundary test with visible answers

Build decisions at 09:30:15, 09:31:00.100, and 09:31:00.300. Under the illustrative availability assumption, only the third can use the completed 09:30 bar. Record which bar identifier each decision receives. This catches off-by-one errors more clearly than inspecting a rolling-average plot.

Apply the same rule before calculating indicators. A moving average built from completed closes inherits their availability. Shifting the final feature by one row can be insufficient when buckets are missing or irregular. Join using explicit temporal boundaries rather than assuming row position equals elapsed time.

Limits of historical bars

A completed OHLC row does not reveal the order of intrabar highs and lows, every intermediate quote, or the consumer's actual delivery path. Exact execution prices require separate observations. Keep feature reconstruction separate from execution modeling. The useful claim is narrower: a specified completed-bar feature was withheld until its stated availability boundary. That is a testable data contract and an essential prerequisite for interpreting a later model evaluation.

Place this check in the agent loop

Our state and memory framework explains how this input contract fits a persistent trading agent. Use the harness-transfer test design to distinguish a data-adapter change from a change in model behavior.

Sources

Related field notes