Event Sampling Bias in Trading-Agent Market Features
By DX Research Group · · Market data
A two-regime example showing how feed activity changes an average even without a price change.
Trading-agent evaluations can overweight busy intervals before a model sees its first feature. We use a synthetic two-regime stream to show how collection frequency changes the population represented by a summary.
A market-data average depends on how observations enter the sample. A message-level mean weights busy intervals by message count. A regular-clock mean weights captured time intervals more evenly. Neither is automatically the correct answer; they describe different populations.
Calculate the difference
Assume a synthetic ten-second window. For the first five seconds, a spread is one unit and the feed sends ninety messages. For the next five seconds, the spread is three units and the feed sends ten messages. The event-weighted spread is (90 × 1 + 10 × 3) / 100 = 1.2. If each regime lasts exactly five seconds, the time-weighted spread is (5 × 1 + 5 × 3) / 10 = 2.
Both arithmetic results are correct under the assumptions. Calling either simply the average spread hides the weighting rule. A chart showing 1.2 may emphasize the busy regime rather than the typical elapsed-time experience.
Coinbase's channel documentation explains that ticker messages follow matches and may batch cascading matches. This is one reason message frequency has a delivery-dependent weighting. Read the specific feed's delivery rules before using message counts as economic weights.
Define the population first
For an event study, event weighting may be intentional. For a question about what a consumer sees at regularly scheduled decisions, sample at those decision times using a documented availability rule. For duration summaries, assign each valid state a bounded duration and account for missing intervals separately.
Bound any carry-forward duration across a collection outage. An illustrative duration cap of two seconds would leave part of an eight-second silence uncovered. Report that uncovered time rather than allowing an old observation to represent the whole interval. The cap is a local assumption that requires justification for the task.
Keep counts and coverage together
Publish event counts, covered duration, and gap duration alongside the statistic. Break them down by instrument and collection segment. A market with many messages and a market with few messages can contribute very different weights to a pooled event mean.
If combining instruments, state whether each instrument receives equal weight, volume weight, event weight, or covered-time weight. Normalizing within each instrument and then averaging is a different operation from pooling all messages first. Show a tiny worked table to make the choice auditable.
A repeatable diagnostic
Use the two-regime fixture and verify the expected values of 1.2 and 2. Then duplicate every message in the first regime. The event mean changes, while a properly constructed duration mean does not. This duplication test reveals whether an implementation's result depends on message multiplicity.
Add a gap and verify that coverage drops. A method that returns the same complete coverage after deleting observations may be silently carrying values forward beyond its stated rule.
What the comparison establishes
A difference between weighted summaries is evidence about the weighting rule. It exposes the estimand, the quantity the calculation targets. Once that target is explicit, researchers can choose the statistic appropriate to their question and prevent a change in feed activity from being mistaken for a change in market conditions.
Place this check in the agent loop
Our state and memory framework explains how this input contract fits a persistent trading agent. Use the harness-transfer test design to distinguish a data-adapter change from a change in model behavior.