Latency Quantiles Reveal Trading-Agent Tail Behavior
By DX Research Group · · Trace evaluation
Compare latency distributions and deadline misses rather than relying on average response speed.
Average latency can hide the delays that determine whether a trading agent completes a useful review. We would compare empirical quantiles, deadline misses and the quality of the affected decisions. A fast average offers weak reassurance when a small tail contains the urgent cases.
Two equal means, different deadlines
Take two illustrative sets of ten completion times. Agent A completes every case in five seconds. Agent B completes eight cases in two seconds and two cases in seventeen seconds. Both means are five seconds: B totals 8 × 2 + 2 × 17 = 50 seconds. With a ten-second deadline, A misses zero cases and B misses two.
Use the nearest-rank empirical quantile for this small example: sort times and select position ceil(q × n). A's median and 90th percentile are both five seconds. B's median is two seconds and its 90th percentile is seventeen seconds. Other software interpolation conventions can give different reported values; the evaluation should record its convention.
The SRE monitoring guidance motivates watching latency distributions. Our proposal joins those measurements to trading task requirements. A delayed long-horizon research summary and a delayed stale-order cancellation occupy different economic contexts.
Preserve the task mix
We would first separate scheduled portfolio reviews, triggered risk checks and background research. Within each class, retain model time, tool waiting and runtime queue time. Then compare matched case distributions. A model assigned more chart calls can look slower even when its own inference stage is faster.
Avoid averaging percentile values across unequal shards. Pool underlying observations when the measurement contract permits it, or present each shard and its count. A median from a five-case quiet period and a median from a five-hundred-case congested period carry different population information.
The continuous record companion provides historical context for agent inference economics. The operating-layer controls companion provides the mandate-to-settlement framing. Neither tells us that the illustrative tail distribution exists in the current product.
For a proposed fixture, replay urgent and ordinary reviews under a controlled queue, then mark each output with correctness and its task-specific deadline. Report the fraction delivering a correct admissible proposal on time. This allows a latency improvement to be interpreted alongside any loss in decision quality from shorter generation or fewer tool calls.
Small samples make extreme percentiles fragile. In the ten-case illustration, each observation represents ten percentage points of the empirical distribution. We would show the sorted times and sample count beside the quantiles, rather than attach false precision to a 99th percentile. Tail investigation then becomes a concrete engineering task: inspect the two seventeen-second traces, identify waiting stages, and test the proposed repair on those stages before claiming a broader speed gain.