Freezing Replay Inputs Before Comparing Trading Models

By DX Research Group · · Trace evaluation

How to define the replay boundary, preserve captured context and avoid silently evaluating a different information set.

A replay comparison answers a narrow question: what would each model propose from the information captured for this decision? If one replay refreshes a quote or adds a later headline, it answers a different question. We treat the captured input boundary as part of the experiment, with the same care as the scoring rule.

Freeze what the agent could read

The continuous record paper companion describes a paired league built from 416 captured production scenarios across seven days. It explicitly scopes replay to behavior under frozen context. Its reported model-quality intervals overlap, which makes small differences especially sensitive to accidental changes in the input population.

For a new experiment, save the exact rendered prompt bytes, the state snapshot and the tool definitions available at the decision timestamp. Store fetched tool responses as fixtures with their observed timestamps and provenance. The manifest should state whether models can request additional information. If they can, every model needs the same captured response service and the same rules for unavailable requests.

Market timestamp and capture timestamp serve different purposes. A candle may close at 12:00 and enter the data service at 12:01. A turn beginning at 12:00:30 could see neither the finished candle nor a later corrected value. The fixture must reproduce the observed version, including its delay.

A quote that changes the experiment

Take an illustrative replay captured when an asset trades at 100. The model saw a recent high of 102 and a permitted maximum entry of 101. During evaluation the asset trades at 104. Refreshing the quote makes the same entry instruction infeasible. A model that declines now might look more disciplined, even though the captured decision allowed an entry.

The repair is a response lookup keyed by scenario and tool request, with explicit unavailable status for requests the capture cannot answer. Record that status in the result. Avoid replacing an absent fixture with the current API response. If a model needs information outside the frozen set, classify the scenario under the declared missing-information rule or evaluate that capability in a separate live test.

Before scoring, run a leakage check over every tool timestamp and every rendered reference. Inspect a sample of actual prompts, because a clean manifest can coexist with a renderer that quietly inserts the current clock or current account balance. Compare the stored input digest to the bytes delivered to each candidate.

Preserve the decision question

Our operating-layer controls account describes repeated scenario testing during pre-launch prompt revision. The useful principle is a stable question with a deliberate treatment change. Model substitution, prompt revision and tool expansion should each have an identified role in the experiment.

DXAP supports model selection within its trading harness. That architecture makes controlled comparison useful, while the replay manifest establishes what the comparison actually held fixed. Publish the population, capture window, response coverage and excluded scenarios beside the score. A frozen replay can reveal decision consistency or policy adherence. Economic claims require an additional execution model and a clearly defined outcome horizon.

Sources

Related field notes