Recording Model and Harness Versions for Trading-Agent Results
By DX Research Group · · Trace evaluation
Why a model label alone cannot identify the configuration that produced a trading decision or evaluation score.
A model name is an incomplete description of a trading agent. Its rendered context, tools and policy can change while the label remains the same. We would publish a compact version manifest beside any result so a reader can identify the configuration that actually produced the decisions.
Follow the configuration that generated the turn
The manifest should identify the provider model reference and observed model revision where available. It should also identify prompt template, tool contract, renderer and policy versions. The execution adapter and evaluator need their own versions because each can change the result after the model has produced an identical action.
Use content identifiers for saved configuration artifacts and retain human-readable release labels for navigation. A digest establishes which bytes were used; it establishes no claim about their quality. Record the effective versions on the turn, rather than reconstructing them from whichever deployment happens to be current when the analysis runs.
Our continuous record describes template lineage from Terminal Pro into the later fleet while preserving differences in venue, cadence and execution. It also restates cross-era P&L at a common fee rate. That is an example of why lineage and economic normalization belong together: a continuous project history can still contain materially different measurement eras.
A rolling change can imitate a model gain
Take an illustrative evaluation in which model A runs during the first week and model B runs during the second. Between weeks the renderer adds funding data, the policy lowers size limits and the execution adapter fixes a rejected-order path. If B's realized losses fall, the result belongs to the combined change in model, harness and calendar conditions.
For a model comparison, replay shared saved scenarios through an identical declared harness, or run a controlled design that explicitly estimates the relevant changes. For a product-release comparison, retain the whole package and describe it as a release effect under the observed population. Those designs answer different questions and can both support useful decisions.
A report spanning several eras should group by effective configuration and observation window. If a provider silently updates a model alias, retain the known alias and state that the underlying revision is unavailable. Avoid inventing precision through a local label that the provider never guaranteed.
Make future evolution inspectable
The operating-layer controls paper records prompt revisions and a frozen historical runtime. That gives its deployment findings a configuration boundary. An evolving alpha product needs the same discipline at a finer cadence, because new data and controls can ship between otherwise similar turns.
DXAP publicly presents the harness as evolving. Our proposed manifest turns that product direction into an evaluable research practice: preserve old configuration artifacts, attach effective versions to decisions and compare releases on an identified population. A reader can then distinguish a historical finding from a current capability and a proposed improvement. The platform can continue changing while its evidence remains interpretable.