Shared Examples Can Contaminate an Agent Benchmark
By DX Research Group · · Trace evaluation
A case-family audit separates copied examples, semantic siblings and genuinely new trading decisions.
A trading benchmark loses interpretability when its evaluation cases also teach the agent how to answer. We would audit the movement of examples through documentation, demonstrations, retrieval and training before reading a score improvement as generalization. The useful unit is the decision family: two traces can use different asset names while testing the same threshold and exception.
Follow the example through the harness
Consider an illustrative benchmark containing twelve position-management cases. Four are literal copies of prompt demonstrations. Another four change the symbol and quantities while preserving the same fee threshold, mandate exception and expected action. Four introduce a different conflict between portfolio state and owner instruction. A model answers eleven correctly: four copied, four rephrased and three new cases. The aggregate is 11/12, or 91.7%; the new-family result is 3/4, or 75%. Neither percentage describes an unseen production population.
The rephrased contamination study motivates examining semantic overlap alongside exact text. Our proposed audit goes further into the agent runtime: a benchmark answer may appear in tool examples even when the model's pretraining exposure is unknown. We can inspect what our own harness supplied without claiming access to a provider's training corpus.
Assign each case a family identifier and an exposure ledger. Record the first public release, every internal demonstration that used it, and whether retrieval could return its expected answer. Hashes detect exact copies; reviewers inspect shared economic structure. Distinguish a common primitive, such as calculating fees, from a near-complete worked solution. Ordinary financial vocabulary alone is weak evidence of contamination.
Split families before publishing solutions
We would place all variants of a decision family in one partition. A case involving a temporary owner restriction should remain alongside its renamed and rescaled siblings. The held-out partition can then test a new combination of restrictions, rather than the ability to repeat an example with another token ticker.
The continuous record companion preserves a historical directional-edge null and scoped replay evidence. That distinction matters here: a cleaner benchmark establishes a more credible behavioral comparison, while executable market value still needs separate evidence. The operating-layer controls companion supplies the reason to inspect harness exposure as well as model exposure.
After publication, move disclosed families into a regression set and build fresh evaluation families. Label the old set as public. Keep release dates attached so an apparent gain after disclosure can be investigated. For a follow-up challenge, generate a privately reviewed mandate conflict with different causal structure, preserve its answer rationale, and reveal it only after the comparison closes. This is a proposed workflow, with no measured contamination estimate for DXAP. Its deliverable is an exposure table that lets readers decide which slice of the score answers their question.