Shared Examples Can Contaminate an Agent Benchmark

By DX Research Group · · Trace evaluation

A case-family audit separates copied examples, semantic siblings and genuinely new trading decisions.

A trading benchmark loses interpretability when its evaluation cases also teach the agent how to answer. We would audit the movement of examples through documentation, demonstrations, retrieval and training before reading a score improvement as generalization. The useful unit is the decision family: two traces can use different asset names while testing the same threshold and exception.

Follow the example through the harness

Consider an illustrative benchmark containing twelve position-management cases. Four are literal copies of prompt demonstrations. Another four change the symbol and quantities while preserving the same fee threshold, mandate exception and expected action. Four introduce a different conflict between portfolio state and owner instruction. A model answers eleven correctly: four copied, four rephrased and three new cases. The aggregate is 11/12, or 91.7%; the new-family result is 3/4, or 75%. Neither percentage describes an unseen production population.

The rephrased contamination study motivates examining semantic overlap alongside exact text. Our proposed audit goes further into the agent runtime: a benchmark answer may appear in tool examples even when the model's pretraining exposure is unknown. We can inspect what our own harness supplied without claiming access to a provider's training corpus.

Assign each case a family identifier and an exposure ledger. Record the first public release, every internal demonstration that used it, and whether retrieval could return its expected answer. Hashes detect exact copies; reviewers inspect shared economic structure. Distinguish a common primitive, such as calculating fees, from a near-complete worked solution. Ordinary financial vocabulary alone is weak evidence of contamination.

Split families before publishing solutions

We would place all variants of a decision family in one partition. A case involving a temporary owner restriction should remain alongside its renamed and rescaled siblings. The held-out partition can then test a new combination of restrictions, rather than the ability to repeat an example with another token ticker.

The continuous record companion preserves a historical directional-edge null and scoped replay evidence. That distinction matters here: a cleaner benchmark establishes a more credible behavioral comparison, while executable market value still needs separate evidence. The operating-layer controls companion supplies the reason to inspect harness exposure as well as model exposure.

After publication, move disclosed families into a regression set and build fresh evaluation families. Label the old set as public. Keep release dates attached so an apparent gain after disclosure can be investigated. For a follow-up challenge, generate a privately reviewed mandate conflict with different causal structure, preserve its answer rationale, and reveal it only after the comparison closes. This is a proposed workflow, with no measured contamination estimate for DXAP. Its deliverable is an exposure table that lets readers decide which slice of the score answers their question.

Sources

Related field notes