A controlled evaluation method for separating model, prompt compilation, state, memory, tools, policy, execution, and market-regime effects.

A harness-transfer test is a controlled comparison that asks whether an observed trading-agent improvement survives a declared change to the system around the model. That system begins with the mandate compiler, state snapshot, and memory policy. It continues through tools, typed actions, deterministic controls, execution, and evaluation. A useful test freezes the rest of that path, changes one named component, and follows the result through settlement or an explicit abstention.
Download the harness-transfer evaluation card or the CSV edition. The first four rows cover model, compiler, state, and memory transfers. The remaining rows cover action schema, policy, execution, and regime transfers. Each row names what stays fixed, what changes, what must be measured, and how far the result can be promoted.
Evidence boundary: This page combines DXRG controlled tests, one bounded historical deployment, current primary-source methods, and our evaluation framework. A passing transfer test supports the named comparison and does not establish alpha, profitability, universal safety, or transfer beyond the registered conditions.
Trading-agent evaluations often collapse a complete decision system into a model label. That shortcut hides the components that shape what the model can observe and what its output can do. Two runs with the same model can differ because the mandate was compiled differently, the portfolio snapshot arrived at a different time, the memory carried a different provenance rule, or the execution adapter handled an ambiguous timeout differently.
The reverse problem matters too. A stronger model can enter a weak harness and inherit its stale state, loose action schema, or unrealistic fill assumptions. A useful result therefore identifies the evaluation subject as a versioned model inside a versioned mandate-to-settlement path.
Recent primary research makes this separation concrete. KTD-Fin holds prompts, tool interfaces, masking, and execution rules constant for its cross-model comparison, then runs multiple seeds.1 Agent Market Arena reports that agent frameworks produced distinct behavior across a live multi-market benchmark while model backbones explained less of the observed outcome variation in that setting.2 Neither paper validates DXRG's results. Together, they show why model identity and harness identity belong in separate fields.
Our general benchmark card asks what a complete evaluation must disclose. This page answers a narrower question: how do we test whether a measured improvement belongs to the component we changed?
Before DX Terminal Pro launched, the fee sentence sat in paragraph eight of the compiled context. Agents cited the 2.3% total fee in 3% of sampled reasoning traces. We moved the same sentence to paragraph one. The model stayed fixed, the wording stayed fixed, and the market data stayed fixed. Fee citation rose to 74%.3
That scene is useful because it leaves little room for a vague attribution. The observed difference belonged to reading order inside the compiled context. It did not prove better returns. We read the diagnostic beside action rate, observation rate, and capital deployment because a model can mention a fee without using it well.
The same discipline applies to a broader transfer. Suppose the paragraph-one intervention passes with a second model. We still need the same scenario population, tool schema, sampling policy, action definitions, and downstream checks. If the second run also changes memory retrieval and execution rules, the result becomes a system comparison. That can be useful, but it answers a different question.
Our trace-derived regression method starts with the smallest reproducible break. The transfer card extends that method across a declared boundary while preserving the trace fields needed for attribution.
We use transfer contract for the frozen specification that makes two evaluation cells comparable. The contract records the baseline cell, candidate cell, intended change, and shared fixtures. It also records seeds, evidence class, and promotion limit. Later references to the contract mean that exact comparison record, not a general promise that an improvement travels.
The minimum contract fits in five parts:
| Part | What we record | Why it matters |
|---|---|---|
| Subject | Model, harness commit, prompt bundle, tool schemas, policy version | Names the system that actually ran |
| State | Market time, portfolio time, open orders, memory cutoff | Prevents different information sets from masquerading as transfer |
| Intervention | One changed component and its expected mechanism | Keeps attribution local |
| Measures | Local diagnostic, typed action, policy result, settlement outcome | Connects the edit to the full path |
| Boundary | Evidence class, failed cases, excluded regimes, promotion limit | Stops a narrow pass from becoming a broad claim |
NIST's draft benchmark-evaluation practices describe comparability and external validity as distinct execution-environment considerations.4 We treat that distinction as operational. First make the two cells comparable. Then test a new environment and label that move as a separate transfer.
A model-swap cell keeps the mandate compiler, state, memory, tools, policy, and execution adapter fixed. The baseline model and candidate model receive the same frozen fixtures under the same sampling contract. We compare action validity, abstention, mandate fidelity, policy outcomes, and downstream settlement measures across repeated runs.
KTD-Fin's identical-harness cross-model design is a useful primary-source example of this logic.1 Its markets and metrics differ from ours, so its results stay in their own evidence class. The reusable method is the controlled model swap.
We also retain the distribution, rather than only the median. STATE-Bench runs tasks repeatedly and reports both completion and a reliability measure for tasks that succeed across all five trials.5 Its enterprise tasks are separate from trading. The design illustrates why one lucky completion cannot describe an agent's operational consistency.
For trading agents, an abstention belongs in the distribution. If the candidate model submits fewer invalid actions because it observes more often, that is a real behavioral change. Whether the change is desirable depends on the mandate, opportunity set, costs, and evaluation objective. A transfer card that records only completed trades would erase the mechanism.
The next cells keep the model fixed and move one harness component. A prompt-compilation cell can change precedence or context order. A state cell can change freshness enforcement. A memory cell can change provenance labels or retrieval policy. Each cell needs a local measure and a downstream measure.
Consider state assembly. The local question is whether the runtime rejects a snapshot whose portfolio balance is stale relative to an acknowledged fill. The downstream question is whether any action created from that snapshot reaches policy validation or submission. A clean local catch with no submitted payload supports the state-control claim. It says nothing about the agent's market judgment.
Memory has a similar shape. STATE-Bench evaluates agents against stateful environments with deterministic final-state assertions for state-mutating tasks.5 We adapt the principle, not the benchmark outcome. A trading memory fixture should name the source, event time, observation time, and expiry rule. The next trace must show whether recalled information entered state as evidence, context, or an expired record.
These cells also help locate negative results. If an intervention passes under one compiler and fails when the compiler moves, the failure can belong to precedence handling instead of the model. A failed transfer is useful when the trace makes that attribution possible.
Every card includes the typed no-trade action. We use abstention for the explicit path that ends a decision without constructing or submitting an order. It carries a reason code, the bound state identity, and the policy result. Later mentions of abstention refer to this recorded action, rather than missing output or a silent crash.
The distinction becomes visible in a tool or action-schema transfer. One schema may force the model to choose among BUY, SELL, and OBSERVE with bounded parameters. Another may accept prose and repair it after generation. If the first schema reduces malformed actions, we still inspect whether valid trades became unsupported repairs or whether uncertainty moved into OBSERVE.
NIST's work on agent evaluation probes stores structured verdicts in a machine-readable audit trail and can apply those probes during the workflow or after it.6 The specific probes focus on factual grounding, not trading decisions. The method supports our requirement that a transfer verdict remain attached to the trace evidence that produced it.
A valid typed action remains upstream from a venue outcome. Execution-transfer cells therefore keep the decision fixed and change the adapter, acknowledgement path, or reconciliation rule. The primary measure can be duplicate prevention, fill-state accuracy, fee capture, or time to authoritative reconciliation.
An ambiguous timeout is the cleanest fixture. The baseline and candidate receive one payload with one stable request identity. The response disappears after submission. A passing adapter queries authoritative order state before any retry and binds the recovered acknowledgement to the original action. The next state snapshot includes the fill or preserves the unresolved order.
Our execution and settlement matrix contains the deterministic fixtures for this path. A model comparison that ends at JSON validity cannot inherit an execution-reliability claim. A harness comparison that ends at acknowledgement cannot inherit a settlement claim. Each denominator stays attached to its stage.
This is where a mandate-to-settlement trace earns its name. The trace links the authenticated mandate to compiled instructions, state, memory, and action. It then links policy, payload, acknowledgement, settlement, and reconciled state. Passing only the early fields produces an early-stage result.
Market, venue, and regime moves are external-validity tests. We schedule them after a component has passed comparable cells. The card changes the environment field while retaining the system versions and evaluation definitions.
KTD-Fin uses regime-stratified windows and distinguishes its contamination controls from its performance comparisons.1 NIST's draft benchmark practices also treats environment similarity to real deployment as a separate design consideration.4 These sources support the need to disclose the move. They do not supply a universal recipe for market transfer.
A result can survive quiet and volatile windows while failing on a second venue because order semantics differ. Another can survive a venue change while failing when borrow, funding, or gas enters the state. We record each break at the earliest supported layer. "Cross-market transfer" is too broad when the actual evidence covers one adapter and two declared regimes.
NIST's analysis of evaluation cheating adds another reason to preserve transcripts and tool restrictions. It distinguishes solution contamination from grader gaming and recommends transcript review plus standardized agent affordances.7 In a trading test, future data access, hidden answer leakage, or a scorer that rewards invalid shortcuts can create a high score without a better decision system.
We read a transfer pass in three lines. First, name the component and boundary that changed. Second, report the local and downstream measures with repeated-run uncertainty. Third, state the highest evidence class the result can support.
The resulting claim should sound narrow: "Under the registered replay fixtures, the candidate memory policy reduced expired-record use while the model, tools, and policy stayed fixed." The next question is whether the same mechanism survives a fresh paper, shadow, or live setting. It is a new evaluation cell.
The same restraint applies to a failure. A candidate that fails a transfer cell may expose a real dependency between components. It may also reveal a broken fixture or scorer. We inspect the trace before assigning the cause. NIST's evaluation-cheating work is a reminder that benchmark implementation can distort the property being measured.7
This discipline protects model-versus-harness attribution. It also makes the result reusable. Another team can reconstruct the declared cell, change one boundary, and report whether the mechanism survives.
The JSON card is the canonical machine-readable version. The CSV edition flattens the eight transfer rows for review. Each row is a template, so the result field begins empty.
Start with the row closest to the claimed improvement. Freeze the comparison contract. Fill the baseline and candidate identifiers, then attach immutable trace references. Record failures and abstentions with the successful cases. Promote the claim only to the evidence class earned by the weakest required stage.
For the wider disclosure envelope, pair this card with the DXRG benchmark card. For component-level regression construction, continue with Using Trading Traces to Improve Agent Behavior. The three pages answer separate questions: what to disclose, how to isolate transfer, and how to turn a failure into the next test.
Version 1.0. Published July 28, 2026. Corrections: poof@dxrg.ai. This article is research and educational material, not financial, investment, legal, compliance, or security advice.
KTD-Fin, From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets (2026), evaluation design and experiments. (arXiv) ↩ ↩2 ↩3
When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents (2025), abstract and benchmark design. (arXiv) ↩
DXRG, Operating-Layer Controls for Onchain Language-Model Agents Under Real Capital (2026), control-loop method and failure-mode table. (arXiv) ↩
NIST, Practices for Automated Benchmark Evaluations of Language Models, Initial Public Draft, January 2026, execution-environment and evaluation-protocol settings. (PDF) ↩ ↩2
Microsoft Research, STATE-Bench: A Benchmark for AI Agent Memory (2026), evaluation loop and production-readiness metrics. (Project overview) ↩ ↩2
NIST, Building Evaluation Probes into Agentic AI (2026), structured audit-trail method. (Project overview) ↩
NIST CAISI, Cheating on AI Agent Evaluations (2025), contamination, grader gaming, and transcript review. (Research note) ↩ ↩2