A release regression corpus for an evolving trading harness

By DX Research Group · · Frontier research

A proposed corpus tests mandate, input, execution and recovery behavior so improvements can be evaluated across DXAP harness releases.

An evolving agentic trading product needs a way to preserve working behavior while fixing failures. Our operating-layer controls paper describes 24 pre-launch prompt revisions tested against replayed scenarios. We describe a PROPOSED release regression corpus that extends this discipline across the complete agent turn. Its purpose is to make improvements reviewable, rather than announce a current DXAP release process or treat a test-suite pass as investment performance.

Each fixture should name the behavior it protects

A fixture binds a mandate, state snapshot, rendered context and expected result. Some expected results are deterministic: an asset outside an allowlist must be rejected under that policy. Others admit several valid model choices: observing or proposing one of several compliant actions may all satisfy the case. The test should encode that acceptable set explicitly rather than require a model to reproduce one historical phrase.

We would include stale data, conflicting instructions, unavailable tools and an uncertain settlement response. Execution cases need a simulated venue state and an expected reconciliation path. A model answer alone cannot pass a recovery test whose failure lives after submission. Every fixture should carry a source incident or a clearly labelled synthetic rationale so reviewers know why it exists.

Keep a discovered failure after repairing it

An illustrative fixture contains a previous agent rationale that invents a named trading rule. The next turn must treat that text according to its context status and the authenticated mandate. The earlier paper reports a compound intervention reducing fabricated-rule traces from 57% to 3% in affected pre-launch populations, with some comparison-arm counts unavailable. We would preserve that bounded historical lesson while evaluating a fresh corpus with disclosed denominators.

Another fixture opens a position but encounters a protective-trigger rejection. The expected outcome depends on the declared recovery policy and reconciled exposure. A prompt revision that improves explanations while leaving that position unresolved should still fail the execution case. This prevents a visible text improvement from masking an unchanged operational fault.

Version the test and the system together

Every run should record model, prompt, tool schema, policy version and fixture version. A changed tool interface can invalidate an old fixture, so the maintenance record must explain whether a case was updated, retired or newly added. We would keep a stable comparison subset alongside expanding incident coverage, avoiding a headline pass rate that changes because the denominator became easier.

The continuous record shows why a second track is needed: historical control and behavioral findings coexist with a directional-edge null. Passing regression cases establishes specified behavior under those cases. Forward forecasting and economic evaluations establish whether decisions improve under later market data. DXAP publicly describes a persistent loop with external policy checks and recorded turns. A versioned corpus gives that evolving architecture a concrete learning mechanism: preserve resolved failures, expose new regressions and attach each release claim to the behavior actually measured.

Sources

Related field notes