A Shared Harness Lineage Still Needs Transfer Tests

By DX Research Group · · DXRG findings

The two historical fleets shared configuration concepts but resolved text-setting conflicts differently. Transfer must name the component and environment.

Our two historical fleets shared five behavioral sliders and a one-action-per-turn design lineage, yet resolved text-versus-slider conflicts in opposite directions. Terminal Pro’s sliders constrained insistent strategy text. In the later fleet, strategy text routinely overrode the frequency slider. Shared interface concepts therefore supplied weak evidence that the underlying instruction precedence had transferred.

The continuous record joins a February 26 to March 18, 2026 Base event with a June 8 to August 15 Hyperliquid fleet. It pools findings at the metric level rather than treating their trades as one exchangeable population. Real capital in one setting and predominantly paper execution in the other further limit a cross-system economic comparison.

Name the boundary before claiming progress

Terminal Pro ran one frozen Qwen model family in a 12-token bounded market. The later fleet used a changing model mix across roughly 99 perpetuals, including HIP-3 synthetics. Schedules, tool manifests and execution paths changed too. A successful behavior in both environments can suggest a useful mechanism, while still leaving its cause unresolved.

The first paper companion reports a sixfold activity gradient from 2.8% to 16.8% of invocations and trade-size settings mapping approximately from 2% to 95% of available ETH. Those are local observations under that harness. Applying those exact coefficients to perpetuals would invent a transfer result.

We start a transfer claim by naming one boundary: the model, mandate compiler, state adapter, memory policy or execution path. The remaining components need immutable references. If several change together, the result belongs to the bundle, and its limitations should say so.

Eight templates, with results still open

Our public harness-transfer card supplies eight controlled comparison templates. It covers model swaps, mandate compilation, state assembly and memory, then action schemas, policy, execution reconciliation and environment changes. Every template’s result field is null. These rows organize proposed evaluations; they report no completed superiority test.

The card’s execution comparison freezes the typed action, policy result, payload and request identity while changing the adapter or reconciliation boundary. Its measures include submission counts, duplicate prevention and fill accuracy. Those choices allow a researcher to locate an execution regression without calling it lost predictive skill.

For mandate transfer, the useful evidence is the compiled instruction paired with reasoning, action and policy records. We would put the same conflicting slider and text fixture through both compilers, score which authority survives, and register the expected behavior first. That experiment could establish precedence transfer for the named fixtures. Economic transfer would still require a separate test.

How an evolving DXAP can earn stronger claims

Current DXAP publicly describes a persistent agent with a model harness, external policy checks and recorded turns. These are concrete dimensions for product comparison. The historical research gives us a method for asking whether each component preserves its intended behavior as markets and models change.

Our claim of frontier work rests on publishing failures and specifying the next comparison. A broader assertion of superior returns or universal portability would require new matched evidence. The platform can keep evolving while every measured result stays bound to the version, market and authority that produced it.

Sources

Related field notes