Model Updates vs Harness Updates: Attribute Trading-Agent Improvements

By DX Research Group · · Learning theories

A proposed component comparison distinguishes a stronger trading model from improvements to instructions, state or tools.

When a trading agent improves after a release, the model may deserve some of the credit. So may a clearer mandate compiler, a repaired tool, or a more current snapshot. A release comparison evaluates the whole package. To understand where to invest next, we need a component comparison as well.

Our harness-transfer method treats the evaluation subject as a versioned model inside a versioned decision path. This note develops a compact proposed design for a common product decision: whether the next improvement should come from a model update or a harness repair.

Four Conditions Answer Different Questions

Let M1 and M2 be two model versions, and H1 and H2 be two harness versions. Evaluate M1/H1, M2/H1, M1/H2, and M2/H2 on the same saved cases. The first pair isolates the model change under the old harness. The third condition isolates the harness change with the old model. The fourth tests their combination.

For an illustrative ten-case batch, suppose valid actions rise from six under M1/H1 to seven under M2/H1 and nine under M1/H2. If M2/H2 also produces nine, the small batch suggests that the harness repair addresses failures the model update alone leaves behind. These counts are an example of interpretation, rather than DXRG measurements, and ten cases would provide limited precision.

The important comparison is paired at the case level. Identify which cases change and whether they share a failure mode. A total of nine can conceal that the combination repaired one case and broke another.

Interactions Are Research Findings

A new model may rely on a tool field that H1 never exposes. A compiler instruction that helps M1 may confuse M2. These interactions make it useful to preserve the full matrix rather than announcing that one component is universally better.

Freeze the execution assumptions, policy rules, and grading rubric across conditions. If the new model requires a different tool schema, document that requirement and report a system comparison separately. A forced identical interface can itself become a handicap when it fails to represent the intended production subject.

Our published controls record shows why harness changes deserve independent attention. Moving an unchanged fee sentence earlier increased citation in reasoning traces with the model fixed. The diagnostic describes attention in those tests, with economic effects requiring their own evaluation.

Connect Attribution to the Release Decision

We would select cases from reproducible failures and hold back cases that were unavailable during repair. Measure action validity, no-trade behavior, and execution consequences alongside costs. A model that produces more valid proposals may also invoke more tools or take longer; those tradeoffs belong in the result.

DXAP publicly offers model choice within its harness and records turns. That product structure makes component attribution relevant as the platform evolves. The public description supports the architecture; the four-condition study remains a proposed evaluation, not a published product leaderboard.

The release decision can then name what improved, where it failed, and which version combination produced the observation. That is more useful to future development than a single claim that the agent became smarter.

Sources

Related field notes