The harness defines the trading decision system

By DX Research Group · · Trading agent theory

A transition-system formulation explains why an identical model and prompt can still describe different trading agents.

A trading harness defines which observations become available, which actions become possible and what the next decision knows about the last one. We therefore treat it as a decision-system object with its own identity. A model name and a prompt string describe only part of that object. This distinction matters whenever builders compare agents, transfer a successful configuration or claim that a model upgrade improved trading.

Our operating-layer controls record follows the path from owner mandate through rendered context, typed action, validation and settlement. The theoretical step is to represent that path as a state transition. Let S contain the account state, outstanding actions and effective owner configuration. An observation function exposes a permitted view of S and market information to the model. The model proposes an action. A checking function evaluates it against the applicable configuration. Execution and reconciliation then produce S for the next turn.

This formulation locates the harness in three places: the information the model receives, the actions the runtime admits and the state that survives execution. Changing any of those functions can change the agent even when its model and wording stay fixed.

Identity requires the transition, not just the answer

Consider an illustrative pair of runtimes receiving the same request to open a new symbol. Both models emit the same typed order. Runtime A includes pending entries when computing an account's open-symbol count. Runtime B counts only filled positions. With two filled symbols, one pending new symbol and a configured cap of three, A rejects another new symbol while B may pass its cap check. Their answer-level agreement conceals a decision-system difference.

The current DXAP reference documents execution checks outside the proposal, including a configured cap on new symbols. Its configuration reference specifies that the count includes manual positions and pending entries. These are concrete semantics a system identity needs to preserve. They also show why changing prose alone can leave the material behavior untouched.

We propose defining equivalence relative to a declared observation and action contract. Two harness versions are equivalent for a position-cap question if every relevant account transition produces the same allowed or rejected action and the same subsequent exposure accounting. They may still differ in wording, latency or another contract. Equivalence belongs to a specified question, rather than a universal claim that the agents are identical.

The environment closes the loop

In a single-answer task, a response can be scored without creating tomorrow's input. A settled trading action changes capital, exposure and often the market other agents observe. Repeated waiting changes the time at which new evidence is sampled. Memory changes what an earlier rationale contributes to later context. The transition must therefore include those consequences whenever they affect the assessed behavior.

The continuous record makes this dependency visible across two historical fleets. A shared five-slider lineage produced different text-versus-control conflict resolution. The inherited interface alone failed to establish inherited semantics. We read that contrast as a reason to test transitions during transfer, rather than assume that familiar controls define the same agent.

For builders, the useful comparison becomes precise: which observation, admission or state-update relationship changed? A model upgrade can improve proposals while a state-accounting regression worsens executed behavior. Recording both gives research a falsifiable object. The harness is the structure that turns an answer into a continuing decision system, and its identity should travel with every result.

Sources

Related field notes