Process Reward vs Economic Reward for Trading Agents

By DX Research Group · · Learning theories

A proposed two-axis evaluation tests whether better trading-agent process scores translate into decision economics under realistic costs.

A trading agent can improve at citing fees while its trading results remain unchanged. That possibility is central to our research. A process score tells us whether a specific behavior occurred. An economic score tells us what the resulting positions did under a defined accounting method. A useful experiment keeps both visible.

Our operating-layer controls paper reports a concrete process intervention: moving an unchanged fee sentence earlier raised fee citation in sampled reasoning traces from 3% to 74%. The comparison held the model, wording, and market data fixed. Exact per-arm counts are absent from the published account. The result establishes a bounded change in trace behavior, with return effects left unresolved.

Put the Mechanism Between the Scores

Suppose a proposed grader rewards every explicit mention of trading fees. An agent can earn that reward while placing the same order as before. The grader has detected attention, but the hypothesized mechanism requires fees to affect the action. We would therefore record three stages: fee recognition, estimated cost relative to the thesis, and the resulting action.

An illustrative trade expects a gross gain of $8 while estimated entry and exit costs total $10. A fee-aware decision might decline it. Another expects $30 against the same costs and may remain admissible under the mandate. Repeating the word fee in both explanations carries less information than showing how the expected margin changes the choice.

The arithmetic is illustrative. Real execution introduces uncertainty in the forecast, costs, fill price, and holding horizon. Each of those inputs needs its own source and timestamp before the example becomes a scored market case.

A Two-Axis Experiment

Compare a frozen baseline against one named intervention on the same saved cases. Report process adherence beside action changes. Then evaluate the economic consequences using one declared fill model, costs, and exit policy. Include no-trade outcomes so a lower action rate remains interpretable.

The informative cases sit off the diagonal. Higher process scores with unchanged actions suggest an attention change that stops before the decision. Higher process scores with worse economics suggest the behavior was applied inappropriately or the economic estimator was weak. Better economics with unchanged process scores may indicate a separate cause that the rubric missed.

Avoid collapsing these cases into a weighted total before inspecting them. A total can hide a material compliance regression behind a fortunate return. Hard policy restrictions belong in admissibility checks, while economic comparisons describe outcomes among admissible decisions.

What Counts as Progress

Our trace feedback framework treats diagnostic measures as evidence for a named failure and repair. The next study would register the expected causal path from attention to action to outcome, then measure where that path succeeds or stops.

DXAP's architecture separates model proposals from policy checks and execution. It offers a concrete place to examine that path as the harness evolves. These public capabilities support inspectable evaluation; they do not establish that process rewards are being used to train the product.

We would call a process intervention successful on its declared diagnostic first. We would describe an economic improvement only after the separate downstream comparison supports it, with costs and uncertainty attached.

Sources

Related field notes