What a Trading-Agent Rationale Can Actually Establish

By DX Research Group · · Trace evaluation

How to test explanation consistency against recorded inputs and actions while keeping causal claims bounded.

An agent's explanation is useful evidence about the text it produced. Its causal role in the trade needs a separate test. We use that distinction to keep an attractive rationale from becoming a substitute for recorded inputs, policy checks or execution outcomes.

Read the explanation against the trace

A proposed rationale audit starts with three questions. Does each factual statement have support in the information delivered to the agent? Does the explanation agree with the typed action? Does its description of a constraint agree with the actual policy result? These checks evaluate observable consistency and can be reviewed without inferring hidden reasoning.

Consider an illustrative rationale that says a position was reduced because an owner imposed a maximum exposure. The submitted action increases exposure, and the policy record rejects it. The text is inconsistent with the proposal even if the final account exposure stays unchanged. Crediting the rationale for restraint would attribute the policy engine's rejection to the model.

The operating-layer controls paper companion records fabricated named rules in sell traces and a compound intervention associated with a reduction from 57% to 3% in the affected pre-launch tests. Exact per-arm denominators remain unavailable in the published account. This is evidence about trace behavior under that intervention, with a clearly bounded setting.

Citation can change without economic improvement

The same account reports that moving an unchanged fee sentence raised fee citation from 3% to 74% in the affected reasoning traces. The reported outcome is citation frequency. A rationale that mentions fees more often can become more informative to an auditor, while the effect on action quality or realized returns remains a separate empirical question.

A proposed extension would score factual fee recall and action consistency on the same replay fixtures. Then it would measure whether the chosen action changes after the placement treatment. Finally, a declared economic evaluator could score those changes under its execution assumptions. Preserve the outcomes as separate columns so improvement in explanation quality remains visible without implying improved trading skill.

For sensitivity testing, alter one supported input feature at a time. A rationale that cites a funding rate should respond coherently when the funding-rate fixture changes. A failure to respond is a useful diagnostic. A response alone still leaves several possible mechanisms, including surface imitation, so the claim should stay at measured sensitivity.

Explanations as a repair interface

The continuous record describes attempts to elicit liquidation-distance statements that failed to improve sizing. That finding illustrates why explicit statements need action-level verification. The record also separates historical model behavior from improvements in the evolving product.

DXAP lets users inspect recorded decisions. Our proposed audit makes that surface more useful: every scored explanation points to its supporting input, typed proposal and policy result. Reviewers can identify unsupported assertions and action contradictions directly. A good rationale becomes a navigable account of the recorded decision, while causal attribution remains attached to the intervention evidence that actually supports it.

Sources

Related field notes