Evaluate a Revised Rationale While Holding Actions Fixed

By DX Research Group · · Learning theories

Clearer explanations deserve their own evidence rather than a trading-performance claim.

A revised rationale can improve factual support and owner understanding even when every trading action stays identical. We would evaluate that change as an explanation intervention with action invariance explicitly checked. It should earn a claim about the record that changed.

Our operating-layer paper distinguishes measured trace behavior from investment outcomes. Trace feedback provides the linked action and evidence record. Preference-learning research motivates evaluating human judgments, while the protocol below is our proposed action-fixed audit.

Freeze more than the action name

In an illustrative set of fifty saved turns, a rewritten explanation improves supported factual claims from 120 of 150 to 140 of 150. That moves claim-level support from 80% to approximately 93.3%. If each turn's typed action remains byte-identical after canonical serialization, the action-change count is zero of fifty.

The improvement describes these explanation claims. It does not establish better prediction, more compliant orders, or greater returns. Moreover, claim-level scoring weights verbose explanations more heavily. We would also report the fraction of turns containing at least one unsupported claim, so a few long texts cannot dominate the aggregate.

Action invariance covers side, quantity, asset, timing, and order conditions. A rationale rewrite that silently changes a stop price or inserts a delayed execution instruction has become an action intervention. The comparison should identify that change before assigning an explanation-only label.

Evaluate the record from the owner's perspective

We would give reviewers the authenticated mandate, decision-time evidence, and fixed action, then show either original or revised explanation under blinded ordering. Reviewers answer concrete questions: can they identify the cited fact, verify the stated constraint, and understand why this action was selected? Preference is useful, but the criteria need to distinguish attractive wording from supported content.

An adversarial subset includes a polished explanation for a flawed fixed action. The revised rationale should acknowledge the inconsistency rather than make the action sound justified. Another subset has a valid action with insufficient evidence to identify the model's internal cause. It should describe the recorded basis and uncertainty without inventing a causal account.

If later turns consume past rationales as memory, explanation changes may eventually change behavior. We would exclude that pathway from this initial fixed-action study and register a separate memory-use experiment. The direct readability result and the downstream policy result have different populations and mechanisms.

The useful outcome is a paired explanation audit with action hashes, factual support counts, and reviewer answers. A successful revision can earn a clear, modest claim: the saved decision became easier to inspect in these cases. That is a worthwhile improvement in a persistent agent's record, even while its trading-performance evidence remains unchanged.

Sources

Related field notes