A persistent agent record creates options for improving the whole system
By DX Research Group · · Data and learning flywheels
A proposed reconstruction test measures whether retained decisions make competing repairs distinguishable.
A persistent trading agent can produce something more useful than a sequence of answers: a record of how the same owner mandate interacts with changing account state. DXAP combines that continuity with configured execution checks and recorded outcomes. We see a system-level learning opportunity in the ability to revisit a failure, distinguish its causes and test a repair against the original conditions.
The public execution documentation separates a model's action request, execution checks and venue outcomes. The distinction creates several possible repair locations. An incorrect quantity could originate in the proposal. A correctly rejected request could be explained badly. A submitted order could have an unresolved execution outcome. Retaining those differences makes the next investigation more precise.
Measure the questions a record can answer
Suppose an illustrative owner asks why the agent opened another symbol when a position cap was configured. Three hypotheses deserve inspection: the configuration changed after the request, the account snapshot omitted another position, or the execution check used the wrong state. A final chat answer cannot distinguish them. A joined record of the effective setting, account snapshot and checked request can.
Our proposed reconstruction test would give independent reviewers two representations of the same consented incidents. One contains the final user-visible explanation. The other contains a minimized, permissioned record of the relevant stages, with the resolution withheld. Reviewers choose among predefined causal hypotheses and identify the missing observation that would resolve uncertainty. Score correct localization and confidently incorrect localization separately.
For an illustrative twelve-incident fixture, imagine the explanation-only representation resolves three incidents, while the joined record resolves nine. The six additional resolutions would measure diagnostic availability on that fixture. They would justify better retention or joining of those fields, rather than establish improved trading decisions. A subsequent repair comparison asks whether correct diagnosis actually reduced recurrence on unseen cases.
Persistence has a cost and a scope
The October 3 release notes describe shipped agent-scoped chat memory when enabled. Saved preferences and adopted theses can carry across conversations, while memory stays separate from trading notes and leaves strategy and settings unchanged. This is useful owner continuity. A research record has a different purpose and requires its own permitted use; a saved preference becomes neither a cross-owner label nor a training entitlement simply because it persists.
We would retain fields according to the competing hypotheses they help distinguish. Keeping every repeated sentence increases review cost. Dropping the setting's effective time can destroy the ability to resolve a race. The design question is which retained information preserves a useful comparison, with account details restricted to the reviewers who need them.
Our historical continuous record already demonstrates the research value of looking across configurations and runtime versions, including findings that failed to support a directional edge. The forward possibility is a platform that becomes faster at identifying its own repair opportunities. The first receipt should be a reconstruction result and the next an unseen-case repair result. Those two measurements turn a large archive into a usable learning asset.