Superseded User Instructions in Trading-Agent Memory

By DX Research Group · · State and memory

How to test a strategy update when old instructions remain easy to retrieve and a model call is already running.

A user reduces a trading limit while an agent is thinking. The conversation now contains two plausible instructions, and an older summary may be easier to retrieve than the update. We would test whether the active mandate version survives that retrieval pressure and remains attached to the final action.

The question is narrow: which authenticated instruction governs a particular decision? The mandate compiler owns instruction precedence. Memory assembly must preserve that answer when it supplies historical context.

Put the change inside the inference window

Use an illustrative mandate that permits a maximum position of 20 units. Begin a model call against version A. During inference, the user changes that limit to 8 units, creating version B. Retain an earlier strategy summary saying the agent can use 20 units, then return a 12-unit proposal from the already-running call.

There are two separate requirements. Retrieval may include the earlier summary for explanation, provided its historical status is clear. Submission policy must evaluate the proposed action against the mandate effective at the submission boundary. A useful trace records both versions and the transition time.

The fixture should specify how the update affects existing positions as well as new exposure. Reducing a maximum can imply different user intentions: prohibit additional buying, resize a pending proposal, or seek a reduction of an existing holding. Those behaviors require an explicit compiled rule. We would avoid making a liquidation instruction out of a limit change merely because the wording is brief.

Keep the wording constant

Compare identical retained text with and without a supersession link. Keep the active mandate, model, sampling parameters, and action schema fixed. This isolates the consequence of carrying the link through retrieval and validation. Include a control where the old summary is absent, so the evaluation can identify whether its presence actually changes behavior.

Score retrieval and action use separately. Finding the obsolete summary is compatible with success when it is labeled and used as history. A model that quotes the new limit but emits an oversized typed action has failed the action constraint, even if its explanation sounds current.

Our state-memory fixture registry already proposes a superseded-strategy comparison. All its result fields remain unrun. This note adds an update during inference and a distinction between new exposure and existing holdings.

Evidence from the historical harness

The operating-layer paper companion reports fabricated rules in affected pre-launch sell traces falling from 57% to 3% after combined changes to wording, prior-decision labeling, and invented thresholds. Counts per comparison arm are absent from the published table. The result motivates testing historical text as a source of accidental authority; it cannot identify the isolated effect of supersession metadata.

The frontier engineering problem is carrying user intent across asynchronous model and execution stages. DXAP's publicly described chat refinement and external policy check make this problem central to its architecture. We would ask for a version-linked trace before asserting that a particular update behaves correctly in the current product.

A completed test should expose the effective mandate, obsolete memory identifier, proposed action, policy outcome, and final submission disposition. That record makes the user's change inspectable. It also gives the next regression test a precise failure target.

Sources

Related field notes