Testing Whether Retrieved Memory Is Relevant to a Trade

By DX Research Group · · State and memory

A relevance fixture that separates semantic similarity from current decision eligibility and downstream usefulness.

A retrieved memory can resemble the current question closely while describing a different account, venue, or mandate. We would evaluate relevance through decision dependencies rather than relying on similarity alone. The key outcome is whether the item improves or distorts the specific action under consideration.

The state-memory framework describes why old regimes and superseded settings can appear plausible in retrieval. This note turns that concern into a small distractor test.

Build distractors that sound useful

Use an illustrative query about buying asset X on venue A for account 1. Prepare three retained items with similar wording. One concerns the same asset on venue B. Another concerns venue A under an obsolete mandate. The third records a recent, source-labeled observation for the correct venue and account.

Preserve the text similarity while varying the eligibility metadata. A retrieval system may rank the wrong item first because its phrasing matches the query more closely. The runtime should still determine whether its subject, time interval, and authority fit the active decision.

Include a control query for which no retained item is eligible. A system forced to return something can make empty memory look like a defect. In a trading decision, a typed “no eligible memory” result can be the accurate answer and a useful signal to fetch current information.

Measure three different stages

We would measure candidate retrieval, eligibility filtering, and downstream use separately. Candidate retrieval asks whether the relevant item appears in the returned set. Eligibility asks whether obsolete or mismatched items are labeled correctly. Downstream use asks which items actually change the proposal or rationale.

Report false inclusion and false exclusion alongside recall. Excluding every memory could prevent contamination while discarding useful context. Including every memory could make retrieval recall look strong while passing obsolete facts directly into the decision.

To test causal usefulness, compare the same current snapshot with the eligible item included and withheld. Keep model, prompt structure, mandate, and sampling fixed. Repeatable changes in a rationale are a behavioral result; improvements in forecast scoring or simulated decision economics require their own measured targets.

The public evaluation fixture registry separates retrieval identifiers, state resolution, typed actions, and policy outcomes. Its comparisons remain unrun. Our distractor construction adds account and venue mismatch controls to that method.

What the published work leaves open

Our operating-layer paper companion documents structured state and linked traces in a bounded deployment. The related state article says open-ended recall was not obviously helpful in that setting. That observation leaves the value of a particular retrieval method open until a matched comparison measures it.

DXAP's public harness description includes indexed data and recorded decisions. Those features can support a useful retrieval study, but the current homepage provides no benchmark result for this distractor fixture. Transparent eligibility rules and published error cases would make the implementation assessable.

Read the benchmark card for the larger evaluation record. A useful memory result names the item, shows why it was eligible, and records the decision change it caused. That gives retrieval a concrete purpose beyond producing a plausible collection of related text.

Sources

Related field notes