Compact State Snapshots Versus Long Context for Trading Agents

By DX Research Group · · State and memory

A proposed representation comparison that preserves modern context capacity while testing what current decisions actually need.

A large context window makes more information available. It also makes it easier to carry expired facts, repeated observations, and old instructions into a current trading decision. We would compare representations of the same evidence rather than assume either maximal history or maximal compression is best.

Our state-memory research describes structured recent state and source-labeled memory in the historical runtime. This note proposes a controlled comparison against a longer event history with the same underlying information.

Preserve the evidence in both arms

Build an illustrative episode with a user limit update, an acknowledged order, a partial fill, and a later fee. The compact arm presents the current mandate, reconciled portfolio, outstanding order state, and references to the events that produced them. The history arm includes the full chronological record and those same authoritative current fields.

Both arms must expose the same decision-relevant facts. Removing the fee from the compact arm would test missing information rather than compact representation. Omitting the active mandate from the history arm would create an avoidable instruction-resolution disadvantage.

Choose a modern model and verify its actual context capacity, serving limit, and truncation behavior for the run. The experiment can use a long context without filling every available token. Record actual input length and any compression method, so another team can understand the comparison.

Inspect the cost of compression

Give the model a proposed trade that depends on both remaining capital and the active position cap. Score whether it identifies the current limits and whether its typed action respects them. Preserve rationale citations to determine whether it used current state or an obsolete historical number.

Then ask an audit question about why the current holding differs from the initial one. A compact state can support the action while omitting the evidence needed to explain it. Event references should let the audit retrieve the original fill and fee rather than force the model to invent a history.

We would report decision correctness, obsolete-fact use, audit reconstructability, input tokens, and latency. A smaller input that passes the action task but fails reconstruction has a different tradeoff from a long input that cites an expired mandate. These are representation outcomes; market skill needs its own scoring target.

The public evaluation registry freezes model and policy components while injecting state transitions. Its fixtures remain unrun. This comparison would add representation as an explicit intervention and retain the same underlying event set.

Interpret a result within its workload

The operating-layer paper companion supports linked state and outcome diagnosis in a bounded historical deployment. It supplies no matched compact-versus-long-context result. A favorable finding here would belong to the selected episode population, model, and task distribution.

DXAP's publicly described persistent agent, recorded decisions, and policy check make both current-state decisions and historical reconstruction relevant. We would assess those needs together when evaluating future context designs. A large window and a structured snapshot can coexist: the snapshot identifies current authority, while linked history supports targeted investigation.

Read the benchmark card before interpreting any comparison. The strongest result states what evidence each representation carried, what each decision required, and which errors remained after the representation changed.

Sources

Related field notes