External State Versus Model Belief in a Trading Runtime

By DX Research Group · · State and memory

How to test a model that insists an order filled when the authoritative venue state says it remains open.

An agent may confidently describe a fill that the venue has yet to confirm. The runtime needs a representation of the external order state that survives the model's interpretation. We would test that boundary by supplying a persuasive prior rationale alongside an authoritative open-order record.

Our state-memory research assigns portfolio and order facts to their own sources. This note asks how a model belief should be recorded when it conflicts with those facts.

Preserve the disagreement

Construct an illustrative order with a stable request identifier and venue status “open, zero filled.” Add an earlier model response saying “the purchase completed.” The current model then receives both records and is asked whether it can spend the remaining account capital on a second purchase.

The runtime should preserve the first order's actual status and any required capital reservation. It may record the model's statement as a mistaken belief, but that statement cannot produce a fill event in the ledger. A derived portfolio based on the imagined fill would already be corrupted before policy checked the next action.

Use explicit status values for proposed, submitted, acknowledged, partially filled, and terminal orders. The exact adapter may use different names, but the trace should show the mapping. “Success” is too broad if it can mean accepted submission, completed execution, or reconciled account state.

Move the evidence after the belief

In a second fixture, deliver a real partial fill after the incorrect completion claim. The agent must update to the partial state while preserving the order's remaining quantity. This tests whether the runtime reconciles newer external events or simply carries forward the stronger-sounding conclusion.

Freeze the model-authored completion sentence across arms. Vary the authoritative order status and compare the next state identifier, reserved capital, proposed action, and policy result. Count fabricated terminal transitions separately from wrong sizing decisions, because they are different failure mechanisms.

A recovery path should query the authoritative venue when acknowledgement is unknown. A completed query can resolve the uncertainty; another model explanation cannot. Record query time and response sequence so an investigator can establish what was known at the next submission boundary.

The unrun state fixture registry includes unknown acknowledgement and next-cycle reconciliation. The belief-conflict fixture is a proposed extension, with no measured current-product result.

State integrity is a distinct research target

The operating-layer paper companion describes linked portfolio snapshots, validation, and chain outcomes in the historical 21-day deployment. That method supports diagnosing where a belief became an account fact. Its published intervention results address their own cases and do not establish this test's outcome.

DXAP publicly separates model proposals from external policy and execution. We would evaluate that architecture by checking whether authoritative order status constrains every subsequent decision, including confident mistaken explanations. Passing would demonstrate state integrity within the fixture; forecasting skill remains a separate question.

Continue with execution and settlement for the complete order path. A useful trace leaves the disagreement visible: what the model believed, what the venue reported, which state governed policy, and how reconciliation resolved the next action.

Sources

Related field notes