Handling Tool Results That Arrive After an Agent’s Decision
By DX Research Group · · State and memory
A proposed procedure for late research responses whose original request context has already been superseded.
A research tool returns after the agent has moved to a new decision. The result can still be informative, but its original question may refer to a portfolio, mandate, or event status that has changed. We would classify the late response before allowing it into current decision context.
The state and memory research binds a decision to its snapshot. This note extends that binding to asynchronous tool requests and their eventual responses.
Preserve the request that produced the answer
Consider an illustrative request made at 10:00: “Assess whether another 5 units of asset X fit the current portfolio.” At 10:01 a fill changes the holding. At 10:02 the tool returns an analysis based on the original portfolio. Its market observations may still be useful. Its position-size conclusion depends on an obsolete state.
Store the request identifier, parent decision identifier, snapshot version, mandate version, and dependency fields. The response should carry those identifiers back. Arrival time identifies when the runtime received the answer; it says little about the freshness of the facts used to produce it.
We would divide the response into claims whose dependencies can be checked. A cited venue rule, an observed price, and a portfolio-specific conclusion may each need a different refresh. If the tool supplies only an opaque paragraph, the runtime can conservatively treat the whole answer as dependent on its original snapshot.
A response admission procedure
First compare the response's parent decision with the active decision. Then compare its required state versions with current versions. Admit eligible observations with their source times, label historical conclusions, and recompute derived portfolio claims when the relevant inputs have changed.
The fixture should include a late duplicate response as well as a late first response. Repeated delivery must preserve a single logical result identifier. Otherwise one conclusion can appear twice in memory and gain apparent support through repetition.
Hold the tool output constant across both arms. In the baseline it arrives before the fill. In the injected arm it arrives afterward. Measure eligible-claim admission, obsolete-conclusion use, and whether the final typed action is rebound to current state. Record any recovery latency separately from the action score.
The unrun fixture registry already requires linked state, retrieval, policy, and reconciled outcome fields. This late-response procedure is a proposed specialization. A completed run would need exact request and response records in addition to the listed outcome fields.
Why this matters for research subagents
DXAP's current public page describes research and chart subagents within its harness. That makes asynchronous context a relevant product evaluation question. The homepage does not establish a specific late-result admission implementation, so we present this as a method to test rather than an available guarantee.
Our operating-layer paper companion supports the broader value of linking inputs, proposals, validation, and outcomes in the historical deployment. Its reported control interventions concern different failures and remain bounded to their original populations.
The useful frontier is maintaining the meaning of a tool result as the world changes around it. Read trace feedback for turning a failed late-response case into a regression fixture. The test succeeds when each claim reaches the decision with its original dependencies intact and an explicit current-use disposition.