Planning when a tool answer may arrive incomplete

By DX Research Group · · Frontier research

A proposed planning study varies tool reliability and tests how agents value information before acting.

A trading agent should plan around the possibility that a research tool returns late, incomplete or inconclusive evidence. We propose varying those response conditions in a fixed decision problem and measuring whether the agent changes its research plan sensibly. The experiment is unrun and addresses planning under uncertainty, beyond merely admitting a late result into context.

The controls paper companion ties a typed action to its observed inputs and subsequent policy check. That structure makes the missing evidence visible. The continuous record companion provides historical motivation for testing the surrounding harness while preserving the limits of its directional results.

Information has a deadline and a distribution

In an illustrative decision, an optional tool can resolve whether a catalyst is confirmed before a thirty-second opportunity expiry. It succeeds in time with probability 0.60. A useful response is estimated, for this fixture only, to improve the chosen action's expected outcome by $8 relative to the available fallback. The request costs $1. Its expected incremental value is 0.60 × $8 − $1 = $3.80, assuming late responses have zero value and the fallback remains available.

Add a second tool with a 0.90 timely success probability, a $5 improvement and a $2 cost. Its corresponding value is $2.50. The first tool has higher expected incremental value under these assumptions despite being less reliable. If waiting for it causes the fallback opportunity to disappear, the calculation changes. The fixture must specify that opportunity cost rather than reward the model for choosing the more reliable tool by habit.

These amounts are hypothetical utilities used to test planning consistency. They establish no market edge. In the initial experiment, supply the reliability distributions directly. A later study can ask whether agents estimate them from prior tool receipts, with a separate held-out assessment of those estimates.

A response tree the grader can inspect

Require a plan to state the intended request, the latest useful receipt time and the action for each response status. Use tool outcomes generated from a fixed saved schedule so all policies encounter comparable failure patterns. Vary correlated failures as well as independent failures: two providers relying on the same upstream source may supply less diversification than their separate names imply.

Score expected-value consistency, deadline compliance, factual support and policy validity. An agent may choose a reasonable request and still act as if an unanswered question were resolved. That should appear as a distinct failure in the trace. Include a no-request baseline, since investigation itself can consume a useful decision window.

We would use the resulting response trees to identify unnecessary requests and fragile plans. A successful planner would preserve an admissible fallback, update beliefs from actual responses and respect the original exposure limits. The next experiment would test whether those planning improvements survive realistic tool distributions, with decision economics measured independently.

Sources

Related field notes