Unsuccessful Tool Calls Are Information-Access Failures

By DX Research Group · · Trace evaluation

Evaluate how agents respond when needed evidence is unavailable instead of grading only the final narrative.

An unsuccessful research tool call changes what the agent could know. We would grade information access and the response to missing evidence separately from the final proposal. An articulate rationale cannot establish that the requested market state was acquired.

The quote request fails before the decision

In an illustrative fixture, the agent needs a current quote to size an order under a maximum dollar exposure. The quote tool times out. The model then uses a remembered price of 100 even though the saved current quote would have been 125. A proposed ten-unit order represents 1,000 dollars at the remembered price and 1,250 dollars at the fixture's current price.

The audit should record three states: evidence attempted, evidence acquired and evidence consumed. The request establishes the first. A usable response establishes the second. A trace dependency showing that the price informed sizing supports the third. These states can diverge when a result arrives but is ignored, or when the agent uses an unsupported fallback.

OpenTelemetry's tracing API supplies status and event recording concepts. Our proposed rubric assigns the economic consequence separately. A span marked failed identifies a technical event; it leaves the adequacy of the agent's recovery unresolved.

Make the missing evidence change the admissible action

Construct a paired fixture with the quote present and absent. In the present arm, the expected size calculation uses 125. In the absent arm, the task contract requires an explicit unresolved-price status and prevents a newly sized order until an acceptable source arrives. This fixture tests behavior under an information restriction, rather than rewarding universal inactivity.

A recovery arm can supply a documented alternative quote with timestamp and instrument identity. Check whether the agent reports the substitution and uses its units correctly. A successful fallback needs evidence coverage, freshness and an authorization-compatible result. Merely calling another tool increases activity without demonstrating restored access.

The operating-layer controls companion explains why reliable behavior requires controls around the model. The continuous record companion supplies historical multi-tool context and its scoped performance boundaries. This note proposes an access-failure audit; it introduces no new current-product reliability figure.

For a worked report of twelve required quote requests, suppose nine return usable evidence and three fail. Access coverage is 9/12, or 75%. If two failed cases are handled according to the missing-evidence contract and one guesses, failure-response compliance is 2/3. Keep these distinct from the accuracy of the nine evidence-supported decisions.

The resulting review names what remained unavailable and whether the agent respected that limitation. It also distinguishes provider outage, connector parsing and unsupported query construction as possible repair locations. A tool failure becomes a concrete information constraint in the evaluation, with a traceable recovery outcome, instead of an invisible inconvenience between the prompt and a polished answer.

Sources

Related field notes