Tool Availability Is Part of a Fair Trading-Model Comparison

By DX Research Group · · Trace evaluation

A procedure for separating tool-use skill from information access in trading-agent model evaluations.

A model that receives a richer tool response has more information to trade with. A model that knows how to ask a useful question has a different capability. Those effects can both matter, but a single leaderboard often mixes them. We would separate access to information from skill in using the available interface before interpreting a trading-model comparison.

Equal access has several meanings

A fixed-context test gives every candidate the same rendered evidence and asks for one typed action. It evaluates the decision from that input. A tool-enabled test gives every candidate the same tool manifest and response service, then lets each candidate choose requests. It evaluates the interaction as well as the final action. Both are useful, provided their titles identify what was tested.

The continuous record reports a frozen-context replay league. Its result belongs to that information boundary. It provides no direct ranking of research workers that can search, inspect charts or assemble new evidence during a turn.

For a proposed tool-enabled extension, normalize the semantic tool contract. A request for recent candles should return the same interval, missing-value policy and timestamp convention across providers. Model-specific function-call syntax can differ while the available economic information remains identical. Save each request and response so an evaluator can determine whether a difference arose before or after retrieval.

Build a coverage table before a score table

An illustrative test has 20 scenarios. In four scenarios the captured tool archive lacks an order-book response. Candidate A requests that response in all four; candidate B never requests it. If the archive silently supplies live order books to A, the test leaks future information. If it silently fails A's requests, the test may favor B's interaction style.

Declare the handling rule in advance. One option is to score the incomplete scenarios separately and report their proportion. Another is to restrict the primary comparison to scenarios whose permitted response set is complete. Either choice changes the population, so retain the all-scenario coverage table and explain which question the primary estimate answers.

Report requested-tool coverage by candidate, response availability by scenario and terminal completion rate. A low final score after frequent missing responses is a harness result until the missingness is resolved. A candidate that receives valid responses and misinterprets them supplies a more direct model-behavior finding.

How this bears on an integrated product

The operating-layer controls paper places tool calls inside a trace joining mandate, validation and execution. That allows a retrieval failure to remain distinguishable from a proposal failure. It also means the evaluator should inspect the exact response the agent consumed, rather than an idealized API response reconstructed afterward.

DXAP publicly describes research workers inside its harness. Evaluating that capability calls for a second track beside frozen-context decisions. The resulting report should say which candidate uses the interface well, which candidate decides well from identical evidence, and where the response service constrains both. That distinction helps us improve an agentic trading system without attributing every infrastructure effect to the model.

Sources

Related field notes