Evaluate Learned Tool Routing Against a Fixed Policy
By DX Research Group · · Learning theories
A learned router earns its complexity by choosing useful information under the same available tools.
A learned tool router should be compared with a clear fixed routing policy under the same tool availability and decision inputs. Calling fewer tools is useful only if the retained information supports the task; calling more tools can improve scores simply by buying additional observations. We propose a budget-matched routing evaluation.
Off-policy evaluation research makes the logging policy and action support relevant when learning from prior choices. Tool routing inherits that issue: an uncalled tool has no observed response in the original trace. Our harness-transfer tests freeze tool boundaries, while trace feedback follows responses into later decisions.
Price the route at the parent turn
Consider an illustrative fixed policy that always calls a chart tool costing 2 units and a news tool costing 3 units. Its cost is 5 per turn. A proposed adaptive router calls both tools on 40% of turns, only the chart tool on 40%, and neither on 20%. Average tool cost is 0.4 times 5 plus 0.4 times 2, or 2.8 units.
That 44% tool-cost reduction omits router inference. If routing itself costs 1 unit per turn, total routing-plus-tool cost becomes 3.8, a 24% reduction relative to the fixed policy's 5. The accounting should also include retries and any parent-model consumption of larger tool outputs.
We would compare those costs alongside answer quality, valid action coverage, and mandate compliance. A cheap router that skips information needed to verify a position limit can save resources while worsening the operational task. A router that sends every difficult case to both tools may still be valuable if it reliably identifies those cases.
Missing responses limit offline claims
The proposed test uses frozen response bundles for cases where every eligible tool response is available at the same information cutoff. This allows each route to reveal only its selected responses. Cases with incomplete bundles enter a separate coverage report rather than receiving invented tool outputs.
Compare the learned router with always-call, never-call, and a registered deterministic condition such as calling news only when a source-age flag exceeds a threshold. The condition and threshold should be selected on development data. Keep the final case set hidden during route choice.
Measure routing errors by their downstream effect: omitted decisive evidence, redundant requests, stale-result admission, and correct economical routing. Tool-selection accuracy alone assumes there is a unique right route, which may be false when two sources provide equivalent facts.
The decision to learn routing should follow a visible cost-quality frontier on the supported cases. It remains a proposed experiment here. A fixed policy can be the appropriate outcome when it matches quality at comparable total cost or when logged data cannot support evaluation of alternative routes.