Measuring the information value of research subagents

By DX Research Group · · Frontier research

A proposed ablation distinguishes better information, changed decisions and net economic value when specialist research workers are added.

A specialist worker can produce useful research while adding little to a trading decision. DXAP currently describes research and chart-analysis subagents within its agent workflow. We want to measure their incremental information value rather than infer it from a persuasive report or low inference cost. The study below is PROPOSED, with no measured return benefit or assertion about which specialist configuration the product should ship next.

Three arms answer three different questions

The reference arm receives the main agent's saved state. The treatment adds the actual specialist output available before the decision. A third arm adds a length-matched summary of information already present in the reference snapshot. That placebo summary helps distinguish new evidence from the effect of reorganizing familiar information or simply increasing attention to a topic.

The specialist must work from a bounded source packet with receipt timestamps. We would save its questions, retrieved evidence and final answer. Claims that cannot be traced to a source receive an explicit status in the evaluation record. A correct conclusion reached with unavailable future information belongs outside a valid information-value comparison.

Inspect the changed decision, then its outcome

An illustrative fixture begins with a main agent considering a long position because price momentum is positive. A specialist identifies a scheduled event whose timing is absent from the initial snapshot. The treatment agent waits; the reference agent enters. We first ask whether the event information was accurate and timely, then whether waiting obeyed the mandate, and finally what happened under the fixed forward evaluation. One favorable fixture demonstrates a mechanism to investigate, rather than the average value of the worker.

Across the saved corpus we would measure factual additions, forecast score changes and action changes separately. A worker may improve factual coverage without moving the decision. It may move decisions while reducing forecast quality. It may improve forecast quality while generating trades whose additional costs consume the benefit. Those outcomes call for different engineering responses.

Include the opportunity cost of waiting

Inference spend is only one cost. Specialist latency can delay a time-sensitive order or cause a trigger condition to disappear. The evaluation should therefore score the complete decision timestamp and record cases where research arrives after its usefulness expires. We would compare specialist value by task type rather than combine chart interpretation, source verification and event research into one uninformative average.

The continuous record reports that adding very large context in a historical paired comparison produced no improvement, reminding us that more information requires an incremental test. The controls paper supplies the trace discipline for connecting rendered inputs to actions. Together they motivate a concrete path for DXAP's evolution: retain specialist outputs that improve defined decisions under held-out evidence, diagnose failures by task and measure latency alongside quality before expanding the research budget.

Sources

Related field notes