Retrieval Recall Under a Fixed Token Budget
By DX Research Group · · State and memory
Compare memory retrieval by required evidence delivered inside the same rendered budget.
Memory retrieval should be compared at the token boundary the parent agent actually receives. A retriever that finds the right item but places it beyond truncation has failed to deliver that evidence for the decision. We would score the final rendered context under a fixed budget.
This extends the provenance requirements in our state and memory framework. It also supplies an inspectable stage for trace feedback: the available corpus, selected items, and delivered spans are separate records.
Finding an item is only one step
Consider an illustrative query with five required evidence items in a frozen corpus. Retriever A selects all five, plus long distractors. Its final 2,000-token context contains only three required items after truncation. Retriever B selects four and renders all four inside the same budget. Delivered item recall is 3/5, or 60%, for A and 4/5, or 80%, for B.
This fixture does not establish a general ranking between the retrievers. It shows why pre-render recall can disagree with evidence actually delivered. We would record tokenizer version, formatting overhead, and the truncation rule, because a nominal budget excluding source labels gives one method extra room.
Item recall is also insufficient when a relevant document contains several needed facts. A summary may retain its title while dropping the account or time qualifier. We would map required facts to source spans and score delivered fact recall alongside item recall. The annotation should identify why each fact is required for the task, rather than treating every semantically related passage as relevant.
Hold the expensive boundary fixed
The proposed comparison freezes corpus version, query set, and maximum rendered tokens. Retrieval methods can use different ranking algorithms, but their compute and latency remain reported. If one method performs an additional model call, that cost belongs in the comparison even when the parent sees the same 2,000 tokens.
A paired query suite should include a short answer supported by one item, a multi-item question, and a question with a highly similar expired record. We would also include a query with no supporting evidence. That last case tests whether the retriever or parent can expose absence without filling the budget with persuasive substitutes.
Downstream action correctness is a second measurement. Higher delivered recall may increase distracting material, while lower recall may still include the one decisive fact. We would inspect disagreements rather than assume retrieval metrics translate directly into agent quality. The full-context reference is useful for diagnosing evidence availability, provided its extra token allowance is clearly labeled.
A budget study ends with a curve across several declared budgets and the marginal cost of each method. The numbers above are a worked example, with no empirical curve claimed. For a frontier runtime, the practical decision is where another retrieval step buys evidence the parent can use. A stable budget makes that question measurable without confusing a larger prompt with a stronger retrieval mechanism.