Aggregate inference tokens do not set a per-turn budget

By DX Research Group · · DXRG findings

Interpret roughly 70 billion tokens across 7.5 million invocations without inventing context or serving limits.

DX Terminal Pro's roughly 70 billion inference tokens across 7.5 million invocations imply an average of about 9,333 tokens per invocation. That is a ratio of aggregate summaries. It cannot establish the maximum context length, the model's native limit or the runtime's per-turn budget. We should keep those configuration questions separate from consumed volume.

The controls paper companion publishes both rounded totals for the 21-day real-capital event. Its frozen runtime used Qwen3-235B-A22B-Thinking-2507 on H100 hardware through SGLang, in a twelve-token market. The scale table leaves the 70-billion total unpartitioned across prompt, completion, reasoning and cached-token categories.

A mean leaves the distribution unknown

An illustrative pair of invocations consumes 2,000 and 18,000 tokens. Its average is 10,000 tokens. A second pair consuming 10,000 tokens each has the same average but a different maximum and tail. A configured ceiling of 256,000 tokens would remain compatible with either pair if actual usage stayed below it. These examples illustrate possible usage distributions; actual Terminal Pro budgets require configuration evidence.

The aggregate average also depends on which events enter the token counter. If failed retries count tokens but are absent from the invocation denominator, the ratio means something different from successful-turn usage. If cached input tokens are counted differently across providers, cross-system comparisons may mix billing units with logical context size. The published summaries are insufficient to resolve those accounting details.

A concrete token artifact would record invocation ID privately, model version, served context ceiling, input tokens, generated tokens and counter semantics. Public reporting could release quantiles and an aggregate reconciliation without releasing prompts or individual identifiers. That proposed artifact would distinguish capacity, actual usage and cost while preserving the privacy boundary.

Consumption and information value need separate measures

The continuous record companion reports that adding 770,000 tokens of context in a paired comparison yielded no demonstrated directional benefit. Its statement belongs to that comparison and its evaluation contract. The result remains specific to that context comparison. Other representations and the earlier tournament's usage require their own evaluations.

Similarly, approximately 70 billion tokens divided by 21 days is about 3.3 billion per day on average. That derived throughput says nothing about peak serving demand or an individual agent's memory. The longest recorded prompt-state-action cycles exceeded 6,000, but cycle count is another unit and cannot be converted into a retained context size without the memory policy.

For operators, the useful question is which token category drives capacity and which additional information changes decisions. Historical aggregate consumption answers scale. A current serving budget needs configuration evidence, and an information-value claim needs a matched evaluation. Treating the average as a ceiling would collapse all three questions into a number that was never designed to answer them.

Sources

Related field notes