Auditing Information Loss in Agent Summaries

By DX Research Group · · State and memory

Measure the decision-relevant facts a summary drops, changes, or merges.

A summary should be evaluated against the facts a later task requires, rather than its readability alone. We would inventory those facts before compression and test whether the shortened record preserves their relationships. A fluent paragraph can discard a qualifier that changes the admissible action.

Our state and memory article explains why retained context needs lineage. The trace feedback method makes a summary transformation inspectable as a distinct stage. This proposed audit focuses on information loss within that transformation.

Count claims, not sentences

Suppose an illustrative tool result contains six atomic facts: a position is long 2 units; its venue is A; available cash is 300; an order reserves 50; a quoted news claim is unverified; and that claim expires after its event resolves. A summary says, “Long 2 units, cash 300, positive news.” It preserves two values but loses venue identity, the reservation relationship, source status, and expiry.

Under an equal-weight inventory, preserved-fact coverage is 2 divided by 6, or one third. That is a narrow transcription measure. It says little about the consequence of each omission. If the later task is closing the position, venue identity may be decisive. If it is allocating cash, the meaning of the reservation matters more.

We would therefore predeclare task-specific critical fields as well as report unweighted coverage. A summary preserving five of six facts still fails a close-position fixture if the missing fact is the venue. Weighting after reading the generated action would make the metric too easy to reshape around the outcome.

A loss audit has several error types

Omission drops a fact. Mutation changes its value. Merge combines claims that should remain separate. Unsupported addition creates a new claim. These classes need different remediation: preserving more text addresses an omission but can leave a mistaken merge intact.

We would retain a mapping from each summary claim to its source span. The audit asks whether the span supports the exact claim, including units and qualifiers. A source saying “cash 300, including reserved 50” supports available cash of 250 under that stated definition. It cannot support “available cash 300.” This difference survives even when every number appears somewhere in the source.

The evaluation then gives a downstream agent either the full source or the summary under otherwise fixed conditions. Summary fidelity and downstream action correctness remain separate outputs. A model may recover a missing field through a permitted tool; that successful recovery should be recorded, along with its extra cost and latency.

The fixture should include a concise faithful summary and a longer misleading one. That comparison prevents token reduction from becoming the whole objective. These are illustrative inputs, with no measured improvement claim. The review artifact is a claim-to-source table and the first task that breaks under each loss. It lets us repair a specific compression behavior instead of asking the model to produce a vaguely better summary.

Sources

Related field notes