Confidence Labels for Unverified Source Facts
By DX Research Group · · State and memory
Separate a fact’s verification status from the model’s confidence in its interpretation.
A source fact needs a verification status that survives retrieval and summarization. Model confidence describes a different quantity. We would keep an unverified claim visibly unverified even when the agent finds it persuasive, repeats it fluently, or uses a high verbal confidence label.
Our state and memory framework treats retained material as source-labeled claims. The trace feedback framework lets us examine whether that label survived into the decision. This note proposes a test of epistemic handling, with no implied result from a deployed agent.
A confident reading of a weak source
In an illustrative case, a social post says a protocol will unlock 10 million tokens tomorrow. The post gives no primary reference. A research worker returns, “Likely accurate, confidence 90%.” That probability may describe the worker's subjective judgment. It does not establish that the unlock schedule was independently checked.
We would store the claim text, source URL, retrieval time, and verification status separately from any probability estimate. The initial status might be “single-source, primary confirmation unavailable.” If the worker later retrieves an official schedule showing 8 million tokens on another date, the verified fact receives its own record and the discrepancy remains linked.
The numbers here are hypothetical. In particular, the 90% label has no calibration evidence. A useful record should say what would resolve the uncertainty: the official schedule and its applicable version. Merely changing “likely” to “high confidence” leaves the evidence unchanged.
Exercise the transformations
We would pass the unverified record through a summary, a retrieval result, and a parent-agent synthesis. At each stage, a deterministic check asks whether the claim retains its status and source identifier. A human review then checks whether the prose implies confirmation despite carrying a mechanically correct label elsewhere.
The paired case supplies the official schedule alongside the same social post. Expected behavior is to state the discrepancy and bind the factual schedule to its primary source. A third case supplies an official page whose date is ambiguous. That case remains unresolved rather than inheriting the appearance of certainty from an official domain.
We would distinguish label preservation from decision sensitivity. A model can preserve the label while sizing a trade as if the claim were confirmed. Conversely, it can abstain for an unrelated reason while losing the label. The scorecard should therefore retain factual status, cited support, and the stated dependency of the proposed action on that fact.
This approach also helps owners review a persistent trading agent. The useful question becomes which source establishes the claim and what remains unverified. It avoids turning a fluent research answer into a silently upgraded market input. A short field describing the missing evidence is often more informative than another confidence adjective. The acceptance target is that an unverified source stays identifiable after every transformation that could carry it into a future turn.