Count Reposts as Distribution, Separate from Independent Claims

By DX Research Group · · Market data

A post-reference fixture keeps social reach visible while preventing copied claims from becoming false corroboration.

Keep circulation and independent evidence as different social features. A hundred reposts can indicate attention, but they may all repeat one underlying claim. An agent interpreting that volume as a hundred corroborating sources receives a distorted evidence summary. We would retain the distribution graph and separately count the claim's identifiable origins.

X's data dictionary provides referenced_tweets for relationships such as reposts, quotes, and replies. These references support a graph of explicit relationships. Copy-pasted text without a reference needs an additional, uncertainty-aware procedure.

Six posts contain two originating claims

Assume six fictional posts captured in a complete observation window. One original says a protocol upgrade is live. Three posts repost that original. A fifth quotes it with the words “excited to see this.” A sixth links an independent official release confirming deployment.

PostsRelationshipNew factual evidence in fixture
1Original claimOne originating claim
2 to 4Reposts of 1Distribution only
5Quote of 1Reaction only
6Independent release linkSeparate supporting document

Raw activity is six posts. Explicit origin count is two in this fixture. Even that count says little about correctness until the documents are inspected. The quote contributes a reaction and possibly reach, while its added text supplies no independent deployment evidence.

Preserve a graph before flattening a summary

Retain each post ID, author, referenced IDs, and relationship type. Build the origin grouping from explicit relationships first. Quote posts can add novel facts, disagreement, or firsthand evidence, so inspect their added content before inheriting the original's classification.

For unlinked copies, normalize only enough to compare text, preserving the original. Near-duplicate matching can group likely copies, but record its threshold and uncertainty. Similar words can also arise from independent descriptions of the same event. A fuzzy match should remain a grouping hypothesis rather than proof of coordination or shared control.

Keep attention measures such as unique observed authors and repost velocity separate from evidence measures such as independently supported claims. Describe the observation window and collection coverage. An incomplete API response can make both measures lower bounds, and account counts should never become a bot classification merely through repetition.

Change one feature channel at a time

A proposed agent fixture would provide the six posts as raw text, then as a graph-aware summary. Freeze the underlying facts and show the same attention count in both conditions. Compare whether the agent falsely describes six independent confirmations, with economic decisions scored separately.

Our market facts and instructions note covers how quoted text enters a prompt without becoming authority. Our research-subagent value test offers a way to evaluate the resulting summary. The contribution here is a concrete origin-versus-distribution contract. Whether attention or corroboration predicts a useful market outcome requires a held-out study with explicit trading costs and information cutoffs.

Sources

Related field notes