Resolve contradictory feedback without averaging away intent

By DX Research Group · · Data and learning flywheels

A proposed adjudication process separates factual disagreement from different owner objectives.

Two owners can watch similar agent behavior and give opposite feedback. One wants fewer entries; another wants faster participation. Averaging their ratings could teach an agent to satisfy neither. We propose resolving contradictions by identifying the mandate, evidence and kind of judgment each report contains.

DXAP lets owners refine strategy through chat while retaining approval for settings and persistent instructions, as described in the chat guide. That owner-specific direction matters. The same action can fit one strategy and contradict another, even when both owners receive identical market information.

First ask whether the cases conflict

In an illustrative pair, owner A says, “Waiting here was correct; my strategy requires confirmation.” Owner B says, “Waiting missed my entry; my strategy permits an early attempt.” If the saved instructions match those descriptions, the reports are compatible conditional labels. Each supplies a desired response under a different mandate. A pooled majority vote would erase the conditioning variable that explains the difference.

A harder case occurs when two reviewers inspect the same turn under the same instruction. One says confirmation was present; the other says the candle was incomplete. This is a factual dispute. The adjudicator should inspect the snapshot's observation time and candle status before reviewing the eventual market outcome. If the record lacks that status, the correct disposition is unresolved evidence, with the missing field identified.

A third disagreement concerns economics. Both reviewers agree the agent followed the strategy, but one calls the losing trade a bad decision. Here we retain instruction compliance and score the forecast or execution under a separately declared question. A loss can motivate deeper analysis without rewriting the historical mandate.

Decide who can resolve which question

We would appoint an adjudicator for factual and annotation disputes while leaving strategy ownership with the authenticated owner. A support reviewer can correct a mistaken label using the trace; that role supplies no authority to replace the owner's preferred strategy. Each disputed case would receive a short adjudication record: the competing labels, evidence examined, resolution and the change that might address it. A preference difference may require clearer strategy configuration. A factual discrepancy may require better state representation. An unsupported explanation may require a harness intervention. A weak forecast may warrant model evaluation against a reference.

The historical controls paper companion shows why that routing matters: published interventions addressed distinct trace failures. Its findings motivate targeted diagnosis, while each new feedback-derived intervention still needs its own test.

Preference aggregation would require a separate, explicit research purpose and permission appropriate to that use. Feedback about one agent should stay attached to that owner's intended strategy unless the owner agrees to contribute it to a broader preference study. Even with that agreement, a pooled preference would become a candidate design choice, with individual agent changes retaining their approval flow.

For a proposed study, independently label a frozen collection before adjudication. Report agreement separately for factual support, mandate compliance and owner preference. Then give a new reviewer either the original report alone or the report plus the adjudication context. Measure whether the added context improves reconstruction and reduces incorrect labels. The test would evaluate label usability, with any downstream model or release benefit requiring a further comparison.

We would also preserve minority cases that are reproducible and materially different. Popularity is a poor substitute for specificity: one precise report about a rare instruction transition can reveal a failure hidden by many satisfied owners. Conversely, repeated unsupported complaints can expose confusing product language even when execution is correct.

A useful learning loop retains these different signals. Owners remain authoritative about their intended strategy within their permissions; saved records establish what happened; research evaluates whether a proposed change improves the behavior. Contradiction becomes an opportunity to find the missing condition, rather than a reason to train on an ambiguous average.

Sources

Related field notes