Use owner comparisons to teach answer order and length

By DX Research Group · · Data and learning flywheels

A proposed pairwise dataset holds facts and actions fixed while owners compare presentation choices.

Owners often recognize a useful answer more easily than they can write an ideal one. We propose asking them to compare two factually equivalent responses, then curating the judgments into a preference objective for answer order and length. Holding the facts and action fixed keeps the learning problem specific.

DXAP’s chat guide supports asking about recent decisions and exploring a strategy with the selected agent. Those interactions create concrete presentation questions: which fact should appear first, how much explanation helps, and when does compression hide something material? Pairwise judgments could answer these questions without asking users to become technical annotators.

Here is a fictional comparison. Both responses describe the same waiting decision. Response A opens with a broad market recap, then identifies the absent condition. Response B opens with “Your entry condition was absent,” gives the supporting observation and follows with a short recap. An owner prefers B because it answers the question sooner. That reason supports an ordering label.

Now offer B against response C, which is even shorter but omits the observation supporting the decision. A preference for B supplies a different lesson: retain enough evidence to inspect the answer. Collapsing both comparisons into “shorter is better” would teach the wrong objective. The curation record should retain the chosen response and the reason category, alongside any free-text explanation.

Make the comparison fair

Randomize left and right placement. Remove product-version cues and keep typography equivalent. Before collecting a judgment, check that the responses state the same action, applicable constraint and uncertainty. If one invents a fact, the pair belongs to factual correctness evaluation, even when the owner says the shorter answer feels better.

Allow a tie and an “insufficient context” response. Forcing a winner converts weak judgments into confident labels. An owner may also prefer different lengths for routine updates and strategy reviews. The question type and explicit preference scope must accompany the pair so the model can learn that conditional relationship.

An illustrative dataset contains thirty comparisons about routine activity questions and thirty about strategy reviews. If routine questions favor concise responses in twenty-four cases while reviews favor expanded evidence in twenty-two, a pooled winner count hides the conditional lesson. We would train and evaluate with the task context attached rather than nominate one global answer length.

The controls paper companion reports that changing the placement of a fee sentence affected citation in historical traces. That bounded observation makes ordering worth studying; it does not validate owner comparisons or show that a preferred layout improves economics.

Inspect the learned behavior

The proposed evaluation would hold out owners and response subjects, then ask whether the candidate better matches declared preferences while preserving required facts. Report preference wins, omitted evidence and unsupported claims separately. A model can win comparisons through agreeable wording while becoming less useful for checking a decision.

Pairwise data become a training objective only after the comparison is interpreted. “B wins because it answers the question first” is a teachable relation between question and response. “B wins because it sounds confident” may require review for unsupported certainty. Our proposed contribution is that interpretation step: user judgments become conditional labels for a bounded communication skill, with evidence preservation tested alongside preference.

Sources

Related field notes