A feedback reward model needs three separate targets

By DX Research Group · · Data and learning flywheels

A proposed feedback model keeps owner satisfaction, mandate compliance and economic outcomes independently inspectable.

A single thumbs-up can describe a clear explanation, a compliant action or a profitable outcome. Those are different training targets. We propose a feedback model with separate outputs for owner satisfaction, mandate compliance and decision economics, so the optimization objective can be chosen explicitly rather than inferred from a blended rating.

DXAP’s activity guide separates decisions from submitted orders and fills. Its policy reference explains that a successful check establishes compliance with applicable execution checks, while outcomes still depend on the venue and market. A feedback record can preserve those distinctions before anyone trains a reward model on it.

Consider three fictional responses to the same owner question. The first is concise and appreciated, but incorrectly claims that settings were approved. The second accurately explains that approval remains pending, but takes too long to answer. The third is clear and correct, and the associated position later loses money. A blended favorable-versus-unfavorable label gives each example an ambiguous lesson.

Instead, ask the owner a focused satisfaction question about the answer. Derive compliance labels from the authenticated mandate and relevant policy record, with expert review where meaning is ambiguous. Attach economics at the appropriate decision episode, including the declared horizon and measured costs. A later loss changes the economic label without rewriting whether the answer was clear or the action authorized.

A vector before a scalar

In an illustrative scoring scheme, satisfaction and compliance each range from zero to one. A preferred but unauthorized response might score 1.0 on satisfaction and 0.0 on compliance. Averaging them gives 0.5, which can look comparable to a mediocre authorized response. A release objective that requires compliance first would reject that tradeoff rather than optimize the average.

The proposed reward model predicts separate dimensions from the inputs relevant to each one. Economics may be better evaluated with an explicit simulator or held-out outcome analysis than predicted from conversational approval. The word “reward” names an optimization signal; it gives no guarantee that a learned score reflects market value.

We would test the satisfaction predictor on new owners and tasks. Test compliance with contrast cases that change one mandate or approval fact while preserving style. Test economic scoring against complete episodes with a fixed action policy and costs. Each test asks whether that dimension responds to its intended evidence.

Users should know what they are judging

A useful feedback prompt is “Did this answer make the decision understandable?” That yields a clearer satisfaction label than “Was this good?” A separate correction field can identify a factual error or missed instruction. Participation incentives should recognize useful, reviewable feedback regardless of whether it praises the product. Otherwise the dataset can teach approval-seeking behavior.

Our controls paper companion documents bounded trace interventions whose performance meaning remains separate. The proposed model follows that discipline by retaining independent outputs through training and evaluation.

A downstream objective could optimize communication subject to preserved compliance, then examine economics separately. Another research objective might target compliant decision quality under a fixed economic evaluator. Naming the objective exposes the choice. User contributions become labels for the dimension they can actually judge, and a tested update has a precise claim rather than an opaque higher score.

Sources

Related field notes