Does a feedback improvement reach the next group of owners?
By DX Research Group · · Data and learning flywheels
A proposed prospective cohort study tests whether feedback-driven fixes generalize beyond their contributors.
An improvement developed from one group's feedback should be evaluated on the next group that encounters the product. Existing contributors know its language, may have learned workarounds, and can recognize the cases developers already repaired. We propose a prospective cohort study for feedback-driven agent releases that makes the next group an independent source of evidence.
The controls paper companion reports historical behavior under a bounded deployment and distinguishes observational language cohorts from controlled interventions. That distinction matters here. A later owner cohort can differ in strategy, familiarity, and market conditions. Comparing its raw outcomes with the earlier group would combine release effects with those differences.
Enroll before the answer is known
An illustrative candidate improves explanations of rejected entry requests. Freeze it after development on contributed cases. Define a prospective enrollment window and a common eligibility rule for owners who consent to the study. Keep contributors who supplied development cases in a separately reported group. New owners supply assessment cases only after the candidate artifact is frozen.
For the offline comparison, render each eligible saved case to both incumbent and candidate. Both receive the same strategy version, policy fields, and point-in-time account state. The evaluator grades whether the explanation identifies the actual rejecting condition and accurately describes the owner's available approval path. No study arm alters the owner's saved policy.
The DXAP policy reference identifies current checks such as entry and projected position notional, rolling entry limits, position caps, and re-entry cooldowns when configured. These provide public categories for assessment. A correct explanation should match the relevant rejection, with opening and closing actions treated according to their different roles.
Separate prevalence from repair quality
Suppose an illustrative new cohort contributes 100 eligible cases: 80 straightforward explanations and 20 involving several interacting checks. The candidate improves the straightforward cases from 72 correct to 76, while complex cases remain at 10 correct for both versions. Overall correctness rises from 82% to 86%. The defensible conclusion is an improvement concentrated in the straightforward group, with the difficult mechanism unresolved.
Now suppose the next cohort contains 50 cases of each type. Even with the same within-group accuracy, the overall score changes because the population changed. We would therefore publish both cohort-specific results and a standardized summary using a declared common case mix. Report paired disagreements and uncertainty grouped by owner, since repeated cases from one person share language and configuration.
The next release cycle needs a fresh boundary. Once researchers inspect this assessment cohort to design another fix, its cases become development evidence for that later candidate. A newly enrolled cohort then supplies prospective assessment. That rolling sequence supports continual improvement without recycling an inspected population as fresh proof.
This proposed study makes new users valuable for discovering unfamiliar failure mechanisms. It also gives them a clearer contribution contract: permitted cases inform evaluation, while their trading authority stays with their configured agent. Instruction compliance and explanation quality are the initial outcomes. Forecast skill and economics would require their own prospective questions and scoring.