Reward feedback that changes a test, rather than repeated activity
By DX Research Group · · Incentives and participation
A proposed feedback rubric values reproducibility and coverage without rewarding trading churn.
The most useful feedback makes a failure testable. A participant who identifies a repeatable contradiction can help improve an agent platform without generating another trade. We would design feedback recognition around the evidence a reviewer can use and the distinct behavior it exposes.
DXAP's documentation describes a configurable agent whose instructions, policies and schedule govern subsequent turns. That creates concrete feedback targets: whether an instruction was carried forward, whether an order met the configured boundary, and whether the decision record explains a wait. A generic statement that the agent should be more active supplies little information about any of those stages.
Two submissions, different value
In an illustrative review queue, submission A says “it traded too little” twenty times across consecutive quiet turns. Submission B identifies one case where a confirmed instruction was absent from the next eligible turn, includes the relevant timestamp, and explains the expected behavior. Counting messages makes A appear larger. Counting distinct, reproducible failure cases makes B the stronger contribution.
We propose a review rubric with three recorded judgments: can the reviewer locate the event, can the expected behavior be established from an authorized instruction, and does the case add coverage beyond an existing report? Each judgment needs a short reason. A contribution can be useful even when review establishes correct behavior, if it exposes a confusing interface or a missing explanation. That result should receive its own classification.
For a worked allocation rule, suppose a future research program has twelve recognition slots. Reserve eight for independently reproducible cases and four for reports that substantially improve owner understanding or documentation. Review duplicate cases together, recognize useful additional evidence, and avoid multiplying awards merely because the same issue appears in many messages. The allocation is a proposed study rule, rather than a statement of current DXAP rewards.
Turn the report into an evaluation case
The controls paper companion describes the historical practice of tracing failures to a stage and testing revisions against replayed scenarios. We would extend that discipline to participant reports by recording the original trigger, observed behavior, expected behavior and review outcome. Sensitive account material belongs in an authorized support channel; public summaries can describe the mechanism without exposing private records.
A verified case can become a regression fixture. A second, withheld case should vary the instruction timing or account state so the repair has to preserve the underlying rule. If a proposed training experiment uses these reports, label provenance and consent need to be established before examples enter a dataset. The training proposal would then require held-out evaluation of compliance and forecasting separately.
The economic question remains distinct. Reduced confusion and fewer policy failures may improve the experience, but return improvement requires its own matched evaluation after costs. Rewarding a report on that basis gives participants a reason to inspect carefully without asking them to churn positions or manufacture a loss. The next useful measurement is reviewer time per resolved, distinct issue, alongside coverage of failure families.