An incentive change also changes the feedback dataset

By DX Research Group · · Incentives and participation

A prospective sampling audit checks whether new rewards alter the cases available for future training and evaluation.

A participation program can change the dataset before anyone changes the model. Owners respond to what receives attention or recognition, and those responses determine which failures reach reviewers. We would audit that sampling shift whenever an incentive changes.

DXAP's public points page links participation to points and referral benefits. Its dated Chapter 0 guide describes additional one-time recognition. These are documented programs. Using resulting reports for training would be a separate proposal requiring consent, provenance and evaluation. The sources establish recorded activity and review opportunities, rather than current online weight updates.

Freeze a reference population before changing the offer

A proposed audit samples 200 eligible recorded turns from each of two four-week periods around an authorized program change. Sample from the eligible turn population independently of user reports, keeping owner limits and privacy permissions intact. In parallel, classify every consented report submitted during those windows. Record whether each report concerns a wait, a proposed action, execution, instruction compliance or an explanation problem.

DXAP's activity guide distinguishes those stages. Preserve that distinction in labels so a confusing submitted-order display cannot become a forecasting-error example. Keep report origin, program version and adjudication outcome attached to the case throughout curation.

In an illustrative pre-change queue, 100 accepted reports contain 20 wait cases and 80 action cases. After the change, another 100 contain 60 wait cases and 40 action cases. The queue's wait share triples from 20% to 60%. If the independently sampled turn population remains 40% waits in both windows, the report shift reflects changing representation of cases rather than a comparable change in runtime behavior.

An unweighted training sample would teach from the new 60/40 mix. For a descriptive summary matching the 40/60 reference mix, illustrative post-change weights are 0.40/0.60, or two-thirds, for waits and 0.60/0.40, or 1.5, for actions. These weights repair that particular observed imbalance. They cannot recover absent case types or correct hidden selection within a type. Reviewers still need to inspect whether rewarded wait reports became easier, more repetitive or more detailed.

Evaluate against two distributions

Before a proposed training run, reserve an owner-disjoint, later-time holdout drawn from the independent turn sample. Keep a second diagnostic set representing the actual submitted-report distribution. The first tests transfer to ordinary eligible activity; the second tests whether the model handles what participants currently surface. Report results on both, with subtype counts and missing labels.

Our controls paper companion supplies a historical precedent for tracing failures and evaluating harness changes. A future dataset audit extends that discipline to participation selection. Compliance labels, predictive targets and execution outcomes retain separate meanings throughout training.

If the incentive change broadens coverage of a previously missing failure family, preserve that gain even when the raw average score falls because cases became harder. The useful release decision asks whether a candidate handles the reference population and the newly exposed failures. A larger feedback queue earns its value through coverage and transfer, rather than through its size alone.

Sources

Related field notes