Why early participation can support an agent platform

By DX Research Group · · Incentives and participation

A proposed participation experiment connects user incentives to usable engineering evidence.

An early agent platform needs people who will inspect what their agents actually do. Incentives can help sustain that attention long enough to expose unclear instructions, missing observations and confusing recovery paths. We would judge the program by whether that participation produces reproducible evidence and tested improvements.

DXAP has a concrete setting for this work. Its documentation describes owner strategies, model selection, schedules and trading policies, with decisions available for review. These create several places where a participant can explain a mismatch between intended behavior and observed behavior. A report about a strategy revision reaching the wrong turn has more engineering value than an undifferentiated request for better performance.

Follow one report through the loop

Consider an illustrative participant who changes an agent's maximum order size from $200 to $100. The next displayed proposal still shows $200. The participant supplies the revision time, relevant decision and expected limit, using an authorized private support channel for account details. An engineer can then determine whether the proposal began under the earlier strategy, whether the new instruction was confirmed, and whether the policy check enforced the current limit.

The proposed path is specific: participation produces a case; review establishes the failure stage; an authorized repair changes that stage; a held-out fixture checks whether the repair transfers to another timing pattern. A duplicate of the original case belongs in regression testing, while a different revision timing belongs in evaluation. That distinction keeps successful reproduction from becoming an inflated improvement claim.

Our controls paper companion describes historical harness revisions tested against replayed scenarios. It supplies a precedent for this engineering loop. It supplies no evidence that current participation updates model weights online or improves current trading returns. Training on consented, reviewed labels would be a further proposal with its own data and evaluation contract.

Measure the missing middle

A proposed four-week study could compare two invitation groups given the same product access and support availability. One receives a structured review task and an explicitly defined recognition program; the other receives the review task alone. Freeze the task instructions before assignment. Measure valid reports per invited participant, distinct reproduced failure families, reviewer time and successful repairs on held-out cases.

For illustration, ten reports that each take an hour to classify may provide less useful coverage than four reports that take ten minutes and expose four different failures. Report count and participation rate explain different parts of that result. We would also retain the participants who submitted nothing in the invitation denominator, rather than comparing only the most enthusiastic reporters.

Public Chapter 0 terms describe a dated recognition program ending October 6, 2026 at 00:00 UTC. That program is existing context; the study above is a proposed addition. A useful incentive experiment earns its place when the evidence becomes easier to review and the resulting repair survives independent cases. More activity alone leaves that question open.

Sources

Related field notes