How participation could become a tested agent release
By DX Research Group · · Data and learning flywheels
A proposed feedback loop connects user participation to reproducible fixes and evaluated releases.
Participation becomes useful research when an experience can be reconstructed, classified and tested against a proposed change. We see a concrete opportunity for DXAP: owners can inspect persistent agents, describe what they intended and help identify where the harness departed from that intention. The valuable output is a release supported by comparative evidence.
The public activity guide describes reviewing a turn, its action details and the account's positions and trades. Our controls paper companion records a historical method of diagnosing traces and testing harness revisions against replayed scenarios. Together they motivate the proposed loop below. They establish recording and bounded intervention work; this proposal extends that work into a systematic participation process.
Follow one report through the loop
Consider an illustrative owner report: “I asked the agent to wait after a full close, but its explanation says it can immediately reopen.” The first step is to recover the relevant instruction, configured cooldown and decision time. A disagreement in the explanation differs from a forbidden order that passed execution. Both deserve investigation, with different expected fixes.
A researcher classifies the case as explanation inconsistency if the policy correctly rejected reopening. The researcher creates a replay containing the saved account state and applicable configuration, removes private identifiers, and writes an expected result: the explanation should identify the active cooldown while the enforcement result remains unchanged. A candidate change might alter how the current cooldown state is represented to the model.
We would then compare the existing and candidate harness on that fixture and on a separate collection of close-and-reopen cases. The development example verifies reproduction. The held-out collection tests whether the change generalizes to partial closes, expired cooldowns and delayed fill receipts. A fix that makes every reopening sound forbidden would fail the expired-cooldown cases.
The release decision belongs to a reviewer who can inspect the difference, affected population and remaining failures. If accepted, the release record should identify the changed component and tested behavior. The reporting owner can receive a precise resolution: the explanation changed, enforcement was already correct, and specified cases were checked. That closes the loop without presenting a wording improvement as a return improvement.
Measure each conversion
An illustrative intake of 40 reports might yield 24 reconstructable cases, 10 distinct failure mechanisms and three candidate changes. Those are separate denominators. Counting 40 messages as 40 independent learning examples would conceal duplicate reports and missing context. A useful dashboard would track reproduction coverage, new-case diversity, held-out improvement and regressions rather than treating participation alone as success.
Owners also need a clear distinction between submitting feedback and authorizing a change to their agent. The chat guide preserves review and approval for settings proposals and persistent instructions. A research report should retain that boundary.
The next measurable advance would be one complete, documented passage from report to tested release, with the test population and unresolved failures visible. Model training could later become one candidate intervention among several. Better state representation, tool behavior or explanation may solve the reported problem more directly. The mechanism earns its value when participation produces a verifiable improvement in the component that actually failed.