When more users add new evidence
By DX Research Group · · Data and learning flywheels
A proposed coverage audit distinguishes new failure contexts from repeated observations of one event.
More users can expose an agent to new mandates, instruments and operating conditions. They can also produce thousands of near-identical observations of one shared market event. We would evaluate a participation-driven data loop by the coverage it adds and the failures it reveals, alongside its raw volume.
The controls paper companion reports historical agents acting differently under a shared model and shared market state, with differences associated with their harness context. It also records concentrated activity. That combination explains why user count alone is a weak measure of independent evidence: owner variation can matter while market dependence remains substantial.
Consider a synthetic collection of 100 reports. Eighty concern one symbol during the same short volatility event under one harness version. Twenty concern distinct situations, including partial closes, interrupted approvals and several instruction styles. All 100 may be valuable to support. For evaluation, the first eighty should carry an event-group identifier so their shared conditions remain visible.
Define diversity before counting it
We propose a coverage table built around mechanisms a candidate change could affect. For an instruction-handling change, relevant dimensions might include explicit numerical conditions, temporary instructions and corrections to prior direction. For an execution change, order state and receipt timing may matter more. The table should follow the engineering question rather than collect demographic attributes for their own sake.
Each cell would show the number of distinct owners, decision opportunities and underlying event groups. A cell with ten owners observing one event differs from ten owners observing ten separate events. Empty cells identify missing evidence; they do not prove missing product capability.
An illustrative next intake adds 50 reports. If 45 repeat the same event and five fill previously empty instruction contexts, report volume grows by half while relevant coverage grows in only five places. That can still be useful. The repeated reports might establish how widespread an interface misunderstanding is, while the five new contexts are more informative for testing generalization.
DXAP's activity guide distinguishes completed turns, proposed actions and fills. We would carry that stage into the coverage table. Otherwise a large collection of completed turns could mask sparse evidence about actual execution outcomes.
Choose the next observation deliberately
A proposed study would compare two equal annotation budgets. One arm samples incoming reports uniformly. The other selects underrepresented mechanism cells using a rule fixed before reviewing outcomes. Both arms retain a common, untouched evaluation set and report the number of newly discovered reproducible failures per annotation hour.
The coverage-directed arm could discover more failure types yet yield weaker estimates of their population frequency. Uniform sampling could estimate common failures more cleanly while missing rare interactions. Report those benefits separately. Neither sampling policy should recruit owners to trade merely to generate difficult cases; synthetic fixtures and historical reconstruction can cover many gaps.
Our aim is a data loop where participation makes the system easier to inspect and test. A release would earn a broader claim only after it succeeds across the additional contexts. Growth becomes a research advantage when it supplies variation that changes the experiment, rather than a larger denominator for the same unresolved question.