Ten support tickets may describe one learning example

By DX Research Group · · Data and learning flywheels

A proposed grouping rule preserves support demand without inflating independent evidence.

Support volume and dataset diversity answer different questions. Ten reports of the same failed turn may justify urgent support work while supplying a single underlying evaluation case. We propose retaining every report for service accounting, then grouping contributions by the event and failure mechanism before constructing training or evaluation splits. The grouping would be a research design, rather than a description of an established DXAP training pipeline.

The controls paper companion describes revisions tested against replayed scenarios. Its trace structure gives a natural anchor for grouping: the mandate, state and action that produced a failure. Counting messages alone discards that anchor and lets one widely shared incident dominate a dataset.

An illustrative twelve-report queue

Suppose twelve authorized reports arrive. Seven refer to the same cooldown failure, with screenshots copied from one owner. Three describe separate occurrences under the same release and instruction pattern. Two describe distinct failures in position-cap interpretation. The service queue still contains twelve reports. The underlying event inventory contains six events: one shared cooldown event, three other cooldown events and two position-cap events.

Even six is an incomplete independence claim. Four cooldown events may share one faulty parser branch, one prompt template and one market regime. For learning, we would record both event groups and mechanism families. The seven repeated reports can add useful perspectives or missing artifacts, but they should strengthen the documentation of one event rather than create seven copies of its target answer.

A curator can attach new evidence to the event record while retaining separate reporter attribution and permission receipts. Combining reports must respect the narrowest applicable permission for each field. One owner's public-use approval cannot authorize another owner's private screenshot. A dataset export selects permitted artifacts individually, then resolves their group membership.

Keep related cases on the same side of a split

We would assign all versions of a shared event to one dataset partition. A stricter generalization test would hold out an entire mechanism family or later release period. Otherwise a model can encounter a paraphrase of the test case during training and appear to generalize. Text similarity helps find candidates for grouping; event identity and curator review determine the final relationship.

Report the number of tickets, unique events and mechanism families alongside evaluation results. If twenty paraphrases of one bug all pass after a revision, readers should see one repaired event with twenty representation checks. The paraphrases still test resilience to wording, but their correlated answers carry less evidence about new failures.

For a proposed learning loop, this distinction changes incentives as well as measurement. Rewarding accepted messages by count encourages resubmission. Crediting a novel reproducible event or a materially useful addition gives contributors a reason to improve the evidence. Neither credit rule establishes a reward program or a return benefit. The first measurable advance would be a dataset whose apparent size can be reconciled to its independent failure coverage.

Sources

Related field notes