Stop buying more examples when the next batch stops helping

By DX Research Group · · Data and learning flywheels

A proposed marginal curation curve allocates research capacity by held-out benefit per added unit of cost.

A sustainable learning loop has to decide when another example is worth its cost. More agent use creates potential evidence, but storing, reviewing and testing that evidence consumes capacity. We propose allocating DXAP's curation budget through a marginal-benefit curve: measure what the next accepted batch adds to a frozen task, rather than assuming that dataset size is the objective.

Our continuous-record companion reports a historical replay comparison where models had overlapping decision-quality intervals and substantially different inference costs. Its practical lesson is that economics can distinguish choices even when assessed quality stays close. For a future data pipeline, collection and review costs deserve the same treatment as model inference costs.

Build a cost curve around one question

Choose a specific repair task, such as identifying whether a missing outcome requires execution recovery or merely a clearer explanation. Begin with a fixed development set, then add prespecified batches of independently reviewed failure families. At each step, evaluate a frozen candidate on the same sealed benchmark through a restricted evaluator. Release only the measures the development protocol permits. Once developers inspect individual benchmark answers, retire those cases from fresh assessment. Pre-register the batch order and decision rule, and reserve a separate untouched final set for the selected candidate. Repeated aggregate feedback can influence development even when individual answers remain hidden.

An illustrative cost ledger starts with 40 correct diagnoses out of 60. A first curated batch costs $120 and raises correctness to 46. The next batch costs $100 and raises it to 48. The third costs $100 and leaves it at 48. The observed marginal costs are $20 per additional correct diagnosis for the first batch and $50 for the second. The third has zero observed gain, so a cost-per-gain ratio is undefined. These fixture calculations express allocation choices; uncertainty may still allow a small benefit or harm.

The dollar ledger includes the actual development resources being compared. We would also report severe-error counts independently. A batch that leaves average correctness unchanged while resolving one serious recovery error could deserve further study under a stated severity objective. Choosing that objective before observing the curve prevents retrospective claims that every batch was valuable in some unspecified way.

Move capacity to the missing mechanism

When a curve flattens, reviewers can inspect coverage and disagreement on development cases. A duplicate-rich intake might benefit from sampling fewer copies. An unresolved label disagreement might need a domain review rather than another model call. A poorly represented state might justify a focused request for consenting examples. Each choice changes where the next unit of capacity goes.

The DXAP execution documentation distinguishes authorization checks from submitted orders and execution outcomes. That distinction gives curation a natural unit: a failure mechanism with enough evidence to locate the stage. Ten extra explanations of an already-resolved issue may cost more than one unfamiliar recovery case with complete lineage.

We would keep the capacity decision reversible. A paused collection stream can reopen for a newly observed mechanism, while a broad stream can narrow once common cases are covered. Sustainable use means that each further batch has a defined question and a plausible route to resolving it. The marginal curve supplies a receipt for continuing, redirecting or stopping that particular collection effort.

Sources

Related field notes