Twenty-four prompt revisions create a selection history
By DX Research Group · · DXRG findings
Read repeated pre-launch repair as development evidence while preserving the trials behind a chosen prompt.
The controls paper records 24 pre-launch prompt revisions tested against replayed scenarios across thousands of agents. That history shows an iterative repair process. It also means the final prompt was selected after repeated feedback, so performance on the development fixtures needs a different interpretation from an untouched holdout. We should preserve both the repairs and the selection path.
Our controls paper companion describes tracing each failure to the stage that produced it and measuring fixes against the same fixtures. The published five-row table covers fee placement, rule fabrication, tokenomics framing, soft numbers becoming quotas and cadence drift. Some intervention arms lack exact published counts, which limits reconstruction of uncertainty from the companion.
Revision history and statistical tests have different counts
The independence and number of statistical tests require the trial ledger behind those twenty-four revisions. A revision may repair parsing, change several phrases or revisit a known failure. Conversely, one revision can contain several explored alternatives. Without a trial ledger, the revision number cannot support a precise multiple-testing correction or an estimate of how much apparent improvement came from selection.
An illustrative sequence has three candidates scored on the same ten failure cases. Candidate A fixes six, B fixes seven and C fixes eight. Choosing C documents improvement on those cases. If C fails on a separately held ten-case set, the original repair result remains true while transfer remains weak. This synthetic fixture separates regression-corpus fitting from generalization; the actual 24-revision outcomes remain in the historical record.
A useful selection artifact would retain the candidate version, changed text, fixture version, metric, outcome and reason for retaining or rejecting it. A protected holdout would have a release rule that prevents repeated inspection from gradually turning it into another development set. The proposed contract should also disclose any manual selection or abandoned metric after results were seen.
Preserve measured interventions at their actual scope
Moving an unchanged fee sentence raised trace citation from 3% to 74% in the reported comparison, a 71 percentage-point increase. The compound rule-fabrication intervention reduced fabricated sell-rule traces from 57% to 3%, a 54-point decrease. These are trace measures from affected pre-launch populations. Trading-return gains require a separate economic result, and missing per-arm counts prevent exact confidence calculations here.
Once live, the Terminal Pro production prompt and harness were frozen for the 21-day, twelve-token Base event on one model family. That freeze helps describe production behavior under a stable runtime. The earlier development selection remains part of the lineage, with out-of-sample status determined by each replay's data role.
The continuous record companion later emphasizes preregistration, complete sweeps and leave-window-out validation after failed claims. Those principles suggest how to strengthen a new repair study. The appropriate reader takeaway is to ask where a result was measured in the selection process, so a successful local fix and a demonstrated transferable improvement remain distinct.