Checkpoint Selection Can Leak the Evaluation Set
By DX Research Group · · Learning theories
Repeatedly choosing checkpoints on a holdout turns that holdout into a development signal.
A holdout loses its final-evaluation role when its scores repeatedly select checkpoints or training changes. The leakage can occur through human selection even when no holdout tokens enter the gradient. We would separate checkpoint selection from the final audit and preserve the complete selection history.
Our harness-transfer framework requires frozen comparison boundaries. Implementation research highlights the influence of design choices; the proposed protocol here treats checkpoint selection as one such choice. Our continuous record remains a historical record rather than a checkpoint-selection benchmark.
The best observed score contains selection noise
Consider five illustrative checkpoints with validation accuracies of 71%, 73%, 72%, 75%, and 74%. Choosing the fourth checkpoint is a legitimate validation decision. Reporting 75% as its final untouched-test accuracy would mislabel the same data's role.
A simple probability illustration shows why repeated looks matter. If five independent null checks each have a 5% false-positive rate, the chance of at least one false positive is 1 minus 0.95 to the fifth power, approximately 22.6%. Real checkpoints are correlated, so this arithmetic does not estimate the actual false-positive rate. It demonstrates the effect under a stated toy assumption.
The selection procedure may also react to compliance failures. If an engineer inspects final-audit cases and modifies the training data to fix them, those cases now provide development feedback. Their status changes even if the modification is entirely sensible.
Register which data makes which decision
We would designate training data for parameter updates, validation data for checkpoint choice, and an untouched audit set for the final registered candidate. The split should respect agent, mandate, and time dependencies so near-duplicate turns cannot cross boundaries unnoticed. Record the criterion used to choose a checkpoint, including how ties and tradeoffs between compliance and forecast quality are handled.
The ledger includes every checkpoint evaluated, all visible scores, and all later changes influenced by those scores. A final candidate should be identified before the audit set is opened. If the audit reveals a failure that triggers repair, the repaired version needs fresh final evidence; the old audit remains useful as a regression fixture under its new role.
Where data are scarce, nested temporal evaluation can separate selection and reporting, but its estimates inherit their window definitions and deployment assumptions. It should be planned for a specific candidate comparison rather than used as permission for uncontrolled repeated searches.
The deliverable is a data-role ledger and one clearly labeled final result. This preserves useful feedback while preventing a selected maximum from posing as independent evidence. We propose the protocol here; no checkpoints or training runs were produced for this note.