Register an agent-release evaluation before its outcomes arrive

By DX Research Group · · Frontier research

A proposed release protocol freezes questions, comparisons and decision rules prospectively.

A prospective evaluation should freeze its primary question before release outcomes arrive. We propose a compact preregistration record for agent releases that preserves the candidate version, comparison and decision rule. This protocol has yet to be exercised here. Its purpose is to make interpretation possible after results become tempting to reframe.

The continuous record companion preserves a historical cutoff and a directional-edge null across studied groups. That bounded record shows why release evaluation needs a defined window and complete trial history. The controls paper companion shows that a compound harness intervention can change a trace measure without establishing return gains. A new release should declare which claim it actually intends to test.

A receipt small enough to use

An illustrative registration states: compare candidate C with baseline B on 200 saved episodes, use the same model and action policy, and make policy-valid episode completion the primary outcome. The candidate changes one state-reconciliation component. Forecast scores and simulated economics remain secondary diagnostics. The manifest hashes the episode inputs, expected-state rules and code revisions before scoring begins.

Suppose the observed completion counts are 184 of 200 for C and 180 of 200 for B. The rates are 92% and 90%, a two-percentage-point difference. A claim of improvement still requires the paired episode table and uncertainty, because four additional successes could arise in several different overlap patterns. The registration should specify the comparison method and what counts as a missing episode. It should also preserve every candidate tried before C was selected.

For a prospective shadow window, register the start condition and end condition rather than choosing a flattering date after inspecting the curve. Describe the evidence class precisely: intended actions without execution support a different question from real order receipts. Exposure, funding and other costs belong to the selected question only where the environment measures them.

Changes deserve their own timestamps

A preregistration is useful when deviations remain visible. If a feed outage removes thirty cases, retain their original assignment and record the exclusion rule, decision time and reason. If the primary metric changes, label the new analysis exploratory and retain the original one. Emergency repairs can proceed while their effect on interpretation is documented.

We would publish a short result table with planned outcomes, observed outcomes and deviations linked to the manifest. This keeps the release decision readable without burying the evidence in a long methods document. A failed primary result should remain attached even when a secondary diagnostic suggests a worthwhile next repair.

The contribution is a temporal boundary on researcher discretion. Registration does not make a grader valid or a dataset representative; those require their own checks. It does make the intended comparison inspectable and separates a prospective release claim from an explanation discovered after the fact. That distinction allows useful exploration while preserving the meaning of the original test.

Sources

Related field notes