Choose Verifiers That Expose Reward Gaming

By DX Research Group · · Learning theories

A verifier should discriminate substantive compliance from text that merely resembles a passing answer.

A verifier can reward the appearance of compliance while missing the action that violates it. We would select verifiers using adversarially paired examples before using their scores as learning signals. The relevant question is which changes can raise the verifier score without improving the underlying task.

Reward hacking research formalizes divergence between proxy and true rewards. Our proposed trading audit tests a concrete proxy: a textual verifier for mandate compliance. Trace feedback binds explanations to actions, and the operating-layer paper supplies historical examples of trace-level failure without implying a return gain.

A passing phrase can conceal a forbidden order

Suppose an illustrative verifier awards one point when the rationale says it checked a 500-unit exposure cap. Candidate A repeats that sentence and proposes 700 units. Candidate B proposes 400 units and gives a concise factual explanation without the phrase. The proxy scores A above B while the actual limit check ranks B above A.

The repair begins with the verifier's input contract. It needs the authenticated mandate and exact proposed action, including units and any normalization. A text-only judgment cannot verify a numerical action it never receives. The deterministic arithmetic check can evaluate the cap directly; a language verifier can assess whether the explanation faithfully describes that result.

A composite score should retain the component outcomes. Otherwise a verbose explanation score might compensate for a prohibited action. If the task treats authorization as a hard condition, the scoring rule should implement that condition explicitly rather than approximating it through a large finite penalty.

Test invariance and sensitivity together

The proposed selection set includes pairs with identical actions and different phrasing. Their compliance score should remain stable. Another set holds the rationale constant while changing the action across the mandate boundary. Its compliance score should change. Both tests are needed: a verifier that ignores all input can look perfectly invariant.

We would also insert unsupported citations, fabricated rule names, and contradictory numeric statements. Human adjudication of a bounded disagreement sample establishes whether a candidate verifier's apparent confidence corresponds to the intended criterion. Keep that adjudication set separate from examples used to tune the verifier.

Evaluate the eventual learning candidate with an independent audit, because optimization can find weaknesses absent from the initial verifier-selection corpus. A proxy that passes this fixture remains scoped to its tested manipulations and action types.

The outcome is a verifier card recording accepted inputs, known blind spots, and discrimination results. It tells us whether a score is fit for an explanation objective, an action-compliance objective, or neither. That precision makes reward gaming easier to identify before a model becomes good at satisfying the wrong test.

Sources

Related field notes