Test Representation While Holding Information Fixed
By DX Research Group · · Learning theories
A representation test needs a fact-level equivalence check before model scores are compared.
A representation ablation should change how facts are expressed while preserving which facts the agent receives. Otherwise a better result may come from additional information, corrected data, or a more helpful instruction. We propose a fact-level equivalence audit followed by a paired decision test.
Our harness-transfer card separates changed components from frozen components. This note narrows that rule to an often-missed problem: two prompts can have similar length while supplying different decision facts. The historical operating-layer paper motivates checking rendered context, but it establishes no result for the proposed representation experiment.
Construct a reversible pair
An illustrative portfolio contains 200 units of cash, a position of 3 contracts, and a mandate allowing at most 5 contracts. Version A presents those facts in prose. Version B uses a typed table with fields for cash, current quantity, and maximum quantity. Both must imply residual capacity of 2 contracts.
Suppose B also supplies a precomputed field saying residual capacity equals 2. That addition makes the arithmetic more accessible, but it introduces an explicit derived fact. To test layout alone, either supply the derived value in both representations or omit it from both. A comparison between raw facts and a calculation aid answers a different question and can be valuable under its own label.
We would define a canonical fact map containing values, units, timestamps, source roles, and authority. Each renderer produces its text from that map. A reverse parser or independent audit reconstructs the map from each output. Missing fields and ambiguous units fail the equivalence check before any agent response is graded.
The map also includes negative space. If one representation says the balance is reconciled and the other leaves reconciliation status unspecified, the second conveys less operational information. A convenient summary can quietly remove this distinction.
Use errors that discriminate between formats
The proposed held-out fixture varies cash, existing quantity, and mandate limits around exact boundaries. Cases should require different actions rather than all sharing the same answer. A quantity of 4 under the same limit leaves capacity 1; a quantity of 5 leaves capacity zero. Record arithmetic errors separately from unauthorized proposals and valid abstentions.
The implementation study in reinforcement learning illustrates how design choices can affect observed performance. Our adaptation is to treat rendering as an experimentally declared component. We would freeze model, tool availability, and decoding configuration, retain both rendered inputs, and score paired cases with the same adjudication rules.
A representation can also improve auditability without changing decisions. Report reconstruction accuracy and review effort as separate outcomes. The final claim should name the fact map and decision population: for example, a table reduced quantity-arithmetic errors on these boundary fixtures. It should leave broader market skill to a separate measurement.