Missing Action Propensities Break an Off-Policy Ratio

By DX Research Group · · Learning theories

A saved action and token log probabilities may still leave the behavior-policy denominator unknown.

Importance weighting requires the probability that the behavior policy assigned to the observed action. A trace containing the chosen action alone leaves that denominator unidentified. We would classify such data by what can actually be estimated before using it to judge a proposed policy.

Doubly robust off-policy evaluation studies estimating a new policy's value from another policy's data. Its assumptions matter at the logging boundary. Our historical record provides observations under recorded systems; the existence of that record does not establish every probability needed by a new estimator. This is a proposed identifiability audit.

The denominator belongs to the action mechanism

In an illustrative one-step problem, the target policy chooses buy with probability 0.4. The behavior policy chose buy with probability 0.2. A logged buy receiving reward 3 contributes an importance-weighted value of (0.4/0.2) times 3, or 6, before averaging with other draws.

If the behavior probability is missing, plausible denominators of 0.1 and 0.5 produce contributions of 12 and 2.4. The observed action and reward are identical. Those observations cannot select the denominator. Substituting the target probability gives a weight of one and changes the estimand into an uncorrected logged-policy average.

An LLM action also has several layers. Multiple output strings can map to the same typed action. A deterministic validator can reject some strings or round quantities. The probability of a particular token sequence therefore differs from the probability of the final executed action. Retries and fallback behavior can change that mapping further.

We would identify the evaluation unit first: generated sequence, parsed proposal, accepted action, or execution. The behavior propensity must correspond to that unit and its conditioning information. A quantity bucket requires the probability mass of the bucket, rather than the density or probability of one observed spelling.

Three useful dispositions

Cases with known behavior probabilities and supported target actions can enter the registered estimator. Cases with probabilities reconstructed from an exact saved policy can enter a separately labeled reconstruction track, with reconstruction error checked on logs where probabilities were retained. Cases with unknown denominators remain observational or support sensitivity calculations with explicit assumptions.

Our harness-transfer tests help freeze the parser and policy boundary during reconstruction. For a sequential estimator, the audit also follows probabilities through time; multiplying ratios can magnify both missing information and variance. Weight clipping changes the estimator and should be reported with its bias tradeoff.

The actionable outcome is a coverage report: how many decisions have an identified denominator at the selected action boundary, which actions lack support, and how much the estimate depends on assumed probabilities. A precise label for missing information is more useful than a numerical policy ranking built from an invented denominator.

Sources

Related field notes