Model Alias Drift Can Break Replay Reproducibility

By DX Research Group · · Trace evaluation

Resolve model identity at request time and quarantine ambiguous alias changes before comparing agent behavior.

A saved model alias identifies a request label, while reproducibility needs evidence about the implementation that served the request. We would attach a routing receipt to each replay and treat an unexplained identity change as a comparison boundary. The same alias string can conceal a changed serving path.

The run label survives while the route changes

Imagine an illustrative two-day evaluation. Monday's run requests research-model-latest through provider A. Tuesday requests the same label through provider B after a fallback. Tuesday produces more valid tool calls. The saved label supports continuity of requested configuration; it leaves the serving identity unresolved. A harness fix, model update and provider adaptation remain competing explanations.

Our receipt would retain requested identifier, returned identifier, provider endpoint, request time and any documented snapshot identifier. It would also preserve observable routing decisions and response metadata. A field that the provider omits should remain unknown. Substituting the alias into the snapshot field creates false precision.

The official model catalog is a primary place to inspect available identifiers. A catalog lookup today establishes today's documentation, whereas a dated response receipt supports a historical request. We would archive the relevant documentation reference with the evaluation rather than infer past service behavior from a current page.

Use a small boundary fixture

Before merging results across the routing boundary, pass a deliberately diagnostic set through each route. Include a nested tool argument, a long mandate with an exception, and a response requiring a supported structured-output feature. These cases probe the interface most likely to alter the trading trace. Keep their expected interface outcomes explicit, and describe any semantic behavior as observed variation.

The operating-layer controls paper companion shows why changes around the model deserve attribution. The continuous record paper companion separates historical fleet evidence from replay comparisons. A route change can affect both kinds of records, but the repair is different: stratify historical traces by known route, and use paired interface fixtures for a prospective release.

If identity remains opaque, report the comparison as alias-and-provider behavior during a dated window. A persistent difference in fixture outcomes establishes an observable boundary, even when its internal cause is unavailable. Conversely, matching fixtures provide limited regression evidence; they cannot prove identical weights or serving internals.

For an illustrative ledger of twenty cases split ten before and ten after fallback, we would retain two ten-case slices until the boundary is explained. A pooled fifteen successes out of twenty gives 75%, but hides whether the slices are ten out of ten and five out of ten. This proposed protocol yields a reviewable uncertainty record. Its value is preserving what the run can actually identify when an apparently stable name changes underneath it.

Sources

Related field notes