Separate Replay Determinism from Model Choice Variation
By DX Research Group · · Trace evaluation
A two-stage replay fixture distinguishes runtime instability from stochastic changes in model proposals.
A replay can preserve its market inputs while the model produces different proposals. We would test deterministic runtime behavior separately from stochastic inference. Otherwise, a changed action can be blamed on the harness when the downstream code is stable, or attributed to model sampling when the runtime is changing the state.
First replay the saved proposal
Take an illustrative captured proposal to buy two units at a limit of 100. Replay that exact proposal through normalization, policy checks and simulated execution against the same saved state. The expected downstream result should be explicit. Divergent results here indicate runtime state, nondeterministic dependencies or a missing fixture input, before a new model generation is involved.
Only after this path is stable should a comparison regenerate proposals from the saved context. Preserve generation settings and any available seed or service metadata, while acknowledging that those controls may provide limited repeatability. A deterministic-looking configuration is a claim to test, rather than proof of identical outputs.
The continuous record companion distinguishes choice stability from measured decision quality in its historical replay analysis. The operating-layer controls companion provides the downstream validation and settlement framing. Our proposed fixture uses that separation without treating action identity as an economic score.
Different proposals can occupy the same economic class
For an illustrative three-output set, the model proposes buying one unit, buying two units and holding. Raw action agreement is low. Whether those differences matter depends on the mandate and the eligible opportunity. If both buy sizes satisfy the allowed range, they can share an admissibility label while retaining different exposure. Holding introduces another policy decision that needs its own justification.
Define equivalence before seeing outputs. Possible dimensions include identical tool identity, same instrument and side, size within a declared tolerance, and equal mandate compliance. Report each dimension separately. A broad equivalence class chosen after inspecting differences can conceal meaningful variation.
OpenTelemetry's tracing API supplies relationships between replay stages; our fixture needs a run identifier that keeps regenerated inference apart from replayed downstream actions. Link the two tracks through the original captured case.
We would plan a bounded variation study with a declared compute allowance and stop rule, authorized before requests. Additional generations answer a variability question and consume a different budget from deterministic regression. Keep all outputs, including failed requests, and avoid expanding the study merely because the first comparison is ambiguous.
This proposed method yields two findings: whether saved actions replay consistently through the tested runtime, and how regenerated proposals vary within the sampled cases. Economic evaluation then asks whether that variation changes executable outcomes under stated assumptions. Stable action text alone supplies little evidence of market skill; varied text can still produce equally admissible behavior. The report should leave those possibilities visible.