Stratify Manual Audits Around the Failures You Need to Understand

By DX Research Group · · Trace evaluation

A weighted review design finds rare severe traces while preserving a population estimate.

A manual trace audit should oversample cases that can change the engineering decision, then account for that sampling when estimating prevalence. We would preserve both the discovery sample and a population-weighted estimate. Reviewing only suspicious traces finds useful defects but says little about how common they are.

Spend reviewer time where ambiguity lives

Consider an illustrative population of one thousand turns: nine hundred ordinary completions and one hundred traces carrying a recovery flag. We sample fifty from each stratum. Reviewers find one defect among ordinary traces and ten among flagged traces. The unweighted sample rate is 11/100, or 11%.

Weighting back to the population gives 0.9 × (1/50) + 0.1 × (10/50) = 3.8%. That estimate relies on random selection within each stratum and correct stratum counts. The audit intentionally enriched the sample for recovery cases, so the raw 11% describes the reviewed set rather than the population.

The tracing API specification supplies concepts for locating related events. Our proposed sampling ledger uses those records to define strata before reading outcomes. Sampling should operate on the intended unit, such as completed decision intents, instead of whichever span is easiest to query.

Keep a discovery lane beside estimation

A known severe trace deserves direct inspection even when it was selected outside the random sample. Label such cases as purposive discoveries. Their findings can populate regression fixtures, while the prevalence estimate stays attached to the probability sample. This gives engineers useful counterexamples without pretending every reviewed trace had a known inclusion probability.

The operating-layer controls companion shows the value of trace-level failure diagnosis. The continuous record companion demonstrates why historical populations and inference units need explicit definitions. This note proposes a new audit design and supplies illustrative arithmetic, rather than a new fleet defect estimate.

We would save population extraction time, stratum rule, counts, random selection seed and inclusion probability for each sampled intent. Audit instructions should name which evidence the reviewer can inspect and how missing records are labeled. If recovery flags themselves are incomplete, the ordinary stratum may contain hidden recoveries; reviewers should record them rather than silently reassign the sampling probabilities.

Estimate uncertainty with a method appropriate to the stratified design and any clustering by agent or day. The weighted 3.8% is a point estimate from a small example, with substantial uncertainty in the one-defect ordinary slice. A simple unweighted interval on one hundred reviewed cases answers the wrong sampling question.

The release decision can then use two complementary outputs: a weighted defect estimate for the declared population and a library of verified severe cases. Each tells the reader why it exists. The first quantifies exposure to the measured failure; the second identifies a concrete repair that can be tested.

Sources

Related field notes