Error Severity and Failure Frequency Need Separate Reports
By DX Research Group · · Trace evaluation
A severity-aware trace audit avoids treating common harmless defects as equivalent to rare unauthorized actions.
Failure frequency and severity answer different engineering questions. We would report both before ranking repairs. A common recoverable formatting error may consume operator time, while a rare authorization breach can change exposure. Combining them into one success percentage hides the decision that matters.
One thousand turns, two kinds of harm
In an illustrative batch of one thousand turns, forty contain a citation formatting defect and two propose an action above the owner's maximum exposure. Formatting frequency is 4%; oversized-proposal frequency is 0.2%. If deterministic validation blocks both oversized proposals, the proposal defect and realized unauthorized exposure remain separate outcomes.
Record the attempted deviation, the control response and the economic state reached. Severity should attach to an explicit consequence or credible reachable state. A blocked proposal can reveal a serious model behavior while leaving the account protected by the harness. That distinction directs repairs toward both proposal generation and the control that prevented execution.
The operating-layer controls companion connects trace behavior to validated actions. The continuous record companion separates observed risk behavior from directional performance. Those historical records motivate a stage-specific severity ledger, while the counts here are illustrative.
Make priorities reviewable
A proposed ledger would store failure class, stage, observed consequence, recovery cost and maximum authorized state deviation. Keep severity definitions concrete: missing rationale citation, delayed review, unresolved order state or executed exposure beyond a mandate. Avoid vague labels whose meaning changes from reviewer to reviewer.
An operator may choose to estimate repair burden. For illustration, forty formatting defects requiring one minute each consume forty minutes. Two recovery investigations requiring thirty minutes each consume sixty minutes. The latter class has lower frequency and greater observed review burden. This arithmetic estimates labor under the stated assumptions; it places no price on capital loss or user trust.
OpenTelemetry's tracing specification helps preserve the stages and relationships needed for that inspection. Our severity interpretation requires economic context beyond span status. A transport error and an unauthorized fill can both produce an error marker while deserving different responses.
We would test a candidate repair against a fixture with both classes. A formatting fix that removes forty defects leaves the two oversized proposals untouched. Aggregate failures drop from forty-two to two, a 95.2% reduction, but the severe class shows zero improvement. The release note should state that directly.
Show class counts and confirmed consequences beside any aggregate. For rare severe events, uncertainty remains large; zero observed executions after a repair supplies bounded sample evidence rather than a guarantee. This proposed audit ends with named residual failures and the control responsible for limiting them. It makes prioritization an explicit consequence-based choice that readers can inspect, rather than an opaque ranking derived from frequency alone.