Weight Agent Errors by Operational Severity Carefully
By DX Research Group · · Learning theories
Severity weights encode an objective and should remain visible beside ordinary error counts.
A rare unauthorized action can deserve more learning emphasis than a frequent cosmetic explanation error. Severity weighting makes that priority explicit, but it also changes the objective and can let noisy high-weight labels dominate. We would publish the weight rationale and evaluate each failure class separately.
Concrete Problems in AI Safety identifies practical problems including reward specification. Our proposal concerns a narrower dataset decision: weights for operational failure classes. The operating-layer paper motivates stage-specific diagnosis, while trace feedback retains the evidence needed to verify labels.
A weighted score can reverse the ranking
Consider an illustrative set with ninety explanation cases and ten unauthorized-exposure cases. Candidate A makes nine explanation errors and one exposure error. Candidate B makes three explanation errors and two exposure errors. Ordinary error rates are 10% for A and 5% for B.
Assign weight 1 to an explanation case and weight 20 to an exposure case. Total weight is 90 plus 200, or 290. A's weighted error is (9 plus 20)/290, or 10%. B's is (3 plus 40)/290, approximately 14.8%. The ranking reverses because the stated objective values exposure failures much more heavily.
That reversal can be reasonable. It should be traceable to the operational consequence, rather than chosen after observing which candidate wins. A severity weight also differs from an estimate of expected monetary loss. A forbidden order that the deterministic policy blocks has different realized consequences from one that reaches the venue, even if both reveal the same model compliance failure.
Check the rare labels first
We would audit every high-weight example before relying on its gradient contribution. Confirm the mandate version, units, and exact action boundary. A mislabeled example with weight 20 can offset many correctly labeled ordinary cases. Where severity is uncertain, retain a range or an explicit unresolved status instead of false precision.
The proposed report includes unweighted counts, class-conditional rates, and the registered weighted score. It also includes effective sample size of the weights: (sum of weights) squared divided by sum of squared weights. For this fixture that is 290 squared divided by (90 plus 10 times 400), approximately 20.6. A hundred rows therefore provide much less balanced weighted evidence than the raw count suggests.
We would inspect whether the proposed emphasis encourages universal inactivity. Boundary cases with permitted exposure distinguish compliance learning from blanket abstention. Severity weights can guide prioritization, while deterministic controls remain a separate enforcement mechanism. The concrete decision is whether the learning objective reflects the operational priority without concealing label fragility or sacrificing valid actions.