Cost-Sensitive False Positives in Agent Decision Screening
By DX Research Group · · Forecast evaluation
Evaluate an internal review screen using declared asymmetric error costs.
A decision screen should be evaluated against the cost of its mistakes, separately from the quality of the underlying probability forecast. False positives can consume review capacity; false negatives can leave a defective trace unreviewed. Our research question here concerns internal screening workload, with illustrative costs that carry no investment recommendation.
Compare two illustrative review screens
Screen A flags 30 traces: 10 need review and 20 are false alarms. It misses two defective traces. Screen B flags 16: eight need review and eight are false alarms. It misses four defective traces. Assign an illustrative cost of one unit per false alarm and five units per missed defect. A costs 20 + 2 × 5 = 30 units. B costs 8 + 4 × 5 = 28 units. If missed defects cost ten units, A costs 40 and B costs 48. The preferred screen changes with the specified consequence.
Bind the cost to a real workflow
A false alarm cost might include reviewer minutes, while a missed-defect cost could reflect later remediation. Use measured workflow quantities when available and retain a sensitivity range where costs are uncertain. The two screens must operate on the same traces with the same adjudicated defect definition. A predicted market event and a malformed action are different targets requiring different labels.
For calibrated probabilities and constant binary mistake costs, flagging is favored when p exceeds C_FP/(C_FP + C_FN), under the simplified assumptions. Costs one and five give a threshold of 1/6. This mathematical illustration assumes correct cost estimates and calibrated defect probabilities. A capacity constraint can require a different selection rule. Evaluate workload, defect capture, and uncertainty separately before changing any runtime behavior.
The saved record that would make this reviewable
The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.
Continue with grader responsibilities for the related question of screening and reviewer workload. Our benchmark card keeps this forecast-level comparison separate from action and execution results.
The screen’s final report should name the adjudication policy and the workload units. Give a cost sensitivity table instead of one unexplained preferred threshold. If the chosen operating point is tested prospectively, retain actual review time and defects found. Those receipts can replace the illustrative costs with measured workflow evidence without making any assertion about investment performance.