AI failure evaluation
Evaluate AI failure modes before use
The evaluated population
The cohort contains 3,100 evaluation records, of which 372 have the defined synthetic outcome: Critical assistance error. The outcome is known by construction here. In production, label uncertainty and selection must be recorded separately.
Figure data and text version
| Outcome | Count |
|---|---|
| Critical assistance error | 372 |
| Other labeled outcomes | 2,728 |
Average fluency does not measure unsupported claims, instruction misuse, missing evidence, or sensitive-data exposure. Evaluation needs a labeled failure taxonomy tied to the use case.
Test critical omissions and unsupported claims across representative and adversarial records.
All amounts, rates, capacity limits, and outcomes in this case are synthetic. The three conditions are separate assumptions for comparison. A better result in the response condition is not measured proof that the proposed control causes that improvement. The figures expose the calculation and its limits; a real deployment needs its own evidence.
Read the result
The rule flags 353 of 3,100 evaluation records. Of those flags, 320 meet the synthetic target, giving 90.65% precision. It misses 52 target events. Under the stated cost assumptions, residual loss and operating friction total $8,077. The important result is the connection between the population, action, capacity, and outcome—not one isolated score.
Model inputs and calculated values
Inputs below are the case-specific values. Each figure states the condition-specific assumptions and units used in its calculation. Calculated values are rounded for display.
| Input | Value |
|---|---|
| population | 3,100 |
| prevalence | 0.12 |
| severity | 125 |
| Calculated value | Result |
|---|---|
| population | 3,100 |
| positive | 372 |
| negative | 2,728 |
| tp | 320 |
| fp | 33 |
| fn | 52 |
| tn | 2,695 |
| loss | 6,500 |
| severity | 125 |
| precision | 90.65 |
| recall | 86.02 |
Four outcomes of the rule
The rule flags 320 synthetic positives and 33 negatives. It misses 52 positives. A flagged item is a decision to intervene; it is not proof of fraud, a legal prohibition, or any other real-world conclusion.
Figure data and text version
| Known outcome | Flagged | Not flagged |
|---|---|---|
| Critical assistance error | 320 | 52 |
| Other outcome | 33 | 2,695 |
Three rates with different denominators
Precision is 90.65%, recall is 86.02%, and the false-positive rate is 1.21%. Changing the denominator changes the meaning. This record keeps each numerator attached to the population from which it came.
Figure data and text version
| Metric | Numerator | Denominator | Result |
|---|---|---|---|
| Precision | 320 | 353 | 90.65% |
| Recall | 320 | 372 | 86.02% |
| False-positive rate | 33 | 2,728 | 1.21% |
Precision changes with prevalence
This sensitivity plot holds recall at 86% and false-positive rate at 1.2%, then changes prevalence. It is an algebraic comparison, not a forecast. Even unchanged detection quality can produce a very different review queue when the base rate changes. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.
Figure data and text version
| Assumed prevalence | Precision % |
|---|---|
| 0.1% | 6.69 |
| 0.5% | 26.48 |
| 1% | 41.99 |
| 2% | 59.39 |
| 5% | 79.04 |
| 10% | 88.84 |
The threshold trade-off
Six illustrative score bands use a stated pair of detection rates. Lower sensitivity can reduce false alarms but miss more target events. These points do not come from a trained model and do not establish the best operating threshold. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.
Figure data and text version
| Score band | True positives | False positives |
|---|---|---|
| Band 1 | 365 | 409 |
| Band 2 | 350 | 218 |
| Band 3 | 320 | 95 |
| Band 4 | 268 | 33 |
| Band 5 | 186 | 11 |
| Band 6 | 93 | 3 |
A transparent loss-and-friction calculation
At $125 severity per missed synthetic positive, residual loss is $6,500. Review costs $1,412; lost contribution on false alarms is $165. The calculation assumes intervention prevents all flagged-positive loss and each false alarm loses the stated contribution. Relax those assumptions before applying it to a real policy.
Figure data and text version
| Cost component | USD |
|---|---|
| Missed-positive loss | 6,500 |
| Review cost | 1,412 |
| False-alarm contribution | 165 |
Observed outcomes mature over time
The final synthetic positive count is 372. Earlier observations reveal only a stated fraction. Comparing a day-1 cohort with a day-30 cohort would confuse label age with control quality. This curve models observation delay only; it does not change the final outcome. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.
Figure data and text version
| Days after event | Observed positives |
|---|---|
| 1 | 67 |
| 3 | 130 |
| 7 | 223 |
| 14 | 305 |
| 30 | 372 |
Review demand and available capacity
The flag count is 353. The comparison capacity is an illustrative 124 reviews per cohort window. A mathematical rule can be coherent while its resulting workload exceeds the operating team’s capacity. Capacity is not permission to ignore an applicable mandatory control.
Figure data and text version
| Queue measure | Items |
|---|---|
| Flagged for review | 353 |
| Available capacity | 124 |
| Excess demand | 229 |
A feature is an observation with provenance
This evidence contract supports evaluate ai failure modes before use. A value needs its event time, arrival time, scope, and source. Keeping unavailable evidence distinct from a measured zero prevents an outage from becoming a falsely reassuring feature.
Figure data and text version
| Field | Example | Meaning |
|---|---|---|
| entity_ref | AI failure evaluation | Subject of this case |
| event_time | 2026-09-18T09:00:00Z | When the event occurred |
| received_time | 2026-09-18T09:00:02Z | When the system learned it |
| signal_status | repaired | Evidence quality, not an outcome |
| label_definition | Critical assistance error | The target used in these calculations |
Missing evidence changes the observed population
The cells show an explicitly constructed completeness profile for three signal groups. The stress condition removes more history and device evidence. Missingness does not prove the target outcome; it changes what the decision process knows.
Figure data and text version
| Signal group | Available | Missing |
|---|---|---|
| Identity evidence | 3,084 | 16 |
| Activity history | 3,054 | 46 |
| Context signal | 3,007 | 93 |
Evidence, score, and action remain separate
The policy can use evaluate ai failure modes before use only within its approved scope. The action record must retain which evidence was available, which model or rule ran, and which action was actually applied. The final action can differ from the score recommendation when a separate constraint applies.
Figure data and text version
| Stage | Record |
|---|---|
| Observe | AI failure evaluation: evidence as of the decision time |
| Evaluate | Rule flags 353 of 3,100 evaluation records |
| Apply | Record action, reason, owner, and expiry |
| Reconcile | Join the action to later outcomes without overwriting history |
What the result cannot establish
Observed classifications do not reveal every counterfactual. The synthetic labels make arithmetic possible, but production decline data is selected by prior policy. Keep measured outcomes, assumed prevention, and unknown alternatives separate when reporting impact.
Figure data and text version
| Claim | Evidence in this case | Limit |
|---|---|---|
| Detected target | 320 known synthetic positives flagged | Production labels may be delayed or wrong |
| Prevented loss | Assumed 40,000 USD | Requires an intervention-effect assumption |
| Customer impact | 33 synthetic negatives flagged | Not every flag causes abandonment |
| Unobserved alternative | Outcome without the action | Needs a valid evaluation design |
Connect the result to the system
Test critical omissions and unsupported claims across representative and adversarial records.
Check the population, evidence, permitted action, and actual effect together. A balanced calculation can still use the wrong population; a successful response can still leave an unknown financial outcome. The case’s numerical result applies only to its stated assumptions.
Sources and further reading
The chapter sources support the concepts and scope. They do not prescribe the synthetic model rates.