Evidence-bound AI assistance
Bound generative AI to evidence-supported tasks
The evaluated population
The cohort contains 6,200 reviewed summaries, of which 403 have the defined synthetic outcome: Materially unsupported case summary. The outcome is known by construction here. In production, label uncertainty and selection must be recorded separately.
Figure data and text version
| Outcome | Count |
|---|---|
| Materially unsupported case summary | 403 |
| Other labeled outcomes | 5,797 |
An assistant can summarize a case while still introducing unsupported facts or omitting decisive evidence. Its output needs links to the source record and a bounded role.
Generated text presents an unverified suspicion as an established fact.
All amounts, rates, capacity limits, and outcomes in this case are synthetic. The three conditions are separate assumptions for comparison. A better result in the response condition is not measured proof that the proposed control causes that improvement. The figures expose the calculation and its limits; a real deployment needs its own evidence.
Read the result
The rule flags 512 of 6,200 reviewed summaries. Of those flags, 193 meet the synthetic target, giving 37.7% precision. It misses 210 target events. Under the stated cost assumptions, residual loss and operating friction total $22,543. The important result is the connection between the population, action, capacity, and outcome—not one isolated score.
Model inputs and calculated values
Inputs below are the case-specific values. Each figure states the condition-specific assumptions and units used in its calculation. Calculated values are rounded for display.
| Input | Value |
|---|---|
| population | 6,200 |
| prevalence | 0.065 |
| severity | 90 |
| Calculated value | Result |
|---|---|
| population | 6,200 |
| positive | 403 |
| negative | 5,797 |
| tp | 193 |
| fp | 319 |
| fn | 210 |
| tn | 5,478 |
| loss | 18,900 |
| severity | 90 |
| precision | 37.7 |
| recall | 47.89 |
Four outcomes of the rule
The rule flags 193 synthetic positives and 319 negatives. It misses 210 positives. A flagged item is a decision to intervene; it is not proof of fraud, a legal prohibition, or any other real-world conclusion.
Figure data and text version
| Known outcome | Flagged | Not flagged |
|---|---|---|
| Materially unsupported case summary | 193 | 210 |
| Other outcome | 319 | 5,478 |
Three rates with different denominators
Precision is 37.7%, recall is 47.89%, and the false-positive rate is 5.5%. Changing the denominator changes the meaning. This record keeps each numerator attached to the population from which it came.
Figure data and text version
| Metric | Numerator | Denominator | Result |
|---|---|---|---|
| Precision | 193 | 512 | 37.7% |
| Recall | 193 | 403 | 47.89% |
| False-positive rate | 319 | 5,797 | 5.5% |
Precision changes with prevalence
This sensitivity plot holds recall at 48% and false-positive rate at 5.5%, then changes prevalence. It is an algebraic comparison, not a forecast. Even unchanged detection quality can produce a very different review queue when the base rate changes. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.
Figure data and text version
| Assumed prevalence | Precision % |
|---|---|
| 0.1% | 0.87 |
| 0.5% | 4.2 |
| 1% | 8.1 |
| 2% | 15.12 |
| 5% | 31.48 |
| 10% | 49.23 |
The threshold trade-off
Six illustrative score bands use a stated pair of detection rates. Lower sensitivity can reduce false alarms but miss more target events. These points do not come from a trained model and do not establish the best operating threshold. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.
Figure data and text version
| Score band | True positives | False positives |
|---|---|---|
| Band 1 | 395 | 870 |
| Band 2 | 379 | 464 |
| Band 3 | 347 | 203 |
| Band 4 | 290 | 70 |
| Band 5 | 202 | 23 |
| Band 6 | 101 | 6 |
A transparent loss-and-friction calculation
At $90 severity per missed synthetic positive, residual loss is $18,900. Review costs $2,048; lost contribution on false alarms is $1,595. The calculation assumes intervention prevents all flagged-positive loss and each false alarm loses the stated contribution. Relax those assumptions before applying it to a real policy.
Figure data and text version
| Cost component | USD |
|---|---|
| Missed-positive loss | 18,900 |
| Review cost | 2,048 |
| False-alarm contribution | 1,595 |
Observed outcomes mature over time
The final synthetic positive count is 403. Earlier observations reveal only a stated fraction. Comparing a day-1 cohort with a day-30 cohort would confuse label age with control quality. This curve models observation delay only; it does not change the final outcome. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.
Figure data and text version
| Days after event | Observed positives |
|---|---|
| 1 | 73 |
| 3 | 141 |
| 7 | 242 |
| 14 | 330 |
| 30 | 403 |
Review demand and available capacity
The flag count is 512. The comparison capacity is an illustrative 248 reviews per cohort window. A mathematical rule can be coherent while its resulting workload exceeds the operating team’s capacity. Capacity is not permission to ignore an applicable mandatory control.
Figure data and text version
| Queue measure | Items |
|---|---|
| Flagged for review | 512 |
| Available capacity | 248 |
| Excess demand | 264 |
A feature is an observation with provenance
This evidence contract supports bound generative ai to evidence-supported tasks. A value needs its event time, arrival time, scope, and source. Keeping unavailable evidence distinct from a measured zero prevents an outage from becoming a falsely reassuring feature.
Figure data and text version
| Field | Example | Meaning |
|---|---|---|
| entity_ref | Evidence-bound AI assistance | Subject of this case |
| event_time | 2026-09-18T09:00:00Z | When the event occurred |
| received_time | 2026-09-18T09:00:02Z | When the system learned it |
| signal_status | late | Evidence quality, not an outcome |
| label_definition | Materially unsupported case summary | The target used in these calculations |
Missing evidence changes the observed population
The cells show an explicitly constructed completeness profile for three signal groups. The stress condition removes more history and device evidence. Missingness does not prove the target outcome; it changes what the decision process knows.
Figure data and text version
| Signal group | Available | Missing |
|---|---|---|
| Identity evidence | 5,456 | 744 |
| Activity history | 4,464 | 1,736 |
| Context signal | 3,720 | 2,480 |
Evidence, score, and action remain separate
The policy can use bound generative ai to evidence-supported tasks only within its approved scope. The action record must retain which evidence was available, which model or rule ran, and which action was actually applied. The final action can differ from the score recommendation when a separate constraint applies.
Figure data and text version
| Stage | Record |
|---|---|
| Observe | Evidence-bound AI assistance: evidence as of the decision time |
| Evaluate | Rule flags 512 of 6,200 reviewed summaries |
| Apply | Record action, reason, owner, and expiry |
| Reconcile | Join the action to later outcomes without overwriting history |
What the result cannot establish
Observed classifications do not reveal every counterfactual. The synthetic labels make arithmetic possible, but production decline data is selected by prior policy. Keep measured outcomes, assumed prevention, and unknown alternatives separate when reporting impact.
Figure data and text version
| Claim | Evidence in this case | Limit |
|---|---|---|
| Detected target | 193 known synthetic positives flagged | Production labels may be delayed or wrong |
| Prevented loss | Assumed 17,370 USD | Requires an intervention-effect assumption |
| Customer impact | 319 synthetic negatives flagged | Not every flag causes abandonment |
| Unobserved alternative | Outcome without the action | Needs a valid evaluation design |
Connect the result to the system
Require evidence references, explicit uncertainty, and authorized human decisions for the defined task.
Check the population, evidence, permitted action, and actual effect together. A balanced calculation can still use the wrong population; a successful response can still leave an unknown financial outcome. The case’s numerical result applies only to its stated assumptions.
Sources and further reading
The chapter sources support the concepts and scope. They do not prescribe the synthetic model rates.