Worked case · Failure and stress · 12 figures

Proxy and data provenance review

Examine proxies and data provenance

Figure 01 / 12

The evaluated population

The evaluated population — Proxy and data provenance review. Count; one closed observation cohort. Exact values are in the figure data below.
Count; one closed observation cohort

The cohort contains 27,500 evaluated applications, of which 2,062 have the defined synthetic outcome: Synthetic credit-performance outcome. The outcome is known by construction here. In production, label uncertainty and selection must be recorded separately.

Figure data and text version
OutcomeCount
Synthetic credit-performance outcome2,062
Other labeled outcomes25,438

A feature can carry information about a prohibited basis or reflect unequal data coverage even when its label looks neutral. The source and use both matter.

A missing-history pattern is treated as low quality without examining who lacks the data.

All amounts, rates, capacity limits, and outcomes in this case are synthetic. The three conditions are separate assumptions for comparison. A better result in the response condition is not measured proof that the proposed control causes that improvement. The figures expose the calculation and its limits; a real deployment needs its own evidence.

Read the result

The rule flags 2,389 of 27,500 evaluated applications. Of those flags, 990 meet the synthetic target, giving 41.44% precision. It misses 1,072 target events. Under the stated cost assumptions, residual loss and operating friction total $1,877,175. The important result is the connection between the population, action, capacity, and outcome—not one isolated score.

Model inputs and calculated values

Inputs below are the case-specific values. Each figure states the condition-specific assumptions and units used in its calculation. Calculated values are rounded for display.

InputValue
population27,500
prevalence0.075
severity1,700
reviewCost20
Calculated valueResult
population27,500
positive2,062
negative25,438
tp990
fp1,399
fn1,072
tn24,039
loss1,822,400
severity1,700
precision41.44
recall48.01
Figure 02 / 12

Four outcomes of the rule

Four outcomes of the rule — Proxy and data provenance review. Counts; rows are actual labels, columns are actions. Exact values are in the figure data below.
Counts; rows are actual labels, columns are actions

The rule flags 990 synthetic positives and 1,399 negatives. It misses 1,072 positives. A flagged item is a decision to intervene; it is not proof of fraud, a legal prohibition, or any other real-world conclusion.

Figure data and text version
Known outcomeFlaggedNot flagged
Synthetic credit-performance outcome9901,072
Other outcome1,39924,039
Figure 03 / 12

Three rates with different denominators

Three rates with different denominators — Proxy and data provenance review. Rates within this cohort. Exact values are in the figure data below.
Rates within this cohort

Precision is 41.44%, recall is 48.01%, and the false-positive rate is 5.5%. Changing the denominator changes the meaning. This record keeps each numerator attached to the population from which it came.

Figure data and text version
MetricNumeratorDenominatorResult
Precision9902,38941.44%
Recall9902,06248.01%
False-positive rate1,39925,4385.5%
Figure 04 / 12

Precision changes with prevalence

Precision changes with prevalence — Proxy and data provenance review. Percent; fixed conditional detection rates. Exact values are in the figure data below.
Percent; fixed conditional detection rates

This sensitivity plot holds recall at 48% and false-positive rate at 5.5%, then changes prevalence. It is an algebraic comparison, not a forecast. Even unchanged detection quality can produce a very different review queue when the base rate changes. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.

Figure data and text version
Assumed prevalencePrecision %
0.1%0.87
0.5%4.2
1%8.1
2%15.12
5%31.48
10%49.23
Figure 05 / 12

The threshold trade-off

The threshold trade-off — Proxy and data provenance review. Count in the same cohort. Exact values are in the figure data below.
Count in the same cohort

Six illustrative score bands use a stated pair of detection rates. Lower sensitivity can reduce false alarms but miss more target events. These points do not come from a trained model and do not establish the best operating threshold. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.

Figure data and text version
Score bandTrue positivesFalse positives
Band 12,0213,816
Band 21,9382,035
Band 31,773890
Band 41,485305
Band 51,031102
Band 651625
Figure 06 / 12

A transparent loss-and-friction calculation

A transparent loss-and-friction calculation — Proxy and data provenance review. Illustrative USD; expected cost under stated intervention assumptions. Exact values are in the figure data below.
Illustrative USD; expected cost under stated intervention assumptions

At $1700 severity per missed synthetic positive, residual loss is $1,822,400. Review costs $47,780; lost contribution on false alarms is $6,995. The calculation assumes intervention prevents all flagged-positive loss and each false alarm loses the stated contribution. Relax those assumptions before applying it to a real policy.

Figure data and text version
Cost componentUSD
Missed-positive loss1,822,400
Review cost47,780
False-alarm contribution6,995
Figure 07 / 12

Observed outcomes mature over time

Observed outcomes mature over time — Proxy and data provenance review. Count; final outcome fixed. Exact values are in the figure data below.
Count; final outcome fixed

The final synthetic positive count is 2,062. Earlier observations reveal only a stated fraction. Comparing a day-1 cohort with a day-30 cohort would confuse label age with control quality. This curve models observation delay only; it does not change the final outcome. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.

Figure data and text version
Days after eventObserved positives
1371
3722
71,237
141,691
302,062
Figure 08 / 12

Review demand and available capacity

Review demand and available capacity — Proxy and data provenance review. Items in the cohort window. Exact values are in the figure data below.
Items in the cohort window

The flag count is 2,389. The comparison capacity is an illustrative 1,100 reviews per cohort window. A mathematical rule can be coherent while its resulting workload exceeds the operating team’s capacity. Capacity is not permission to ignore an applicable mandatory control.

Figure data and text version
Queue measureItems
Flagged for review2,389
Available capacity1,100
Excess demand1,289
Figure 09 / 12

A feature is an observation with provenance

A feature is an observation with provenance — Proxy and data provenance review. Illustrative data contract. Exact values are in the figure data below.
Illustrative data contract

This evidence contract supports examine proxies and data provenance. A value needs its event time, arrival time, scope, and source. Keeping unavailable evidence distinct from a measured zero prevents an outage from becoming a falsely reassuring feature.

Figure data and text version
FieldExampleMeaning
entity_refProxy and data provenance reviewSubject of this case
event_time2026-09-18T09:00:00ZWhen the event occurred
received_time2026-09-18T09:00:02ZWhen the system learned it
signal_statuslateEvidence quality, not an outcome
label_definitionSynthetic credit-performance outcomeThe target used in these calculations
Figure 10 / 12

Missing evidence changes the observed population

Missing evidence changes the observed population — Proxy and data provenance review. Count; each row is the same cohort. Exact values are in the figure data below.
Count; each row is the same cohort

The cells show an explicitly constructed completeness profile for three signal groups. The stress condition removes more history and device evidence. Missingness does not prove the target outcome; it changes what the decision process knows.

Figure data and text version
Signal groupAvailableMissing
Identity evidence24,2003,300
Activity history19,8007,700
Context signal16,50011,000
Figure 11 / 12

Evidence, score, and action remain separate

Evidence, score, and action remain separate — Proxy and data provenance review. Decision lifecycle. Exact values are in the figure data below.
Decision lifecycle

The policy can use examine proxies and data provenance only within its approved scope. The action record must retain which evidence was available, which model or rule ran, and which action was actually applied. The final action can differ from the score recommendation when a separate constraint applies.

Figure data and text version
StageRecord
ObserveProxy and data provenance review: evidence as of the decision time
EvaluateRule flags 2,389 of 27,500 evaluated applications
ApplyRecord action, reason, owner, and expiry
ReconcileJoin the action to later outcomes without overwriting history
Figure 12 / 12

What the result cannot establish

What the result cannot establish — Proxy and data provenance review. Interpretation boundary. Exact values are in the figure data below.
Interpretation boundary

Observed classifications do not reveal every counterfactual. The synthetic labels make arithmetic possible, but production decline data is selected by prior policy. Keep measured outcomes, assumed prevention, and unknown alternatives separate when reporting impact.

Figure data and text version
ClaimEvidence in this caseLimit
Detected target990 known synthetic positives flaggedProduction labels may be delayed or wrong
Prevented lossAssumed 1,683,000 USDRequires an intervention-effect assumption
Customer impact1399 synthetic negatives flaggedNot every flag causes abandonment
Unobserved alternativeOutcome without the actionNeeds a valid evaluation design

Connect the result to the system

Review provenance, missingness, predictive purpose, and appropriate outcome comparisons.

Check the population, evidence, permitted action, and actual effect together. A balanced calculation can still use the wrong population; a successful response can still leave an unknown financial outcome. The case’s numerical result applies only to its stated assumptions.

Sources and further reading

The chapter sources support the concepts and scope. They do not prescribe the synthetic model rates.

  1. Regulation B, 12 CFR 1002.6: evaluation of applications
  2. Regulation B, 12 CFR 1002.9: notifications
  3. NIST: AI Risk Management Framework