Worked case · Failure and stress · 12 figures

Identity-preserving normalization

Normalize without destroying identity

Figure 01 / 12

A match is an identity hypothesis

A match is an identity hypothesis — Identity-preserving normalization. Count; identity match does not itself state legal disposition. Exact values are in the figure data below.
Count; identity match does not itself state legal disposition

The synthetic population contains 47 known identity matches to fictional list records and 16753 nonmatches. This ground truth is supplied for the example. A real list match still requires correct identity resolution and a separate analysis of the applicable restriction.

Figure data and text version
Known synthetic statusRecords
Same subject as fictional list record47
Different subject16,753

Case, spacing, word order, transliteration, and punctuation can affect matching. Normalization should improve comparison while retaining the original evidence.

Aggressive normalization removes distinctions needed to separate unrelated entities.

All amounts, rates, capacity limits, and outcomes in this case are synthetic. The three conditions are separate assumptions for comparison. A better result in the response condition is not measured proof that the proposed control causes that improvement. The figures expose the calculation and its limits; a real deployment needs its own evidence.

Read the result

The matcher returns 500 candidates from 16,800 records. It identifies 31 of 47 known fictional identity matches and misses 16. After scoped suppressions and stale-evidence returns, review demand is 343. The example keeps identity resolution, control availability, and the final legal disposition separate.

Model inputs and calculated values

Inputs below are the case-specific values. Each figure states the condition-specific assumptions and units used in its calculation. Calculated values are rounded for display.

InputValue
population16,800
knownMatches47
capacity210
Calculated valueResult
population16,800
actual47
tp31
fn16
fp469
tn16,284
candidates500
suppressed225
stale68
review343
pending133
capacity210
Figure 02 / 12

Candidate generation changes reviewer workload

Candidate generation changes reviewer workload — Identity-preserving normalization. Count; constructed identity-matching outcome. Exact values are in the figure data below.
Count; constructed identity-matching outcome

The matcher returns 500 candidates: 31 known synthetic matches and 469 unrelated records. It misses 16 known matches. A low candidate count can mean precision improved or coverage failed; the truth set and data pipeline are needed to distinguish those explanations.

Figure data and text version
MeasureRecords
Candidates500
Known matches found31
Unrelated candidates469
Known matches missed16
Figure 03 / 12

Evaluate matching against a defined truth set

Evaluate matching against a defined truth set — Identity-preserving normalization. Count; identity-matching evaluation. Exact values are in the figure data below.
Count; identity-matching evaluation

The four cells sum to 16,800. Candidate precision is 6.2% and matching recall is 65.96%. The labels refer only to fictional identity matches in this test set, not to whether a real payment is lawful or a customer is suspicious.

Figure data and text version
Actual identityCandidateNo candidate
Same fictional subject3116
Different subject46916,284
Figure 04 / 12

A toy token score exposes a limitation

A toy token score exposes a limitation — Identity-preserving normalization. Percent; explicitly defined toy text metric. Exact values are in the figure data below.
Percent; explicitly defined toy text metric

The example query is “Marin Trade”. The score is the Jaccard overlap of lowercase space-separated token sets, multiplied by 100. Reversing token order leaves this toy score unchanged. It ignores spelling distance, transliteration, dates, addresses, and legal identity, so it is not a production screening method.

Figure data and text version
Candidate stringToken overlap %
Marin Trade100
MARIN TRADE100
Trade Marin100
Marin Trade Holdings66.67
Marin Services33.33
North Harbor Logistics0
Figure 05 / 12

More attributes can change the interpretation

More attributes can change the interpretation — Identity-preserving normalization. Fictional entity-matching record. Exact values are in the figure data below.
Fictional entity-matching record

A name score is a lead. A reviewer compares reliable identifiers and relevant context, while preserving uncertainty when attributes are missing. A disagreement can support exclusion only under the approved method and facts; missing information is not a reliable contradiction.

Figure data and text version
AttributeIllustrative evidenceInterpretation
NameMarin TradeCandidate-generation input
Registration identifierCase-specific identifierPotentially strong identity evidence
CountryDeclared and document valuesContext, not a universal exclusion
Effective dateList and customer record datesFacts must refer to the relevant period
OwnershipSeparate graph recordNot inferred from text similarity
Figure 06 / 12

List freshness is an operational dependency

List freshness is an operational dependency — Identity-preserving normalization. Illustrative minutes; not a required legal interval. Exact values are in the figure data below.
Illustrative minutes; not a required legal interval

The stated ingestion lag is 180 minutes. A download, successful parse, validated release, activation, and rescreen are separate events. “Latest file received” is not proof that the production matcher used it for the relevant decision.

Figure data and text version
EventRelative minuteEvidence
Publisher release0Source release identifier
Ingestion180Retrieved bytes and checksum
Validation182Schema and population checks
Activation184Version used by production decisions
Rescreen batch194Affected population and outcomes
Figure 07 / 12

A clearance has scope and expiry

A clearance has scope and expiry — Identity-preserving normalization. Count; review = candidates − suppressed + stale. Exact values are in the figure data below.
Count; review = candidates − suppressed + stale

225 candidate references meet a prior suppression condition, but 68 no longer meet its evidence requirements and return to review. Suppression applies to a defined subject, list record, and evidence basis; it is not a general exemption for similar names.

Figure data and text version
StateCandidate references
Raw candidates500
Suppression matches225
Stale suppression evidence68
Net review demand343
Figure 08 / 12

Review demand must reach an owned queue

Review demand must reach an owned queue — Identity-preserving normalization. Cases in one review window. Exact values are in the figure data below.
Cases in one review window

Net review demand is 343 cases and the illustrative capacity is 210. That leaves 133 pending. Staffing can change waiting time, but it does not change the applicable disposition or justify an automatic release when a required control is unresolved.

Figure data and text version
MeasureCases
Net review demand343
Available capacity210
Completed within capacity210
Pending133
Figure 09 / 12

Unavailable screening is not a clear result

Unavailable screening is not a clear result — Identity-preserving normalization. Count; pending is distinct from clear. Exact values are in the figure data below.
Count; pending is distinct from clear

3024 records enter a pending-control state in this scenario, while 13776 have a recorded screening result. The status distinction prevents an unavailable dependency from being represented as no match. The actual action follows approved scope-specific policy and applicable duties.

Figure data and text version
Control stateRecords
Screening result recorded13,776
Required result pending3,024
Figure 10 / 12

Match resolution and legal disposition differ

Match resolution and legal disposition differ — Identity-preserving normalization. Illustrative screening-to-disposition sequence. Exact values are in the figure data below.
Illustrative screening-to-disposition sequence

The workflow first establishes whether the candidate is the relevant subject, then applies the appropriate legal analysis and disposition. A reviewer can document a false identity match without deciding that every possible restriction on the transaction has been resolved.

Figure data and text version
StageDecision scope
CandidatePotential similarity or identifier relationship
Identity resolutionSame subject, different subject, or unresolved
Scope analysisRelevant jurisdiction, activity, program, and authority
DispositionApproved handling and any required records or reporting
Release controlAuthorized evidence-linked transition
Figure 11 / 12

Keep the list, matcher, and evidence versions

Keep the list, matcher, and evidence versions — Identity-preserving normalization. Versioned decision trace. Exact values are in the figure data below.
Versioned decision trace

The record explains which inputs supported normalize without destroying identity. Replaying an old payment against a current list answers a current-screening question. Reconstructing the original decision requires the list and logic versions actually available at that time.

Figure data and text version
FieldIllustrative value
entity_refIdentity-preserving normalization
list_releasefictional-release-2026-09-18
matcher_versionteaching-v2
observed_lag_minutes180
candidate_count500
review_case_count343
evidence_ownerrisk operations
Figure 12 / 12

False clears and false alerts have different costs

False clears and false alerts have different costs — Identity-preserving normalization. Diagnostic control view. Exact values are in the figure data below.
Diagnostic control view

A complete assessment includes missed identity matches, unnecessary review, stale clearances, and unavailable controls. These observations cannot be collapsed into one accuracy percentage without losing their different operational and legal implications.

Figure data and text version
Failure modeObserved in this exampleRequired response
Missed known match16Investigate matching and data coverage
Unrelated candidate469Improve evidence and resolution
Stale suppression68Reassess the clearance basis
Pending required result3,024Apply the approved failure policy

Connect the result to the system

Keep original and normalized values, document the transform, and evaluate known matches and lookalikes.

Check the population, evidence, permitted action, and actual effect together. A balanced calculation can still use the wrong population; a successful response can still leave an unknown financial outcome. The case’s numerical result applies only to its stated assumptions.

Sources and further reading

The chapter sources support the concepts and scope. They do not prescribe the synthetic model rates.

  1. OFAC: Sanctions List Service
  2. OFAC: A Framework for Compliance Commitments
  3. OFAC FAQ 5: resolving matches and choosing a disposition