Worked case · Controlled response · 12 figures

Identity-preserving normalization

Normalize without destroying identity

Figure 01 / 12

A match is an identity hypothesis

A match is an identity hypothesis — Identity-preserving normalization. Count; identity match does not itself state legal disposition. Exact values are in the figure data below.
Count; identity match does not itself state legal disposition

The synthetic population contains 47 known identity matches to fictional list records and 16753 nonmatches. This ground truth is supplied for the example. A real list match still requires correct identity resolution and a separate analysis of the applicable restriction.

Figure data and text version
Known synthetic statusRecords
Same subject as fictional list record47
Different subject16,753

Case, spacing, word order, transliteration, and punctuation can affect matching. Normalization should improve comparison while retaining the original evidence.

Keep original and normalized values, document the transform, and evaluate known matches and lookalikes.

All amounts, rates, capacity limits, and outcomes in this case are synthetic. The three conditions are separate assumptions for comparison. A better result in the response condition is not measured proof that the proposed control causes that improvement. The figures expose the calculation and its limits; a real deployment needs its own evidence.

Read the result

The matcher returns 146 candidates from 16,800 records. It identifies 45 of 47 known fictional identity matches and misses 2. After scoped suppressions and stale-evidence returns, review demand is 128. The example keeps identity resolution, control availability, and the final legal disposition separate.

Model inputs and calculated values

Inputs below are the case-specific values. Each figure states the condition-specific assumptions and units used in its calculation. Calculated values are rounded for display.

InputValue
population16,800
knownMatches47
capacity210
Calculated valueResult
population16,800
actual47
tp45
fn2
fp101
tn16,652
candidates146
suppressed18
stale0
review128
pending0
capacity210
Figure 02 / 12

Candidate generation changes reviewer workload

Candidate generation changes reviewer workload — Identity-preserving normalization. Count; constructed identity-matching outcome. Exact values are in the figure data below.
Count; constructed identity-matching outcome

The matcher returns 146 candidates: 45 known synthetic matches and 101 unrelated records. It misses 2 known matches. A low candidate count can mean precision improved or coverage failed; the truth set and data pipeline are needed to distinguish those explanations.

Figure data and text version
MeasureRecords
Candidates146
Known matches found45
Unrelated candidates101
Known matches missed2
Figure 03 / 12

Evaluate matching against a defined truth set

Evaluate matching against a defined truth set — Identity-preserving normalization. Count; identity-matching evaluation. Exact values are in the figure data below.
Count; identity-matching evaluation

The four cells sum to 16,800. Candidate precision is 30.82% and matching recall is 95.74%. The labels refer only to fictional identity matches in this test set, not to whether a real payment is lawful or a customer is suspicious.

Figure data and text version
Actual identityCandidateNo candidate
Same fictional subject452
Different subject10116,652
Figure 04 / 12

A toy token score exposes a limitation

A toy token score exposes a limitation — Identity-preserving normalization. Percent; explicitly defined toy text metric. Exact values are in the figure data below.
Percent; explicitly defined toy text metric

The example query is “Marin Trade”. The score is the Jaccard overlap of lowercase space-separated token sets, multiplied by 100. Reversing token order leaves this toy score unchanged. It ignores spelling distance, transliteration, dates, addresses, and legal identity, so it is not a production screening method.

Figure data and text version
Candidate stringToken overlap %
Marin Trade100
MARIN TRADE100
Trade Marin100
Marin Trade Holdings66.67
Marin Services33.33
North Harbor Logistics0
Figure 05 / 12

More attributes can change the interpretation

More attributes can change the interpretation — Identity-preserving normalization. Fictional entity-matching record. Exact values are in the figure data below.
Fictional entity-matching record

A name score is a lead. A reviewer compares reliable identifiers and relevant context, while preserving uncertainty when attributes are missing. A disagreement can support exclusion only under the approved method and facts; missing information is not a reliable contradiction.

Figure data and text version
AttributeIllustrative evidenceInterpretation
NameMarin TradeCandidate-generation input
Registration identifierCase-specific identifierPotentially strong identity evidence
CountryDeclared and document valuesContext, not a universal exclusion
Effective dateList and customer record datesFacts must refer to the relevant period
OwnershipSeparate graph recordNot inferred from text similarity
Figure 06 / 12

List freshness is an operational dependency

List freshness is an operational dependency — Identity-preserving normalization. Illustrative minutes; not a required legal interval. Exact values are in the figure data below.
Illustrative minutes; not a required legal interval

The stated ingestion lag is 3 minutes. A download, successful parse, validated release, activation, and rescreen are separate events. “Latest file received” is not proof that the production matcher used it for the relevant decision.

Figure data and text version
EventRelative minuteEvidence
Publisher release0Source release identifier
Ingestion3Retrieved bytes and checksum
Validation5Schema and population checks
Activation7Version used by production decisions
Rescreen batch17Affected population and outcomes
Figure 07 / 12

A clearance has scope and expiry

A clearance has scope and expiry — Identity-preserving normalization. Count; review = candidates − suppressed + stale. Exact values are in the figure data below.
Count; review = candidates − suppressed + stale

18 candidate references meet a prior suppression condition, but 0 no longer meet its evidence requirements and return to review. Suppression applies to a defined subject, list record, and evidence basis; it is not a general exemption for similar names.

Figure data and text version
StateCandidate references
Raw candidates146
Suppression matches18
Stale suppression evidence0
Net review demand128
Figure 08 / 12

Review demand must reach an owned queue

Review demand must reach an owned queue — Identity-preserving normalization. Cases in one review window. Exact values are in the figure data below.
Cases in one review window

Net review demand is 128 cases and the illustrative capacity is 210. That leaves 0 pending. Staffing can change waiting time, but it does not change the applicable disposition or justify an automatic release when a required control is unresolved.

Figure data and text version
MeasureCases
Net review demand128
Available capacity210
Completed within capacity128
Pending0
Figure 09 / 12

Unavailable screening is not a clear result

Unavailable screening is not a clear result — Identity-preserving normalization. Count; pending is distinct from clear. Exact values are in the figure data below.
Count; pending is distinct from clear

67 records enter a pending-control state in this scenario, while 16733 have a recorded screening result. The status distinction prevents an unavailable dependency from being represented as no match. The actual action follows approved scope-specific policy and applicable duties.

Figure data and text version
Control stateRecords
Screening result recorded16,733
Required result pending67
Figure 10 / 12

Match resolution and legal disposition differ

Match resolution and legal disposition differ — Identity-preserving normalization. Illustrative screening-to-disposition sequence. Exact values are in the figure data below.
Illustrative screening-to-disposition sequence

The workflow first establishes whether the candidate is the relevant subject, then applies the appropriate legal analysis and disposition. A reviewer can document a false identity match without deciding that every possible restriction on the transaction has been resolved.

Figure data and text version
StageDecision scope
CandidatePotential similarity or identifier relationship
Identity resolutionSame subject, different subject, or unresolved
Scope analysisRelevant jurisdiction, activity, program, and authority
DispositionApproved handling and any required records or reporting
Release controlAuthorized evidence-linked transition
Figure 11 / 12

Keep the list, matcher, and evidence versions

Keep the list, matcher, and evidence versions — Identity-preserving normalization. Versioned decision trace. Exact values are in the figure data below.
Versioned decision trace

The record explains which inputs supported normalize without destroying identity. Replaying an old payment against a current list answers a current-screening question. Reconstructing the original decision requires the list and logic versions actually available at that time.

Figure data and text version
FieldIllustrative value
entity_refIdentity-preserving normalization
list_releasefictional-release-2026-09-18
matcher_versionteaching-v3
observed_lag_minutes3
candidate_count146
review_case_count128
evidence_ownerrisk operations
Figure 12 / 12

False clears and false alerts have different costs

False clears and false alerts have different costs — Identity-preserving normalization. Diagnostic control view. Exact values are in the figure data below.
Diagnostic control view

A complete assessment includes missed identity matches, unnecessary review, stale clearances, and unavailable controls. These observations cannot be collapsed into one accuracy percentage without losing their different operational and legal implications.

Figure data and text version
Failure modeObserved in this exampleRequired response
Missed known match2Investigate matching and data coverage
Unrelated candidate101Improve evidence and resolution
Stale suppression0Reassess the clearance basis
Pending required result67Apply the approved failure policy

Connect the result to the system

Keep original and normalized values, document the transform, and evaluate known matches and lookalikes.

Check the population, evidence, permitted action, and actual effect together. A balanced calculation can still use the wrong population; a successful response can still leave an unknown financial outcome. The case’s numerical result applies only to its stated assumptions.

Sources and further reading

The chapter sources support the concepts and scope. They do not prescribe the synthetic model rates.

  1. OFAC: Sanctions List Service
  2. OFAC: A Framework for Compliance Commitments
  3. OFAC FAQ 5: resolving matches and choosing a disposition