Attributes beyond the name
Use more than name similarity
A match is an identity hypothesis
The synthetic population contains 62 known identity matches to fictional list records and 22338 nonmatches. This ground truth is supplied for the example. A real list match still requires correct identity resolution and a separate analysis of the applicable restriction.
Figure data and text version
| Known synthetic status | Records |
|---|---|
| Same subject as fictional list record | 62 |
| Different subject | 22,338 |
A name candidate becomes useful when reliable identifiers and context help establish whether it is the same subject. Missing attributes leave uncertainty.
The reference case starts with the stated population and a functioning evidence path. The owner is risk operations.
All amounts, rates, capacity limits, and outcomes in this case are synthetic. The three conditions are separate assumptions for comparison. A better result in the response condition is not measured proof that the proposed control causes that improvement. The figures expose the calculation and its limits; a real deployment needs its own evidence.
Read the result
The matcher returns 236 candidates from 22,400 records. It identifies 57 of 62 known fictional identity matches and misses 5. After scoped suppressions and stale-evidence returns, review demand is 190. The example keeps identity resolution, control availability, and the final legal disposition separate.
Model inputs and calculated values
Inputs below are the case-specific values. Each figure states the condition-specific assumptions and units used in its calculation. Calculated values are rounded for display.
| Input | Value |
|---|---|
| population | 22,400 |
| knownMatches | 62 |
| capacity | 270 |
| Calculated value | Result |
|---|---|
| population | 22,400 |
| actual | 62 |
| tp | 57 |
| fn | 5 |
| fp | 179 |
| tn | 22,159 |
| candidates | 236 |
| suppressed | 47 |
| stale | 1 |
| review | 190 |
| pending | 0 |
| capacity | 270 |
Candidate generation changes reviewer workload
The matcher returns 236 candidates: 57 known synthetic matches and 179 unrelated records. It misses 5 known matches. A low candidate count can mean precision improved or coverage failed; the truth set and data pipeline are needed to distinguish those explanations.
Figure data and text version
| Measure | Records |
|---|---|
| Candidates | 236 |
| Known matches found | 57 |
| Unrelated candidates | 179 |
| Known matches missed | 5 |
Evaluate matching against a defined truth set
The four cells sum to 22,400. Candidate precision is 24.15% and matching recall is 91.94%. The labels refer only to fictional identity matches in this test set, not to whether a real payment is lawful or a customer is suspicious.
Figure data and text version
| Actual identity | Candidate | No candidate |
|---|---|---|
| Same fictional subject | 57 | 5 |
| Different subject | 179 | 22,159 |
A toy token score exposes a limitation
The example query is “Marin Trade”. The score is the Jaccard overlap of lowercase space-separated token sets, multiplied by 100. Reversing token order leaves this toy score unchanged. It ignores spelling distance, transliteration, dates, addresses, and legal identity, so it is not a production screening method.
Figure data and text version
| Candidate string | Token overlap % |
|---|---|
| Marin Trade | 100 |
| MARIN TRADE | 100 |
| Trade Marin | 100 |
| Marin Trade Holdings | 66.67 |
| Marin Services | 33.33 |
| North Harbor Logistics | 0 |
More attributes can change the interpretation
A name score is a lead. A reviewer compares reliable identifiers and relevant context, while preserving uncertainty when attributes are missing. A disagreement can support exclusion only under the approved method and facts; missing information is not a reliable contradiction.
Figure data and text version
| Attribute | Illustrative evidence | Interpretation |
|---|---|---|
| Name | Marin Trade | Candidate-generation input |
| Registration identifier | Case-specific identifier | Potentially strong identity evidence |
| Country | Declared and document values | Context, not a universal exclusion |
| Effective date | List and customer record dates | Facts must refer to the relevant period |
| Ownership | Separate graph record | Not inferred from text similarity |
List freshness is an operational dependency
The stated ingestion lag is 5 minutes. A download, successful parse, validated release, activation, and rescreen are separate events. “Latest file received” is not proof that the production matcher used it for the relevant decision.
Figure data and text version
| Event | Relative minute | Evidence |
|---|---|---|
| Publisher release | 0 | Source release identifier |
| Ingestion | 5 | Retrieved bytes and checksum |
| Validation | 7 | Schema and population checks |
| Activation | 9 | Version used by production decisions |
| Rescreen batch | 19 | Affected population and outcomes |
A clearance has scope and expiry
47 candidate references meet a prior suppression condition, but 1 no longer meet its evidence requirements and return to review. Suppression applies to a defined subject, list record, and evidence basis; it is not a general exemption for similar names.
Figure data and text version
| State | Candidate references |
|---|---|
| Raw candidates | 236 |
| Suppression matches | 47 |
| Stale suppression evidence | 1 |
| Net review demand | 190 |
Review demand must reach an owned queue
Net review demand is 190 cases and the illustrative capacity is 270. That leaves 0 pending. Staffing can change waiting time, but it does not change the applicable disposition or justify an automatic release when a required control is unresolved.
Figure data and text version
| Measure | Cases |
|---|---|
| Net review demand | 190 |
| Available capacity | 270 |
| Completed within capacity | 190 |
| Pending | 0 |
Unavailable screening is not a clear result
45 records enter a pending-control state in this scenario, while 22355 have a recorded screening result. The status distinction prevents an unavailable dependency from being represented as no match. The actual action follows approved scope-specific policy and applicable duties.
Figure data and text version
| Control state | Records |
|---|---|
| Screening result recorded | 22,355 |
| Required result pending | 45 |
Match resolution and legal disposition differ
The workflow first establishes whether the candidate is the relevant subject, then applies the appropriate legal analysis and disposition. A reviewer can document a false identity match without deciding that every possible restriction on the transaction has been resolved.
Figure data and text version
| Stage | Decision scope |
|---|---|
| Candidate | Potential similarity or identifier relationship |
| Identity resolution | Same subject, different subject, or unresolved |
| Scope analysis | Relevant jurisdiction, activity, program, and authority |
| Disposition | Approved handling and any required records or reporting |
| Release control | Authorized evidence-linked transition |
Keep the list, matcher, and evidence versions
The record explains which inputs supported use more than name similarity. Replaying an old payment against a current list answers a current-screening question. Reconstructing the original decision requires the list and logic versions actually available at that time.
Figure data and text version
| Field | Illustrative value |
|---|---|
| entity_ref | Attributes beyond the name |
| list_release | fictional-release-2026-09-18 |
| matcher_version | teaching-v1 |
| observed_lag_minutes | 5 |
| candidate_count | 236 |
| review_case_count | 190 |
| evidence_owner | risk operations |
False clears and false alerts have different costs
A complete assessment includes missed identity matches, unnecessary review, stale clearances, and unavailable controls. These observations cannot be collapsed into one accuracy percentage without losing their different operational and legal implications.
Figure data and text version
| Failure mode | Observed in this example | Required response |
|---|---|---|
| Missed known match | 5 | Investigate matching and data coverage |
| Unrelated candidate | 179 | Improve evidence and resolution |
| Stale suppression | 1 | Reassess the clearance basis |
| Pending required result | 45 | Apply the approved failure policy |
Connect the result to the system
Compare reliable attributes under a defined resolution method and preserve conflicting evidence.
Check the population, evidence, permitted action, and actual effect together. A balanced calculation can still use the wrong population; a successful response can still leave an unknown financial outcome. The case’s numerical result applies only to its stated assumptions.
Sources and further reading
The chapter sources support the concepts and scope. They do not prescribe the synthetic model rates.