Versioned list ingestion
Version the list pipeline
A match is an identity hypothesis
The synthetic population contains 84 known identity matches to fictional list records and 27416 nonmatches. This ground truth is supplied for the example. A real list match still requires correct identity resolution and a separate analysis of the applicable restriction.
Figure data and text version
| Known synthetic status | Records |
|---|---|
| Same subject as fictional list record | 84 |
| Different subject | 27,416 |
A list release must be retrieved, parsed, validated, activated, and used by production. Each step can fail independently.
The monitoring dashboard reports download success while the matcher still uses the old version.
All amounts, rates, capacity limits, and outcomes in this case are synthetic. The three conditions are separate assumptions for comparison. A better result in the response condition is not measured proof that the proposed control causes that improvement. The figures expose the calculation and its limits; a real deployment needs its own evidence.
Read the result
The matcher returns 823 candidates from 27,500 records. It identifies 55 of 84 known fictional identity matches and misses 29. After scoped suppressions and stale-evidence returns, review demand is 564. The example keeps identity resolution, control availability, and the final legal disposition separate.
Model inputs and calculated values
Inputs below are the case-specific values. Each figure states the condition-specific assumptions and units used in its calculation. Calculated values are rounded for display.
| Input | Value |
|---|---|
| population | 27,500 |
| knownMatches | 84 |
| capacity | 330 |
| Calculated value | Result |
|---|---|
| population | 27,500 |
| actual | 84 |
| tp | 55 |
| fn | 29 |
| fp | 768 |
| tn | 26,648 |
| candidates | 823 |
| suppressed | 370 |
| stale | 111 |
| review | 564 |
| pending | 234 |
| capacity | 330 |
Candidate generation changes reviewer workload
The matcher returns 823 candidates: 55 known synthetic matches and 768 unrelated records. It misses 29 known matches. A low candidate count can mean precision improved or coverage failed; the truth set and data pipeline are needed to distinguish those explanations.
Figure data and text version
| Measure | Records |
|---|---|
| Candidates | 823 |
| Known matches found | 55 |
| Unrelated candidates | 768 |
| Known matches missed | 29 |
Evaluate matching against a defined truth set
The four cells sum to 27,500. Candidate precision is 6.68% and matching recall is 65.48%. The labels refer only to fictional identity matches in this test set, not to whether a real payment is lawful or a customer is suspicious.
Figure data and text version
| Actual identity | Candidate | No candidate |
|---|---|---|
| Same fictional subject | 55 | 29 |
| Different subject | 768 | 26,648 |
A toy token score exposes a limitation
The example query is “Marin Trade”. The score is the Jaccard overlap of lowercase space-separated token sets, multiplied by 100. Reversing token order leaves this toy score unchanged. It ignores spelling distance, transliteration, dates, addresses, and legal identity, so it is not a production screening method.
Figure data and text version
| Candidate string | Token overlap % |
|---|---|
| Marin Trade | 100 |
| MARIN TRADE | 100 |
| Trade Marin | 100 |
| Marin Trade Holdings | 66.67 |
| Marin Services | 33.33 |
| North Harbor Logistics | 0 |
More attributes can change the interpretation
A name score is a lead. A reviewer compares reliable identifiers and relevant context, while preserving uncertainty when attributes are missing. A disagreement can support exclusion only under the approved method and facts; missing information is not a reliable contradiction.
Figure data and text version
| Attribute | Illustrative evidence | Interpretation |
|---|---|---|
| Name | Marin Trade | Candidate-generation input |
| Registration identifier | Case-specific identifier | Potentially strong identity evidence |
| Country | Declared and document values | Context, not a universal exclusion |
| Effective date | List and customer record dates | Facts must refer to the relevant period |
| Ownership | Separate graph record | Not inferred from text similarity |
List freshness is an operational dependency
The stated ingestion lag is 180 minutes. A download, successful parse, validated release, activation, and rescreen are separate events. “Latest file received” is not proof that the production matcher used it for the relevant decision.
Figure data and text version
| Event | Relative minute | Evidence |
|---|---|---|
| Publisher release | 0 | Source release identifier |
| Ingestion | 180 | Retrieved bytes and checksum |
| Validation | 182 | Schema and population checks |
| Activation | 184 | Version used by production decisions |
| Rescreen batch | 194 | Affected population and outcomes |
A clearance has scope and expiry
370 candidate references meet a prior suppression condition, but 111 no longer meet its evidence requirements and return to review. Suppression applies to a defined subject, list record, and evidence basis; it is not a general exemption for similar names.
Figure data and text version
| State | Candidate references |
|---|---|
| Raw candidates | 823 |
| Suppression matches | 370 |
| Stale suppression evidence | 111 |
| Net review demand | 564 |
Review demand must reach an owned queue
Net review demand is 564 cases and the illustrative capacity is 330. That leaves 234 pending. Staffing can change waiting time, but it does not change the applicable disposition or justify an automatic release when a required control is unresolved.
Figure data and text version
| Measure | Cases |
|---|---|
| Net review demand | 564 |
| Available capacity | 330 |
| Completed within capacity | 330 |
| Pending | 234 |
Unavailable screening is not a clear result
4950 records enter a pending-control state in this scenario, while 22550 have a recorded screening result. The status distinction prevents an unavailable dependency from being represented as no match. The actual action follows approved scope-specific policy and applicable duties.
Figure data and text version
| Control state | Records |
|---|---|
| Screening result recorded | 22,550 |
| Required result pending | 4,950 |
Match resolution and legal disposition differ
The workflow first establishes whether the candidate is the relevant subject, then applies the appropriate legal analysis and disposition. A reviewer can document a false identity match without deciding that every possible restriction on the transaction has been resolved.
Figure data and text version
| Stage | Decision scope |
|---|---|
| Candidate | Potential similarity or identifier relationship |
| Identity resolution | Same subject, different subject, or unresolved |
| Scope analysis | Relevant jurisdiction, activity, program, and authority |
| Disposition | Approved handling and any required records or reporting |
| Release control | Authorized evidence-linked transition |
Keep the list, matcher, and evidence versions
The record explains which inputs supported version the list pipeline. Replaying an old payment against a current list answers a current-screening question. Reconstructing the original decision requires the list and logic versions actually available at that time.
Figure data and text version
| Field | Illustrative value |
|---|---|
| entity_ref | Versioned list ingestion |
| list_release | fictional-release-2026-09-18 |
| matcher_version | teaching-v2 |
| observed_lag_minutes | 180 |
| candidate_count | 823 |
| review_case_count | 564 |
| evidence_owner | risk operations |
False clears and false alerts have different costs
A complete assessment includes missed identity matches, unnecessary review, stale clearances, and unavailable controls. These observations cannot be collapsed into one accuracy percentage without losing their different operational and legal implications.
Figure data and text version
| Failure mode | Observed in this example | Required response |
|---|---|---|
| Missed known match | 29 | Investigate matching and data coverage |
| Unrelated candidate | 768 | Improve evidence and resolution |
| Stale suppression | 111 | Reassess the clearance basis |
| Pending required result | 4,950 | Apply the approved failure policy |
Connect the result to the system
Record the active list version on each decision and test activation and rescreen coverage.
Check the population, evidence, permitted action, and actual effect together. A balanced calculation can still use the wrong population; a successful response can still leave an unknown financial outcome. The case’s numerical result applies only to its stated assumptions.
Sources and further reading
The chapter sources support the concepts and scope. They do not prescribe the synthetic model rates.