Worked case · Reference condition · 12 figures

Outcome comparisons in context

Define fairness in the product context

Figure 01 / 12

Start with the assigned population

Start with the assigned population — Outcome comparisons in context. Count; synthetic experiment. Exact values are in the figure data below.
Count; synthetic experiment

This teaching experiment assigns 7,000 eligible units to comparison and 7,000 to treatment. The units are eligible application units. Mandatory controls are outside the experimental choice. Eligibility and assignment are defined before observing outcomes.

Figure data and text version
Assigned armUnits
Comparison7,000
Treatment7,000

Approval, pricing, limits, servicing, and exceptions can each produce different customer outcomes. A single aggregate metric cannot describe the entire credit journey.

The reference case starts with the stated population and a functioning evidence path. The owner is risk operations.

All amounts, rates, capacity limits, and outcomes in this case are synthetic. The three conditions are separate assumptions for comparison. A better result in the response condition is not measured proof that the proposed control causes that improvement. The figures expose the calculation and its limits; a real deployment needs its own evidence.

Read the result

The comparison arm has 840/7000 adverse outcomes (12.00%) and the treatment arm has 773/7000 (11.04%). The absolute difference is -0.96 percentage points, with an illustrative large-sample 95% interval from -2.01 to 0.10. Interpretation depends on assignment integrity, outcome maturity, independence, and the actual decision being evaluated.

Model inputs and calculated values

Inputs below are the case-specific values. Each figure states the condition-specific assumptions and units used in its calculation. Calculated values are rounded for display.

InputValue
population14,000
rate0.12
clusterSize5
Calculated valueResult
n14,000
nc7,000
nt7,000
yc840
yt773
effect-0.0096
lower-0.0201
upper0.001
clusterSize5
icc0.01
effectiveN13,461.5385
Figure 02 / 12

Retain events and nonevents in both arms

Retain events and nonevents in both arms — Outcome comparisons in context. Count; binary outcome over one fixed window. Exact values are in the figure data below.
Count; binary outcome over one fixed window

The defined adverse outcome occurs 840 times in comparison and 773 times in treatment. Each row includes every assigned unit under this complete-outcome illustration. Omitting nonevents or unresolved cases changes the denominator and the estimand.

Figure data and text version
Assigned armAdverse outcomeOther outcome
Comparison8406,160
Treatment7736,227
Figure 03 / 12

Absolute and relative effects answer different questions

Absolute and relative effects answer different questions — Outcome comparisons in context. Observed synthetic rates. Exact values are in the figure data below.
Observed synthetic rates

Treatment minus comparison is -0.96 percentage points. The relative change is -7.98% when the comparison rate is nonzero. A percentage-point difference and a percentage change must not share the same label.

Figure data and text version
MeasureValueUnit
Comparison rate12Percent
Treatment rate11.043Percent
Absolute difference-0.957Percentage points
Relative change-7.976Percent of comparison rate
Figure 04 / 12

Each arm’s rate has sampling uncertainty

Each arm’s rate has sampling uncertainty — Outcome comparisons in context. Percent; Wilson score intervals. Exact values are in the figure data below.
Percent; Wilson score intervals

The 95% Wilson intervals are [11.26, 12.78]% and [10.33, 11.80]%. These intervals use independent Bernoulli sampling assumptions. Shared customers or merchants can violate that independence and require an appropriate clustered analysis.

Figure data and text version
ArmObserved %Lower 95%Upper 95%
Comparison1211.2612.782
Treatment11.04310.3311.799
Figure 05 / 12

Uncertainty in the difference

Uncertainty in the difference — Outcome comparisons in context. Percentage points; treatment minus comparison. Exact values are in the figure data below.
Percentage points; treatment minus comparison

The normal-approximation 95% interval for treatment minus comparison is [-2.01, 0.10] percentage points. It is a teaching large-sample interval, not a replacement for the analysis appropriate to the design, low counts, sequential monitoring, or multiple comparisons.

Figure data and text version
Difference measurePercentage points
Lower bound-2.015
Point estimate-0.957
Upper bound0.101
Figure 06 / 12

A fixed outcome horizon prevents premature comparison

A fixed outcome horizon prevents premature comparison — Outcome comparisons in context. Observed adverse events. Exact values are in the figure data below.
Observed adverse events

Both arms use the same explicitly constructed maturation fractions. Earlier data can undercount the final outcome even when assignment is correct. Stopping at the first favorable partial result also changes the statistical interpretation unless the monitoring design accounts for it. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.

Figure data and text version
Observation ageComparison eventsTreatment events
Day 1168155
Day 3336309
Day 7546502
Day 14714657
Day 30840773
Figure 07 / 12

Assignment and received treatment differ

Assignment and received treatment differ — Outcome comparisons in context. Count; received-action counts are illustrative. Exact values are in the figure data below.
Count; received-action counts are illustrative

7000 units are assigned to treatment, while 6580 receive the modeled action. The primary assignment-based comparison retains all assigned units. Restricting to action recipients selects a post-assignment subgroup and can introduce bias.

Figure data and text version
PopulationUnits
Assigned comparison7,000
Comparison action observed6,860
Assigned treatment7,000
Treatment action observed6,580
Figure 08 / 12

Correlated units reduce independent information

Correlated units reduce independent information — Outcome comparisons in context. Effective units under stated approximation. Exact values are in the figure data below.
Effective units under stated approximation

The design-effect illustration is 1 + (m − 1)ρ with cluster size m = 5. At the case’s assumed ρ = 0.01, the design effect is 1.040 and n/design-effect is about 13,461.5. This approximation assumes a simple equal-cluster design; it is not a universal effective-sample-size formula. Horizontal positions are the labeled observations or scenarios; equal spacing does not imply equal numerical increments.

Figure data and text version
Intracluster correlationApproximate effective n
0.0014,000
0.0113,461.5
0.0312,727.3
0.0511,666.7
0.1010,000
0.207,777.8
Figure 09 / 12

Segment totals must reconcile to the whole

Segment totals must reconcile to the whole — Outcome comparisons in context. Count; constructed disjoint segments. Exact values are in the figure data below.
Count; constructed disjoint segments

The two synthetic segments partition each arm. Their outcome counts add back to the original totals. Segment effects can differ from the overall effect, but small cells, selection, and multiple comparisons must be considered before interpreting a subgroup result.

Figure data and text version
SegmentComparison nComparison eventsTreatment nTreatment events
Segment A3,5005463,500309
Segment B3,5002943,500464
Figure 10 / 12

The primary outcome is only one decision input

The primary outcome is only one decision input — Outcome comparisons in context. Experiment decision contract. Exact values are in the figure data below.
Experiment decision contract

The primary target is Defined adverse decision outcome. A release also considers customer friction, queue capacity, required controls, and the uncertainty of the result. A favorable loss metric cannot justify an action that violates a separate obligation.

Figure data and text version
DimensionRequired definition
Primary outcomeDefined adverse decision outcome
Customer impactCompletion, delay, complaints, and access
Operating costReview effort and capacity in the same window
Mandatory controlsOutside randomized relaxation
Decision ownerrisk operations
Figure 11 / 12

Specify the analysis before observing the result

Specify the analysis before observing the result — Outcome comparisons in context. Synthetic experimental workflow. Exact values are in the figure data below.
Synthetic experimental workflow

The protocol fixes assignment, exclusions, metrics, outcome maturity, and the comparison method. Changing these after seeing results can make an ordinary noise fluctuation look like an improvement. Amendments need a documented reason and a clear distinction between confirmatory and exploratory work.

Figure data and text version
StageTimingRecord
Define estimandBefore assignmentdefine fairness in the product context
Assign unitsExperiment startStable unit and arm
Collect outcomesFixed windowSame definition for both arms
AnalyzeAt planned maturityChosen method and uncertainty
DecideAfter reviewNet effect, guardrails, and limitations
Figure 12 / 12

What this example establishes and omits

What this example establishes and omits — Outcome comparisons in context. Interpretation boundaries. Exact values are in the figure data below.
Interpretation boundaries

The arithmetic shows how the stated counts become rates, differences, and intervals. It does not prove a causal effect in production. Causal interpretation requires a valid assignment process, appropriate handling of interference and missing outcomes, and analysis consistent with the design.

Figure data and text version
PropertyIn this exampleProduction requirement
AssignmentAssumed validVerify implementation and unit integrity
OutcomesComplete at final windowResolve delayed and missing outcomes
IndependenceSimple interval assumptionAccount for clustering or interference
CostNot included in rate differenceMeasure net economics separately

Connect the result to the system

Define the population, decision, comparison, uncertainty, and relevant governance review.

Check the population, evidence, permitted action, and actual effect together. A balanced calculation can still use the wrong population; a successful response can still leave an unknown financial outcome. The case’s numerical result applies only to its stated assumptions.

Sources and further reading

The chapter sources support the concepts and scope. They do not prescribe the synthetic model rates.

  1. Regulation B, 12 CFR 1002.6: evaluation of applications
  2. Regulation B, 12 CFR 1002.9: notifications
  3. NIST: AI Risk Management Framework
  4. SciPy: binomial proportion confidence intervals