Unit 07 · Chapter 4 · 15 min read

Experiments, causal effects, and risk tradeoffs

Measure whether a control improves outcomes rather than only changing them.

The concept at a glance

Compare cohorts at equal maturity

Eligible accounts are randomly assigned by a stable account key into control and treatment. Each arm has an enrollment point, the same outcome window, and a mature assessment. An interim reporting line lies inside both windows, showing why early observations are not final loss outcomes.

Enlarge to read every label and explore the connections

Stable random assignment helps identify an intervention’s effect. Compare risk outcomes after equivalent follow-up.

  1. Define the intervention and assignment unit.

  2. Keep assignment stable through repeat activity.

  3. Compare mature outcomes and show uncertainty.

The new rule launches on Monday. Loss falls on Tuesday. On the same day, the largest risky merchant leaves the platform. The chart shows a change. It does not yet show what caused it.

Define the intervention and estimand

Define the intervention and estimand — the flow
Define the intervention and estimand Define the intervention and estimand — the flow Follow the sequence. Specify the effect and observation window. Intervention Name the changed control Population Define eligible participants Outcome Specify the effect and observation window
  1. InterventionName the changed control
  2. PopulationDefine eligible participants
  3. OutcomeSpecify the effect and observation window
Follow the sequence. Specify the effect and observation window. Chapter sources · Open image

An experiment needs a specific intervention and a quantity it aims to estimate. For example, compare two permitted authentication flows on completion, confirmed loss, and customer support cost. The estimand states whose outcomes, over what period, under which assignment.

Keep mandatory legal controls outside an experiment that would bypass them. Use an approved test population and guardrails. The goal is to learn within permitted activity. A vague objective such as improve risk cannot specify a sample, a stopping rule, or a useful result.

An intervention is the change applied to the system. An estimand is the particular effect the analysis intends to measure. “Does the new rule help” is incomplete until help has a unit, population, time horizon, and outcome. The effect on all eligible accounts can differ from the effect on accounts that actually received a challenge, because challenge receipt itself depends on the policy and customer behavior.

Write the analysis contract before reading the result. Define assignment, exclusions, outcome maturity, primary measures, and material customer or operational limits. Otherwise, a team can unintentionally select the time window or subgroup that makes a weak result look convincing. A clear contract also makes an inconclusive result useful by showing which uncertainty remains.

Inside the mechanism. An estimand states the population, intervention, comparison, outcome, and period whose effect is being estimated. Assignment to a new policy and receipt of a particular action are different variables. An intention-to-treat analysis follows assignment and can preserve the benefit of randomization; an analysis restricted to those who received an action can reintroduce selection. Define the primary comparison before inspecting outcomes.

A concrete example. The experiment must state whose outcome changes, over which period, and under which assignment. Approval rate and net mature loss answer different questions. The comparison arm has 312/6000 adverse outcomes (5.20%) and the treatment arm has 287/6000 (4.78%). The absolute difference is -0.42 percentage points, with an illustrative large-sample 95% interval from -1.20 to 0.36. Interpretation depends on assignment integrity, outcome maturity, independence, and the actual decision being evaluated.

When the assumption fails. The analysis changes its primary outcome after seeing an attractive early result. Freeze the eligible population, assignment, outcome, and analysis window before release. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.

Follow a worked case3 conditions · 36 figures

The experiment must state whose outcome changes, over which period, and under which assignment. Approval rate and net mature loss answer different questions.

Define the intervention and estimand — the distinction
Define the intervention and estimand Define the intervention and estimand — the distinction These concepts answer different questions. Read each definition in the context of the section. Observed difference Groups have different outcomes Causal effect Difference attributable to the intervention under the design
Observed difference
  • Groups have different outcomes
Causal effect
  • Difference attributable to the intervention under the design
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Experiment contract
Define the intervention and estimand Experiment contract Fictional teaching record. Customer and financial protection. Experiment contract Illustrative data; not a real customer record or a prescribed policy. Change clearer step-up prompt Permitted intervention Primary outcome completed legitimate purchases Defined measure Guardrail confirmed loss and complaints Customer and financial protection The experiment needs a clear success definition
Fictional educational excerpt / Not for execution

Experiment contract

Illustrative data; not a real customer record or a prescribed policy.

  1. Changeclearer step-up prompt

    Permitted intervention

  2. Primary outcomecompleted legitimate purchases

    Defined measure

  3. Guardrailconfirmed loss and complaints

    Customer and financial protection

The experiment needs a clear success definition

Fictional teaching record. Customer and financial protection. Chapter sources · Open image
Define the intervention and estimand — control and failure modes
Define the intervention and estimand Define the intervention and estimand — control and failure modes The experiment needs a clear success definition. The branches show why alternative designs fail. Control design Specify the effect and guardrails before launch. The experiment needs a clear success definition. Failure mode 1 Disable mandatory controls to learn faster. That exceeds permitted experimentation. avoid Failure mode 2 Change many unrelated things at once. Attribution becomes difficult. avoid Failure mode 3 Choose the metric after seeing results. That increases selection bias. avoid
Control design

Specify the effect and guardrails before launch. The experiment needs a clear success definition.

Failure mode 1avoid
Disable mandatory controls to learn faster. That exceeds permitted experimentation.
Failure mode 2avoid
Change many unrelated things at once. Attribution becomes difficult.
Failure mode 3avoid
Choose the metric after seeing results. That increases selection bias.
The experiment needs a clear success definition. The branches show why alternative designs fail. Chapter sources · Open image

Randomize at the right level

Randomize at the right level — the flow
Randomize at the right level Randomize at the right level — the flow Follow the sequence. Verify balance and actual exposure. Unit Choose account merchant or transaction assignment Assign Use a stable randomized rule Check Verify balance and actual exposure
  1. UnitChoose account merchant or transaction assignment
  2. AssignUse a stable randomized rule
  3. CheckVerify balance and actual exposure
Follow the sequence. Verify balance and actual exposure. Chapter sources · Open image

Randomization helps balance confounders on average, but the assignment unit matters. If one customer sees both treatments across retries, behavior can spill across groups. Merchant-level interventions may require merchant-level assignment.

Use a stable assignment key and preserve it through the relevant lifecycle. Check sample-ratio mismatch and implementation errors. Randomization does not guarantee balance in every small sample, nor does it solve missing outcomes. Document exclusions and whether they were decided before or after treatment, because post-treatment exclusions can bias the estimate.

Randomization must account for interference. If one customer can make many payments, assigning each payment independently can expose that customer to both policies and change later behavior. Shared merchants, devices, or counterparties can also connect observations. Choose an assignment unit that fits the intervention and use an analysis that respects the resulting dependence. Larger assignment groups may reduce the number of independent observations, which changes uncertainty even when the raw transaction count is large.

Inside the mechanism. Randomization should respect interference and shared state. Transactions within one account can affect a shared limit; related accounts can share devices or reviewers. Cluster assignment may reduce contamination but changes precision and sample-size needs. A simple design-effect illustration uses 1 plus cluster size minus one times within-cluster correlation. That approximation does not replace an analysis suited to the actual assignment and dependence structure.

A concrete example. Related transactions can influence each other through shared limits, devices, and review decisions. Row-level assignment may contaminate both treatment groups. The comparison arm has 346/7200 adverse outcomes (4.81%) and the treatment arm has 318/7200 (4.42%). The absolute difference is -0.39 percentage points, with an illustrative large-sample 95% interval from -1.07 to 0.30. Interpretation depends on assignment integrity, outcome maturity, independence, and the actual decision being evaluated.

When the assumption fails. One account receives both policies and its shared limit changes the control group. Assign at a defensible entity boundary and account for dependence in uncertainty estimates. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.

Follow a worked case3 conditions · 36 figures

Related transactions can influence each other through shared limits, devices, and review decisions. Row-level assignment may contaminate both treatment groups.

Randomize at the right level — the distinction
Randomize at the right level Randomize at the right level — the distinction These concepts answer different questions. Read each definition in the context of the section. Assignment Intended treatment group Exposure Treatment the participant actually received
Assignment
  • Intended treatment group
Exposure
  • Treatment the participant actually received
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Assignment defect
Randomize at the right level Assignment defect Fictional teaching record. Avoid cross-treatment contamination. Assignment defect Illustrative data; not a real customer record or a prescribed policy. Customer three retries Same decision journey Groups A then B then A Inconsistent exposure Fix stable journey assignment Avoid cross-treatment contamination Retries and spillovers can contaminate the comparison
Fictional educational excerpt / Not for execution

Assignment defect

Illustrative data; not a real customer record or a prescribed policy.

  1. Customerthree retries

    Same decision journey

  2. GroupsA then B then A

    Inconsistent exposure

  3. Fixstable journey assignment

    Avoid cross-treatment contamination

Retries and spillovers can contaminate the comparison

Fictional teaching record. Avoid cross-treatment contamination. Chapter sources · Open image
Randomize at the right level — control and failure modes
Randomize at the right level Randomize at the right level — control and failure modes Retries and spillovers can contaminate the comparison. The branches show why alternative designs fail. Control design Keep assignment stable at the relevant unit. Retries and spillovers can contaminate the comparison. Failure mode 1 Randomize every page render. One user can receive mixed treatments. avoid Failure mode 2 Ignore sample-ratio mismatch. It can reveal implementation defects. avoid Failure mode 3 Exclude inconvenient outcomes after treatment. That can bias the result. avoid
Control design

Keep assignment stable at the relevant unit. Retries and spillovers can contaminate the comparison.

Failure mode 1avoid
Randomize every page render. One user can receive mixed treatments.
Failure mode 2avoid
Ignore sample-ratio mismatch. It can reveal implementation defects.
Failure mode 3avoid
Exclude inconvenient outcomes after treatment. That can bias the result.
Retries and spillovers can contaminate the comparison. The branches show why alternative designs fail. Chapter sources · Open image

Wait for mature outcomes

Wait for mature outcomes — the flow
Wait for mature outcomes Wait for mature outcomes — the flow Follow the sequence. Use the planned analysis and harm guardrails. Immediate Observe completion and latency Delayed Wait for comparable loss maturity Decide Use the planned analysis and harm guardrails
  1. ImmediateObserve completion and latency
  2. DelayedWait for comparable loss maturity
  3. DecideUse the planned analysis and harm guardrails
Follow the sequence. Use the planned analysis and harm guardrails. Chapter sources · Open image

Risk outcomes often arrive after the customer interaction. Conversion can be measured quickly; disputes, defaults, and recoveries take longer. A test can show a short-term benefit before its loss cost becomes visible.

Define leading and final outcomes with separate reporting. Compare groups at equal maturity and keep uncertainty visible. Avoid stopping as soon as one metric looks favorable. Sequential monitoring requires an appropriate analysis plan if repeated looks influence the decision. The operational team also needs stop conditions for clear harm, independent of the final statistical conclusion.

Inside the mechanism. Specify the outcome window and the required follow-up before comparing arms. An equally sized but newer treatment cohort can have fewer observed losses simply because reports are delayed. Track enrollment, exposure, censoring, and label maturity separately. Interim operational signals can support safety monitoring, but they should not be mislabeled as mature financial outcomes.

A concrete example. Returns, disputes, and credit losses appear on different schedules. Equal calendar dates do not imply equal time at risk or label maturity. The comparison arm has 339/5300 adverse outcomes (6.40%) and the treatment arm has 312/5300 (5.89%). The absolute difference is -0.51 percentage points, with an illustrative large-sample 95% interval from -1.42 to 0.40. Interpretation depends on assignment integrity, outcome maturity, independence, and the actual decision being evaluated.

When the assumption fails. The treatment cohort is newer and appears safer because fewer losses have arrived. Compare equally mature cohorts and report incomplete follow-up explicitly. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.

Follow a worked case3 conditions · 36 figures

Returns, disputes, and credit losses appear on different schedules. Equal calendar dates do not imply equal time at risk or label maturity.

Wait for mature outcomes — the distinction
Wait for mature outcomes Wait for mature outcomes — the distinction These concepts answer different questions. Read each definition in the context of the section. Leading metric Early signal of behavior Mature outcome Result after the relevant observation window
Leading metric
  • Early signal of behavior
Mature outcome
  • Result after the relevant observation window
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Experiment timeline
Wait for mature outcomes Experiment timeline Fictional teaching record. Not a final loss conclusion. Experiment timeline Illustrative data; not a real customer record or a prescribed policy. Day 1 conversion improves Early result Day 30 disputes still developing Incomplete cost Decision provisional evidence Not a final loss conclusion Fast conversion data cannot establish final net benefit
Fictional educational excerpt / Not for execution

Experiment timeline

Illustrative data; not a real customer record or a prescribed policy.

  1. Day 1conversion improves

    Early result

  2. Day 30disputes still developing

    Incomplete cost

  3. Decisionprovisional evidence

    Not a final loss conclusion

Fast conversion data cannot establish final net benefit

Fictional teaching record. Not a final loss conclusion. Chapter sources · Open image
Wait for mature outcomes — control and failure modes
Wait for mature outcomes Wait for mature outcomes — control and failure modes Fast conversion data cannot establish final net benefit. The branches show why alternative designs fail. Control design Separate early signals from mature outcomes. Fast conversion data cannot establish final net benefit. Failure mode 1 Stop at the first favorable chart. Repeated peeking can distort inference. avoid Failure mode 2 Compare unequal follow-up periods. Groups have different opportunity for loss. avoid Failure mode 3 Ignore clear harm while waiting. Operational guardrails still apply. avoid
Control design

Separate early signals from mature outcomes. Fast conversion data cannot establish final net benefit.

Failure mode 1avoid
Stop at the first favorable chart. Repeated peeking can distort inference.
Failure mode 2avoid
Compare unequal follow-up periods. Groups have different opportunity for loss.
Failure mode 3avoid
Ignore clear harm while waiting. Operational guardrails still apply.
Fast conversion data cannot establish final net benefit. The branches show why alternative designs fail. Chapter sources · Open image

Account for selection and counterfactuals

Account for selection and counterfactuals — the flow
Account for selection and counterfactuals Account for selection and counterfactuals — the flow Follow the sequence. Use a justified method and state uncertainty. Policy Determines which outcomes become observable Counterfactual Outcome under an alternative action Estimate Use a justified method and state uncertainty
  1. PolicyDetermines which outcomes become observable
  2. CounterfactualOutcome under an alternative action
  3. EstimateUse a justified method and state uncertainty
Follow the sequence. Use a justified method and state uncertainty. Chapter sources · Open image

A declined payment does not reveal whether it would have become fraud if approved. A reviewed customer may behave differently because of the review. Observed labels are shaped by the policy that produced them.

Use randomized evidence where appropriate, carefully designed observational analysis where necessary, and explicit uncertainty where the counterfactual is unavailable. Do not label every decline a prevented loss. In a teaching example, 1,000 declines with a model estimate of 5 percent fraud are not 1,000 confirmed fraud events. Even the 50-event estimate depends on model validity and population assumptions.

Inside the mechanism. Observed outcomes depend on prior policy. A declined payment does not reveal the loss that would have occurred if approved. Comparing only approvals can change the population in each arm and distort the intended effect. State the missing counterfactual and the assumptions of any correction method. Historical replay can estimate action differences on recorded inputs; it cannot manufacture unobserved customer outcomes.

A concrete example. A declined transaction has no observed approved outcome. Comparing only approved transactions can confuse changes in selection with changes in underlying risk. The comparison arm has 640/7900 adverse outcomes (8.10%) and the treatment arm has 589/7900 (7.46%). The absolute difference is -0.65 percentage points, with an illustrative large-sample 95% interval from -1.48 to 0.19. Interpretation depends on assignment integrity, outcome maturity, independence, and the actual decision being evaluated.

When the assumption fails. The analysis drops all declines and treats the remaining cohorts as exchangeable. Define the counterfactual and use an evaluation design whose assumptions support the claim. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.

Follow a worked case3 conditions · 36 figures

A declined transaction has no observed approved outcome. Comparing only approved transactions can confuse changes in selection with changes in underlying risk.

Account for selection and counterfactuals — the distinction
Account for selection and counterfactuals Account for selection and counterfactuals — the distinction These concepts answer different questions. Read each definition in the context of the section. Declined amount Value of blocked attempts Prevented loss Estimated or evidenced loss avoided by the control
Declined amount
  • Value of blocked attempts
Prevented loss
  • Estimated or evidenced loss avoided by the control
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Counterfactual example
Account for selection and counterfactuals Counterfactual example Fictional teaching record. Not equal to all declines. Counterfactual example Illustrative data; not a real customer record or a prescribed policy. Declines 1000 Observed actions Estimated fraud 5 percent Model assumption Confirmed prevented events unknown Not equal to all declines The rejected outcome is often unobserved
Fictional educational excerpt / Not for execution

Counterfactual example

Illustrative data; not a real customer record or a prescribed policy.

  1. Declines1000

    Observed actions

  2. Estimated fraud5 percent

    Model assumption

  3. Confirmed prevented eventsunknown

    Not equal to all declines

The rejected outcome is often unobserved

Fictional teaching record. Not equal to all declines. Chapter sources · Open image
Account for selection and counterfactuals — control and failure modes
Account for selection and counterfactuals Account for selection and counterfactuals — control and failure modes The rejected outcome is often unobserved. The branches show why alternative designs fail. Control design Distinguish observed actions from estimated avoided loss. The rejected outcome is often unobserved. Failure mode 1 Count every decline as fraud stopped. That overstates effectiveness. avoid Failure mode 2 Treat model estimates as confirmed labels. They remain estimates. avoid Failure mode 3 Ignore policy-driven selection. It shapes the available training data. avoid
Control design

Distinguish observed actions from estimated avoided loss. The rejected outcome is often unobserved.

Failure mode 1avoid
Count every decline as fraud stopped. That overstates effectiveness.
Failure mode 2avoid
Treat model estimates as confirmed labels. They remain estimates.
Failure mode 3avoid
Ignore policy-driven selection. It shapes the available training data.
The rejected outcome is often unobserved. The branches show why alternative designs fail. Chapter sources · Open image

Report net effect and uncertainty

Report net effect and uncertainty — the flow
Report net effect and uncertainty Report net effect and uncertainty — the flow Follow the sequence. Record the product action and rationale. Report Show design outcomes and uncertainty Interpret Connect effects to costs and constraints Decide Record the product action and rationale
  1. ReportShow design outcomes and uncertainty
  2. InterpretConnect effects to costs and constraints
  3. DecideRecord the product action and rationale
Follow the sequence. Record the product action and rationale. Chapter sources · Open image

A useful experiment report includes the intervention, population, assignment, sample, duration, outcomes, uncertainty, guardrails, and limitations. Show both the main effect and important segments without turning every noisy subgroup into a firm conclusion.

Translate the result into the product decision. A small approval gain can be valuable if loss and support costs remain acceptable; a large gain may be unacceptable if it creates harm or violates constraints. Keep the decision owner and review date explicit. The report should allow a later reader to understand why the change was adopted or rejected.

Inside the mechanism. Report absolute changes with denominators and uncertainty, then add separately supported financial and customer effects. A large relative improvement can describe a small absolute change when the baseline is rare. The worked cases show binomial intervals and a simple difference interval under stated assumptions. Clustering, repeated looks, multiple outcomes, and noncompliance can require different methods. A confidence interval is not a guarantee of future performance.

A concrete example. A favorable point estimate can coexist with a wide interval and higher operating cost. Decisions need absolute impact, uncertainty, and material customer effects. The comparison arm has 351/9000 adverse outcomes (3.90%) and the treatment arm has 323/9000 (3.59%). The absolute difference is -0.31 percentage points, with an illustrative large-sample 95% interval from -0.87 to 0.24. Interpretation depends on assignment integrity, outcome maturity, independence, and the actual decision being evaluated.

When the assumption fails. A relative percentage improvement hides a small absolute difference and a large review bill. Report denominators, absolute differences, uncertainty, and separately measured cost components. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.

Follow a worked case3 conditions · 36 figures

A favorable point estimate can coexist with a wide interval and higher operating cost. Decisions need absolute impact, uncertainty, and material customer effects.

Report net effect and uncertainty — the distinction
Report net effect and uncertainty Report net effect and uncertainty — the distinction These concepts answer different questions. Read each definition in the context of the section. Statistical signal Evidence of a measured difference Business decision Choice considering magnitude cost harm and constraints
Statistical signal
  • Evidence of a measured difference
Business decision
  • Choice considering magnitude cost harm and constraints
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Net-effect example
Report net effect and uncertainty Net-effect example Fictional teaching record. Before uncertainty and omitted costs. Net-effect example Illustrative data; not a real customer record or a prescribed policy. Contribution gain 12000 USD Illustrative benefit Added loss and support 9000 USD Defined costs Net 3000 USD Before uncertainty and omitted costs A significant result is not automatically a good decision
Fictional educational excerpt / Not for execution

Net-effect example

Illustrative data; not a real customer record or a prescribed policy.

  1. Contribution gain12000 USD

    Illustrative benefit

  2. Added loss and support9000 USD

    Defined costs

  3. Net3000 USD

    Before uncertainty and omitted costs

A significant result is not automatically a good decision

Fictional teaching record. Before uncertainty and omitted costs. Chapter sources · Open image
Report net effect and uncertainty — control and failure modes
Report net effect and uncertainty Report net effect and uncertainty — control and failure modes A significant result is not automatically a good decision. The branches show why alternative designs fail. Control design Report magnitude costs and uncertainty together. A significant result is not automatically a good decision. Failure mode 1 Publish only the winning metric. Tradeoffs disappear. avoid Failure mode 2 Treat every subgroup fluctuation as real. Small samples can be noisy. avoid Failure mode 3 Omit the tested population. Readers may generalize beyond the evidence. avoid
Control design

Report magnitude costs and uncertainty together. A significant result is not automatically a good decision.

Failure mode 1avoid
Publish only the winning metric. Tradeoffs disappear.
Failure mode 2avoid
Treat every subgroup fluctuation as real. Small samples can be noisy.
Failure mode 3avoid
Omit the tested population. Readers may generalize beyond the evidence.
A significant result is not automatically a good decision. The branches show why alternative designs fail. Chapter sources · Open image

Chapter connections

This chapter builds on Risk models, calibration, and delayed outcomes. Continue with Model governance and AI-assisted risk work to follow the next part of the system. Use the glossary for terminology and risk mathematics for formulas and worked calculations.

Sources

Reviewed 2026-09-17
  1. NIST: AI Risk Management Framework
  2. scikit-learn: model evaluation metrics
  3. SciPy: binomial proportion confidence intervals