Unit 07 · Chapter 5 · 15 min read

Model governance and AI-assisted risk work

Keep model use bounded, reviewable, and accountable.

The concept at a glance

An assistant can draft; authority stays explicit

Source documents cross an evidence boundary into an AI-assisted draft with linked records. A reviewer verifies facts and authority before a separately permissioned action. A version and evaluation record captures purpose, limits, tests, owner, release, and rollback. Untrusted text is data, not a source of new instructions or authority.

Enlarge to read every label and explore the connections

AI-generated text must stay tied to evidence and a bounded purpose. Source material cannot grant permission for consequential action.

  1. Treat retrieved material as untrusted evidence.

  2. Verify supported facts and limitations.

  3. Keep action permissions and change records explicit.

A model can be wrong in a sophisticated way. An AI assistant can write a fluent case note with an invented fact. Governance makes those failures visible before the output becomes an unchecked decision.

Use the current model-risk framework

Use the current model-risk framework — the flow
Use the current model-risk framework Use the current model-risk framework — the flow Follow the sequence. Assign proportionate review and ownership. Inventory Identify models and actual uses Assess Evaluate consequence complexity and limitations Govern Assign proportionate review and ownership
  1. InventoryIdentify models and actual uses
  2. AssessEvaluate consequence complexity and limitations
  3. GovernAssign proportionate review and ownership
Follow the sequence. Assign proportionate review and ownership. Chapter sources · Open image

The federal banking agencies issued revised model-risk guidance in April 2026 through SR 26-2, replacing SR 11-7 and SR 21-8. The letter states its expected relevance to Federal Reserve-regulated banking organizations over $30 billion in assets and emphasizes a risk-based approach. Applicability must be assessed for the actual institution.

For this textbook’s engineering design, maintain an inventory of models and their uses, owners, limitations, dependencies, and review evidence. A small rule-based model can still be important if it controls a large exposure. Governance effort should reflect the consequence and complexity of use rather than the prestige of the algorithm.

Inside the mechanism. Governance should reflect the actual model use, institution, materiality, and applicable framework. Inventory models and material decision components by function rather than relying on a team’s preferred label. Record owner, purpose, inputs, dependencies, limitations, validation status, and approved uses. Current guidance must be checked for scope and effective changes before treating an older supervisory document as the governing reference.

A concrete example. Governance begins with the actual model use, materiality, institution, and applicable supervisory framework. An inventory entry needs a decision owner and current evidence. The case identifies 1,428 eligible records from a source population of 1,700. The required workflow completes for 1,385, but 21 completed records miss the illustrative internal target. Another 43 remain incomplete. Communication evidence covers 1,371 generated notices. Scope, completion, timeliness, and delivery are four separate properties of the customer outcome.

When the assumption fails. A material decision model is omitted because the team calls it a rule or vendor score. Inventory decision components by function and apply the relevant review and change process. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.

Follow a worked case3 conditions · 36 figures

Governance begins with the actual model use, materiality, institution, and applicable supervisory framework. An inventory entry needs a decision owner and current evidence.

Use the current model-risk framework — the distinction
Use the current model-risk framework Use the current model-risk framework — the distinction These concepts answer different questions. Read each definition in the context of the section. Algorithm complexity Technical sophistication of the method Use risk Consequence of relying on its output
Algorithm complexity
  • Technical sophistication of the method
Use risk
  • Consequence of relying on its output
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Model inventory
Use the current model-risk framework Model inventory Fictional teaching record. Complexity alone is insufficient. Model inventory Illustrative data; not a real customer record or a prescribed policy. Model simple reserve estimator Low algorithm complexity Exposure large merchant portfolio High consequence Review proportionate to use Complexity alone is insufficient A simple model can support a consequential decision
Fictional educational excerpt / Not for execution

Model inventory

Illustrative data; not a real customer record or a prescribed policy.

  1. Modelsimple reserve estimator

    Low algorithm complexity

  2. Exposurelarge merchant portfolio

    High consequence

  3. Reviewproportionate to use

    Complexity alone is insufficient

A simple model can support a consequential decision

Fictional teaching record. Complexity alone is insufficient. Chapter sources · Open image
Use the current model-risk framework — control and failure modes
Use the current model-risk framework Use the current model-risk framework — control and failure modes A simple model can support a consequential decision. The branches show why alternative designs fail. Control design Assess model risk in its actual use. A simple model can support a consequential decision. Failure mode 1 Treat SR 11-7 as the current sole guidance. The 2026 letter replaced it. avoid Failure mode 2 Apply the banking letter identically to every startup. Institutional scope matters. avoid Failure mode 3 Govern only machine-learning models. Other quantitative tools can be important. avoid
Control design

Assess model risk in its actual use. A simple model can support a consequential decision.

Failure mode 1avoid
Treat SR 11-7 as the current sole guidance. The 2026 letter replaced it.
Failure mode 2avoid
Apply the banking letter identically to every startup. Institutional scope matters.
Failure mode 3avoid
Govern only machine-learning models. Other quantitative tools can be important.
A simple model can support a consequential decision. The branches show why alternative designs fail. Chapter sources · Open image

Separate development from effective challenge

Separate development from effective challenge — the flow
Separate development from effective challenge Separate development from effective challenge — the flow Follow the sequence. Track findings and permitted limitations. Develop Document assumptions and evidence Challenge Test the design and actual use Resolve Track findings and permitted limitations
  1. DevelopDocument assumptions and evidence
  2. ChallengeTest the design and actual use
  3. ResolveTrack findings and permitted limitations
Follow the sequence. Track findings and permitted limitations. Chapter sources · Open image

Developers understand the model deeply, but they also know the assumptions they intended. Independent challenge tests whether those assumptions hold and whether the use is appropriate. The structure should fit the organization and applicable expectations.

Review conceptual soundness, data, implementation, outcomes, limitations, and controls. Track findings through remediation and retest. A reviewer’s signature is not evidence that every concern was resolved. The model owner should know which limitations remain and what use is permitted while they remain.

Effective challenge examines whether the model is suitable for its stated use, including limitations that the development team may not have emphasized. It can inspect the target, data, assumptions, validation evidence, implementation, and downstream policy. Independence is useful because the people responsible for delivery may face pressure to interpret ambiguous results favorably. The form and depth of review should fit the model’s materiality and the applicable framework.

A limitation becomes operational when it changes the permitted use. If a model has weak evidence for a new product segment, the response may be restricted deployment, more review, additional validation, or another approved control. Recording the limitation in a document while using the model without restriction does not manage it.

Inside the mechanism. Effective challenge needs access to the relevant data, methods, assumptions, and outcomes, plus authority to require a response. The reviewer should test what could invalidate the development conclusion. Keep findings, management responses, unresolved limitations, and release conditions explicit. Organizational separation alone is not evidence of a substantive review, and a repeated developer summary is not independent validation.

A concrete example. A reviewer needs sufficient evidence, expertise, and authority to question the development conclusion. Repeating the developer summary does not test its assumptions. The case identifies 673 eligible records from a source population of 740. The required workflow completes for 653, but 10 completed records miss the illustrative internal target. Another 20 remain incomplete. Communication evidence covers 646 generated notices. Scope, completion, timeliness, and delivery are four separate properties of the customer outcome.

When the assumption fails. Validation signs off without access to the evaluation population or known limitations. Document challenges, test evidence, unresolved limitations, and release conditions. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.

Follow a worked case3 conditions · 36 figures

A reviewer needs sufficient evidence, expertise, and authority to question the development conclusion. Repeating the developer summary does not test its assumptions.

Separate development from effective challenge — the distinction
Separate development from effective challenge Separate development from effective challenge — the distinction These concepts answer different questions. Read each definition in the context of the section. Developer explanation Why the model was built this way Independent challenge Evidence that tests the explanation
Developer explanation
  • Why the model was built this way
Independent challenge
  • Evidence that tests the explanation
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Review finding
Separate development from effective challenge Review finding Fictional teaching record. Finding remains tracked. Review finding Illustrative data; not a real customer record or a prescribed policy. Issue weak new-merchant performance Identified limitation Use restriction known merchants only Bounded permitted use Retest scheduled with new evidence Finding remains tracked Approval alone does not remove limitations
Fictional educational excerpt / Not for execution

Review finding

Illustrative data; not a real customer record or a prescribed policy.

  1. Issueweak new-merchant performance

    Identified limitation

  2. Use restrictionknown merchants only

    Bounded permitted use

  3. Retestscheduled with new evidence

    Finding remains tracked

Approval alone does not remove limitations

Fictional teaching record. Finding remains tracked. Chapter sources · Open image
Separate development from effective challenge — control and failure modes
Separate development from effective challenge Separate development from effective challenge — control and failure modes Approval alone does not remove limitations. The branches show why alternative designs fail. Control design Connect review findings to use restrictions and retesting. Approval alone does not remove limitations. Failure mode 1 Treat a signature as full proof. The underlying findings matter. avoid Failure mode 2 Let unresolved issues disappear at launch. They remain part of the risk. avoid Failure mode 3 Review only code style. Conceptual and data weaknesses can dominate. avoid
Control design

Connect review findings to use restrictions and retesting. Approval alone does not remove limitations.

Failure mode 1avoid
Treat a signature as full proof. The underlying findings matter.
Failure mode 2avoid
Let unresolved issues disappear at launch. They remain part of the risk.
Failure mode 3avoid
Review only code style. Conceptual and data weaknesses can dominate.
Approval alone does not remove limitations. The branches show why alternative designs fail. Chapter sources · Open image

Bound generative AI to evidence-supported tasks

Bound generative AI to evidence-supported tasks — the flow
Bound generative AI to evidence-supported tasks Bound generative AI to evidence-supported tasks — the flow Follow the sequence. Verify claims before consequential action. Retrieve Provide approved relevant evidence Draft Generate a bounded supported output Review Verify claims before consequential action
  1. RetrieveProvide approved relevant evidence
  2. DraftGenerate a bounded supported output
  3. ReviewVerify claims before consequential action
Follow the sequence. Verify claims before consequential action. Chapter sources · Open image

Generative AI can help summarize records, draft narratives, or retrieve relevant policy passages. It can also invent facts, omit context, or follow instructions embedded in untrusted documents. Treat retrieved customer material as evidence to analyze, not instructions to the system.

Use constrained inputs, source references, and a reviewable output. Require factual claims to point to supporting records. Keep sensitive reporting information within approved access boundaries. A model-generated narrative should not automatically file a report, release funds, or change a customer restriction without the authorized decision process.

A generative system can help summarize retained evidence, but fluent wording is not evidence. Each material factual statement should be traceable to the permitted source material, and the workflow needs a response when support is absent or contradictory. Treat retrieved documents and customer submissions as data rather than trusted instructions. Separate the assistant’s proposal from the authorized decision, especially when the action changes funds, account access, or a regulated process. Evaluation should include fabricated facts, missing evidence, private-data exposure, and misleading instructions embedded in source material.

Inside the mechanism. Bound an AI assistant by task and authority. A case-summary tool should distinguish quoted source facts, supported synthesis, and unresolved questions, and link material claims to evidence. It should not acquire the authority to move money or make a legal determination merely because it can produce fluent text. Treat retrieved documents and customer messages as data that can contain misleading instructions. Preserve the original sources for review.

A concrete example. An assistant can summarize a case while still introducing unsupported facts or omitting decisive evidence. Its output needs links to the source record and a bounded role. The rule flags 409 of 6,200 reviewed summaries. Of those flags, 322 meet the synthetic target, giving 78.73% precision. It misses 81 target events. Under the stated cost assumptions, residual loss and operating friction total $9,361. The important result is the connection between the population, action, capacity, and outcome—not one isolated score.

When the assumption fails. Generated text presents an unverified suspicion as an established fact. Require evidence references, explicit uncertainty, and authorized human decisions for the defined task. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.

Follow a worked case3 conditions · 36 figures

An assistant can summarize a case while still introducing unsupported facts or omitting decisive evidence. Its output needs links to the source record and a bounded role.

Bound generative AI to evidence-supported tasks — the distinction
Bound generative AI to evidence-supported tasks Bound generative AI to evidence-supported tasks — the distinction These concepts answer different questions. Read each definition in the context of the section. Fluent narrative Text reads plausibly Grounded narrative Claims are supported by the supplied records
Fluent narrative
  • Text reads plausibly
Grounded narrative
  • Claims are supported by the supplied records
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
AI drafting record
Bound generative AI to evidence-supported tasks AI drafting record Fictional teaching record. Do not accept fluent invention. AI drafting record Illustrative data; not a real customer record or a prescribed policy. Claim customer admitted intent Generated statement Source none Unsupported claim Action remove and investigate Do not accept fluent invention Fluency does not establish truth
Fictional educational excerpt / Not for execution

AI drafting record

Illustrative data; not a real customer record or a prescribed policy.

  1. Claimcustomer admitted intent

    Generated statement

  2. Sourcenone

    Unsupported claim

  3. Actionremove and investigate

    Do not accept fluent invention

Fluency does not establish truth

Fictional teaching record. Do not accept fluent invention. Chapter sources · Open image
Bound generative AI to evidence-supported tasks — control and failure modes
Bound generative AI to evidence-supported tasks Bound generative AI to evidence-supported tasks — control and failure modes Fluency does not establish truth. The branches show why alternative designs fail. Control design Require claim-level evidence and authorized review. Fluency does not establish truth. Failure mode 1 Let customer documents instruct the assistant. They are untrusted task data. avoid Failure mode 2 Auto-release funds from a generated summary. Consequential authority needs controls. avoid Failure mode 3 Copy restricted reports into broad AI tools. Access and data rules still apply. avoid
Control design

Require claim-level evidence and authorized review. Fluency does not establish truth.

Failure mode 1avoid
Let customer documents instruct the assistant. They are untrusted task data.
Failure mode 2avoid
Auto-release funds from a generated summary. Consequential authority needs controls.
Failure mode 3avoid
Copy restricted reports into broad AI tools. Access and data rules still apply.
Fluency does not establish truth. The branches show why alternative designs fail. Chapter sources · Open image

Evaluate AI failure modes before use

Evaluate AI failure modes before use — the flow
Evaluate AI failure modes before use Evaluate AI failure modes before use — the flow Follow the sequence. Use thresholds and human review appropriate to consequence. Task Define the permitted output and harm model Evaluate Test representative and adversarial fixtures Gate Use thresholds and human review appropriate to consequence
  1. TaskDefine the permitted output and harm model
  2. EvaluateTest representative and adversarial fixtures
  3. GateUse thresholds and human review appropriate to consequence
Follow the sequence. Use thresholds and human review appropriate to consequence. Chapter sources · Open image

An AI evaluation should reflect the actual task: factual accuracy, omissions, unsupported claims, confidentiality, instruction handling, and consistency. Include difficult cases and deliberate misleading content in controlled fixtures.

Measure the harm of errors, not only an average quality score. A rare invented admission can be more consequential than several awkward sentences. Compare against a useful baseline and define when a human must intervene. Version prompts, retrieval rules, models, and evaluation sets so a vendor update does not silently invalidate the evidence.

Inside the mechanism. Evaluate failures that matter to the task: unsupported claims, omitted decisive facts, wrong entities, incorrect amounts, confidentiality errors, and susceptibility to instructions embedded in evidence. Use representative records and targeted difficult cases with a stated scoring rubric. Measure critical errors separately from style. A high average score can hide an unacceptable failure in a small but consequential class of cases.

A concrete example. Average fluency does not measure unsupported claims, instruction misuse, missing evidence, or sensitive-data exposure. Evaluation needs a labeled failure taxonomy tied to the use case. The rule flags 339 of 3,100 evaluation records. Of those flags, 298 meet the synthetic target, giving 87.91% precision. It misses 74 target events. Under the stated cost assumptions, residual loss and operating friction total $10,811. The important result is the connection between the population, action, capacity, and outcome—not one isolated score.

When the assumption fails. A polished summary passes review while missing the fact that changes the case disposition. Test critical omissions and unsupported claims across representative and adversarial records. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.

Follow a worked case3 conditions · 36 figures

Average fluency does not measure unsupported claims, instruction misuse, missing evidence, or sensitive-data exposure. Evaluation needs a labeled failure taxonomy tied to the use case.

Evaluate AI failure modes before use — the distinction
Evaluate AI failure modes before use Evaluate AI failure modes before use — the distinction These concepts answer different questions. Read each definition in the context of the section. Average fluency General writing quality Critical error rate Frequency of consequential unsupported or unsafe output
Average fluency
  • General writing quality
Critical error rate
  • Frequency of consequential unsupported or unsafe output
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
AI evaluation
Evaluate AI failure modes before use AI evaluation Fictional teaching record. Average style score is insufficient. AI evaluation Illustrative data; not a real customer record or a prescribed policy. Cases 200 synthetic records Controlled test set Invented facts 3 Critical errors Decision not ready for autonomous use Average style score is insufficient A good average can hide unacceptable rare failures
Fictional educational excerpt / Not for execution

AI evaluation

Illustrative data; not a real customer record or a prescribed policy.

  1. Cases200 synthetic records

    Controlled test set

  2. Invented facts3

    Critical errors

  3. Decisionnot ready for autonomous use

    Average style score is insufficient

A good average can hide unacceptable rare failures

Fictional teaching record. Average style score is insufficient. Chapter sources · Open image
Evaluate AI failure modes before use — control and failure modes
Evaluate AI failure modes before use Evaluate AI failure modes before use — control and failure modes A good average can hide unacceptable rare failures. The branches show why alternative designs fail. Control design Evaluate consequential errors separately. A good average can hide unacceptable rare failures. Failure mode 1 Score only writing style. Truth and confidentiality remain untested. avoid Failure mode 2 Reuse old evaluations after major changes. The system behavior may differ. avoid Failure mode 3 Use real confidential cases without approval. Test data also has access constraints. avoid
Control design

Evaluate consequential errors separately. A good average can hide unacceptable rare failures.

Failure mode 1avoid
Score only writing style. Truth and confidentiality remain untested.
Failure mode 2avoid
Reuse old evaluations after major changes. The system behavior may differ.
Failure mode 3avoid
Use real confidential cases without approval. Test data also has access constraints.
A good average can hide unacceptable rare failures. The branches show why alternative designs fail. Chapter sources · Open image

Maintain a complete change record

Maintain a complete change record — the flow
Maintain a complete change record Maintain a complete change record — the flow Follow the sequence. Monitor rollout and preserve rollback. Change Identify all behavior-affecting components Evidence Attach evaluation and approval Operation Monitor rollout and preserve rollback
  1. ChangeIdentify all behavior-affecting components
  2. EvidenceAttach evaluation and approval
  3. OperationMonitor rollout and preserve rollback
Follow the sequence. Monitor rollout and preserve rollback. Chapter sources · Open image

Model behavior can change through data, features, code, thresholds, prompts, retrieval sources, or vendor versions. Keep a release record that connects the change to evaluation, approval, rollout, monitoring, and rollback.

Define what counts as a material change and who decides. A prompt edit that adds a new tool can be more consequential than a model patch. Preserve the affected population and outputs for review under the retention policy. Governance is complete when the organization can explain what changed, why it was allowed, and how the result was checked.

Inside the mechanism. A complete change record includes data, transformations, model or prompt version, tools, retrieval sources, thresholds, and downstream policy. Hosted vendor changes can alter behavior without a local code change. Record available version controls and evaluate material changes within the approved process. Rollback needs a known configuration and a way to identify decisions made during the affected interval.

A concrete example. A model change can include data, features, prompts, vendor versions, thresholds, and downstream policy. The release record should identify the full decision configuration. The case identifies 874 eligible records from a source population of 930. The required workflow completes for 848, but 13 completed records miss the illustrative internal target. Another 26 remain incomplete. Communication evidence covers 840 generated notices. Scope, completion, timeliness, and delivery are four separate properties of the customer outcome.

When the assumption fails. A vendor changes a hosted model without a corresponding internal review record. Pin available versions, assess material changes, and preserve approval and rollback evidence. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.

Follow a worked case3 conditions · 36 figures

A model change can include data, features, prompts, vendor versions, thresholds, and downstream policy. The release record should identify the full decision configuration.

Maintain a complete change record — the distinction
Maintain a complete change record Maintain a complete change record — the distinction These concepts answer different questions. Read each definition in the context of the section. Code version One implementation component System version Model data policy prompts and dependencies together
Code version
  • One implementation component
System version
  • Model data policy prompts and dependencies together
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
AI release record
Maintain a complete change record AI release record Fictional teaching record. Bounded capability. AI release record Illustrative data; not a real customer record or a prescribed policy. Prompt v5 Changed instructions Model provider version recorded Dependency identity Tools read-only evidence search Bounded capability Behavior can change outside application code
Fictional educational excerpt / Not for execution

AI release record

Illustrative data; not a real customer record or a prescribed policy.

  1. Promptv5

    Changed instructions

  2. Modelprovider version recorded

    Dependency identity

  3. Toolsread-only evidence search

    Bounded capability

Behavior can change outside application code

Fictional teaching record. Bounded capability. Chapter sources · Open image
Maintain a complete change record — control and failure modes
Maintain a complete change record Maintain a complete change record — control and failure modes Behavior can change outside application code. The branches show why alternative designs fail. Control design Version the whole decision system. Behavior can change outside application code. Failure mode 1 Review only model weights. Prompts and data can alter outcomes. avoid Failure mode 2 Treat tool access as a minor wording edit. Capabilities change the risk. avoid Failure mode 3 Omit rollback and monitoring. The release lacks an operational response. avoid
Control design

Version the whole decision system. Behavior can change outside application code.

Failure mode 1avoid
Review only model weights. Prompts and data can alter outcomes.
Failure mode 2avoid
Treat tool access as a minor wording edit. Capabilities change the risk.
Failure mode 3avoid
Omit rollback and monitoring. The release lacks an operational response.
Behavior can change outside application code. The branches show why alternative designs fail. Chapter sources · Open image

Chapter connections

This chapter builds on Experiments, causal effects, and risk tradeoffs. Use the glossary for terminology and risk mathematics for formulas and worked calculations.

Sources

Reviewed 2026-09-17
  1. Federal Reserve SR 26-2: revised model-risk guidance (2026)
  2. NIST: AI Risk Management Framework
  3. NIST: Generative AI Profile, AI 600-1