Wealth firms increasingly use AI to do the heavy lifting in Know-Your-Product (KYP) work: reading offering documents, extracting fees and terms, detecting changes in filings, scoring products and drafting summaries. Each of those tools produces outputs the firm then relies on. If an output is wrong and nobody notices, the firm's understanding of a product is wrong too.
That is model risk, and banks have managed it for years. What is new is that the tools have changed faster than the guidance. Canada's banking regulator has written AI into its model risk guideline. US banking regulators rewrote theirs in April 2026 and left generative and agentic AI out of scope. Securities regulators in both countries expect AI tools to be fit for purpose, tested and overseen, without prescribing how.
This article sets out a practical model risk framework for the AI tools used in product due diligence and monitoring. It covers what the standards say, why AI needs different treatment, how to inventory and rate AI tools, how to validate and monitor them, what to require from vendors, and responsibilities and an example written process.
It reflects publications available as of the date above.
Model risk is the chance that a firm makes a bad decision because a model's output was wrong, or was used for something it wasn't built for.
| Source | Applies To | Position on AI |
|---|---|---|
| OSFI Guideline E-23, effective May 1, 2027[1] | Federally regulated financial institutions in Canada, including banks and insurers that own dealers and advisers | AI included. A model is defined to include "AI/ML methods"; monitoring standards should cover "model drift, autonomous decision making, autonomous re-parametrization." |
| Revised interagency guidance on model risk management, April 17, 2026[2][3] | US banking organizations, most relevant above $30 billion in assets; replaces SR 11-7 | Generative and agentic AI excluded: they "are novel and rapidly evolving. As such, they are not within the scope of this guidance."[3] |
| CSA Staff Notice 11-348, December 2024[4] | Canadian registrants, including dealers and advisers | Firms should be "satisfied that the AI system is fit for purpose and that robust testing prior to deployment has taken place," and should test "before and after its adoption." |
| FINRA 2026 Annual Regulatory Oversight Report[5] | US broker-dealers | Describes firms establishing "a supervision, governance or model risk management framework" for generative AI, with robust testing and ongoing monitoring of outputs. |
The gap. For a US firm, the banking agencies' guidance no longer covers generative AI at all, and it states that non-compliance "will not result in supervisory criticism."[3] Securities regulators still expect governance and testing. The result is that firms need to build their own framework for AI tools rather than wait for one. The NIST AI Risk Management Framework, a voluntary US standard organized around four functions (Govern, Map, Measure and Manage), is a common starting point.[6]
Canada is more specific. For firms inside a bank or insurance group, OSFI's definition of model risk is broad: "risk of adverse financial impact arising from the design, development, deployment, and/or use of a model."[1] An AI tool that screens or monitors products is likely to fall within the group's model inventory from May 2027.
| Characteristic | Traditional Model | AI Tool | What It Means for Controls |
|---|---|---|---|
| Consistency | Same input, same output | The same question can produce different answers | Test on repeated runs, not single examples |
| Failure mode | Errors are usually visible (a wrong number, a crash) | Errors can be fluent and confident, including invented facts | Require sources for every output so errors can be caught |
| Transparency | Logic can be inspected | Reasoning is hard to inspect directly | Explain outputs through their sources rather than the model's internals |
| Change | Changes when the firm changes it | A vendor can update the underlying model without the firm doing anything | Contractual change notice; revalidate on version change |
| Scope | Does one defined task | Will attempt tasks it wasn't built for | Define approved uses; block or flag others |
| Inputs | Structured data | Documents, prompts and instructions, any of which can change results | Treat prompts and instructions as part of the model, under change control |
Definitions differ between standards, so the practical test is simpler: does the firm rely on the tool's output in its assessment of a product? If it does, the tool belongs in the inventory and under the framework, whatever it is called.
| Tool | In Scope? | Why |
|---|---|---|
| AI extraction of fees, terms and risk factors from offering documents | Yes | Outputs populate the product file |
| AI change detection across filings and data | Yes | Decides what the firm is alerted to, and what it isn't |
| Product scoring or screening model | Yes | Influences which products are approved or reviewed |
| AI drafting of product summaries for advisors | Yes | Shapes advisors' understanding of the product |
| AI research assistant answering questions about products | Yes, at lower risk if outputs are treated only as inputs | Can introduce errors into analysis |
| Fixed rule that compares a field against a threshold | Usually no | Deterministic logic; covered by ordinary system testing |
| General productivity tools (email drafting, meeting notes) | No, unless used for product analysis | Not relied on for product assessment |
A model risk framework for AI tools doesn't need to be large. It needs an inventory, a way to rate risk, a lifecycle with independent review, and monitoring that catches problems early.
The inventory records, for each AI tool: its purpose and approved uses, owner, developer or vendor, underlying model and version, inputs, outputs, where the outputs are used, risk tier, validation date and status, and next review date. OSFI expects the inventory to be comprehensive for models with non-negligible risk and kept current as a system of record.[1]
The risk rating sets how much control each tool gets. OSFI lists factors such as business use, "model complexity or autonomy, data reliability, customer impacts, or regulatory risk."[1] For KYP tools, an illustrative three-tier approach:
| Tier | When | Examples | Controls |
|---|---|---|---|
| High | Outputs feed approvals or status changes with limited human review, or the tool acts autonomously | Product scoring used in approval; agents that update records | Full independent validation before use; quarterly performance review; annual revalidation |
| Medium | Outputs populate product files or reach advisors after human review | Document extraction; change detection; summary drafting | Independent validation before use; monthly sample testing; revalidation on material change |
| Low | Outputs are inputs to a person's own analysis, fully reviewed | Research assistant with cited answers | Pre-use testing; periodic spot checks; usage rules |
Autonomy is the factor that moves a tool up fastest. The same extraction tool is medium risk when a person reviews its output and high risk when its output flows straight into the product register.
Independence matters. OSFI expects the review process to be "independent from model development" and to validate that models are "properly specified, working as intended, and fit-for-purpose."[1] For a firm without a model risk team, independence can mean the people who test the tool are not the people who built or bought it.
Prompts are part of the model. For AI tools, the instructions given to the model can change its behaviour as much as a code change. Prompt and configuration changes should go through the same change control as a new model version.
The core of validating an AI tool used in KYP is a test set: a collection of real documents where the correct answers are already known and have been checked by people. The tool is run against the test set and its outputs compared field by field. The test set should include the hard cases, such as scanned documents, unusual structures and documents with amendments, not just clean examples.
| Measure | What It Tests |
|---|---|
| Field accuracy | Share of extracted values that match the known answer |
| Critical field accuracy | Accuracy on the fields that matter most, such as fees, barriers, redemption terms and leverage limits |
| Omission rate | Share of values present in the document that the tool missed |
| Unsupported output rate | Share of outputs that don't appear in the source at all |
| Source accuracy | Share of outputs whose cited source actually supports them |
| Consistency | Whether repeated runs on the same document give the same answer |
| Change detection rate | For monitoring tools, share of known changes in the test set the tool flagged |
An illustrative validation record for a hypothetical tool:
The finding is the most useful part. An average accuracy of 98% hid a weak spot in exactly the product type where errors matter most. Validation should always break results down by document type.
A tool that passed validation can still degrade: a vendor updates its model, document formats change, or new product types arrive. Monitoring should track a few measures continuously:
| Measure | How | Illustrative Trigger |
|---|---|---|
| Sample accuracy | People check a random sample of outputs against sources each month | Below validated level by more than one point |
| Reviewer override rate | Share of outputs that reviewers correct | Rising for two months in a row |
| Unprocessed rate | Share of documents the tool couldn't process | Above 2%, or any silent failure |
| Known-change catch rate | Changes found by other means that the tool should have flagged | Any missed Critical change |
| Version changes | Vendor notices and the firm's own change log | Any change to model version, prompts or settings |
Reviewer corrections are the most valuable monitoring data the firm has. If analysts are fixing the tool's output, those fixes should be captured and counted, not just made.
Most firms will buy rather than build their AI tools. Buying moves the development work to the vendor, but not the model risk.
OSFI's framework covers "models or data sourced from external sources like foreign offices or third-party vendors."[1] FINRA has said its rules apply when firms use a third party's AI technology, "including through embedded features in existing third-party products."[7] In practice, the firm needs enough from the vendor to validate and monitor the tool itself:
| Requirement | Why |
|---|---|
| Documentation of what the tool does, its underlying model and its known limitations | The firm can't rate or describe a tool it doesn't understand |
| Sources for every output | Makes outputs checkable by reviewers, validators and examiners |
| Advance notice of model, prompt or data changes | Triggers revalidation before the change reaches production |
| Version pinning or a test period for changes | Lets the firm test a new version before it is used |
| The vendor's own testing results, broken down by document or product type | Supports, but doesn't replace, the firm's validation |
| Error and failure reporting | Silent failures are the hardest to catch |
| Access for the firm to run its own test sets | Independent validation needs independent testing |
| Data handling terms, including whether firm data trains the vendor's models | Confidentiality and data reliability |
| Audit and record retention support | The firm must be able to show what the tool produced, and when |
Embedded AI counts. Many existing platforms now include AI features switched on by default. The inventory should capture these, not just tools bought as AI products.
Model risk is mainly the firm's to manage, but advisors are often the first to see an AI output that is wrong.