Platform Why Features Security Score AI Engine AI Coding KYP Hub Pricing Company About Buckler News Contact Français Book Demo →
Part One

Model Risk and AI

Model risk is the chance that a firm makes a bad decision because a model's output was wrong, or was used for something it wasn't built for.

1
What the Standards Say
Four sources, and a gap
SourceApplies ToPosition on AI
OSFI Guideline E-23, effective May 1, 2027[1]Federally regulated financial institutions in Canada, including banks and insurers that own dealers and advisersAI included. A model is defined to include "AI/ML methods"; monitoring standards should cover "model drift, autonomous decision making, autonomous re-parametrization."
Revised interagency guidance on model risk management, April 17, 2026[2][3]US banking organizations, most relevant above $30 billion in assets; replaces SR 11-7Generative and agentic AI excluded: they "are novel and rapidly evolving. As such, they are not within the scope of this guidance."[3]
CSA Staff Notice 11-348, December 2024[4]Canadian registrants, including dealers and advisersFirms should be "satisfied that the AI system is fit for purpose and that robust testing prior to deployment has taken place," and should test "before and after its adoption."
FINRA 2026 Annual Regulatory Oversight Report[5]US broker-dealersDescribes firms establishing "a supervision, governance or model risk management framework" for generative AI, with robust testing and ongoing monitoring of outputs.

The gap. For a US firm, the banking agencies' guidance no longer covers generative AI at all, and it states that non-compliance "will not result in supervisory criticism."[3] Securities regulators still expect governance and testing. The result is that firms need to build their own framework for AI tools rather than wait for one. The NIST AI Risk Management Framework, a voluntary US standard organized around four functions (Govern, Map, Measure and Manage), is a common starting point.[6]

Canada is more specific. For firms inside a bank or insurance group, OSFI's definition of model risk is broad: "risk of adverse financial impact arising from the design, development, deployment, and/or use of a model."[1] An AI tool that screens or monitors products is likely to fall within the group's model inventory from May 2027.

2
Why AI Is Different
Traditional model controls assume behaviour that AI doesn't have
CharacteristicTraditional ModelAI ToolWhat It Means for Controls
ConsistencySame input, same outputThe same question can produce different answersTest on repeated runs, not single examples
Failure modeErrors are usually visible (a wrong number, a crash)Errors can be fluent and confident, including invented factsRequire sources for every output so errors can be caught
TransparencyLogic can be inspectedReasoning is hard to inspect directlyExplain outputs through their sources rather than the model's internals
ChangeChanges when the firm changes itA vendor can update the underlying model without the firm doing anythingContractual change notice; revalidate on version change
ScopeDoes one defined taskWill attempt tasks it wasn't built forDefine approved uses; block or flag others
InputsStructured dataDocuments, prompts and instructions, any of which can change resultsTreat prompts and instructions as part of the model, under change control
3
Which KYP Tools Are Models
Deciding what goes in the inventory

Definitions differ between standards, so the practical test is simpler: does the firm rely on the tool's output in its assessment of a product? If it does, the tool belongs in the inventory and under the framework, whatever it is called.

ToolIn Scope?Why
AI extraction of fees, terms and risk factors from offering documentsYesOutputs populate the product file
AI change detection across filings and dataYesDecides what the firm is alerted to, and what it isn't
Product scoring or screening modelYesInfluences which products are approved or reviewed
AI drafting of product summaries for advisorsYesShapes advisors' understanding of the product
AI research assistant answering questions about productsYes, at lower risk if outputs are treated only as inputsCan introduce errors into analysis
Fixed rule that compares a field against a thresholdUsually noDeterministic logic; covered by ordinary system testing
General productivity tools (email drafting, meeting notes)No, unless used for product analysisNot relied on for product assessment
Part Two

The Framework

A model risk framework for AI tools doesn't need to be large. It needs an inventory, a way to rate risk, a lifecycle with independent review, and monitoring that catches problems early.

1
Inventory and Risk Rating
Know every tool, and scale the controls to its risk

The inventory records, for each AI tool: its purpose and approved uses, owner, developer or vendor, underlying model and version, inputs, outputs, where the outputs are used, risk tier, validation date and status, and next review date. OSFI expects the inventory to be comprehensive for models with non-negligible risk and kept current as a system of record.[1]

The risk rating sets how much control each tool gets. OSFI lists factors such as business use, "model complexity or autonomy, data reliability, customer impacts, or regulatory risk."[1] For KYP tools, an illustrative three-tier approach:

TierWhenExamplesControls
HighOutputs feed approvals or status changes with limited human review, or the tool acts autonomouslyProduct scoring used in approval; agents that update recordsFull independent validation before use; quarterly performance review; annual revalidation
MediumOutputs populate product files or reach advisors after human reviewDocument extraction; change detection; summary draftingIndependent validation before use; monthly sample testing; revalidation on material change
LowOutputs are inputs to a person's own analysis, fully reviewedResearch assistant with cited answersPre-use testing; periodic spot checks; usage rules

Autonomy is the factor that moves a tool up fastest. The same extraction tool is medium risk when a person reviews its output and high risk when its output flows straight into the product register.

2
The Lifecycle
From design to retirement
1
Design
Define the use case, approved uses, inputs, outputs and success measures. Rate the risk.
2
Validate
Independent testing against a known set of answers. Confirm it is fit for purpose.
3
Approve
Approval with any conditions on use, recorded with the validation results.
4
Deploy
Release with the approved model version, prompts and settings locked.
5
Monitor
Sample testing, performance metrics and revalidation triggers.
6
Retire
Notify users, keep documentation, and check nothing downstream still depends on it.

Independence matters. OSFI expects the review process to be "independent from model development" and to validate that models are "properly specified, working as intended, and fit-for-purpose."[1] For a firm without a model risk team, independence can mean the people who test the tool are not the people who built or bought it.

Prompts are part of the model. For AI tools, the instructions given to the model can change its behaviour as much as a code change. Prompt and configuration changes should go through the same change control as a new model version.

3
Validating an AI Tool
Testing against known answers, with a worked example

The core of validating an AI tool used in KYP is a test set: a collection of real documents where the correct answers are already known and have been checked by people. The tool is run against the test set and its outputs compared field by field. The test set should include the hard cases, such as scanned documents, unusual structures and documents with amendments, not just clean examples.

MeasureWhat It Tests
Field accuracyShare of extracted values that match the known answer
Critical field accuracyAccuracy on the fields that matter most, such as fees, barriers, redemption terms and leverage limits
Omission rateShare of values present in the document that the tool missed
Unsupported output rateShare of outputs that don't appear in the source at all
Source accuracyShare of outputs whose cited source actually supports them
ConsistencyWhether repeated runs on the same document give the same answer
Change detection rateFor monitoring tools, share of known changes in the test set the tool flagged

An illustrative validation record for a hypothetical tool:

AI Model Validation Record: Example
Hypothetical
Tool
Offering document extraction tool, version 2.3. Vendor-supplied model; firm-configured prompts v14. Risk tier: Medium.
Use
Extracts fees, terms, parties and risk factors from fund facts, offering memoranda and structured product term sheets into draft product files, for analyst review.
Test set
200 documents (120 fund facts, 50 offering memoranda, 30 term sheets, including 18 scanned), 4,800 fields with answers checked by two analysts.
Results
Field accuracy 98.1%. Critical fields 99.4%. Omission rate 0.8%. Unsupported outputs 0.2%. Source accuracy 99.6%. Consistency across three runs 99.1%.
Findings
Scanned term sheets: critical field accuracy 95.2%, below the 99% threshold. Errors concentrated in tables of barrier and call levels.
Decision
Approved with conditions: all fields extracted from scanned documents require full analyst verification; tool must flag scanned inputs.
Revalidate
On any change of vendor model version or prompts, a new document type, or monthly sample accuracy below 98%. Otherwise in 12 months.
Validated by
Model risk analyst, independent of the product and technology teams. Approved by the AI governance committee.

The finding is the most useful part. An average accuracy of 98% hid a weak spot in exactly the product type where errors matter most. Validation should always break results down by document type.

4
Ongoing Monitoring
Catching drift before it reaches a product file

A tool that passed validation can still degrade: a vendor updates its model, document formats change, or new product types arrive. Monitoring should track a few measures continuously:

MeasureHowIllustrative Trigger
Sample accuracyPeople check a random sample of outputs against sources each monthBelow validated level by more than one point
Reviewer override rateShare of outputs that reviewers correctRising for two months in a row
Unprocessed rateShare of documents the tool couldn't processAbove 2%, or any silent failure
Known-change catch rateChanges found by other means that the tool should have flaggedAny missed Critical change
Version changesVendor notices and the firm's own change logAny change to model version, prompts or settings

Reviewer corrections are the most valuable monitoring data the firm has. If analysts are fixing the tool's output, those fixes should be captured and counted, not just made.

Part Three

Vendor Models

Most firms will buy rather than build their AI tools. Buying moves the development work to the vendor, but not the model risk.

1
What to Require
What a firm needs from a vendor to manage the risk itself

OSFI's framework covers "models or data sourced from external sources like foreign offices or third-party vendors."[1] FINRA has said its rules apply when firms use a third party's AI technology, "including through embedded features in existing third-party products."[7] In practice, the firm needs enough from the vendor to validate and monitor the tool itself:

RequirementWhy
Documentation of what the tool does, its underlying model and its known limitationsThe firm can't rate or describe a tool it doesn't understand
Sources for every outputMakes outputs checkable by reviewers, validators and examiners
Advance notice of model, prompt or data changesTriggers revalidation before the change reaches production
Version pinning or a test period for changesLets the firm test a new version before it is used
The vendor's own testing results, broken down by document or product typeSupports, but doesn't replace, the firm's validation
Error and failure reportingSilent failures are the hardest to catch
Access for the firm to run its own test setsIndependent validation needs independent testing
Data handling terms, including whether firm data trains the vendor's modelsConfidentiality and data reliability
Audit and record retention supportThe firm must be able to show what the tool produced, and when

Embedded AI counts. Many existing platforms now include AI features switched on by default. The inventory should capture these, not just tools bought as AI products.

Part Four

Responsibilities

Model risk is mainly the firm's to manage, but advisors are often the first to see an AI output that is wrong.

1
Firm and Advisor Duties
Who does what
What the Firm Needs to Do
  • Inventory every AI tool. Including embedded vendor features used in product work.
  • Rate the risk. By use, complexity, autonomy and data reliability; scale controls to the tier.
  • Validate independently. Against a checked test set, broken down by document and product type.
  • Approve with conditions. Record approved uses and any limits found in validation.
  • Control change. Treat model versions, prompts and settings as part of the model.
  • Monitor continuously. Sample accuracy, reviewer corrections, failures and missed changes.
  • Hold vendors to account. Documentation, sources, change notice and testing access in the contract.
  • Keep people accountable. Product decisions stay with named individuals, not tools.
What the Individual Advisor Needs to Do
  • Use tools as approved. Only for the uses the firm has approved.
  • Check the source. Before relying on an AI summary of a product, confirm it against the cited document.
  • Report errors. Every wrong or unsupported output reported is monitoring data for the firm.
  • Keep their own understanding. An AI output supports, but doesn't replace, the advisor's own knowledge of a product.
2
Example Written Process
What the firm writes down, as numbered clauses
Example: Model Risk Management for AI Tools in KYP
Illustrative
1
Scope. This process applies to every AI tool whose outputs the firm relies on in assessing, approving or monitoring products, whether built internally, bought, or embedded in a third-party platform.
2
Inventory. Each tool is recorded in the model inventory with its purpose, approved uses, owner, vendor, model version, inputs, outputs, risk tier, validation status and next review date.
3
Risk rating. Each tool is rated High, Medium or Low based on business use, complexity, autonomy, data reliability and regulatory impact. The rating sets the validation and monitoring required.
4
Validation. Before use, each High and Medium tool is validated by a person independent of its development or purchase, against a test set of checked answers, with results broken down by document and product type. Approval records any conditions of use.
5
Change control. Changes to a tool's model version, prompts, settings or data sources are logged and assessed before use. Material changes require revalidation.
6
Monitoring. Outputs of High and Medium tools are sample-checked monthly. Reviewer corrections, processing failures and missed changes are recorded. Breach of a monitoring trigger requires review by the model owner and, where set, revalidation.
7
Vendors. Contracts for AI tools require documentation, sources for outputs, advance notice of changes, testing access and record retention support.
8
Accountability. AI tools do not approve products or change product status. Decisions that rely on AI outputs are made and recorded by named individuals.
Five Questions to Test AI Model Risk Management
  1. Is every AI tool used in product work in the inventory, including features inside vendor platforms?
  2. Was each tool tested against known answers before use, with results broken down by document type?
  3. Would the firm know if a vendor changed the underlying model tomorrow?
  4. Are reviewer corrections captured and counted, or just made?
  5. Can every AI output in a product file be traced to its source?
A note on scope: This article describes practical approaches to managing model risk in AI tools used for product due diligence and monitoring, based on publicly available guidance as of its date. Applicability of each source depends on the firm's regulatory status. It is general information, not legal, risk or compliance advice. The risk tiers, lifecycle, validation measures, thresholds, vendor requirements and written process are illustrations, not prescribed requirements; the tool and figures are hypothetical.
References
  1. Office of the Superintendent of Financial Institutions. Guideline E-23 - Model Risk Management (2027), published September 11, 2025, effective May 1, 2027. Source document
  2. Board of Governors of the Federal Reserve System. SR 26-2, Revised Guidance on Model Risk Management, April 17, 2026 (issued with the OCC and FDIC; supersedes SR 11-7 and SR 21-8). Source document
  3. Office of the Comptroller of the Currency. News Release 2026-29, OCC Issues Updated Model Risk Management Guidance, April 17, 2026. Source document
  4. Canadian Securities Administrators. CSA Staff Notice and Consultation 11-348, Applicability of Canadian Securities Laws and the use of Artificial Intelligence Systems in Capital Markets, December 5, 2024. Source document (PDF)
  5. FINRA. 2026 Annual Regulatory Oversight Report, December 2025. Generative AI, pp.24-28. Source document (PDF)
  6. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023. Source document
  7. FINRA. Regulatory Notice 24-09, FINRA Reminds Members of Regulatory Obligations When Using Generative Artificial Intelligence and Large Language Models, June 27, 2024. Source document