T20 · AI & Frontier Tech · Financial Governance

AI Governance and Model Risk in Regulated Finance

An evidence-gated framework for governing models, generative AI and agentic systems across family offices and institutional allocators.

Concentric glass governance rings protecting an artificial-intelligence core while an investment committee observes the control boundary
Quick answer

Controlled AI adoption in regulated finance begins with one complete inventory, consequence-led materiality, explicit authority, system-level evaluation, independent challenge, third-party controls, current monitoring and a tested exit.

Abstract

Background. AI is entering research, risk, compliance, operations and investment workflows while legal, supervisory and standards sources retain different scopes and authority.

Objective. This paper develops an evidence-gated AI governance and model-risk framework for A2 family-office CIOs and A1 international institutional allocators.

Approach. The analysis reviews 40 primary laws, regulators, standard setters and recognised standards available through 1 August 2026.

Findings. A complete inventory, consequence-led materiality, explicit authority, system-level evaluation, independent challenge, third-party controls, monitoring and tested exit form one coherent operating model.

Implications. Institutions can tailor governance to scale while preserving control objectives, jurisdictional status and decision evidence. All worked inputs are unverified illustrative management assumptions; attributed Matchpoint or client revenue, cash cost reduction and loss reduction remain USD 0 until approved observed evidence exists.

JEL Classification: C53, C58, G20, G23, G28, G38, M15, O33

Keywords: artificial intelligence, AI governance, model risk, regulated finance, family office, institutional allocator, validation, generative AI, agentic AI, third-party risk, operational resilience, data governance

This Matchpoint Insight presents the web edition of Matchpoint Partners' research. The supporting paper contains the regulatory map, governance architecture, lifecycle, evaluation framework, third-party controls, worked operating cases, board pack, roadmap and source register.

Read the full research paper   Explore AI & Technology Advisory

Introduction

Artificial intelligence has moved from isolated experimentation into investment research, client service, compliance, operations, risk analysis and technology delivery. In regulated finance, the central management question is whether each use can be governed in proportion to its consequence. The answer requires an inventory, a defensible classification, accountable decision rights, independent challenge, reliable evaluation, traceable operation and credible exit. A policy statement alone cannot provide that evidence.

The regulatory baseline changed materially in 2026. The US federal banking agencies issued revised model risk management guidance on 17 April 2026. The guidance supersedes SR 11-7, emphasises a risk-based and tailored approach, and is expected to be most relevant to Federal Reserve regulated banking organisations with more than USD 30 billion in total assets [1]. The OCC also stated that generative AI and agentic AI are outside the scope of that revised guidance because they are novel and rapidly evolving [2]. That scope boundary matters. It means an institution cannot assume that conventional model risk management alone answers every generative or agentic AI question.

In the UK, the April 2026 version of PRA Supervisory Statement SS1/23 applies model risk management principles across models used to inform business decisions, including vendor models and artificial intelligence where the principles apply [4]. The Bank of England, FCA and HM Treasury separately called for board understanding and active mitigation of frontier-AI cyber and operational-resilience risks in May 2026 [9]. In June 2026, the Financial Stability Board proposed 12 sound practices covering organisation-wide governance, lifecycle controls, cyber, information and communication technology, and third-party risk [10]. That FSB document remained a consultation at the 1 August 2026 evidence cut-off; its final report was scheduled for October 2026 [10].

The UAE position also became more specific. The Central Bank of the UAE issued consumer-protection guidance for responsible AI and machine-learning use by licensed financial institutions in February 2026 [16]. The DFSA published regulatory expectations on AI risk management in the DIFC in June 2026 [18]. These materials sit alongside the UAE regulators' 2021 enabling-technology guidelines, the UAE personal-data framework and DIFC rules for personal data processed by autonomous and semi-autonomous systems [17,22,23].

This paper converts those overlapping sources into an operating framework for A2 family-office CIOs and heads of alternatives and A1 international institutional allocators. The framework is designed for investment institutions that may be regulated directly, may allocate to regulated firms, may procure AI-enabled services, or may need institutional-grade governance because of fiduciary, contractual or reputational exposure. It distinguishes law, binding rules, supervisory guidance, consultation material, standards and voluntary practice. Applicability must be confirmed by qualified legal and compliance owners for the entity, activity and jurisdiction.

The contribution is practical. It defines a common inventory covering models, AI systems, assistants, agents and embedded vendor capabilities; a materiality method that separates decision consequence from technical novelty; a lifecycle with evidence gates; evaluation requirements for deterministic, predictive, generative and agentic systems; third-party and concentration controls; and board information that supports actual challenge. Two worked operating cases are included. Every numerical input in those cases is explicitly unverified and illustrative. Attributed Matchpoint or client revenue, cash cost reduction and loss reduction remain USD 0 until approved observed evidence exists.

The paper makes five propositions.

PropositionDecision implication
P1. Inventory completeness precedes risk classification.Unrecorded use cannot be governed, monitored or retired.
P2. Materiality depends on consequence, authority and reversibility.A familiar model can be high consequence; a novel model can be low consequence.
P3. Validation must match system behaviour.Generative and agentic systems require evaluation beyond conventional point-estimate accuracy.
P4. Third-party AI remains the institution's governance problem.Procurement evidence, contractual rights, monitoring and exit are control requirements.
P5. Board oversight needs decision evidence.Counts, incidents, exceptions, residual risk and exit readiness are more useful than adoption narratives.

Scope, Definitions And Evidence Boundaries

Scope

The subject is governance and model risk for AI used in regulated finance and institutionally governed investment activity. Covered uses include research, screening, portfolio analytics, valuation support, forecasting, financial crime controls, client communications, suitability support, operations, coding, document analysis and workflow agents. The paper addresses internally developed systems, configured third-party systems, embedded AI features and externally supplied model outputs.

The paper does not determine legal status, regulatory perimeter, fiduciary compliance, suitability, data-law compliance or accounting treatment for a particular institution. It does not certify a model, a vendor or an AI system. Each institution must establish the applicable requirements through its authorised legal, compliance, risk, information-security, privacy and business owners.

Working definitions

Definitions vary across rules and standards. The institution should preserve the applicable legal definition in its obligations register and maintain a broader operational perimeter for governance.

TermOperational definition used in this paper
ModelA quantitative, statistical, economic or mathematical method that transforms inputs into estimates, decisions or information used in business decisions.
AI systemA machine-based system that infers outputs such as predictions, content, recommendations or decisions from inputs for explicit or implicit objectives.
Generative AIAI that produces content, including text, code, images, audio or structured data.
Foundation modelA model trained broadly and capable of adaptation across multiple downstream tasks.
Agentic systemAn AI-enabled system that plans or selects actions, invokes tools or services, and may continue across multiple steps.
Decision supportOutput informs a human decision while documented authority remains with the human.
Automated decisionThe system determines or executes an outcome within delegated authority.
Model riskPotential adverse consequences from incorrect or misused model outputs and decisions [1].
AI riskRisk arising from the design, development, procurement, deployment, operation, interaction, misuse or retirement of an AI system.
MaterialityThe significance of a use based on consequence, exposure, scale, authority, rights impact, reversibility and interconnectedness.

The operational perimeter is intentionally inclusive. A spreadsheet formula, vendor score, retrieval assistant or workflow agent may not satisfy every legal definition of a model. It can still create consequential dependency. Registration routes the item to the correct governance owner; it does not predetermine regulatory classification.

Evidence classes

Evidence has different authority. Mixing these classes creates false certainty.

ClassExamplesPermitted use
Law and binding regulationEU AI Act, DORA, GDPR, UAE data lawsDetermine obligations after applicability is confirmed.
Supervisory rule or statementPRA SS1/23, CBUAE rulebook guidance, DFSA rulesTranslate supervisory expectations within stated scope.
Supervisory guidanceUS interagency revised MRM guidance, FINMA guidanceDesign risk-based practices while preserving non-binding status where stated.
ConsultationFSB June 2026 sound practicesAnticipate direction; label as proposed and monitor finalisation.
Consensus standardISO/IEC 42001 and 23894Structure management systems and risk processes; confirm certification claims separately.
Voluntary frameworkNIST AI RMF, OECD AI PrinciplesCreate common language and operational outcomes.
Market evidenceRegulator surveysInform priors; preserve sample, date and population.
Local evidenceInventory, tests, incidents, logs, approvalsSupport the institution's own decision.

Evidence cut-off and current-state limitations

Sources were reviewed through 1 August 2026. The EU implementation timetable changed in July 2026, including extended dates for high-risk AI obligations [25]. Article 50 transparency obligations were scheduled to apply from 2 August 2026, one day after the evidence cut-off [26]. The FSB sound practices remained a consultation [10]. NIST stated that AI RMF 1.0 was being revised [29]. The obligations register must therefore record source date, current status, next review date and accountable legal owner.

Research method

The paper prioritises primary laws, regulators, standard setters and recognised standards bodies. It triangulates model risk, AI governance, operational resilience, cybersecurity, data protection and third-party risk. It does not treat a survey percentage as a universal adoption rate. It does not treat a voluntary standard as law. It does not interpret consultation language as final policy.

A2 And A1 Decision Map

A2 family-office CIOs and heads of alternatives

The Topic Tracker defines A2 as investment professionals at single-family offices, multi-family offices, private-wealth organisations and external asset managers across the GCC, UK, Switzerland, Singapore and the EU. Their AI governance problem often spans a small internal team, confidential family information, external managers, private-market documents, bank and administrator data, and multiple outsourced technology providers.

A2 decisionMaterial risk questionMinimum evidence
Use AI for manager researchCan unsupported content influence selection or rejection?Controlled corpus, citation coverage, reviewer protocol, exception log.
Summarise fund and deal documentsCan omitted or altered terms affect a commitment?Clause-level traceability, material-term checklist, legal review boundary.
Generate investment-committee materialCan generated statements enter an approved record?Source map, named preparer, reviewer, version and approval evidence.
Analyse portfolio dataAre data rights, lineage and reconciliation sufficient?Data register, purpose, lineage, completeness and reconciliation tests.
Automate instructions or communicationsCan the system bind, disclose or transmit value?Authority limits, dual control, allow-listed tools, audit log and recovery.
Use embedded vendor AIIs the capability visible and contractually governed?Vendor inventory, change notice, data terms, security evidence and exit plan.

A2 institutions should avoid copying a bank's organisational scale. Proportionality can preserve the same control objectives with fewer committees. A named accountable executive, an independent challenger and a complete evidence pack may be more effective than multiple forums with unclear ownership.

A1 international institutional allocators

The Topic Tracker defines A1 as pensions, insurers, endowments and fund-of-funds in the US, EU and UK evaluating or scaling allocations to UAE and GCC private markets. Their AI governance problem adds delegation, public or beneficiary accountability, regulated outsourcing, committee process, manager oversight and cross-border data flows.

A1 decisionMaterial risk questionMinimum evidence
Screen managers or opportunitiesDoes AI change access, scoring or escalation?Feature rationale, bias analysis where relevant, override and appeal process.
Produce due-diligence questionsAre gaps and adverse facts preserved?Population coverage, source retention, reviewer sign-off and sampling.
Monitor funds and portfoliosCan model drift or data delay change risk signals?Data timeliness, benchmark, drift limits, incident and escalation rules.
Support valuation or risk estimatesIs the method inside the model-risk perimeter?Classification, development evidence, independent validation and limitations.
Outsource AI-enabled processingCan the allocator audit, transition and continue service?Contract, subcontractor map, service levels, testing, portability and exit.
Use AI in regulated reportingCan provenance and control evidence withstand review?Reconciliation, lineage, maker-checker approval, retention and reproducibility.

Shared governance outcome

Both ICPs need a compact decision contract.

QuestionRequired answer
What is the system?Registered scope, version, owner, provider, data and interfaces.
What decision does it affect?Use case, user, beneficiary, process and consequence.
What authority does it have?Read, draft, recommend, approve, execute, communicate or transfer.
What can fail?Error, bias, unsupported output, misuse, security, privacy, concentration and resilience.
How is it tested?Representative evaluation, thresholds, challenger and decision record.
How is it operated?Monitoring, exceptions, changes, incidents and user competence.
How is it stopped?Kill switch, fallback, data export, transition and retirement evidence.

Regulatory And Supervisory Landscape

United States banking model-risk guidance

The April 2026 interagency guidance describes model risk as the potential for adverse consequences from decisions based on incorrect or misused model outputs. It highlights effective development and use, validation, and governance and controls [1]. It expects tailoring to the banking organisation's model risk profile, size and complexity. Federal Reserve applicability is described as most relevant above USD 30 billion in total assets [1]. The OCC stated that the guidance is not an enforceable standard and that generative and agentic AI models are outside its scope [2].

This creates a two-part governance task. Conventional models should be mapped to the revised guidance where applicable. Generative and agentic systems need an adjacent AI-risk method that retains useful model-risk disciplines while adding content, interaction, tool-use, security, data-rights and autonomy controls. Institutions should record which regime or policy owns each control objective.

US guidance pointOperating translation
Risk-based tailoringSet validation depth and frequency from materiality and exposure.
Development and useRecord purpose, design, data, assumptions, limitations and intended users.
ValidationApply independent, effective challenge with scope matched to risk.
Governance and controlsMaintain board or senior oversight, policies, inventory, issues and audit.
Vendor productsObtain enough information and testing to understand and manage dependency.
GenAI and agentic exclusionUse a documented adjacent AI framework; do not imply coverage by the MRM guidance.

United Kingdom

PRA SS1/23 sets five principles: model identification and classification; governance; development, implementation and use; independent validation; and model-risk mitigants [4]. It applies to vendor models as well as internally developed models and addresses AI in modelling techniques where the general principles apply [4]. Bank and FCA materials also emphasise governance, accountability, data, model and third-party risks [5,6].

The May 2026 UK joint statement adds a distinct frontier-AI cyber-resilience concern. It calls for board and senior-management understanding, investment and resourcing, access management, network security, data protection, protective and detective capabilities, threat containment and response [9]. The July 2026 Financial Stability Report described four channels: core financial decision-making, financial markets, service-provider operational risk and the changing external cyber threat [8].

The governance implication is that AI risk cannot sit only within model validation. It needs coordination across business ownership, model risk, operational risk, information security, privacy, compliance, procurement and resilience.

UAE and DIFC

The CBUAE's February 2026 guidance focuses on consumer protection and market conduct for licensed financial institutions using AI and machine learning. It addresses transparency, bias, ethics, accountability, explainability and data privacy [16]. The joint UAE enabling-technology guidelines address AI, big-data analytics, cloud, application programming interfaces, biometrics and distributed ledger technology [17].

The DFSA's June 2026 SEO letter records regulatory expectations for AI risk management in the DIFC [18]. Its 2025 survey covered 661 authorised firms with 88 per cent participation. The DFSA reported that 52 per cent used AI, 60 per cent had some form of AI governance structure and 21 per cent lacked clear accountability or oversight even where use was critical [19]. Those figures describe that survey population and date; they are not universal adoption estimates.

DIFC Data Protection Regulation 10 addresses personal data processed through autonomous and semi-autonomous systems [22]. UAE Federal Decree-Law No. 45 of 2021 establishes the federal personal-data framework [23]. The institution's obligations register should identify the governing data regime, controller and processor roles, legal basis, sensitive-data constraints, transfer requirements, individual rights, retention and incident obligations.

European Union

Regulation (EU) 2024/1689 establishes a risk-based AI framework [24]. The July 2026 AI Omnibus changed elements of implementation and extended the start of high-risk AI rules to 2 December 2027, with rules for AI embedded in regulated physical products scheduled for 2 August 2028 [25]. Transparency obligations for specified AI interactions and generated or manipulated content were scheduled to apply from 2 August 2026 [26].

DORA requires financial entities in scope to identify and document ICT assets, dependencies and critical interconnections; manage ICT risk; govern third-party arrangements; test resilience; and maintain contractual and exit capabilities [27]. GDPR continues to govern personal-data processing and automated decision issues where applicable [28]. AI classification should therefore connect to existing ICT, outsourcing, privacy, conduct and resilience registers.

International coordination and standards

The FSB's June 2026 consultation proposes 12 sound practices. The first four address organisation-wide governance, the next six address lifecycle risk management and the final two address cyber, ICT and third-party risk [10]. The document states that the practices are not intended to create an international standard or prescribe whether institutions should adopt AI [10].

Basel Committee materials link digitalisation to governance, model risk, third-party dependency, operational risk and resilience [12-14]. BIS FSI research observed that most financial authorities address AI through existing technology-neutral frameworks while identifying gaps around governance, skills, model risk, data, third parties and new business models [15].

NIST AI RMF organises outcomes under Govern, Map, Measure and Manage [29]. Its Generative AI Profile identifies risks and suggested actions specific to or amplified by generative AI [30]. ISO/IEC 42001 specifies an AI management-system standard [33]. ISO/IEC 23894 provides AI-risk-management guidance [34]. These are useful control-design references. Their use does not establish legal compliance.

Applicability matrix

SourceStatus at 1 Aug 2026Illustrative relevanceRequired owner action
EU AI Act and OmnibusBinding EU law; staged applicationEU providers, deployers or affected activityLegal applicability and role assessment.
DORABinding EU regulation for in-scope financial entitiesICT and AI service dependencyIntegrate AI services into ICT and third-party registers.
CBUAE AI guidanceRegulatory guidance for UAE licensed financial institutionsConsumer-facing or market-conduct AIMap transparency, bias, accountability and privacy controls.
DFSA expectationsSupervisory communication for DIFC authorised firmsDIFC AI useEvidence risk management and accountable oversight.
PRA SS1/23Supervisory statement for in-scope banksModels used in decisionsMap inventory, governance, development, validation and mitigants.
US interagency MRMSupervisory guidance with stated scopeIn-scope banking modelsTailor model-risk practices; preserve GenAI and agentic exclusion.
FSB sound practicesConsultationDirectional cross-border practiceTrack finalisation; use only as proposed practice.
NIST and ISOVoluntary framework and standardsControl architectureAdopt selected outcomes with an evidence map.

Inventory And Classification

One perimeter, multiple taxonomies

An institution should maintain one discoverable AI and model register with linked classifications. Separate unconnected inventories for models, SaaS, end-user computing, privacy, vendors and operational processes create gaps. The master record should link to specialist registers rather than duplicate every field.

Inventory fieldPurpose
Unique ID and versionPrevent ambiguity across deployments and changes.
Business purpose and processConnect technology to an accountable outcome.
Owner, operator and usersEstablish responsibility and competence.
Provider and hostingIdentify internal and third-party dependency.
Model or service componentsReveal foundation models, tools, retrieval and deterministic services.
Data sources and rightsEstablish provenance, purpose, sensitivity and transfer.
Outputs and affected partiesIdentify decision, conduct and rights exposure.
Authority levelRecord whether the system reads, drafts, recommends, approves or acts.
Jurisdictions and entitiesRoute applicability assessment.
Materiality tierDetermine control depth and escalation.
Evaluation and approvalLink evidence, thresholds, conditions and expiry.
Monitoring and incidentsMaintain current operating risk.
Change historyTrigger proportionate re-evaluation.
Exit and fallbackProve recoverability and portability.

Discovery routes

Inventory completeness needs multiple detection routes. Self-attestation alone is weak because embedded features and shadow use can be invisible to central teams.

Discovery routeExamplesEvidence
ProcurementContracts, renewals, vendor questionnairesProduct, features, provider and data terms.
Identity and accessApplication catalogue, single sign-on, privileged accessUsers, roles and service access.
Network and cloudDomains, application programming interfaces, cloud servicesService interaction and data flow.
Expense and cardIndividual subscriptionsShadow procurement candidates.
Code and model platformsRepositories, registries, endpointsInternally built and configured components.
Business process reviewResearch, investment, risk, compliance and operationsEmbedded use and decisions.
User declarationPeriodic attestation and campaignUnmanaged use with accountability.
Vendor change noticeRelease notes and feature togglesNewly embedded AI functionality.
Incident and support dataService tickets, data events and control exceptionsPreviously unregistered use.

Materiality model

Materiality should be assessed before controls are selected. The proposed score uses eight dimensions, each rated from zero to four. The score supports judgement; it does not replace it.

DimensionLow endHigh end
Decision consequenceConvenience or formattingCapital, client, legal, regulatory or rights outcome
AuthorityRead-only or draftExecute, communicate, transfer or bind
Population and scaleFew internal usersLarge or externally affected population
Data sensitivityPublic or syntheticPersonal, confidential, market-sensitive or restricted
ReversibilityImmediate correctionIrreversible or costly remediation
Explainability needNo material relianceDecision must be reconstructed or explained
DependencyEasy substituteConcentrated provider or critical integration
Change velocityFixed and controlledProvider-driven or adaptive behaviour

Proposed initial bands are 0-7 for Tier 1, 8-15 for Tier 2, 16-23 for Tier 3 and 24-32 for Tier 4. These bands are management design parameters, not regulatory thresholds. A mandatory override should elevate uses involving legal rights, material client communications, autonomous value movement, regulatory reporting, safety, credit or suitability decisions, or critical service dependencies.

Authority classification

Authority frequently drives risk more directly than model type.

LevelCapabilityDefault control position
A0Search or retrieve approved informationSource controls and access logging.
A1Draft or summariseHuman review before external or decision use.
A2Recommend or rankDefined decision owner, reason record and override.
A3Prepare a transaction or instructionSegregation, limit and independent release.
A4Execute within a bounded mandateExplicit delegation, allow list, dual controls where material, monitoring and kill switch.
A5Self-directed planning across toolsExceptional approval, sandboxing, least privilege, transaction limits, continuous monitoring and tested containment.

Classification output

The classification record should state facts, conclusions and uncertainty separately.

Record componentExample format
FactsSystem X summarises investment documents using provider Y; data is hosted in region Z.
Legal classificationLegal owner conclusion, source, date and scope.
Model-risk classificationMRM owner conclusion and policy basis.
AI-risk tierScore, overrides and rationale.
ConditionsHuman review, approved corpus, prohibited data and authority limit.
Residual uncertaintyProvider training-data details unavailable; compensating restriction applied.

Governance And Decision Rights

Accountable ownership

Each registered use needs a business owner who owns the outcome and residual risk. Technology ownership does not substitute for business accountability. Model risk, compliance, privacy and security owners provide independent or specialist challenge within their mandates.

RoleCore accountability
Board or governing bodyRisk appetite, oversight, major exposures and challenge.
Senior accountable executiveOrganisation-wide framework, resources and escalation.
Business ownerPurpose, outcome, users, controls, residual risk and value.
System ownerTechnical design, configuration, access, operation and change.
Model-risk ownerClassification, validation standards, issues and aggregation.
Compliance and legalApplicability, conduct, regulatory obligations and legal risk.
Privacy and dataPurpose, rights, minimisation, transfers, retention and data governance.
Information securityThreat model, testing, access, logging and incident response.
Procurement and third-party riskDue diligence, contract, concentration, monitoring and exit.
Internal auditIndependent assurance over design and operating effectiveness.

Proportional forum design

A2 institutions may use an existing investment, risk or operating committee with a standing AI agenda. A1 institutions may use a central AI risk committee with business-level approvals. The required outcome is clear authority and documented challenge.

DecisionMinimum forumRequired record
Register and classifyBusiness owner plus risk coordinatorInventory record and tier rationale.
Approve low-risk internal useDelegated ownerConditions, expiry and monitoring.
Approve material useCross-functional risk forumEvidence pack, challenge, decision and residual risk.
Accept significant exceptionSenior accountable executiveException, compensating controls, duration and remediation.
Approve autonomous authorityGoverning or expressly delegated forumAuthority boundary, limits, containment and accountability.
Retire or suspendBusiness and system owners; risk notificationTrigger, service plan, evidence preservation and transition.

Policy architecture

The institution should avoid a stand-alone AI policy that conflicts with established obligations. A concise AI policy should define perimeter, principles, roles, prohibited uses, approval, monitoring, incidents and exceptions. Supporting standards should connect to model risk, data, privacy, cybersecurity, outsourcing, operational resilience, conduct, records, change and audit.

Risk appetite

Risk appetite should specify prohibited activities, quantitative or categorical limits, approval authorities and escalation. Examples include:

Appetite statementEvidence measure
No unapproved confidential data enters public AI services.Data-loss events, blocked attempts and approved exceptions.
No material investment decision relies solely on generated output.Decision records with named human authority and source evidence.
No autonomous system moves value outside approved limits.Tool permissions, transaction limits, release logs and incidents.
Tier 4 systems operate only within current approval.Approval expiry, overdue conditions and deployment state.
Critical vendor dependency has a tested fallback.Exit test date, recovery time and unresolved gaps.

Competence and challenge

Training should match role. General users need data, source, hallucination, phishing, prompt-injection and escalation awareness. Owners need classification, evidence and incident responsibilities. Validators need technical and domain competence. Boards need enough understanding to challenge risk concentration, autonomy, limitations and resourcing [9,10]. Training completion is weak evidence on its own; scenario testing and observed behaviour provide stronger evidence.

Lifecycle And Evidence Gates

Lifecycle

The proposed lifecycle contains ten stages: idea, registration, classification, design, development or procurement, evaluation, approval, controlled deployment, monitoring and retirement. Iteration returns to the appropriate earlier gate when the model, prompt, data, tool, purpose, population or provider changes.

GateDecision questionMinimum evidence
G0 IntakeIs the use permissible to explore?Purpose, owner, data class and sandbox conditions.
G1 ClassificationWhat governs the use?Inventory, authority, materiality and applicability.
G2 DesignAre controls designed into the workflow?Architecture, data flow, threat model and human role.
G3 Build or buyIs the component understood enough?Development record or vendor due diligence.
G4 EvaluationDoes it meet defined thresholds?Representative tests, limitations and independent challenge.
G5 ApprovalIs residual risk accepted by the right authority?Decision pack, conditions, expiry and accountable sign-off.
G6 DeploymentCan release be controlled and reversed?Access, change, fallback, logging and runbook.
G7 OperationDoes current evidence remain inside appetite?Monitoring, incidents, exceptions, drift and user feedback.
G8 ChangeDoes the change require re-evaluation?Change classification, regression and approval.
G9 RetirementCan dependency end safely?Export, archive, access removal, vendor termination and lessons.

Design evidence

The design pack should show the full system, including the user's task, data, model, retrieval, deterministic controls, tools, interfaces, human review and downstream action. A model card without workflow context is incomplete for a material AI use.

Design artefactRequired content
Purpose statementIntended outcome, user, population, exclusions and prohibited use.
ArchitectureComponents, providers, versions, data flows, tools and boundaries.
Rights mapData owners, legal basis, licences, confidentiality and transfers.
Decision mapRecommendations, approvals, execution and communication authority.
Failure analysisHazards, causes, controls, detection, consequence and recovery.
Threat modelAssets, actors, attack paths, controls and residual risk.
Evaluation planPopulation, metrics, thresholds, samples and challenger.
Operating modelOwners, monitoring, changes, incidents, fallbacks and support.
Exit planPortability, replacement, data deletion, continuity and cost.

Approval pack

The approval pack should answer one question: does the available evidence support the proposed bounded use? It should identify gaps rather than convert absence of evidence into a positive conclusion.

SectionApproval content
Executive decisionApprove, approve with conditions, pilot, redesign, pause or reject.
ApplicabilityEntity, jurisdiction, activity and authoritative sources.
ClassificationModel, AI, authority, materiality and criticality.
EvidenceDesign, data, evaluation, security, privacy and vendor results.
LimitationsKnown gaps, unsupported populations and prohibited use.
Residual riskSeverity, likelihood, owner and acceptance authority.
ConditionsAction, owner, due date, evidence and consequence of breach.
MonitoringMetrics, limits, frequency and escalation.
ExpiryDate or change trigger requiring renewal.

Change control

AI services can change through model versions, prompts, system instructions, retrieval content, tools, policies, safety filters, providers, data, user populations and hosting. The change standard should classify each change and define required regression.

Change classExampleDefault action
RoutineNon-material interface correctionRecord and test affected function.
ModeratePrompt or retrieval updateTargeted regression and owner approval.
MaterialModel, provider, purpose, authority or population changeReclassification, independent evaluation and approval.
EmergencySecurity mitigation or service withdrawalContain, document, approve temporary state and complete retrospective review.

Retirement

Retirement should remove access and dependency while preserving required records. For critical systems, the institution should test data export, replacement operation and recovery before approving production. Vendor termination rights without an operational transition plan are weak evidence.

Validation And Evaluation

Validation objective

Validation provides effective challenge over conceptual soundness, implementation, performance, use and limitations. The depth depends on materiality, complexity and exposure [1,4]. The validator should be independent enough to challenge the owner and competent across the relevant domain, data, model and workflow.

Evaluation by system type

System typeCore evaluation
Deterministic rulesRequirement coverage, boundary cases, reconciliation and change tests.
Predictive modelDiscrimination, calibration, stability, sensitivity, bias where relevant and benchmark.
Generative assistantFactual support, completeness, harmful output, robustness, privacy, source use and human acceptance.
Retrieval-augmented systemRetrieval recall, precision, permissions, freshness, citation fidelity and answer support.
Code generatorCorrectness, security, dependency, licence, test coverage and reviewer competence.
Agentic workflowGoal compliance, tool selection, authority, state, recovery, attack resistance and cumulative error.
Vendor black boxOutcome testing, limitations, change controls, monitoring and compensating restrictions.

Representative evaluation set

The evaluation population should reflect normal, difficult, adverse and prohibited cases. It should include material subgroups and periods, known failure modes, boundary conditions and out-of-distribution examples. Test-set governance should prevent contamination and untracked changes.

Evaluation stratumPurpose
RoutineEstablish normal performance and unit economics.
ComplexTest long, ambiguous, cross-document or exception-heavy cases.
AdverseTest misleading, conflicting or malicious inputs.
SensitiveTest confidentiality, personal data and restricted information.
Authority boundaryTest attempted prohibited action or escalation.
TemporalTest current and older regimes, terms or market states.
SubgroupTest legally or operationally relevant populations.
RecoveryTest timeout, tool failure, corrupted state and provider outage.

Metrics

A single accuracy percentage is insufficient for most material AI uses.

MetricDefinition
Material factual-support rateOutputs with every material claim supported by an approved source divided by reviewed outputs.
Material omission rateOutputs omitting at least one required material item divided by reviewed outputs.
Unsupported-action rateAttempts to act outside delegated authority divided by relevant tests.
First-pass acceptanceOutputs accepted without material correction divided by reviewed outputs.
Severe-defect rateOutputs with a defect meeting the severity definition divided by reviewed outputs.
Escalation precisionCorrect escalations divided by all escalations.
Escalation recallCorrect escalations divided by all cases requiring escalation.
Citation fidelityCitations that support the associated claim divided by citations reviewed.
Recovery successFailure scenarios restored within defined limits divided by scenarios tested.
Human override rateDecisions changed by authorised reviewers divided by AI-assisted decisions.

Thresholds must be tied to consequence. A drafting assistant and an automated client instruction should not share the same tolerance. For zero-tolerance event types, the evaluation objective should be absence in the tested population plus strong preventative and detective controls; it should not be described as proof that the event cannot occur.

Generative AI evaluation

The NIST Generative AI Profile identifies risks including confabulation, data privacy, harmful bias, human-AI configuration, information integrity, information security, intellectual property and value-chain dependency [30]. Evaluation should cover the system as configured, including prompts, retrieval, tools, safety controls and workflow. Provider benchmark scores cannot establish fitness for a local investment or regulated-finance use.

Agentic evaluation

Agentic systems add sequence, state, tool and authority risk. A step may be individually acceptable while the sequence produces an unauthorised outcome. Evaluation should record each action, tool input, tool output, permission decision, human intervention and final state.

Agent testQuestion
Goal adherenceDoes the system remain within the approved objective?
Tool allow listDoes it select only approved tools for the task?
Parameter limitsAre amounts, recipients, destinations and data bounded?
Prompt injectionCan external content override system policy or exfiltrate data?
Cumulative errorDo small intermediate errors create a material final outcome?
State integrityCan stale or corrupted state change action?
Human checkpointIs approval requested at the right stage with complete information?
Kill and recoveryCan operation stop safely and resume from a known state?

Independent challenge record

The validation report should state scope, methods, evidence, limitations, findings, severity, conditions and conclusion. A qualified conclusion can support a bounded pilot. Missing evidence should remain a limitation or issue. Management acceptance should name the accountable owner and authority.

Data, Privacy, Security And Third Parties

Data governance

AI controls inherit from the data lifecycle. Each material use should establish provenance, ownership, permission, purpose, quality, lineage, minimisation, retention, transfer, deletion and incident response.

Data questionEvidence
May the institution use the data for this purpose?Contract, legal basis, consent or other authorised basis.
May the provider process or retain it?Data-processing terms, configuration and verified service behaviour.
Is the corpus complete and current?Source register, reconciliation, refresh and exception logs.
Can users access only authorised material?Identity, permissions, retrieval filters and adversarial tests.
Can output reveal sensitive input?Privacy testing, logging, redaction and incident process.
Can the data be exported and deleted?Portability test, deletion evidence and contract.

Secure development and operation

NIST CSF 2.0 and the NCSC secure-AI guidance provide governance and lifecycle security references [31,32]. OWASP's LLM application risk work provides a technical threat catalogue [36]. The institution should select controls from its applicable security framework and record their implementation.

ThreatIllustrative control objectives
Prompt injectionTreat external content as untrusted; isolate instructions; restrict tools and data.
Sensitive disclosureMinimise data, enforce access, filter output, monitor and respond.
Supply-chain compromiseVerify components, provenance, dependencies, updates and provider controls.
Data or model poisoningControl contribution, validate sources, detect anomalies and preserve rollback.
Excessive agencyLeast privilege, transaction limits, approvals, allow lists and kill switch.
Insecure output useValidate before execution, encode safely and segregate duties.
Model theft or extractionAccess controls, rate limits, monitoring and response.
Denial or provider outageCapacity plan, fallback, recovery and alternate process.

Third-party AI

Financial institutions remain responsible for risks created by services they procure or rely on. DORA, Basel materials, FSB work, FINMA guidance and UAE regulator materials all reinforce the importance of third-party dependency, governance and resilience [10,12,15-17,27,37].

Due-diligence domainEvidence request
Service definitionComponents, versions, hosting, subprocessors and support.
DataCollection, use, retention, training, location, transfer and deletion.
Model governanceDevelopment, evaluation, limitations, monitoring and change.
SecurityControl assurance, incident history, testing and notification.
ResilienceAvailability, recovery, capacity, continuity and dependency map.
AuditabilityLogs, evidence access, audit and supervisory cooperation.
Legal and rightsLicence, intellectual property, confidentiality and prohibited use.
ChangeNotice, version control, regression evidence and customer options.
ExitExport, assistance, deletion, transition time and fees.

Concentration and fourth parties

Foundation models, cloud infrastructure, data services and specialist tooling can create common dependencies across apparently different applications. The inventory should identify ultimate providers and material subprocessors. Scenario analysis should examine simultaneous failure, degraded service, unilateral term change, region outage, model withdrawal and loss of audit access.

Contract control schedule

The contract should define service scope, approved data use, confidentiality, security, incident notice, business continuity, audit, regulator access where applicable, subcontracting, model changes, performance, intellectual property, record retention, deletion, exit assistance and liability allocation. Contract language must be reviewed by qualified counsel. A right without operational testing remains incomplete evidence.

Human Oversight, Conduct And Accountability

Meaningful human oversight

Human review is meaningful when the reviewer has competence, time, authority, information and an effective ability to challenge or stop the outcome. A mandatory click with no supporting evidence is weak control.

Review conditionEvidence
CompetenceRole-specific training, experience and assessment.
InformationSource material, system output, limitations and uncertainty.
TimeWorkflow capacity and service-level design.
AuthorityAbility to reject, amend, escalate or stop.
IndependenceSeparation where consequence requires it.
TraceabilityNamed reviewer, time, action, reason and final outcome.
FeedbackCorrections routed into monitoring and improvement.

Automation bias and workload

Oversight can fail when users defer to polished output, review too many items, or lack access to source evidence. Monitoring should measure correction, override, review time, exception and disagreement patterns. A very low override rate may indicate high quality; it may also indicate weak challenge. Investigation requires sampled evidence.

Fairness and affected parties

Where AI affects access, pricing, suitability, employment, credit, insurance or other consequential treatment, the institution should identify legally and operationally relevant groups, evaluate differential outcomes, document legitimate objectives, provide appropriate explanation and maintain contest or escalation routes. Applicable law determines specific obligations. OECD principles, NIST outcomes, CBUAE guidance and FINMA expectations provide supporting governance references [16,29,35,37].

Records and explanations

The record should preserve the input, relevant source version, system configuration, output, reviewer action, final decision and reason at the level required for reconstruction. Explanation should be appropriate to the audience and decision. Technical interpretability does not automatically satisfy client or legal explanation requirements.

Prohibited and restricted uses

Each institution should define prohibited uses based on law, mandate and risk appetite. Restricted uses can operate only under stated conditions. Examples for management consideration include unapproved confidential-data entry, unsupervised legal or investment advice, autonomous transfer of value beyond limits, fabricated source attribution, covert client interaction, and use of personal data outside approved purpose. These are proposed policy examples; legal owners must establish the actual prohibitions.

Monitoring, Incidents And Assurance

Monitoring design

Monitoring should connect system behaviour to risk appetite and approval conditions.

Monitoring domainExample measureEscalation trigger
PerformanceMaterial support, omission and severe-defect ratesThreshold breach or adverse trend.
UseVolume, users, processes and out-of-scope attemptsUnapproved population or purpose.
AuthorityTool calls, limits, approvals and blocked actionsAny unauthorised or unexplained action.
DataFreshness, completeness, permissions and sensitive eventsStale critical source or access failure.
ChangeProvider, model, prompt, retrieval and configuration changesMaterial change without approval.
SecurityInjection, exfiltration, vulnerabilities and anomalous accessDefined severity event.
ConductComplaints, overrides, affected outcomes and explanationsMaterial client or rights impact.
ResilienceAvailability, latency, recovery and fallbackCritical-service limit breach.
VendorService, assurance, incidents and concentrationEvidence lapse or control deterioration.

Incident taxonomy

An AI incident can be a model, data, security, privacy, conduct, operational, legal or third-party event. The incident process should avoid routing everything to a specialist AI team. The severity assessment should consider actual and plausible consequence, exposure, containment, regulatory notification, affected parties and recurrence.

Incident stageRequired action
DetectPreserve alert, output, input, logs, versions and context.
ContainStop or restrict the use, tools, access or data flow.
AssessDetermine consequence, population, obligations and root causes.
NotifyFollow legal, regulatory, contractual and internal escalation.
RemediateCorrect outcome, system, control and affected records.
RecoverRestore within approved conditions and validate stability.
LearnUpdate evaluation, monitoring, training, risk and vendor management.

Control assurance

First-line testing should confirm controls operate. Second-line monitoring should challenge classification, approval, issues and aggregate risk. Internal audit should assess framework design and operating effectiveness using a risk-based plan. External assurance may support vendor or technical evidence, but the institution must assess scope, period, exclusions and relevance.

Aggregate risk

Portfolio reporting should identify concentration by provider, model, cloud, data source, business process, geography and authority level. Individually low-tier assistants can become material when they share a provider, process a large volume, or embed common error into multiple decisions.

Illustrative A2 Operating Case

Case definition

[Unverified] A family-office investment team is assumed to use a retrieval-augmented assistant to summarise private-market fund documents, compare material terms and draft investment-committee questions. The system has no authority to approve investments, communicate with managers or transmit instructions. All numerical values and operating assumptions in this section are unverified illustrations.

Item[Unverified] illustrative assumption
Users12 investment and operations professionals
Annual document sets180
Typical source length120-450 pages
OutputSource-linked summary, term table and question draft
AuthorityA1 draft only
DataConfidential manager and fund documents
ProviderHosted foundation model plus controlled retrieval
DecisionHuman investment team and investment committee retain authority

Classification

The proposed materiality assessment is Tier 3 because the use handles confidential data and can influence investment analysis, while output remains draft and reversible. A mandatory override would apply if the output entered committee records without review, generated legal conclusions, or initiated external communications.

Evaluation design

Test set[Unverified] illustrative sizeAcceptance focus
Routine fund documents120Material-term support and completeness
Complex structures40Waterfalls, side letters, conflicts and exceptions
Conflicting sources25Source hierarchy and escalation
Access-control cases25Retrieval isolation across entities and deals
Prompt-injection cases30Instruction separation and data protection
Historical documents20Version and effective-date handling

[Unverified] Illustrative proposed thresholds are 100 per cent support for defined material terms in the tested set, zero cross-deal access failures, zero autonomous actions, at least 98 per cent citation fidelity and documented escalation for every unresolved source conflict. These values are proposed management assumptions and require approval.

Control design

Control objectiveProposed control
Complete sourcesReconcile uploaded corpus to the authorised deal-room manifest.
Correct accessEnforce matter-level permissions before retrieval.
Traceable outputRequire clause-level citation for material terms.
Human authorityBlock external communication and record named review.
Legal boundaryRoute legal interpretation and unresolved terms to authorised counsel.
ChangeRegress after model, retrieval, prompt or provider change.
IncidentPreserve input, corpus, version, output, reviewer and downstream use.
ExitExport corpus metadata, evaluation set, logs and approved outputs.

Economic boundary

[Unverified] The use may release analyst capacity. Released time is not a cash saving unless an approved financial consequence is observed. Attributed Matchpoint or client revenue, cash cost reduction and loss reduction remain USD 0. A value ledger can record accepted outcomes, review time, rework and full operating cost without converting theoretical hours into realised benefit.

Illustrative A1 Operating Case

Case definition

[Unverified] An institutional allocator is assumed to use a predictive and generative system to monitor external managers, flag reporting anomalies and draft due-diligence follow-up questions. The model can prioritise review; it cannot change an allocation, issue a breach notice or communicate externally. All numerical inputs and thresholds in this section are unverified illustrations.

Item[Unverified] illustrative assumption
External managers85
Funds and mandates140
Reporting frequencyMonthly or quarterly by mandate
InputsManager reports, administrator data and internal exposure records
OutputRisk flags, supporting evidence and draft questions
AuthorityA2 recommend or rank
Material consequenceReview priority and potential escalation

Model and workflow boundary

The predictive component detects anomalies against approved features and benchmarks. The generative component explains the flag using controlled records and drafts questions. A deterministic rules layer enforces exposure limits and required escalations. The investment-risk owner decides whether to escalate. This separation allows different evaluation methods and reduces authority assigned to probabilistic output.

Evaluation

LayerProposed evidence
DataReconciliation to administrator and internal books; lateness and correction history.
PredictiveBack-testing, stability, sensitivity, benchmark, false-positive and missed-event review.
GenerativeClaim support, omission, source conflict, explanation and prohibited-content tests.
WorkflowCorrect routing, owner decision, override, service levels and record retention.
ResilienceProvider outage, stale data, fallback report and recovery test.
VendorChange notice, model dependency, audit evidence, subprocessors and exit.

[Unverified] An illustrative pilot uses 18 months of historical records, 60 adjudicated events and a three-month shadow period. These values are management assumptions. The institution would need to assess whether the sample supports the intended claims and material subgroups.

Decision and escalation

System outputHuman decisionProhibited system action
Low-confidence anomalyRequest analyst reviewSuppress or close without review
Supported material anomalyEscalate to investment riskNotify manager externally
Conflicting source dataRoute for reconciliationSelect one source silently
Limit-rule breachImmediate mandatory escalationAlter threshold or exposure record
Provider or model changeSuspend affected output pending regressionContinue under stale approval

Portfolio and third-party aggregation

The allocator should aggregate common providers across manager monitoring, research, reporting and client-service tools. A single foundation-model or cloud dependency may support multiple nominal vendors. The board pack should show this look-through concentration and the tested recovery path.

Economic boundary

[Unverified] No financial benefit is attributed in the illustrative case. A later evaluation could measure review throughput, accepted flags, material omissions, false escalation, time to resolution and operating cost. Revenue, cash-cost and loss claims require approved observed evidence and a defensible counterfactual.

Board Pack, Roadmap And Implementation

Board information

Board reporting should support direction and challenge. It should separate exposure from activity and report exceptions plainly.

Board viewDecision evidence
PortfolioUses by tier, authority, business, provider and status.
ConsequenceClient, capital, regulatory, rights and critical-service exposure.
AssuranceEvaluation current, overdue validation, audit and open findings.
Risk appetiteBreaches, exceptions, conditions and remediation.
IncidentsSeverity, affected population, cause, containment and recurrence.
ChangeMaterial provider, model, purpose and authority changes.
ConcentrationFoundation model, cloud, data and fourth-party dependency.
ResilienceFallback coverage, exit tests, recovery gaps and critical services.
PeopleCompetence, capacity, challenge and key-person dependency.
ValueApproved observed benefits, full cost and unverified hypotheses separated.

Twelve-month roadmap

PeriodCore outcomesGate
Days 0-30Sponsor, perimeter, interim prohibited uses, discovery and incident routeFramework charter approved.
Days 31-60Inventory, authority map, materiality method and obligations registerHigh-risk unknowns escalated.
Days 61-90Lifecycle, evidence templates, vendor controls and monitoring minimumPilot governance approved.
Months 4-6Evaluate material uses, remediate access, logging, data and contractsTier 3 and 4 decisions current.
Months 7-9Concentration analysis, resilience and exit tests, board reportingCritical dependencies tested.
Months 10-12Independent assurance, lessons, policy refresh and next-year planResidual risks accepted or remediated.

First 30 days

The first month should establish authority and visibility. Name the accountable executive, issue interim data and authority restrictions, create a single intake route, reconcile known applications and subscriptions, and connect AI events to existing incident management. Prioritise uses that can affect clients, capital, reporting, rights, confidential data or movement of value.

Days 31-90

Build the obligations register from applicable legal and supervisory sources. Complete classification for material uses. Create evaluation, approval, monitoring, change and retirement templates. Review critical vendors and embedded AI features. Establish an exception process with expiry and compensating controls.

Months four to twelve

Evaluate and approve or suspend material uses. Test access isolation, failure recovery and vendor exit. Reconcile provider concentration. Implement board metrics and independent assurance. Refresh the framework when the FSB consultation is finalised, NIST updates AI RMF, or jurisdictional rules change.

Control maturity model

LevelDescriptionEvidence
0 UncontrolledUses are unknown or unmanagedIncidents and ad hoc declarations.
1 VisibleInventory and interim restrictions existRegistered use and accountable owner.
2 DefinedClassification, lifecycle and standards existApproved policy, templates and roles.
3 OperatingMaterial controls are implemented and monitoredTests, approvals, logs and issues.
4 AssuredIndependent challenge and portfolio evidence operateValidation, audit, board challenge and remediation.
5 AdaptiveChanges, incidents and external developments update controls promptlyMeasured learning, timely refresh and tested resilience.

Limitations, Research Agenda And Conclusion

This paper synthesises cross-jurisdictional materials; it does not establish applicability for a particular entity. Requirements can overlap and change. The FSB sound practices were consultative at the evidence cut-off. The EU timetable changed shortly before publication. The US revised model-risk guidance expressly excludes generative and agentic AI from scope. Standards and voluntary frameworks do not establish legal compliance.

The operating framework requires empirical testing. Future research should compare classification consistency across institutions; measure validation reliability for generative and agentic workflows; evaluate whether citation and human-review controls reduce material errors; assess fourth-party concentration; and test exit plans under provider withdrawal, model change and cyber disruption. Institutions should publish or share incident taxonomies and evaluation methods where confidentiality permits.

The central conclusion is that controlled AI adoption is an evidence problem. A complete inventory shows exposure. Materiality and authority determine proportional control. Evaluation establishes bounded fitness. Governance assigns the decision. Monitoring keeps approval current. Incident and exit capability contain failure. For A2 family offices and A1 institutional allocators, this structure supports innovation while preserving investment authority, confidentiality, traceability and resilience.

Attributed Matchpoint or client revenue, cash cost reduction and loss reduction remain USD 0 until approved observed evidence supports attribution.

Source Register

The full paper records the evidence classification, scope and limitations applied to these sources.

  1. [1] Board of Governors of the Federal Reserve System, Federal Deposit Insurance Corporation and Office of the Comptroller of the Currency (2026). *SR 26-2: Revised Guidance on Model Risk Management*. Open source
  2. [2] Office of the Comptroller of the Currency (2026). *OCC Issues Updated Model Risk Management Guidance*. Open source
  3. [3] Federal Deposit Insurance Corporation (2026). *FIL-16-2026: Agencies Revise the Interagency Model Risk Management Guidance*. Open source
  4. [4] Prudential Regulation Authority (2026). *SS1/23: Model Risk Management Principles for Banks, April 2026*. Open source
  5. [5] Bank of England and Financial Conduct Authority (2023). *FS2/23: Artificial Intelligence and Machine Learning*. Open source
  6. [6] Bank of England and Financial Conduct Authority (2024). *Artificial Intelligence in UK Financial Services, 2024*. Open source
  7. [7] Bank of England and Financial Conduct Authority (2026). *The Bank of England and FCA's 2026 AI Survey*. Open source
  8. [8] Bank of England (2026). *Financial Stability Report, July 2026*. Open source
  9. [9] Bank of England, Financial Conduct Authority and HM Treasury (2026). *Joint Statement on Frontier AI Models and Cyber Resilience*. Open source
  10. [10] Financial Stability Board (2026). *Sound Practices for Financial Institutions' Responsible AI Adoption: Consultation Report*. Open source
  11. [11] Financial Stability Board (2024). *The Financial Stability Implications of Artificial Intelligence*. Open source
  12. [12] Basel Committee on Banking Supervision (2024). *Digitalisation of Finance*. Open source
  13. [13] Basel Committee on Banking Supervision (2021). *Principles for Operational Resilience*. Open source
  14. [14] Basel Committee on Banking Supervision (2013). *Principles for Effective Risk Data Aggregation and Risk Reporting*. Open source
  15. [15] Financial Stability Institute (2024). *Regulating AI in the Financial Sector: Recent Developments and Main Challenges*. Open source
  16. [16] Central Bank of the United Arab Emirates (2026). *Guidance Note on the Consumer Protection and Responsible Adoption and Use of Artificial Intelligence and Machine Learning by Licensed Financial Institutions in the U.A.E.* Open source
  17. [17] Central Bank of the UAE, Securities and Commodities Authority, DFSA and FSRA (2021). *Guidelines for Financial Institutions Adopting Enabling Technologies*. Open source
  18. [18] Dubai Financial Services Authority (2026). *DFSA Regulatory Expectations on Artificial Intelligence Risk Management in the DIFC*. Open source
  19. [19] Dubai Financial Services Authority (2025). *Artificial Intelligence Survey 2025*. Open source
  20. [20] Dubai Financial Services Authority (2025). *Cyber and Artificial Intelligence Risk in Financial Services: Strengthening Oversight Through International Dialogue*. Open source
  21. [21] Financial Services Regulatory Authority of Abu Dhabi Global Market (current at 2026). *Supplementary Guidance: Digital Investment Management*. Open source
  22. [22] Dubai International Financial Centre (current at 2026). *Data Protection Regulations; Regulation 10: Personal Data Processed through Autonomous and Semi-Autonomous Systems*. Open source
  23. [23] United Arab Emirates Government (2021). *Federal Decree-Law No. 45 of 2021 Regarding the Protection of Personal Data*. Open source
  24. [24] European Union (2024). *Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence*. Open source
  25. [25] European Commission (2026). *AI Omnibus Enters into Force*. Open source
  26. [26] European Commission (2026). *Guidelines on Transparency Obligations for Providers and Deployers of Certain AI Systems*. Open source
  27. [27] European Union (2022). *Regulation (EU) 2022/2554 on Digital Operational Resilience for the Financial Sector*. Open source
  28. [28] European Union (2016). *Regulation (EU) 2016/679, General Data Protection Regulation*. Open source
  29. [29] National Institute of Standards and Technology (2023). *Artificial Intelligence Risk Management Framework 1.0*. Open source
  30. [30] National Institute of Standards and Technology (2024). *Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1*. Open source
  31. [31] National Institute of Standards and Technology (2024). *Cybersecurity Framework 2.0*. Open source
  32. [32] UK National Cyber Security Centre and partners (2023). *Guidelines for Secure AI System Development*. Open source
  33. [33] International Organization for Standardization (2023). *ISO/IEC 42001:2023, Artificial Intelligence Management Systems*. Open source
  34. [34] International Organization for Standardization (2023). *ISO/IEC 23894:2023, Artificial Intelligence: Guidance on Risk Management*. Open source
  35. [35] Organisation for Economic Co-operation and Development (2024). *OECD AI Principles*. Open source
  36. [36] OWASP Foundation (2025). *OWASP Top 10 for LLM Applications v2.0*. Open source
  37. [37] Swiss Financial Market Supervisory Authority (2024). *FINMA Guidance 08/2024: Governance and Risk Management When Using Artificial Intelligence*. Open source
  38. [38] Swiss Financial Market Supervisory Authority (2025). *FINMA Survey: Artificial Intelligence Gaining Traction at Swiss Financial Institutions*. Open source
  39. [39] Hong Kong Monetary Authority (2024). *Research Paper on Generative Artificial Intelligence in the Financial Services Sector*. Open source
  40. [40] International Organization of Securities Commissions (2025). *AI Use Cases in Capital Markets*. Open source
Questions, answered

AI governance and model risk: frequently asked questions

The perimeter should cover models, AI systems, assistants, agents and embedded vendor capabilities that can affect decisions, records, clients, capital, rights, operations or resilience. Each registered use should link to the applicable legal, model-risk, data, security, conduct, outsourcing and operational-resilience controls.

The OCC stated that generative and agentic AI are outside the scope of the revised guidance because they are novel and rapidly evolving. Institutions should record the scope boundary and apply an adjacent AI-risk framework that retains useful model-risk disciplines where applicable.

A complete record should identify purpose, owner, users, provider, model and service components, data and rights, outputs, affected parties, authority level, jurisdictions, materiality, evaluation, approval, monitoring, incidents, changes, fallback and exit.

Materiality should be driven by decision consequence, authority, population and scale, data sensitivity, reversibility, explanation need, dependency and change velocity. Management-designed bands can support judgement, while legal rights, autonomous value movement and other high-consequence uses require mandatory escalation.

A family office can use an existing investment, risk or operating committee with a standing AI agenda, a named accountable executive, independent challenge and a complete evidence pack. The control objectives should remain intact while forums and documentation are scaled to the institution.

Evaluation should cover the configured system and workflow, including prompts, retrieval, tools, permissions, data, human checkpoints and recovery. Representative tests should include normal, difficult, adverse, sensitive, authority-boundary, temporal, subgroup and failure-recovery cases.

The institution should evidence service components, data terms, evaluation, security, resilience, auditability, legal rights, change notice, subcontractors and tested exit. Vendor assurance supports the decision; it does not transfer institutional accountability.

Board information should show uses by tier and authority, material consequences, current evaluation, open issues, risk-appetite exceptions, incidents, material changes, provider concentration, resilience, competence and finance-approved observed value.

The A2 and A1 worked cases use unverified illustrative management assumptions. Attributed Matchpoint or client revenue, cash cost reduction and loss reduction remain USD 0 because no approved observed attribution evidence was supplied.

This publication is general research for professional audiences. It is not investment, legal, regulatory, accounting, audit, tax, privacy, cybersecurity, employment, technology or valuation advice, and it is not an offer, solicitation, recommendation or promise of results. Readers should verify current requirements and decisions with qualified advisers.

Build an evidence-gated AI governance operating model

Discuss inventory, materiality, authority, evaluation, third-party controls, monitoring, assurance and implementation with a Matchpoint partner.

WhatsApp