1. Treat AI expenditure as a governed portfolio
An AI portfolio contains more than projects. It usually includes use cases that are expected to change revenue, cost, risk or service; shared data and integration capabilities; model and application subscriptions; cloud or data-centre capacity; specialist teams; control functions; and change activity required to alter how work is performed. Each category has a different economic life, dependency profile and degree of reversibility. A single innovation budget can hide these differences.
The board needs a view that reconciles commitments across accounting classifications. Some expenditure may be capitalised infrastructure, some may be operating expense, and some may be embedded in business-unit payroll or supplier contracts. The economic question is the total incremental cash and capacity committed to the portfolio, including operating cost after launch. An initiative can appear affordable during a pilot because enterprise integration, security, data remediation, process redesign, training, monitoring and run cost remain outside the pilot ledger.
Official disclosures show why commitment discipline matters. Alphabet reported USD 91.4 billion of 2025 capital expenditure, primarily technical infrastructure, and said it expected a significant increase in 2026 investment in servers, network equipment and data centres.[4] Microsoft's 2026 annual filing states that AI investments are being made at significant scale and ahead of fully developed revenue streams.[5] Meta reported USD 72.22 billion of 2025 capital expenditure including finance-lease principal payments and anticipated USD 115 billion to USD 135 billion in 2026 capital expenditure to support AI efforts and its core business.[6] Amazon reported USD 96.3 billion of cash capital expenditure in the first six months of 2026, mainly technology infrastructure supporting AWS growth and fulfilment capacity.[7] These figures have different definitions and mixed business purposes. Their relevance is the magnitude, long lead times and demand uncertainty described in the filings.
The portfolio should therefore disclose five totals: approved cumulative spend; contractual and lease commitments; current annual run rate; forecast operating cost at scale; and uncommitted capacity requests. It should also identify which use cases consume each shared capability. This prevents a platform from being approved as a strategic necessity while its dependent benefits remain unowned.
2. Give the board a defined governance mandate
Board oversight should focus on material exposure, portfolio coherence, risk appetite and evidence quality. It should not require directors to choose models or manage delivery sprints. Management remains responsible for proposing and executing investments. The board establishes which commitments require approval, what evidence must accompany them, how independent challenge operates and what happens when a use case misses its conditions.
Approval thresholds should combine financial materiality, risk and reversibility. A low-cost application affecting employment, credit, health, safety, critical infrastructure or sensitive personal data may deserve greater oversight than a larger productivity tool. A multi-year capacity reservation may require board attention because it reduces strategic flexibility even if near-term cash is modest. Thresholds should therefore consider cumulative spend, non-cancellable commitments, affected decisions and users, data sensitivity, autonomy, geographic reach and exit cost.
The board also needs an accountable executive. A chief executive, chief financial officer or named transformation sponsor should own portfolio value. Technology leaders own architecture and delivery evidence. Business owners own the operating baseline, adoption and realised benefits. Risk, security, privacy, legal and compliance leaders own independent standards within their mandates. Internal audit can test design and operation without becoming a first-line approver.
Table 1. AI portfolio decision rights and evidence ownership
| Decision or artefact | Accountable owner | Independent challenge | Minimum evidence for approval |
|---|---|---|---|
| portfolio appetite and annual envelope | board | finance, risk and internal audit | strategy, affordability, concentration, risk appetite and downside capacity |
| use-case problem and baseline | business executive | finance and process owner | current volume, time, quality, cost, risk and customer outcome |
| platform or infrastructure commitment | technology executive | architecture, procurement and finance | demand schedule, capacity utilisation, alternatives, contract terms and exit route |
| data readiness and rights | data owner | privacy, legal and security | access, quality, lineage, permitted use, retention, resilience and monitoring |
| model and control design | product or model owner | model risk, compliance and security | testing, human oversight, limits, incident response and residual risk |
| pilot release | delegated investment committee | finance and control functions | bounded scope, baseline, test plan, budget, stop rules and named users |
| production release and scale | executive committee or board by threshold | finance, risk and affected control owners | realised pilot evidence, adoption, control effectiveness, unit economics and capacity plan |
| benefits close-out | business owner and finance | portfolio office and internal audit by risk | verified benefit, total cost, residual obligations, lessons and asset disposition |
The allocation is a model. Actual responsibilities and delegations should follow the organisation's constitution, regulation and risk framework.
3. Make the use case the unit of value
A model, platform or data lake is rarely the direct unit of business value. Value appears when a defined group of users changes a decision or workflow and the change produces a measurable outcome. The use case should therefore specify the user, operating process, decision or task, AI contribution, human responsibility, baseline, target outcome and affected controls. This definition creates a line from technology to cash flow, service or risk.
The use-case boundary needs discipline. A claim such as "improve customer service with AI" is too broad for investment approval. A governed case might define assistance for a specified contact type, user population and channel; measure handling time, first-contact resolution, quality, complaints and escalation; identify the approved data and knowledge sources; and retain human authority for defined exceptions. Scope becomes testable and costs become attributable.
Shared capability should be allocated transparently. If ten use cases require a retrieval layer, identity controls and monitoring, the platform case should show minimum shared capacity, incremental capacity by use case and unused headroom. Allocating every shared cost to the first use case can make a viable capability appear uneconomic. Treating all shared cost as strategic overhead can make each use case appear artificially attractive. The portfolio should show economics both before and after a consistent allocation method.
Use cases also need stable identifiers and version control. A pilot that starts as drafting assistance may later generate customer communications or make recommendations. Its risk, data and benefit profiles have changed. The approval record should show the new boundary, cumulative commitment and control impact rather than allowing the initiative to retain its original low-risk label.
4. Require an investment-grade business case
An investment-grade case is a decision record that can withstand finance, operating, technology and risk challenge. It begins with the operating problem and business-as-usual counterfactual. It then compares options, identifies the preferred route, quantifies the full-life cost and benefit, demonstrates delivery and adoption, and states the controls required for use. Its depth should be proportionate to commitment and consequence.
HM Treasury's 2026 Green Book describes five connected dimensions: strategic, economic, commercial, financial and management. It also requires options appraisal and proportionality.[17] The related business-case guidance develops proposals through strategic outline, outline and full business cases.[18] These public-sector frameworks are not corporate approval rules. Their value for boards is the separation of strategic rationale, value, supplier arrangements, affordability and delivery evidence.
The counterfactual deserves particular attention. The case should compare the preferred AI option with business as usual, process redesign without AI, a rules-based or conventional analytics option, outsourcing, buying and building. If baseline performance is deteriorating, maintaining today's result may itself require investment. If an existing transformation already captures part of the benefit, the AI case should claim only the incremental effect.
The case must state uncertainty without hiding it in a single point estimate. Cost ranges should reflect estimate maturity. Benefit ranges should reflect measurement confidence and adoption. Scenario analysis should show demand, price, utilisation, quality, productivity and control outcomes. Dependencies on vendor availability, power, data rights, specialist skills and process change should remain explicit.
Table 2. Minimum evidence for an AI business case
| Business-case field | Required content | Evidence owner | Decision test |
|---|---|---|---|
| problem and counterfactual | process boundary, users, current performance, cause and business-as-usual path | business owner | baseline is current, attributable and decision-relevant |
| options | do minimum, process-only, buy, build, partner and defer routes | sponsor and architecture | alternatives use consistent assumptions and include exit implications |
| benefits | revenue, cash cost, capacity, quality, service and risk outcomes | business owner and finance | metric, baseline, timing, owner and verification method are defined |
| full-life cost | build, licence, compute, data, integration, controls, change, run and retirement | finance, technology and procurement | price date, volume driver, range and contractual basis are traceable |
| data readiness | access, quality, lineage, rights, representativeness, security and monitoring | data owner | limitations and remediation have cost, owner and deadline |
| control design | risk classification, testing, oversight, limits, logging, incident and exit controls | risk and control owners | residual risk is within approved appetite and evidence is reproducible |
| adoption and operating model | process change, roles, training, incentives, support and fallback | business owner | accountable users and minimum adoption threshold are named |
| delivery and dependencies | milestones, capacity, suppliers, integrations, decisions and critical path | programme and technology leads | commitments are matched to evidence and contingency |
| economics and scenarios | cash flows, allocation, discount rate, sensitivity, switching values and downside | finance | decision remains transparent under stated assumptions |
Evidence should mature with commitment. A discovery proposal can use bounded ranges; production and scale decisions require controlled, traceable and current inputs.
5. Segment the portfolio before ranking it
Portfolio comparison works only when unlike investments are identified. One segment may contain revenue use cases, another labour-capacity or cost cases, another risk and control applications, and another strategic capability whose direct benefits accrue through several downstream cases. Infrastructure and mandatory compliance work should remain visible as separate layers rather than competing on an unsupported common return measure.
The portfolio can also be segmented by maturity. Discovery validates the problem and available data. Experiment tests technical feasibility within a bounded environment. Pilot tests end-to-end process, user behaviour and controls. Production establishes a controlled service. Scale increases users, volume, geography or autonomy. Benefits review verifies realised outcomes and determines whether capability should be expanded, corrected, maintained or retired.
Risk segmentation should influence the approval path. NIST's AI Risk Management Framework organises activity around govern, map, measure and manage, with governance operating across the AI lifecycle.[10] Its generative AI profile identifies actions for risks that may be new or amplified by generative systems.[11] The GAO accountability framework groups practices around governance, data, performance and monitoring.[12] These frameworks support a common inventory while allowing controls to vary by use case, context and consequence.
The resulting map should show both strategic and financial exposure. A small experiment can remain in the discovery pool. A production service with recurring model and cloud cost belongs in the committed portfolio. A shared platform should show the use cases and demand that justify it. A mandatory control investment should show the exposure it mitigates and the operating consequences of delay.
6. Build a use-case portfolio map
The portfolio map compares expected economic value with execution confidence. Economic value can combine incremental cash flow, avoided cash cost, usable capacity, service improvement and risk reduction, with separate labels where monetary conversion is weak. Execution confidence combines problem clarity, benefit evidence, data readiness, control design, adoption and delivery capability. Bubble size can represent annual run-rate cost or cumulative commitment.
The map guides action. High-value, high-confidence cases are candidates for controlled scale. High-value, low-confidence cases deserve targeted evidence investment rather than automatic commitment. Lower-value, high-confidence cases may remain useful when they enable strategic capability, satisfy a duty or provide a fast route to adoption. Lower-value, low-confidence cases should normally stop, merge or return to problem definition.
A portfolio office should prevent the map becoming a beauty contest. Scores need anchored definitions, source dates and owners. The committee should see confidence ranges and red-line failures. It should also view concentration by vendor, model, cloud provider, data domain, business process and control function. Ten individually attractive use cases can create a material shared dependency.

Use cases, scores and annual run-rate costs are hypothetical modelling assumptions. Bubble size represents annual cost; the map supports governance and does not establish a forecast or recommendation.
[/FIGURE]
7. Advance benefits through a confidence ladder
Benefit confidence should be earned through evidence. A management hypothesis may justify discovery spend. It should not support a large production commitment. A controlled pilot can establish whether the technology changes task performance under defined conditions. Production evidence can show whether the wider process, users and controls sustain the result. Finance verification can confirm whether the change created cash, capacity, service or risk outcomes after costs and displacement effects.
The ladder avoids false precision. A benefit can have a high estimated amount and low confidence. Another can have a smaller amount and stronger evidence. The portfolio should report both. Confidence should fall when the use case changes population, data, model, process or scale because the earlier evidence may not transfer.
OECD research on AI adoption reports that uncertainty over return on investment and insufficient data maturity are important barriers, and that managers can underestimate enterprise-wide process and cultural implications.[8] A separate OECD review of experimental research finds that performance depends on the task and user experience, while long-term organisational evidence remains limited.[9] These findings support cautious transfer from individual-task experiments to enterprise cash-flow claims.
Finance should define the evidence required at each level. Time saved is not automatically cash saved. It may create capacity, improve service, reduce overtime, avoid hiring or permit headcount reduction; each route has different evidence and consequences. Revenue lift requires a credible control group or other attribution method. Risk reduction may require expected-loss or control-cost analysis and should preserve severe non-financial outcomes that are not suitable for simple monetisation.

The ladder defines evidence maturity. It does not assign a universal probability or valuation uplift to a confidence level.
[/FIGURE]
8. Establish the baseline and counterfactual
The baseline should record performance before the AI intervention. Depending on the use case, it may include transaction volume, handling time, rework, error, defect, loss, conversion, margin, complaint, cycle time, service level, control exceptions and employee effort. Source systems, measurement period, exclusions and seasonality should be documented. A baseline assembled after a successful pilot can introduce selection bias.
The counterfactual defines what would occur without the proposed investment. Business as usual may include existing productivity trends, contractual price changes, headcount plans, process improvement and other technology releases. The case should avoid claiming benefits already expected from these factors. Where a randomised control is impractical, a phased rollout, matched group, interrupted time series or defined pre-post comparison may provide useful evidence if limitations are stated.
Benefits need a chain from operational metric to financial statement or governed outcome. Minutes saved per task become annual capacity only after multiplying by eligible volume, adoption, sustained time reduction and productive redeployment. Capacity becomes avoided cost only when an approved hiring plan changes or a third-party expense is removed. Revenue becomes contribution after delivery cost, cannibalisation, discounting, churn and tax where relevant. The board should see each bridge.
The UK Digital and Data Benefits framework was developed to help quantify benefits from digital and data programmes and is intended to be used with Green Book guidance.[19] The Government Efficiency Framework also links identified savings to benefits management and wider governance arrangements.[20] Their public-sector context differs from a corporate portfolio. The general discipline remains useful: define benefit types, measurement methods, ownership and the relationship between operational change and value.
9. Make data readiness a capital gate
Data readiness has direct economic value because remediation consumes time, specialist effort and operating attention. It also affects performance and control risk. A use case should identify the data required, data owner, legal basis or contractual right, access route, quality, completeness, lineage, representativeness, retention, security classification, resilience and monitoring. Unknowns should carry cost and schedule consequences.
The readiness assessment should separate a dataset's existence from usable access. Data may be distributed across systems, locked in supplier contracts, available only through manual extraction or unsuitable for the required frequency. Labels may be inconsistent. Historical outcomes may reflect past policies that have changed. Training data may not represent current users or edge cases. A high-level claim that the organisation has data does not resolve these issues.
The heatmap should use anchored scores. A red score means a defined condition is absent or unacceptable. Amber means remediation is funded and bounded. Green means current evidence satisfies the gate. Grey means the dimension has not been assessed. Averages can hide red-line failures, so rights, security or prohibited-use issues should remain visible even when other dimensions are mature.
The use-case economics should include data remediation and ongoing control. A data product may support several use cases, allowing shared cost allocation. The board should see the use cases that justify the investment, the sequencing required and the risk that anticipated demand does not materialise.

Scores are hypothetical modelling assumptions: 1 means absent or unacceptable, 3 means bounded remediation, and 5 means gate-ready evidence.
[/FIGURE]
10. Reconcile full-life economics
The cost model should include discovery, build, integration, data, security, validation, licences, model inference, cloud or infrastructure, change, training, support, monitoring and retirement. It should also include internal people where they represent scarce capacity or incremental hiring. Commitments under minimum volumes, reserved capacity, leases and take-or-pay contracts should be shown separately from variable usage.
The benefit model should distinguish cash, capacity, revenue, service and risk. These categories should not be added into a single monetary total unless the conversion basis is credible. A case may show a positive cash return and a separate customer or control outcome. A mandatory compliance use case may have limited conventional return while preserving permission to operate. Its cost still competes for capital and delivery capacity.
The hypothetical example below uses a five-year nominal model, an 11 per cent discount rate, no terminal value and simplified tax treatment. It assumes that implementation costs occur before scale, benefits ramp with adoption, and annual run cost rises with volume. It excludes financing structure, foreign exchange, working capital, accounting capitalisation and interactions with other programmes. These are modelling assumptions, not market inputs.
Table 3. Hypothetical use-case economics from pilot to scale
| Economic component | Base case | Downside case | Evidence and decision implication |
|---|---|---|---|
| discovery, pilot and control build | 1.2 | 1.5 | approved scope and estimate basis before pilot |
| integration and data remediation | 2.8 | 4.2 | release in tranches as data gates are met |
| scale and process change | 2.0 | 3.0 | dependent on adoption and production evidence |
| annual run cost at maturity | 1.8 | 2.6 | volume, price and model-routing assumptions require monitoring |
| annual gross cash benefit at maturity | 5.4 | 2.9 | requires finance-verified staffing, supplier or contribution effect |
| five-year NPV | 5.7 | -2.6 | base case supports conditional scale; downside crosses the stop threshold |
| discounted payback | year 4 | not achieved | scale should remain staged until benefit confidence improves |
Values are hypothetical USD millions. NPV uses an 11 per cent nominal discount rate over five years with no terminal value; totals are rounded and do not represent a forecast.
Switching values improve the decision. The case should show the adoption level, unit price, error rate, benefit per transaction or volume at which NPV reaches zero or a board threshold. It should also show cash-at-risk before the next reversible decision. This focuses debate on the assumptions that determine value.
The portfolio view should avoid double counting. A shared data platform cannot claim the full benefit of every dependent use case while each use case also claims the same value. A productivity benefit cannot be counted as cost reduction, hiring avoidance and additional revenue unless the organisation has a credible allocation of released capacity. Finance should maintain a benefit register with one owner and one treatment for each outcome.
11. Price infrastructure, capacity and energy explicitly
AI economics can change with model choice, token or compute consumption, latency, context length, caching, utilisation, reserved capacity, storage, data transfer, hardware depreciation and power availability. The board does not need a technical tariff sheet. It needs the unit drivers, contractual commitments, sensitivity and mitigation.
The International Energy Agency reports that AI-focused data centres grew rapidly and that energy use per simple AI task has improved substantially while more intensive video, reasoning and agentic uses can consume far more energy.[3] It also reports large and uncertain growth in data-centre capacity and electricity demand. The implication for a corporate use case is that cost per task should be measured under the actual workload rather than assumed from a demonstration.
Capacity approval should follow demand evidence. An enterprise can reserve infrastructure for strategic reasons, supply protection or latency, but the case should show committed utilisation, option value and downside. Long-lead equipment, energy contracts and data-centre leases can reduce reversibility. Alphabet's filing, for example, describes multi-year data-centre projects and material technical-infrastructure commitments.[4] The fact pattern is company-specific; it illustrates how infrastructure economics extend beyond a short software pilot.
Model routing can preserve value. Lower-cost models may handle bounded tasks, with larger models used for exceptions. Retrieval, caching, prompt design and process constraints can reduce consumption. The portfolio should monitor cost per completed and accepted business outcome, not merely cost per request. Failed, repeated or manually corrected outputs belong in the denominator.
12. Design controls as part of the investment
Controls are product and operating requirements with cost, schedule and ownership. They should be designed with the use case rather than added after technical success. The control set may include approved purpose, access, data minimisation, testing, human review, segregation, output restrictions, logging, monitoring, incident response, supplier oversight, record retention, fallback and retirement.
Control proportionality matters. NIST's framework is voluntary and use-case agnostic, and it emphasises continuous risk management across the lifecycle.[10] ISO/IEC 42001 specifies an AI management system for organisations providing or using AI and uses a continual-improvement structure.[13] The Federal Reserve's revised 2026 model-risk guidance emphasises a risk-based approach tailored to a banking organisation's model profile, size and complexity.[14] The banking guidance applies within its stated regulatory scope. Its broader governance lesson is that validation, limitations, oversight and controls should be proportionate to material use.
Legal obligations require jurisdiction-specific analysis. The EU AI Act uses a risk-based structure and applies progressively; the European Commission reported that enforcement powers and certain transparency obligations applied from August 2026, with other high-risk provisions subject to later dates.[15][16] A board paper should state the organisation's role, location, use-case classification and legal advice. A generic "AI compliant" label is insufficient.
Table 4. Control design linked to investment evidence
| Risk domain | Minimum control evidence | Operating indicator | Gate consequence |
|---|---|---|---|
| purpose and accountability | approved use, owner, users, limits and prohibited actions | unauthorised-use events and scope changes | stop release until boundary and ownership are clear |
| data and privacy | source, rights, minimisation, retention, quality and access controls | exceptions, drift, access failure and deletion performance | hold or restrict population |
| performance and validation | representative test, acceptance criteria, error analysis and limitations | quality, false outcome, override and rework | revalidate after material model, data or process change |
| human oversight | decision authority, review point, competence and escalation | review rate, override quality and missed escalation | reduce autonomy or suspend use |
| security and resilience | threat model, supplier controls, testing, logging, fallback and recovery | incidents, latency, availability and recovery time | invoke fallback and incident authority |
| fairness and affected parties | impact assessment, subgroup testing, complaint and remedy route | outcome difference, complaints and remediation | narrow scope, correct or stop |
| supplier and concentration | due diligence, service levels, change notice, audit and exit rights | price, availability, model change and concentration | renegotiate, dual-source or cap exposure |
The table is a governance model and does not replace legal, regulatory, security, privacy or sector-specific requirements.
13. Use control gates to match authority with commitment
Gates should authorise a defined next commitment. Gate 0 accepts the problem into discovery. Gate 1 approves a bounded experiment after the baseline, data access and test plan are clear. Gate 2 approves a pilot when the process, users, controls and benefit method are ready. Gate 3 approves production release after end-to-end performance and control evidence passes. Gate 4 approves scale after adoption and unit economics are demonstrated. Gate 5 closes benefits and determines expansion, maintenance or retirement.
Each gate should have four outcomes: advance; advance with time-bound conditions; hold for evidence; or stop. Conditions require an owner, evidence item, due date and consequence. Repeated conditional approvals should be reported because they can convert missing evidence into permanent exposure.
Approval authority should increase with commitment and consequence. A delegated team may approve a small, isolated experiment. An investment committee may approve a production release. The board may approve a large infrastructure commitment, a material change in automation or a high-consequence use case. Emergency action remains possible through a defined exception and retrospective review.

Gate evidence is cumulative. Approval authority should reflect cumulative commitment, risk, reversibility and affected users.
[/FIGURE]
14. Test adoption before claiming scale value
Technical acceptance shows that a system can perform. Adoption evidence shows that people use it in the intended process and that the operating result persists. A scale case should define eligible users and transactions, active use, frequency, completion, acceptance, override, rework, fallback and support. It should also identify work that moved outside the measured system.
Adoption should be measured against a process denominator. The number of licensed users is weak evidence if only a small share of eligible work passes through the system. High usage can also be misleading if output is discarded or corrected. Accepted outcomes per eligible transaction provide a stronger bridge to value.
Process design and incentives matter. Users may bypass a slow or unreliable control. Managers may encourage usage without changing workload or targets. Employees may save time but receive more work, affecting quality and retention. A business case should state how roles, performance measures, training, supervision, support and workforce consultation change. Where employment rights or collective arrangements apply, specialist advice is required.
The pilot should represent the intended operating population. Enthusiastic volunteers, clean data and low volume can overstate scale performance. A production trial should include ordinary users, peak periods, exceptions and failure recovery. Adoption thresholds should therefore be paired with quality and control thresholds.
15. Make stop-or-scale decisions explicit
A stop-or-scale review combines benefit confidence, unit economics, adoption, data readiness, control effectiveness and delivery capacity. It should occur at a stated date or volume and before the next material commitment. The decision is a portfolio action, not a retrospective project score.
Scale can mean more users, transactions, countries, business units, data sources, autonomy or model capability. Each form changes risk and cost differently. Approval should specify the scale dimension and the evidence that remains valid. A use case that expands into a new legal jurisdiction or affected population may require reclassification and a fresh control review.
Stopping should have an operating plan. The team needs authority to terminate a contract, release reserved capacity, archive data, remove access, preserve records, notify users, revert the process and capture learning. A sunk-cost narrative should not keep an uneconomic or uncontrolled use case alive.
Table 5. Stop-or-scale decision matrix
| Decision dimension | Scale signal | Hold or correct signal | Stop signal |
|---|---|---|---|
| benefit confidence | controlled production effect and finance bridge | pilot effect exists but attribution or ramp is incomplete | no material effect under representative conditions |
| unit economics | accepted outcome cost within approved range | cost above range with bounded remediation | downside remains uneconomic at credible switching values |
| adoption | sustained eligible-process use above threshold | uneven use with accountable change plan | persistent bypass or rejection undermines benefit case |
| data readiness | gate-ready rights, quality, access and monitoring | funded remediation with time-bound conditions | unresolved rights, security or severe quality failure |
| controls | design and operating effectiveness evidenced | non-critical exceptions with owner and deadline | residual risk outside appetite or failed critical control |
| delivery and capacity | scale resources, suppliers and support committed | dependency remains conditional | critical capacity, supplier or operating model unavailable |
Thresholds are hypothetical examples. Each organisation should set materiality, risk and evidence thresholds within its governance framework.

Values and thresholds are hypothetical modelling assumptions. Status colours summarise underlying evidence and do not independently establish readiness or value.
[/FIGURE]
16. Control vendors, concentration and exit
Vendor selection can determine cost, performance, data use, auditability, service resilience and exit. The business case should distinguish model provider, application provider, cloud or infrastructure provider, integrator and data supplier. A single contract may combine several roles. Concentration should be assessed across the portfolio, because separate business units can create a large shared exposure without a central view.
Due diligence should cover service scope, price drivers, minimums, capacity, security, privacy, intellectual property, training-data use, model and feature changes, subcontractors, incident notification, data location, audit, regulatory cooperation, portability, transition and termination. Supplier claims should be tested against the intended use and internal evidence. A certification or standard can support due diligence while remaining insufficient for use-case acceptance.
Exit economics should be quantified. Data export, workflow rebuilding, user retraining, parallel operation and record retention can be material. A model change can alter performance without a conventional software release. Contracts and operations should therefore support notice, testing and rollback. Where dual sourcing is uneconomic, the board should see the accepted dependency and contingency.
The portfolio office should maintain a commitment calendar. It should show renewal dates, price reviews, minimum-volume steps, capacity reservations and termination windows. A stop decision made after an automatic renewal provides less economic protection than a review scheduled before the commitment date.
17. Give the board an evidence-led dashboard
The board dashboard should be compact and traceable. It needs portfolio totals, movement since the prior meeting, decisions required and material exceptions. It should show approved spend, commitments, annual run rate, forecast scale cost, benefit by confidence level, adoption, critical control failures, data gates, concentration and use cases approaching stop-or-scale review.
Aggregates require drill-down. A green portfolio average can hide a red high-consequence use case. The board should see exceptions by cumulative commitment and risk. It should also see ageing conditions, repeated gate deferrals and benefits that have remained at hypothesis or pilot level while spending continued.
Forecast and realised value should remain separate. The dashboard can show expected benefit weighted by confidence for scenario purposes, but this should not be reported as realised value. Finance-verified cash, capacity, revenue and service outcomes need their own lines. Avoided cost should identify the approved plan or contract that changed.
The committee pack should state material assumptions and changes. A lower model price, higher user volume, revised data licence or new regulation can change the case. Version control should preserve the basis approved at each gate and explain the current movement. This supports accountability without creating a static business case that becomes obsolete after launch.
18. Mobilise the governance system in ninety days
The first month should establish the inventory and decision boundary. Management should identify all AI use cases, platforms, infrastructure commitments, subscriptions, suppliers and data domains. Each item receives an owner, maturity, cumulative spend, run rate, next commitment date and risk classification. Finance reconciles the ledger and contracts to the inventory.
The second month should establish the common evidence standard. Business owners define baselines and benefit mechanisms. Technology and data teams assess architecture and readiness. Control functions define proportionate gates. Procurement identifies concentration and renewal exposure. The portfolio office maps dependencies and identifies cases whose next commitment should pause pending evidence.
The third month should run the first stop-or-scale cycle. Management selects the highest-exposure and nearest-commitment cases. The investment committee tests business cases, conditions and exit routes. The board approves the portfolio envelope, material exceptions and future reporting cadence. Teams then operate the system as a recurring decision process.
Table 6. Ninety-day AI capex governance mobilisation
| Period | Primary work | Required output | Board or committee decision |
|---|---|---|---|
| days 1-15 | inventory use cases, platforms, infrastructure, contracts and owners | controlled portfolio register with cumulative commitments | confirm scope, accountable executive and immediate red flags |
| days 16-30 | reconcile spend, run rate, renewals, capacity and shared dependencies | affordability and concentration baseline | pause unsupported material commitments where authorised |
| days 31-45 | define use-case, benefit, data and control evidence standards | business-case template and gate criteria | approve proportional paths and delegated limits |
| days 46-60 | assess highest-exposure cases and shared capabilities | portfolio map, confidence ladder and readiness heatmap | set remediation owners, dates and stop rules |
| days 61-75 | complete representative stop-or-scale reviews | decision papers with economics, adoption and controls | advance, condition, hold or stop each reviewed case |
| days 76-90 | establish dashboard, assurance and recurring calendar | board pack, benefits register and commitment calendar | approve portfolio envelope and reporting cadence |
The sequence is indicative. Timing should reflect portfolio scale, regulatory obligations, data availability and existing governance maturity.
The office should begin with material decisions, not a complete historical reconstruction. Lower-value experiments can be placed on a simplified path. High-commitment, high-consequence or near-renewal items deserve immediate attention. The inventory can mature while the governance system begins to protect the next decision.
Conclusion
AI capex governance connects technology progress to capital discipline. The board needs a reconciled portfolio of use cases, shared capabilities and commitments. Each material use case needs a defined operating problem, counterfactual, full-life economics, data readiness, control design, adoption plan and accountable benefit owner. Evidence should mature through gates before commitment increases.
The system described in this paper uses five linked instruments: a use-case portfolio map, benefit-confidence ladder, data-readiness heatmap, control gates and stop-or-scale dashboard. Together they show which cases deserve discovery funding, which need targeted remediation, which are ready for controlled scale and which should close. They also reveal shared infrastructure and supplier exposures that cannot be governed one experiment at a time.
Good governance preserves learning while protecting capital. A stopped case can produce useful evidence when commitments are bounded and lessons are captured. A scaled case earns greater authority through representative performance, user adoption, effective controls and finance-verified value. The board's role is to make those standards explicit and ensure that management applies them before the portfolio becomes difficult to reverse.
Appendix A. Board questions for a material AI commitment
- What operating problem and affected process are within scope?
- What happens under business as usual, and what non-AI options were tested?
- Which benefits are cash, capacity, revenue, service or risk outcomes?
- What is the current confidence level for each benefit and what evidence raises it?
- What are the full-life cash, capacity and contractual commitments?
- Which data rights, quality, lineage, security and monitoring conditions remain open?
- What human decisions, affected parties and jurisdictions are involved?
- Which controls are required before pilot, production and scale?
- What adoption and quality thresholds trigger scale, hold or stop?
- What is the next irreversible commitment and the exit plan before it?
Appendix B. Minimum stop-or-scale evidence pack
- Current use-case definition and version history.
- Baseline, counterfactual and benefit measurement method.
- Actual and forecast spend, commitments and annual run cost.
- Pilot or production results across representative users and conditions.
- Adoption, acceptance, rework, override and fallback evidence.
- Data-readiness status and open remediation.
- Control test results, incidents, limitations and residual risk acceptance.
- Supplier performance, concentration, renewal and exit position.
- Scenario economics, switching values and cash at risk before the next gate.
- Management recommendation with conditions, owner, date and consequence.
References
- Bank for International Settlements, Financing the AI Boom: From Cash Flows to Debt, BIS Bulletin No. 120, published 7 January 2026. https://www.bis.org/publ/bisbull120.htm
- Bank for International Settlements, The AI Investment Race, BIS research paper No. 1367, published 14 July 2026. https://www.bis.org/publ/work1367.htm
- International Energy Agency, Key Questions on Energy and AI, published 2026. https://www.iea.org/reports/key-questions-on-energy-and-ai
- Alphabet Inc., Annual Report on Form 10-K for the year ended 31 December 2025, filed 5 February 2026. https://www.sec.gov/Archives/edgar/data/1652044/000165204426000018/goog-20251231.htm
- Microsoft Corporation, Annual Report on Form 10-K for the year ended 30 June 2026, filed 30 July 2026. https://www.sec.gov/Archives/edgar/data/789019/000119312526323660/msft-20260630.htm
- Meta Platforms, Inc., Annual Report on Form 10-K for the year ended 31 December 2025, filed 29 January 2026. https://www.sec.gov/Archives/edgar/data/1326801/000162828026025534/meta-12312025x10kars.htm
- Amazon.com, Inc., Quarterly Report on Form 10-Q for the quarter ended 30 June 2026, filed 31 July 2026. https://www.sec.gov/Archives/edgar/data/1018724/000101872426000026/amzn-20260630.htm
- OECD, Boston Consulting Group and INSEAD, The Adoption of Artificial Intelligence in Firms: New Evidence for Policymaking, published 2 May 2025. https://doi.org/10.1787/f9ef33c3-en
- OECD, The Effects of Generative AI on Productivity, Innovation and Entrepreneurship, OECD Artificial Intelligence Papers No. 39, published 20 June 2025. https://doi.org/10.1787/b21df222-en
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework 1.0, published 26 January 2023. https://doi.org/10.6028/NIST.AI.100-1
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, published 26 July 2024. https://doi.org/10.6028/NIST.AI.600-1
- U.S. Government Accountability Office, Artificial Intelligence: An Accountability Framework for Federal Agencies and Other Entities, GAO-21-519SP, published 30 June 2021. https://www.gao.gov/products/gao-21-519sp
- International Organization for Standardization, ISO/IEC 42001:2023 Information Technology - Artificial Intelligence - Management System, published December 2023. https://www.iso.org/standard/81230.html
- Board of Governors of the Federal Reserve System, Revised Guidance on Model Risk Management, SR 26-2, published 17 April 2026. https://www.federalreserve.gov/supervisionreg/srletters/SR2602.htm
- European Commission, AI Act: Regulatory Framework for Artificial Intelligence, updated 3 August 2026. https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
- European Union, Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence, published 12 July 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
- HM Treasury, The Green Book 2026, published 5 February 2026. https://www.gov.uk/government/publications/the-green-book-appraisal-and-evaluation-in-central-government/the-green-book-2026
- HM Treasury, Guidance on Developing Business Cases, updated 30 June 2026. https://www.gov.uk/government/publications/guidance-on-developing-business-cases
- Department for Science, Innovation and Technology and Government Digital Service, Digital and Data Benefits Framework, published 7 April 2026. https://www.gov.uk/government/publications/digital-and-data-benefits-framework
- Cabinet Office and HM Treasury, The Government Efficiency Framework, published 2025. https://www.gov.uk/government/publications/the-government-efficiency-framework/the-government-efficiency-framework--2

