Introduction
Investment banking combines standardised production with concentrated professional judgement. A team repeatedly collects information, compares businesses, constructs financial models, prepares questions and assembles presentations. Each mandate also contains facts, conflicts, assumptions and positioning choices that require accountable judgement. Generative AI therefore enters a workflow with substantial technical potential and substantial control requirements.
Anthropic describes Claude for Financial Services as supporting financial analysis, due diligence, market research, modelling and presentation workflows through financial-data and internal-platform connections. Its finance-agent templates extend this description to pitchbooks, market research, due diligence and modelling. These publications establish product intent and workflow patterns. Customer statements in vendor publications remain customer-reported evidence and require attribution.
The paper asks which origination, diligence and pitch tasks suit retrieval, extraction, calculation, drafting, checking or connected execution; which controls preserve source traceability, confidentiality and professional accountability; and how a team should measure capacity effects when model output creates review and exception-handling work.
The controlled task
The unit of analysis is a task with a defined input, output, evidence standard, permission scope and approver. A transaction workflow can be decomposed into seven task classes: retrieval, extraction, calculation, drafting, checking, judgement and approval. Each class receives its own evidence requirement and control position.
A model can draft a comparable-company rationale. The sector lead owns the comparable set. A model can reconcile figures across a spreadsheet and presentation. The engagement lead owns the released valuation and client positioning. This task-level structure makes decision authority explicit.
Evidence Base and Propositions
Anthropic's public materials describe a connected environment for market data, internal knowledge and office-document workflows. The Claude API citations feature can tie response text to passages in user-provided documents. A citation establishes that a supplied passage was linked to an output. The reviewer still assesses authority, date, context and analytical relevance.
The Model Context Protocol distinguishes prompts, resources and tools. Resources supply context. Tools can retrieve information or perform actions. The specification recommends a human in the loop who can deny tool invocations and interfaces that reveal available and invoked tools. These distinctions support a risk-based permission model.
The Bank of England and Financial Conduct Authority's 2024 survey records broad AI use among responding financial-services firms and identifies data privacy, data quality, data security, third-party dependency and model-complexity risks. The Financial Stability Institute identifies governance, skills, model-risk management, data governance and third-party providers as areas requiring attention. NIST's Generative AI Profile identifies confabulation, privacy, information-security, intellectual-property and human-AI configuration risks.
Four operating propositions
Proposition 1. Source-bound tasks with explicit schemas and deterministic checks should reach an acceptable control state earlier than open-ended judgement tasks.
Proposition 2. Realised capacity depends on reviewer correction and exception-handling effort as well as gross generation speed.
Proposition 3. Connected tools increase utility and operational exposure together; least privilege and approval gates become more important as write authority expands.
Proposition 4. Adoption quality should become visible in source coverage, correction rates, exception closure, audit completeness and user adherence before headline productivity.
These propositions are operational hypotheses. The paper does not claim that they have been tested across a representative sample of banks.
The Controlled Reference Architecture
The reference architecture contains six layers. Identity and policy establish named users, role-based entitlements, mandate membership and information barriers. Approved sources include market-data services, public filings, internal research, CRM history and deal documents. A retrieval and citation layer provides bounded context with source identifiers. The model and orchestration layer contains approved models, task instructions and validation routines. Controlled tools cover spreadsheets, documents, presentations and approved internal systems. The evidence and approval layer retains logs, exceptions, reviewer decisions and released versions.
The information path should carry a mandate identifier, user identity, source identifier, document classification, task version and output status. A released output should be reproducible from an archived source set, a recorded instruction version and a named review record.
Source hierarchy
Level A evidence includes executed agreements, audited accounts, regulator filings and data obtained directly from an authorised provider. Level B includes management information and client-provided working data, labelled by provenance and review status. Level C includes reputable secondary research. Level D includes unverified web content or generated material. Decision-critical claims should trace to the highest available level. Generated text does not acquire source authority through repetition.
Four review gates
The input gate covers classification, entitlement, malware control and source registration. The analysis gate covers citation coverage, calculation reconciliation and exception review. The mandate gate covers conflicts, client facts, positioning and senior judgement. The release gate covers version control, disclosure, credentials, formatting and authorised delivery. An output remains a working draft until all applicable gates are recorded.
Origination Workflow
Origination begins with a defined thesis. The team specifies geography, sector, size, ownership, transaction trigger and exclusion criteria. Retrieval can collect source-bound facts from approved databases and filings. Extraction can place companies into a standard schema. Drafting can create short hypotheses for senior review.
A target record should separate verified facts from hypotheses. Verified fields can include legal name, reporting currency, filing date and disclosed ownership. Hypothesis fields can include likely capital need, strategic trigger and relevance to a buyer or investor. Every hypothesis should carry a rationale and review date.
Meeting preparation can retrieve public information, relationship history, recent correspondence and prior proposals from entitled systems. The resulting brief should label public facts, internal relationship notes and analyst hypotheses separately. The relationship owner approves the meeting objective, participants, sensitivities and outreach language.
A controlled pipeline can flag missing fields, overdue follow-ups and unsupported stage changes. It can prepare a daily exception list. The relationship owner remains accountable for client contact, stage assessment and forecast judgement.
Origination sequence
The redesigned sequence is thesis definition; approved-source retrieval; schema extraction; evidence and uncertainty review; senior hypothesis selection; conflicts process; authorised outreach; and CRM recording. The control objective is a visible boundary between collected fact and commercial hypothesis.
Diligence Workflow
Diligence begins with the document population. The system should inventory filenames, versions, dates, owners, page counts and document types before substantive extraction. Duplicate, corrupt, password-protected and scanned files enter an exception queue. The inventory creates a denominator for completeness.
Extraction should use a task-specific schema. A debt schedule can require facility, lender, currency, commitment, drawn amount, rate, maturity, security and change-of-control terms. Empty fields remain empty and become exceptions. The system should not complete missing facts through prediction.
Reconciliation presents the same fact across sources. Revenue can differ across audited accounts, management reporting and the operating model because of period, scope, accounting treatment or error. The workflow should display each value with source and date. The reviewer determines the explanation and materiality.
Missing, conflicting or unusual observations can become draft questions. Each question should carry the triggering evidence, issue category, materiality and owner. The deal team selects which questions go to the client or counterparty. The issue register records open, answered, partially answered and closed states with supporting evidence.
Prompt-injection control
File content should be treated as untrusted input. Connected implementations should isolate document content from system instructions, restrict tool authority and require approval for state-changing actions. Tool-call logs and destination restrictions help the team investigate anomalous behaviour.
Pitch and Execution Workflow
Financial work should separate data movement, formula construction, assumption selection and valuation judgement. Claude can assist with mapping source data, explaining formulas, generating checks and documenting assumptions. The approved spreadsheet remains subject to deterministic tests covering balance-sheet balance, cash-flow reconciliation, signs, units, period consistency, circularity and scenario boundaries.
Comparable-company analysis requires judgement in peer selection, metric definition and treatment of outliers. The model can assemble candidate data and draft a rationale. The sector lead approves the peer set, dates, adjustments and interpretation. Discounted cash-flow analysis also requires human approval of forecast logic, discount rate, terminal value and sensitivities.
A pitchbook should be generated from an approved fact and assumption register. Slide text can be drafted from that register with claim-level source links. The presentation workflow should check that figures, dates and units agree with the financial model and source register. Credentials, transaction claims and market statements require separate verification.
The release package contains the approved presentation, supporting model, source register, exception log and sign-off record. Draft outputs carry a visible status. External delivery permissions sit outside the model's default authority.
Productivity Economics
Productivity analysis should begin with observed task baselines. A team records elapsed and active hours for a representative sample, broken down by retrieval, extraction, calculation, drafting, checking, judgement, approval and rework. The AI-enabled pilot records the same categories plus prompt preparation, exception handling and model review.
The paper's numerical example is Matchpoint scenario analysis. It is illustrative and does not represent observed performance at Matchpoint, Anthropic or an investment bank. The fictional mandate contains 420 baseline team hours. The model limits eligibility to retrieval, structured extraction, first-draft preparation and defined checks. Judgement and approval hours are excluded from time-reduction assumptions.
The disclosed inputs produce 60.8 gross hours released. Incremental review and exception handling consume 38 hours. The resulting 22.8 realised hours equal 5.4 per cent of the 420-hour baseline. These values are model outputs from illustrative assumptions. They are not forecasts or measured results.
A live pilot should measure correction minutes per output, material error rate, exception volume, reviewer seniority, source coverage and audit completeness. Management can set an exit rule for each task and assign released analyst time deliberately to deeper client preparation, diligence, sector work or reduced peak-hour pressure.
Risk, Governance and the Ninety-Day Roadmap
The governance model assigns an accountable executive sponsor, workflow owner, technology owner, risk owner, data owners and named reviewers. The workflow owner defines intended use, excluded use, evidence standards and exit criteria. The technology owner maintains model, connector and instruction versions. The risk owner reviews controls and incidents. Data owners approve sources and access.
Core risks include confabulation, weak source authority, data leakage, prompt injection, spreadsheet error, permission escalation, model or vendor change, automation bias, record failure and third-party concentration. Preventive controls include source-bound tasks, source hierarchies, contract and configuration evidence, least privilege, scoped credentials, controlled templates and exit planning. Detective controls include source sampling, access and tool-call logs, reconciliation tests, correction analysis and audit-completeness checks.
Days 1 to 20: govern and map
Select one bounded workflow and define intended use, exclusions, users, sources, permissions, outputs and review rules. Establish a benchmark from completed or public-material cases. Record contract, retention and security evidence.
Days 21 to 45: read-only pilot
Use read-only sources and create working drafts in a segregated environment. Suitable tasks include document inventory, schema extraction, evidence-linked summaries and draft question lists. Review every output and record source coverage, material errors, correction time, exceptions and user adherence.
Days 46 to 70: governed production
Move approved tasks into a controlled live workflow. Produce evidence objects, exception registers and versioned outputs. Spreadsheet and presentation assistance can enter scope where formula and release controls exist.
Days 71 to 90: connected tools
Consider tightly scoped write actions. Give each action an allowlist, destination restriction, confirmation rule and log. Keep external communications and publication behind explicit human release. Rerun benchmarks after model or connector changes.
Conclusion
Claude can support investment-banking origination, diligence and pitch production through retrieval, structured extraction, calculation assistance, evidence-linked drafting and defined checks. Current Anthropic materials describe these workflow capabilities. Current financial-sector and NIST sources identify data, model, governance, security and third-party risks that require active management.
The controlled task provides a practical adoption unit. Each task receives approved inputs, bounded permissions, an evidence standard, deterministic checks where available, exception handling and a named approver. Judgement and release remain accountable human actions.
The implementation sequence begins with read-only, source-bound work. Measurement focuses on source coverage, correction effort, material errors, exceptions, audit completeness and user adherence. Connected actions enter scope after the control evidence supports them. This sequence gives management a traceable basis for deciding which workflows to expand, restrict or retire.


