On April 17, 2026, the Federal Reserve, OCC, and FDIC issued SR 26-2 — the first rewrite of model risk management guidance since SR 11-7 in 2011. It superseded fifteen years of supervisory expectations. It clarified scope. It introduced a risk-based, proportional approach. And buried inside its scope statement was a single sentence that is already reshaping how financial institutions think about AI governance.

The sentence, confirmed by Federal Reserve Governor Michelle Bowman on May 1, 2026: SR 26-2 “does not apply to generative or agentic AI.” The revised guidance “now applies narrowly to traditional models and basic AI applications.”

That sentence is the SR 26-2 gap.

For CISOs and model risk officers whose institutions are deploying LLMs in customer service, underwriting support, fraud detection, and document processing, this guidance change created a governance condition that most programs have not fully absorbed: the AI systems generating the most institutional exposure are now the ones your model risk framework does not cover.

What SR 26-2 Actually Changed

SR 26-2 did not start from scratch. The Federal Reserve’s interagency guidance retained the core governance pillars that financial institutions have built programs around since 2011: model inventory, independent validation, conceptual soundness review, ongoing monitoring, governance policies and controls, and an effective challenge process. Banks with mature model risk programs did not need to rebuild — they needed to read the scope statement carefully.

The key change was architectural. SR 26-2 moved model risk management from a single prescriptive framework covering all quantitative systems to a risk-based, proportional approach explicitly tailored to an institution’s model risk profile, size, and operational complexity. It is principles-based guidance, not a rulebook. The agencies themselves stated that it “emphasizes a risk-based approach to model risk management that is tailored to a banking organization’s model risk profile” and that practices “appropriately vary among banking organizations based on their specific risk profiles and model usage.”

That principles-based framing matters, because it means the governance bar is not a checklist — it is a demonstrated alignment with sound risk management practice. Examiners assess whether programs reflect sound principles, not whether specific artifacts were produced. A program that ignores those principles entirely, as The Data Experts analysis of SR 26-2 notes, “may draw questions about the adequacy of its model risk management.”

SR 26-2 also significantly expanded expectations around vendor and third-party model risk. Banks relying on externally sourced models — which now includes foundation model APIs from providers such as OpenAI, Anthropic, Google, and AWS — face enhanced due diligence, contracting, and ongoing monitoring requirements for those traditional models. The irony is not lost: third-party risk expectations for SR 26-2-covered systems got stricter at precisely the moment the agency excluded the largest new category of third-party AI from coverage.

One other structural note: SR 26-2 is explicitly most relevant to banking organizations with over $30 billion in total assets. Smaller institutions have proportional applicability where model risk exposure is material. The governance obligation scales with exposure, not with the size of the guidance document.

The Carve-Out, Explained

The agencies characterized generative and agentic AI systems as “novel and rapidly evolving” and stated that other risk-management practices will govern them. That characterization is accurate — these systems are architecturally different from the statistical and econometric models SR 11-7 was designed to govern. They are non-deterministic. Their outputs are not reproducible from fixed inputs. They can take multi-step actions, call external tools, and modify their own behavior based on context retrieved at runtime. Applying model validation standards designed for a credit scoring model to a customer-service LLM produces governance theater, not governance.

The agencies recognized that problem and responded by moving generative and agentic AI outside the model risk guidance perimeter entirely. The Fed, OCC, and FDIC have announced intent to issue a dedicated request for information addressing AI — including generative AI, agentic AI, and AI-based models — as a future issuance. That RFI had not been formally published as of the date of this post.

What banks cannot do is treat this carve-out as permission to govern these systems less rigorously. As The Data Experts stated in their June 21, 2026 analysis: “One critical point: the carve-out is not a regulatory safe harbor.”

SR 26-2 excluded generative and agentic AI from model risk guidance. It did not exclude them from regulatory examination. Consumer protection rules, fair lending laws, and third-party risk guidance have no generative AI exception. An institution that deploys an LLM in a lending workflow and produces discriminatory outputs does not escape ECOA examination because the system was out of SR 26-2 scope. An institution that deploys a customer-facing AI agent and suffers a data breach does not escape GLBA scrutiny because the system was characterized as “novel and rapidly evolving.”

The governance obligation exists. What is missing is the prescribed framework defining how to satisfy it.

Why the Gap Creates Material Audit Risk

The governance gap is not theoretical. It is operational. And it shows up most clearly when an examiner asks a direct question: which governance authority applies to your generative AI systems?

A bank that has quietly assumed its LLM customer service platform was covered by the model risk program — because it used to debate that internally — is now, in The Data Experts’ precise framing, “governing a system the primary guidance has explicitly set aside, and is also leaving itself without a documented framework when an examiner asks which governance authority applies.”

That is the audit exposure. Not non-compliance with SR 26-2 — the guidance itself states that non-compliance does not result in supervisory criticism. The exposure is the inability to answer the examiner’s question with a documented, defensible governance framework. Examiners may not cite SR 26-2 violations for generative AI systems. They can, and will, ask whether any governance framework applies, and whether the institution can demonstrate it.

The risk surfaces in five specific areas where most financial institutions’ current governance programs are operating without coverage:

  1. The inventory problem. SR 26-2 maintains a model inventory requirement for covered systems. Generative AI deployments — customer service chatbots, document summarization tools, underwriting assistants, fraud narrative generators — are frequently not in that inventory and are not subject to an alternative inventory process. The institution cannot govern what it has not catalogued.
  2. The validation problem. Independent validation under SR 26-2 tests conceptual soundness and outcomes analysis. LLMs and agentic systems cannot be validated using the same methodology — their outputs are non-deterministic, their “decisions” are not traceable to discrete model logic, and their behavior changes with context. Most institutions have not built an alternative validation process for these systems.
  3. The evidence problem. When an examiner, regulator, or litigant asks for the documentation proving that an AI system operated within its intended parameters on a specific date, most financial institutions cannot produce it. Traditional model monitoring generates performance metrics. It does not generate the prompt-level, session-level audit trail that proves an LLM behaved appropriately at the moment a specific customer interaction occurred.
  4. The agent-specific problem. AI agents — systems that take multi-step actions, call external tools, access data sources, and execute workflows — introduce risks that go beyond what a language model poses on its own. A prompt injection attack embedded in an external data source can redirect an agent’s behavior without any change to the system’s configuration. As a June 2026 case study published in Developers Digest documented, security researchers at Blue41 demonstrated that a €0.02 SEPA transfer carrying a malicious payload in the free-text description field was sufficient to manipulate a banking AI assistant into generating a realistic phishing message inside the bank’s own interface — using real customer account data to make it more credible than any external phishing email. The attack required zero interaction with the agent’s code. It required only that the agent retrieve a data source the attacker had corrupted. “The vulnerability is not exotic,” the analysis noted. “It is one of the most predictable failure modes in agent architecture.”
  5. The ownership problem. SR 26-2’s governance requirements include a documented owner for every covered system — the person accountable for the model’s performance and appropriate use. Generative AI systems frequently have no equivalent. They are deployed by business units, supported by IT, and owned by nobody in the governance sense. When something goes wrong, the accountability gap becomes an examination finding.

What Sound Governance Looks Like in the Absence of a Framework

Banks cannot wait for the GenAI RFI. The regulatory environment is moving in real time. Fair lending examinations, UDAP/UDAAP scrutiny, third-party risk reviews, and consumer complaint investigations are all live channels through which AI governance gaps surface today.

Sound governance in the absence of a prescribed framework starts with the five components SR 26-2 would require if it applied — adapted for systems it does not cover.

Inventory: A documented registry of all generative AI and agentic AI deployments, separate from the SR 26-2 model inventory. Each entry should include the system’s intended use, data sources accessed, outputs produced, and a named human owner. This is the foundation. Without it, none of the other governance components are credible.

Risk classification: Each system classified by use case risk — customer-facing, credit-relevant, data-intensive, agent-enabled — with the classification driving the depth of governance applied. A low-risk document summarization tool and a high-risk lending recommendation tool should not be governed at the same depth. Risk-proportional governance is the principle SR 26-2 itself establishes for traditional models.

Testing documentation: Testing appropriate to the system type. For LLMs: adversarial prompt testing, output boundary testing, bias assessment across relevant demographic groups. For AI agents: injection resistance testing, tool call boundary testing, and behavioral profiling under adversarial input conditions. The documentation proves the testing happened and what it found.

Monitoring and evidence: Continuous monitoring with session-level logging sufficient to reconstruct what the system did at any point in time. This is the governance component most institutions are currently missing and most regulators will eventually require. The audit trail is not a performance dashboard. It is the document that answers the examiner’s question about a specific interaction.

Ownership and accountability: A named human owner for every system — the person who approved its deployment, understands its risk profile, receives its monitoring reports, and has documented authority to suspend it. This is the human accountability layer that connects AI governance to the institution’s existing accountability structures.

For institutions asking how to operationalize this governance in practice, the frameworks that most naturally bridge from SR 26-2 into generative AI territory are the NIST AI Risk Management Framework and ISO/IEC 42001. Neither is a perfect fit. Both provide more structure than a blank page. Both are frameworks that examiners recognize and that institutions can point to when asked which governance authority applies.

See also: [Closing the AI Proof Gap: Shadow AI Governance Blueprint] for a detailed treatment of how to build the evidence layer for shadow AI deployments.

The External Enforcement Boundary: How Purpose-Built Tooling Closes the Audit Trail Gap

The governance components described above can be built internally, but one component consistently exceeds what internal engineering can deliver at acceptable cost and speed: the session-level audit trail for LLM and agent interactions.

The fundamental problem is architectural. Traditional application logging captures what a user requested and what the system returned. It does not capture what the LLM reasoned over, which version of the model processed the request, what data entered the context window, whether any policy constraint was triggered, or what the model would have returned if given different input. For audit purposes — and for the forensic work that follows an incident or examination — that gap is consequential.

A category of purpose-built tooling operates specifically at this layer: between the institution’s applications and the LLM, intercepting every prompt and response, applying policy controls in real time, and generating the immutable session record that an audit trail requires.

Two vendors working at this layer are representative of the category:

Prompt Security operates as an AI security platform positioned between enterprise applications and the LLMs they call. The platform defends against prompt injection — including indirect injection of the type demonstrated in the Bunq case — data leakage, and harmful LLM outputs. It provides real-time risk assessment and enforcement for agentic AI systems, which the company describes as “demanding real-time, machine-level security for visibility, risk assessment, and enforcement beyond traditional analysis boundaries.” Prompt Security’s deployment at 10x Banking, a UK digital banking infrastructure company, provides a direct financial services reference point. The 10x Banking CISO noted the platform “allowed us to be enablers for the business — we’re able to expand AI use across the company while keeping it safe and proportionate.” The platform supports cloud and self-hosted deployment, which matters for institutions with data residency requirements.

SafePrompts.ai (TVR Labs) operates at the prompt and session execution layer with an explicit focus on the audit trail gap that SR 26-2 leaves unaddressed. The platform generates session-level logs covering what each AI system sent, received, what data entered the context window, and whether any policy constraint was triggered — producing the immutable, reconstructable record that satisfies an audit requirement. For institutions building toward examination readiness, the SafePrompts.ai architecture addresses the evidence problem directly: not by describing what the system should do, but by documenting what it actually did, at the session level, for every interaction. The platform applies to LLM deployments and AI agent deployments, covering both the conversational AI surface and the agentic workflow surface where the audit trail gap is most acute.

Both platforms address what internal logging cannot: the evidence that proves a system operated within its intended parameters at a specific moment in time, with a specific version of the model, in response to specific input. That evidence is what transforms a governance program from a policy document into an audit-defensible record.

Neither platform replaces the governance framework. An institution without an AI inventory, risk classification, ownership documentation, and testing records has a governance problem that tooling cannot solve. But an institution with sound governance infrastructure and no audit trail is producing documentation that describes its program and cannot prove its program operated as described. The tooling closes that specific gap.

See also: [Zero Trust for AI Agents: A 2026 Blueprint] for the architectural foundation that external enforcement boundary tools operate within.

What CISOs and Model Risk Officers Should Do Before the RFI Arrives

The Fed, OCC, and FDIC have signaled that dedicated generative AI guidance is in development. When the RFI is published, it will ask financial institutions to describe their current practices. The institutions that respond from a documented program — with an AI inventory, a governance framework mapped to a recognized standard, testing records, and an operational audit trail — will be in a materially different position than those that respond from a blank page.

The practical sequence, starting now:

Build the GenAI inventory first. Separate from the SR 26-2 model inventory, document every generative AI and agentic AI deployment: what it does, who deployed it, who owns it, what data it accesses, and what outputs it produces. This is the governance foundation. Without it, risk classification, testing, and monitoring have no anchor.

Establish the governance authority. Map each system to a recognized framework — NIST AI RMF, ISO 42001, or a documented internal framework that the institution can defend. The answer to “which governance authority applies” cannot be “our model risk program” for generative AI systems, and it cannot be “none.” Document the answer before the examiner asks.

Implement testing appropriate to the system type. For each system in the inventory: adversarial prompt testing, output boundary documentation, bias assessment where the system touches credit or customer decisions. For AI agents: injection resistance testing and behavioral profiling. Testing that is not documented did not happen for governance purposes.

Close the audit trail gap. Implement session-level logging for every LLM and agent deployment. Whether through purpose-built tooling or internal engineering, the audit trail must be capable of answering the examiner’s question about any specific interaction. Aggregate monitoring dashboards are not sufficient.

Assign human ownership. Every system in the inventory needs a named human owner with documented accountability. That owner approves the system’s deployment, receives monitoring reports, and has authority to suspend the system. This is the accountability layer that connects AI governance to the institution’s broader risk governance structures.

The SR 26-2 gap is real. The governance obligation it leaves unresolved is also real. The institutions building their GenAI governance programs now — before the RFI, before an incident, before an examination — are building the evidentiary foundation that will define their regulatory posture for the next several years.

The guidance that would have told them exactly how to do it has not been published. That is the gap. The response to a gap is not to wait. It is to govern.

TechVision Research publishes practitioner-grade analysis for CISOs, model risk officers, and enterprise security architects.

See also:

We can help

If you want to find out more detail, we're happy to help. Just give us your business email so that we can start a conversation.

Thanks, we'll be in touch!

Subscribe

Join our mailing list to receive the latest announcements and offers.

You have Successfully Subscribed!

Stay in the know!

Keep informed of new speakers, topics, and activities as they are added. By registering now you are not making a firm commitment to attend.

Congrats! We'll be sending you updates on the progress of the conference.