It’s hard to open a newspaper or read a tech article today without seeing at least some mention of large language models (LLMs) and their usage in just about any industry. The auto lending sector is no exception, with LLMs and artificial intelligence in general remaining a hot topic for lenders and consumers alike.

In fact, over the past 14 months, the guidelines for handling AI-related adverse actions were changed twice, even though the underlying legislation has not changed.

In 2022 and again in 2023, the CFPB established that companies are not exempt from their responsibilities because they use a complex algorithm they may not be able to understand or explain. The CFPB also clarified that a company citing a broad list of reasons also fails to meet Regulation B’s requirements.

To clarify ongoing confusion, on May 12, 2025, the CFPB removed 67 guidance documents, which the bureau’s leadership positioned as an effort to align policy with regulatory language. The bottom line for lenders, however, is that even after the CFPB effectively amended Regulation B, the Equal Credit Opportunity Act’s legal requirements remain in effect.

The regulatory framework is the starting point for any lender that is considering where LLM technology belongs in an underwriting workflow.

The industry’s default instinct points the model at the wrong layer.

While LLMs are technological marvels, they do not, in fact, “reason.” An LLM writing a well-organized letter in seconds is using probability to place words in a specific order because training on millions of word patterns taught it to do so, not because the LLM actually understands what the words it just used actually mean.

Whether a lender uses technologies like LLMs to accelerate underwriting, speed up adverse-action letter drafting, or write better credit-committee memos faster, LLMs can deliver on the promise of speed, but it’s important to understand that there are limitations.

A lender using an LLM to write an adverse-action letter, for example, is not actually using the LLM to “reason” through the process. Because LLMs can create plausible, human-sounding language patterns, well-written LLM-generated text can be misinterpreted as “reasoned” thought when the text actually had no basis in reason at all. An adverse-action letter that describes a reason for a denial, but that is not based on a verifiable, specific score, does not satisfy the requirements of Regulation B.

In fact, riskier, more confident language might slip past a reviewer who could mistake well-written wording for sound, data-backed decision-making, when the model may not have used any data for the decision at all.

What the separation actually looks like in production

Solving how to benefit from LLM technology means knowing where it fits best within the architecture lenders already have in place. Deterministic, traceable scoring models are already widely used to generate decisions and to provide specific reason codes for each decision. The output from these “risk” models must meet the strict requirements of Regulation B. LLMs, by contrast, sit further downstream and translate reason codes into plain language for consumer letters, credit committee memos, and examiner packets.

The scenario described earlier, however, oversimplifies where lending is headed. LLM technology can go far beyond just “writing content” and is already being used in the form of agents, such as automatic decisioning systems, that carry out complex processes in various situations.

In auto lending, this means deploying multi-agent pipelines where agents handle data collection, credit scoring, and result interpretation. These multi-agent workflows, however, should be orchestrated to route exceptions to humans rather than relying on automated resolution, creating clear boundaries for each agent’s role and providing a well-defined audit trail.

Having clear boundaries and auditabillity for agents is extremely important. Decision pipelines that are built without guardrails can result in a series of black-box decisions that become even more problematic, complicating compliance and making it harder for examiners to walk through the decision-making process during an audit. The solution is not to avoid a multi-agent process completely, but instead to build explainability and traceability into the workflow from the start, something already in place in anti-money laundering and know-your-customer workflows.

Validating the dual-layer pipeline: An MRM approach

For risk officers, the primary concern with LLM adoption is the auditability of the decision logic. In a traditional risk model, Model Risk Management (MRM) is well-defined where input data is validated, a variable is selected, the output is weighted, and a final score is calculated.

The LLM introduces a “probabilistic” element that standard MRM workflows cannot accommodate.

To maintain compliance, lenders should decouple the validation process, making sure that the model undergoes traditional, rigid validation for bias, variable stability, and statistical significance. In contrast, LLM validation should focus on maintaining semantic integrity. “Compliance teams should implement rationale audits as a critical step, where reasons provided by the LLM are verified against pre-existing, easily traceable variables that were generated by a risk model.

By restricting the LLM to a well-defined library of “truth-aligned” language patterns, lenders can reduce the risk of LLMs hallucinating a justification or inventing reasons that never existed in the underlying data.

Where regulators are drawing the newest line

The clearest current signal on how supervisors are thinking is in SR 26-2. The letter places generative and agentic AI outside its formal scope as a category still evolving too quickly for a settled framework, while directing institutions to apply their existing risk-management principles — materiality, ongoing monitoring, independent challenge — to those systems regardless.

A framework that names agentic AI explicitly, only to say the rules are not in writing yet, simply tells institutions that the underlying discipline still applies to each agent in the pipeline, ahead of paperwork and regulations catching up. The compliance risk in a multi-agent underwriting system lies in the risk that decisioning authority can leak into the orchestration layer itself, through a series of individually defensible steps, with no single agent accountable for the overall outcome, but complicity in some way.

Instead of just calling AI risks a mystery or a ‘black box,’ a new risk guide gives a clearer name to a specific danger: communication breakdowns and domino-effect failures across networks of connected AI tools. This description is much more accurate and gives compliance teams a practical map of exactly what to look for and fix.

The next frontier of lending architecture is the multi-agent pipeline—where specific agents handle data retrieval, scoring, interpretation, and exception routing. While this structure offers immense operational efficiency, it complicates the traditional “single model” view of risk.

Guardrails for multi-agent orchestration

The framework for these pipelines is called the “deterministic hand-off.” Within an agentic system, every transition between agents must be secured by a deterministic gate. For example, when an interpretation agent creates an adverse action letter, it should not be able to finalize or send that letter. A checkpoint must verify that the reason codes the scoring agent generates match the codes in the interpretation the agent received.

Any mismatches should automatically trigger an automated human validation loop. Enforcing “Human-in-the-Loop” (HITL) gates at every point where the system moves from data processing to customer-facing output ensures no agentic decision is finalized without rigid, deterministic approval. This strategy changes the multi-agent design from a “black box” of interacting agents into a transparent, verifiable chain of command.

Navigating the regulatory ‘double jeopardy’

Regulatory uncertainty around AI in lending is often seen as a barrier to innovation, but for sophisticated institutions, it calls for operational discipline. As agencies like the CFPB and the Federal Reserve refine their stance — shifting toward process-based audits rather than model-based audits — the focus is moving toward “coordination failure.”

Recent AI risk taxonomies like the Multi-Agent Risk from Advanced AI report (MARAI) highlight that failure modes in multi-agent systems are rarely a single “rogue” model, but involve a series of failures in which multiple agents interact in unexpected ways, creating a “cascading failure” across the network.

Examiners are increasingly trained to look for these systemic failures and are moving away from asking “what does this one model do?” to “what happens at the interface between your agents?” Compliance teams must adapt by documenting the “interface risk” between agents.

Showing regulators that you have identified failure modes at these critical junctions is the sign of a mature, compliant, and defensible AI governance framework.

A composite deployment that gets the boundary right

Assume you are a mid-sized lender and you have deployed an LLM for two reasons. First, to draft adverse-action letters based on reason codes produced by risk scoring models, and second, to assemble documentation from the same risk model meant for examiners. The scoring layer remains separate from the LLM layer and untouched so that it can be audited independently of any LLM layer.

Any decision based on this approach does not change simply because an LLM wrote a better letter. In this example, the LLM component only accelerates the process and clarifies the explanation drafted for the ultimate recipient.

Once interpretation and decisioning are architecturally separated, the tension in the premise dissolves. There’s no real conflict between the operational efficiency generative AI offers and the specificity Regulation B requires. The conflict exists only for institutions that let fluent output stand in for traceable output.

The guidance will likely change again before this decade is out; it has already changed twice in three years. Institutions that keep blurring the line between an interpretation layer and a decisioning layer — in a single model or across a whole agent pipeline — are building their next adverse-action enforcement exposure directly into the workflow they believed was solving compliance, not creating it.

Joshua Gordon is a data scientist specializing in statistical modeling, machine learning, and deep learning technologies. While earning his Ph.D. in statistics from the University of California, Los Angeles, he published six articles focused on geospatial time series algorithms for earthquake prediction as well as residual analysis for spatial-temporal models. Joshua has over 10 years of experience working as a data scientist in the insurance and manufacturing industries before joining dotData as a senior data scientist. At dotData, he leads customer-facing data science projects and is responsible for customers’ success with dotData technologies.