The Exam Is Being Written Before the Rule Is
What carriers should know about the NAIC AI Systems Evaluation Tool, and why the calendar matters more than the debate
Most of the attention paid to AI regulation in insurance has gone to a question that has not been answered yet: what will the rule require?
That is the wrong thing to be watching right now. The more immediate development is that the examination instrument is arriving ahead of the rule, and it is arriving on a published schedule.
The NAIC's AI Systems Evaluation Tool (aka AI Risk Evaluation Supplement) entered a multistate pilot in early 2026, running through September. Twelve states are participating: California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin. Feedback from the pilot is being folded into a revised version this fall, which will be re-exposed for public comment, with adoption anticipated at the Fall National Meeting in November.
The practical consequence is worth stating plainly. Examiners are on track to have a standardized AI evaluation instrument before the NAIC has decided whether a model law is needed, or what one would require. Carriers will describe their AI governance in a structure regulators designed rather than one the industry proposed. And that description becomes the record.
What the tool actually asks
The framework is organized into four exhibits.
Exhibit A quantifies AI usage across the organization. This is a scoping exercise, and it is the one worth reading first.
Exhibit B is a governance risk assessment framework.
Exhibit C requests detail on the systems a carrier identifies as high risk.
Exhibit D covers the data used as inputs.
Exhibit A deserves the early attention because an inventory question asked cold produces a very different answer than an inventory question an organization has already answered for itself. The gap between what a compliance function believes is deployed and what is actually deployed is where most carriers will find their surprises, and the largest part of that gap is usually AI embedded inside third-party platforms. The carrier did not build it, may have limited visibility into it, and cannot outsource its responsibility for outcomes affecting policyholders.
Participating states are running monthly coordination calls to avoid duplicative requests and to share what they are learning. Information requested is protected under the confidentiality rules of the administering state. Reviews are being folded into existing work: market conduct exams, financial analysis, and financial examinations, across property and casualty, life, and health.
This is not a new regulatory regime bolted on the side. It is the existing examination apparatus acquiring a new set of questions.
The posture underneath it
The NAIC's stated position has been consistent and it is not complicated: existing insurance law applies to AI-driven decisions the same way it applies to decisions made by human underwriters and adjusters. Where AI touches the policyholder, regulatory interest follows.
The Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted by the NAIC in December 2023 and since adopted across numerous states, sets out what regulators expect a carrier to have: a written AI systems program covering governance, risk management and internal controls, model validation and testing for bias and error, consumer notice, oversight of third-party systems, and the ability to respond to a regulatory inquiry about any of it.
None of that is novel as principle.
What changes with the Evaluation Tool is practical rather than philosophical. Expectations that could previously live comfortably in policy will increasingly have to survive structured regulatory questioning. And structured questioning has a way of exposing the difference between a program that is described and one that can be evidenced.
That distinction is the entire subject.
A policy is a document describing what should happen. A control is a mechanism that produces a record when it happens.
Most AI governance programs written in the past two years are policies. Comparatively few are controls. The difference is invisible right up until someone asks to see the record.
What is worth having ready
Not a program overhaul. A set of answers, decided internally, before they are requested externally.
An AI inventory that survives Exhibit A. Including vendor-embedded AI, where responsibility and visibility diverge most sharply.
A written internal definition of what counts as a high-risk AI system. Decide this yourself. Deciding it in the week Exhibit C arrives means deciding it under pressure, and the definition you produce under pressure is the one you will be held to.
A documented reporting path for AI matters. Where they go, how often, and to whom. Most carriers have this well established for cybersecurity, typically quarterly CISO reporting into audit or a risk committee. Fewer have an equivalent for AI. The asymmetry between those two paragraphs is visible to anyone reading both.
Evidence of model testing. Meaning artifacts with dates and named reviewers, not the existence of a testing policy.
A defensible position on third parties. If a vendor model produces an adverse outcome for a policyholder, the carrier answers for it. The contract does not transfer that.
The record most organizations don't keep
There is a sixth item that is rarely part of the standard AI governance record, and it is the one that turns a governance description into something an examiner can verify.
Every organization can tell you what was reviewed. Very few can tell you what was auto-approved, by what role, and on whose authority.
That second record is the one that describes where decisions went out into the world with no human judgment attached. It is also the record that determines who is accountable when one of those decisions is challenged. An examiner asking how oversight works is, functionally, asking for it.
Building it is not a technology project. It is a decision to start counting something that is already happening.
Why the timing is the point
There is a window between now and November in which a carrier can decide what its own answers are, at its own pace, with its own counsel, on its own terms.
After adoption, the first time a carrier answers these questions, it answers them to an examiner.
The instrument is being built in public, on a published schedule, with draft exhibits that are readable today. That is unusual and it is a gift. Most regulatory change does not announce itself this clearly this far in advance.
The carriers that answer calmly next year will be the ones building the evidence this year.
Good Faith EI helps carriers turn AI governance policies into evidence-producing controls. If you want to know what your organization could actually produce if Exhibit A arrived tomorrow, book a 20-minute governance architecture consult at goodfaithei.com.
Regulatory details current as of August 2026 and drawn from public NAIC materials. Nothing here is legal advice.