← Insights

AI Hallucinations in Insurance: When Confident Answers Become Compliance Risks

Generative AI is genuinely good at making insurance easier to understand. It can translate dense policy language, compare options side by side, and help someone figure out what a denial letter actually means. That's real value, and it's why the industry has moved so fast: in a 2025 NAIC survey of 93 health insurers, 84% reported using AI or machine learning in some capacity.

8 min read
AI Hallucinations in Insurance: When Confident Answers Become Compliance Risks

Generative AI is genuinely good at making insurance easier to understand. It can translate dense policy language, compare options side by side, and help someone figure out what a denial letter actually means. That's real value, and it's why the industry has moved so fast: in a 2025 NAIC survey of 93 health insurers, 84% reported using AI or machine learning in some capacity.

But insurance creates a problem that ordinary chatbot use does not. A plausible wrong answer can change what someone does with their money, their coverage, or their healthcare. "Your plan covers this," "you have until Friday to appeal," "this provider is in network" — these aren't harmless conversational slips when they're wrong. They're representations a customer will act on.

That's the line worth paying attention to. Not the moment an AI system makes a mistake, but the moment an unsupported answer travels far enough to become someone's coverage assumption.

How one wrong answer becomes a real problem

Say a customer asks whether a specialist is in network. The system retrieves information for the right carrier but the wrong network, and answers confidently: yes, this doctor is in the network. She books the appointment. The claim processes out of the network.

Technically, that was a retrieval error. From the customer's side, it was a surprise bill. From the business's side, it may now be a complaint, an inaccurate representation, remediation work, and — depending on the facts and the jurisdiction — a regulatory question.

The same chain runs through the whole customer journey. Before a sale, "Plan A has a $1,500 deductible" is only reliable if the system has the current plan document. During enrollment, an incorrect statement about a documentation deadline occupies one sentence but can cost someone a coverage year. After a purchase, questions arrive attached to a concrete event — a denial, an appeal window, a required form — which is exactly where a system should be least willing to fill a gap with a plausible guess.

It also runs internally. An AI-generated policy summary pasted into a service agent's notes can propagate an error even when the customer never spoke to the model at all.

A useful test cuts across all of it: how expensive is it to act on this answer before discovering it was wrong? A mistaken definition of a non-material term is low risk. A false coverage confirmation, a wrong appeal deadline, or an incorrect eligibility conclusion is not, because the customer may spend money, miss a window, or delay care before anyone catches it.

"The AI said it" doesn't move the responsibility

Regulators have been clear about the direction, even where the specific rules still vary by state.

The NAIC's 2023 Model Bulletin tells insurers that AI-supported consumer outcomes remain subject to existing insurance law, and expects a written AI program proportionate to risk — governance, documentation, testing, validation, and oversight of third parties. It's guidance rather than federal law; adoption by individual states is what gives it force. By late 2025 the NAIC said more than half of states had adopted it or something similar, and the organization has continued the work, including an AI Systems Evaluation Tool pilot reported in July 2026.

Vendor technology doesn't transfer the risk to the vendor. New York's 2024 circular letter — formally scoped to underwriting and pricing — says insurers should conduct due diligence over third-party AI and remain responsible for the outcomes of its use. It also expects procedures to investigate and remove incorrect information, and says a vendor's proprietary algorithm isn't an excuse for failing to explain an adverse action. Texas makes a similar point about third-party data: regulated entities stay responsible for accuracy in rating, underwriting, and claims handling even when the data came from somewhere else.

Requirements are also getting more specific. Colorado extended its governance rules for external data, algorithms, and predictive models from individual life to private passenger auto and health benefit plans, effective October 2025, with a July 2026 compliance milestone for newly covered insurers. California's SB 1120 goes further in one narrow but high-stakes place: AI can assist utilization management, but it cannot displace the qualified professional responsible for medical-necessity determinations.

None of this means every chatbot error violates insurance law. It means introducing AI doesn't create an exemption from the underlying obligation to communicate accurately.

There's a parallel principle outside insurance regulation. The FTC has repeatedly pursued unsupported claims about what AI products can do, including a finalized order against DoNotPay over claims that its chatbot could substitute for a lawyer. The lesson isn't that regulators dislike AI. It's closer to the opposite: what you say your AI does needs to match what it actually does.

Retrieval helps. It isn't a truth machine.

Retrieval-augmented generation — RAG — is a real part of the answer. Instead of generating from what a model absorbed during training, the system first looks up relevant external information and works from that. The original research showed clear factuality gains over a comparable model relying only on its own parameters.

For insurance that's appealing, because the system can be pointed at an approved policy or benefit summary instead of trying to remember a rule.

But retrieval fails in ordinary ways. The retriever finds the wrong document. The knowledge base holds last year's version. The customer is on Plan A and the system pulls Plan B. The right document comes back and gets read wrong. Two accurate passages combine into an inaccurate conclusion. Research continues to document hallucinations in retrieval-grounded systems, including in financial domains where the underlying facts are time-sensitive.

So the question isn't whether a system has RAG. It's whether it can show that it used the right source, for the right customer, as of the right date — and recognize when that evidence isn't enough.

Citations help with the first half of that, but only when they're substantive. A citation pointing at a document that doesn't actually support the claim creates false reassurance rather than less of it. The more useful function is reconstruction: if a customer later says "your AI told me this was covered," the organization should be able to establish which plan version was live, what the customer supplied, what was retrieved, what was generated, and whether a person reviewed it. In a regulated workflow, logs stop being an engineering detail and start being part of how you investigate and remediate.

Sometimes the right answer is that there isn't one

This is the design principle that matters most, and it runs against the instinct to make an assistant maximally helpful.

For a general question, answer it. For an ambiguous one, ask. For a plan-specific question without the documentation to support an answer, say so. For a disputed claim, an eligibility question, an appeal deadline, or anything with a medical consequence, the bar for escalation should be low.

Research on current language models supports this more than it supports optimism about the next release. Work published in Nature found that newer, more capable, more instruction-following systems became more likely in some settings to produce plausible wrong answers rather than declining to answer — and that human reviewers often didn't catch them. Other experimental work has found people continuing to rely on incorrect AI advice even when shown explanations, which is why making an answer look transparent doesn't reliably prevent overreliance.

The goal shouldn't be maximizing the share of questions an AI answers. It should be maximizing the share it handles appropriately.

What this means if you're the one asking

Treat an AI explanation as a guide, not as the policy. When an answer materially affects cost, coverage, care, or a deadline, ask what document it came from, which version, and whether it applies to your plan. A well-built system should make those questions easier to answer, not discourage them.

Five categories deserve extra care: coverage confirmations, network status, eligibility, claims and appeal instructions, and anything touching a medical determination. In those cases, "please confirm this with a representative" isn't an AI failure. It's often evidence of a system that knows its own limits.

Keep the important exchanges. If an answer influenced a decision, the record makes a later discrepancy much easier to resolve.

And when an answer conflicts with your policy document, EOB, or carrier statement, don't rephrase the question until you get the answer you want. Flag it. A good system should either explain the difference or route it to someone who can — and the discrepancy should become a test case, not just a corrected conversation.

The standard worth holding

It's tempting to treat hallucination as a problem the next model version solves. The evidence argues otherwise. Retrieval improves grounding. Domain specialization reduces the irrelevant ground a system has to cover. Version-controlled sources cut down on stale answers. Testing catches recurring failures. Human review contains the expensive ones. None of them, alone or together, guarantees that every future response is correct.

Which changes the question. Not "can this system promise never to be wrong," but: what happens when it doesn't know enough to answer safely? Does it use a current source? Does it separate an explanation from a confirmation? Does it show uncertainty instead of smoothing over it? Does it escalate what it should? Can the interaction be reconstructed afterward?

Insurance is hard partly because people are constantly translating between contracts, benefits, carriers, providers, bills, and actual decisions. AI that makes those relationships legible, quickly and in plain language, removes a great deal of friction. But trust won't come from making every answer sound certain.

It comes from making the certainty earned.

Share this article

Keep reading