Insurance converts uncertain future events into contracts, prices, reserves, and claims decisions. Its work is rich in language and documents, but the consequences are financial and regulated. That combination favors specialist AI systems with explicit evidence, thresholds, and human authority.
The most valuable portfolio is not a universal insurance chatbot. It is a set of task-specific models across acquisition, underwriting, service, claims, fraud, and compliance.
Adoption and maturity
The European Insurance and Occupational Pensions Authority's 2026 market survey covered 347 insurance undertakings in 25 EU and EEA countries. It found:
- 65% were actively using generative AI.
- A further 23% planned to use it within three years.
- 64% of reported use cases were still at proof-of-concept or experimentation stage, while 32% had reached production.
- Of 957 submitted use cases, 64% were back-office and 36% customer-facing.
- Current use was most common in customer service at 40%, claims at 32%, and sales and distribution at 22%.
- Planned adoption was highest for fraud detection at 64%, claims at 59%, and sales and distribution at 54%.
- 49% had a dedicated AI policy, twice the reported 2023 proportion.
EIOPA's survey also found hallucinations were the top-cited GenAI risk and said, "Human oversight remains dominant."
US regulatory surveys show similarly broad interest. The National Association of Insurance Commissioners reports that 88% of surveyed auto insurers, 70% of homeowners insurers, 58% of life insurers, and 92% of health insurers use, plan to use, or are exploring AI and machine learning. These categories combine current use with planning and exploration, so they should not be read as production adoption.
The model stack across the lifecycle
The model stack across the lifecycle
Acquisition and distribution
Speech and language models can transcribe agent and customer conversations, extract needs, prepare compliant follow-up, and identify missing application fields.
Underwriting
Document models extract evidence from applications, medical files, financial statements, inspection reports, and loss runs. RAG retrieves current guidelines.
Policy issuance and servicing
Document generation, comparison, and retrieval can support endorsements, renewals, certificates, billing questions, and coverage explanations.
First notice of loss
A voice agent can capture the event, parties, location, injuries, damage, coverage identifiers, evidence, and contact preference.
Claims triage and handling
Models can classify severity, extract documents, summarize files, detect missing evidence, estimate routing, and draft communication.
Fraud investigation
Graph, anomaly, and supervised models can identify suspicious relationships, behavior, documents, and claims patterns.
Acquisition and distribution
Speech and language models can transcribe agent and customer conversations, extract needs, prepare compliant follow-up, and identify missing application fields. Voice agents can handle appointment and status tasks.
Measure quote completion, field accuracy, conversion, handle time, disclosure compliance, complaint rate, and subgroup outcomes.
Underwriting
Document models extract evidence from applications, medical files, financial statements, inspection reports, and loss runs. RAG retrieves current guidelines. An SLM summarizes risk factors and missing evidence. Predictive models score defined risks.
Zurich's underwriting proof of concept combines RAG over underwriting guidance with a similarity engine for relevant historical cases. The architecture is important: guidance retrieval and case comparison support the underwriter rather than hiding the evidence behind one score.
EIOPA recorded examples of insurers using internal LLMs to search underwriting manuals and extract information from medical documents to accelerate review.
Policy issuance and servicing
Document generation, comparison, and retrieval can support endorsements, renewals, certificates, billing questions, and coverage explanations. The system must identify the exact policy version, endorsement, jurisdiction, effective date, and insured party.
First notice of loss
A voice agent can capture the event, parties, location, injuries, damage, coverage identifiers, evidence, and contact preference. It should adapt questions to the loss type while escalating injury, vulnerability, fraud, dispute, and coverage ambiguity.
Branch and Liberate report 43% adoption of digital or voice FNOL, average call duration falling from 12 minutes 23 seconds to 7 minutes 10 seconds, a 42% reduction, and an expected 70% reduction in handling cost. This is a vendor-published customer case and the cost figure is an expectation, not an audited outcome.
Claims triage and handling
Models can classify severity, extract documents, summarize files, detect missing evidence, estimate routing, and draft communication. Image models can assess visible damage for review. Decision support should separate coverage, liability, severity, and fraud signals.
Measure cycle time, touch time, reopen rate, supplement rate, leakage, litigation, complaint, and claimant outcomes. A faster incorrect denial is not a successful automation.
Fraud investigation
Graph, anomaly, and supervised models can identify suspicious relationships, behavior, documents, and claims patterns. An SLM can assemble an investigator brief with linked evidence. It should not state that a claimant committed fraud.
Measure precision among referred cases, investigator yield, prevented loss, false-positive burden, and subgroup disparities.
Compliance and quality
Speech analytics can review required disclosures and vulnerable-customer signals. RAG can map regulation to policy and control evidence. Document models can test issued artifacts. Human compliance teams decide breaches and remediation.
Reference architecture
- Party and policy graph: customer, insured, policy, endorsement, risk, claim, provider, and payment.
- Evidence layer: calls, documents, images, telematics, third-party data, and source provenance.
- Specialist models: ASR, document extraction, classifier, forecast, anomaly, graph, and SLM.
- Knowledge layer: product wording, underwriting manual, claims guidance, regulation, and internal policy.
- Decision services: deterministic coverage rules, predictive scores, and optimization.
- Workflow layer: policy, billing, claims, contact center, CRM, and investigation systems.
- Governance: consent, purpose, permissions, human approval, explanation, monitoring, and audit.
Keep final adverse or consequential decisions under applicable law, policy, and human authority. A generative model should not be the hidden rules engine.
Evaluation scorecard
| Use case | Model metric | Insurance outcome | Guardrail |
|---|---|---|---|
| FNOL voice | entity accuracy, escalation recall | handle time, completion | injury and vulnerability routing |
| Underwriting RAG | retrieval recall, citation precision | review time, referral quality | current guideline only |
| Document AI | field precision and recall | straight-through rate | source traceability |
| Claims summary | factual consistency, omission | adjuster time, reopen | human file review |
| Fraud | precision, calibration | investigator yield | fairness and appeal |
| Service agent | task completion | first-contact resolution | correct policy and identity |
| Compliance | detection precision and recall | confirmed breach, review time | legal interpretation by experts |
A portfolio rollout
Begin with back-office assistance, where EIOPA found the largest current concentration. Choose one measurable workflow such as underwriting-guideline retrieval, claims-file summarization, or document intake. Build a test set from closed cases and current policy.
Run shadow evaluation, then assisted use. Measure corrections and downstream outcomes. Expand toward customer-facing and agentic work only after permissions, escalation, monitoring, complaint handling, and version control are proven.
The strategic conclusion
Insurers do not need one model to understand insurance. They need a governed model portfolio where each component performs a bounded job and every output can be traced to policy, evidence, and accountable review.
The durable moat is the insurer's connected policy and claims data, expert corrections, and evaluation set. Models can change. That institutional layer compounds.
Research note
Research is current through September 5, 2026. Regulatory survey categories and denominators are stated to avoid overstating production use. Vendor metrics are labeled. Requirements vary by product and jurisdiction. This article is not insurance, actuarial, or legal advice.
Continue the research
- AI first notice of loss for insurance claims
- Insurance policy language AI
- Why enterprises should use small language models
- How enterprise model routing balances quality and cost
- The private AI bank model stack
Building a Production-Ready System
Conscious Engines builds insurance AI solutions around policy language, customer conversations, claim evidence, underwriting workflows, and regulated actions. Specialized speech, document intelligence, private RAG, and task-specific agents can improve cycle time and consistency while keeping consequential decisions reviewable.