Blog
September 14, 2026

How to Evaluate an AI Vendor for Claims Document Analysis

par
Andrej Evtimov

Evaluating an AI vendor for claims document analysis comes down to six requirements.

Evaluating an AI vendor for claims document analysis comes down to six requirements: a production-ready system rather than a prototype, current SOC 2 Type 2 security certification, a business that serves the defence only, analysis that supports the adjuster's decision rather than a shorter document, coverage across every injury line the carrier writes, and completeness across both medical and non-medical documents.

A vendor that misses any one of these tends to fall out of a serious evaluation, often before a technical review even starts.

Why Choosing an AI Claims Vendor Got Harder

The market caught up to demand fast. By 2025, roughly 84% to 90% of insurers had adopted or were actively deploying AI somewhere in claims operations, and the global AI-in-insurance market was on a path from roughly $10 billion in 2025 toward $26 billion or more in 2026. That growth pulled in vendors of every kind: incumbents adding a feature, document-AI startups pointing a general model at claims, medical-record specialists extending into adjacent work, and a wave of new entrants built on the latest large language models.

For a buyer, the result is noise. Vendor pitches sound alike: a clean demo on a clean file, and a similar time-savings number. Accuracy, security, and speed have become claims everyone makes, so they no longer help a buyer choose. The category is also young enough that vendors don't share definitions: one vendor's “medical summary” is a structured, cited chronology, and another's is a paragraph of generated prose describing the same file.

The Six Requirements Carriers Use to Choose

1. Production-ready, not a prototype to build together

The first check is whether the solution already runs in production at organizations like the buyer's, on real files, at real volume, not a research project the carrier would have to help finish. Ask how many carriers run the system in production today, at what volume, and for two references willing to talk about what works and what does not.

2. Security treated as the entry price

For any vendor handling medical records, current SOC 2 Type 2 certification has become the baseline. A Type 1 report describes whether controls are designed properly at a single point in time; a Type 2 report tests whether those controls actually operated effectively over a review period, typically three to twelve months . Buyers also ask where data is stored and processed, how it is encrypted, who can access it, and, with generative models in particular, whether claim data is ever used to train the vendor's models without explicit consent.

3. Built for the defence, and only the defence

Carriers increasingly ask who else a vendor sells to. A vendor serving both carriers and plaintiff firms has divided incentives: the same document-analysis capability that helps an adjuster find weaknesses in a demand can help a plaintiff firm build a stronger one. Buyers ask directly whether the vendor sells to plaintiff-side legal services and whether anything learned from the carrier's data could ever benefit a plaintiff-side product.

4. Analysis that supports the adjuster's decision

This requirement separates a general-purpose summarizer from a tool built for claims work. A summary tells the adjuster what the file says; analysis tells the adjuster what the file means for this claim: the treatment gaps, the prior conditions, the billing that does not match the documented care, the missing records, each traceable to a source page. Carriers test this by running the tool on files where they already know the answer and checking whether it finds what an experienced adjuster would find, without inventing findings the record does not support.

5. One platform that works across every injury line

Motor liability, general liability, workers' compensation, medical malpractice, and long-term care or disability files all turn on different questions: causation and mechanism, premises and responsibility, compensability and return to work, standard of care, and functional capacity over time, respectively. A platform that returns the same generic read across all of them has shown it understands none of them. Carriers test this by running the tool on files from several lines and checking whether the findings actually reflect what matters in each one.

6. The whole claim, medical and non-medical

Liability and causation usually live outside the medical record: in the police report, recorded statements, and the demand package itself. A tool that ingests medical records and ignores everything else gives the adjuster half a file. Carriers test completeness by handing the vendor a full claim file, including the messy non-medical bundle, and checking whether findings connect across both halves.

Four Tells That Separate Production-Grade Vendors From Demos

Claims leaders who run frequent evaluations described a consistent set of signals that surface in the first few conversations, before any formal testing begins:

  • Whose files the demo runs on. A confident vendor will run on the carrier's own files, including hard ones the carrier picks. A vendor that insists on its own curated examples is showing only its best case.
  • How the vendor talks about accuracy. A vendor that can explain where it is strong, where it struggles, and how it handles the files it finds hard has met real data. A single accuracy percentage with no context usually has not.
  • What happens when the system is wrong. Every AI system makes mistakes; a vendor that can explain how it catches and surfaces errors has run in production.
  • Whether references exist. A vendor with real production deployments can connect a buyer to carriers running the system on live claims today.

Questions to Ask Before You Sign

A compact version of the questions above, usable as a working checklist during vendor calls:

  1. How many carriers run this in production, at what volume, and can we call two references?
  2. Do you hold current SOC 2 Type 2 certification, and is our data ever used to train your models?
  3. Do you sell to plaintiff firms or plaintiff-side legal services, in any form?
  4. Can you show findings on our own files, with citations to the source page, not just a summary?
  5. Which of our specific lines do you support in production today, with line-appropriate analysis?
  6. Which document types do you ingest, and can you connect findings across medical and non-medical sources?

The category is also under closer regulatory watch than it was even two years ago. The NAIC's model bulletin on insurers' use of AI, adopted in December 2023, has now been adopted by more than half of U.S. states, and a multistate pilot of a formal AI Systems Evaluation Tool for market conduct examinations is running across twelve participating states from January through September 2026 . Because carriers remain accountable for what a vendor's system does, a vendor's governance, documentation, and explainability have become the buyer's concern too, not an afterthought to raise after a contract is signed .

Frequently Asked Questions

What is SOC 2 Type 2 and why does it matter for claims AI?

SOC 2 is an AICPA audit standard evaluating a service provider's controls around security, availability, processing integrity, confidentiality, and privacy. A Type 2 report confirms an independent auditor tested those controls over a period of months and found they actually worked, not just that they were designed well at one point in time. For any vendor handling protected health information in claim files, Type 2 certification has become the baseline expectation before a security review even begins .

What's the difference between AI summarization and AI analysis in claims?

Summarization condenses a file into a shorter version of the same text: it tells the adjuster what the file says. Analysis goes further: it finds the treatment gaps, prior conditions, billing inconsistencies, and causation issues that change a reserve or a negotiation, and traces each finding to a specific page in the record. First-generation claims AI tools stopped at summarization, which is why buyers now test specifically for analysis.

How do I evaluate an AI vendor for insurance claims?

Screen first on the fast-disqualifying requirements: current SOC 2 Type 2 certification and whether the vendor also serves plaintiff firms. Then confirm the solution runs in production at similar carriers today. For the vendors that clear those gates, run the software on your own production files, including the hard ones, against success criteria and a baseline you define in advance, rather than relying on a demo.

Should claims AI vendors also sell to plaintiff firms?

In practice, the answer is usually no. A vendor serving both sides has divided incentives, and the same analytical capability that helps a defense adjuster find weaknesses in a demand can help a plaintiff firm build a stronger one. Carriers typically ask vendors directly about their customer base before moving forward with a technical evaluation.

amaise builds AI for bodily injury claims and underwriting, built for the defense and only the defense, and holds current SOC 2 Type 2 certification. To test it against these six requirements on your own files, contact amaise at hello@amaise.com.