Blog
September 27, 2026

Build or Buy: What Carriers Learn When They Build Claims AI In-House

by
Andrej Evtimov

Cost, time to go live, upkeep and adjuster adoption, compared across both paths.

Both paths are defensible, and they fail for different reasons at different points. This compares building and buying claims AI on the four dimensions that decide the outcome: cost, time to go live, ongoing upkeep, and adjuster adoption.

‍

The question is not which one is better

‍

Carriers that built claims AI in-house and carriers that bought it both made defensible decisions. The two paths fail for different reasons, and they fail at different points in the programme.

‍

This is an attempt to set out what each path actually costs, on four dimensions that decide the outcome. Cost, time to go live, ongoing upkeep, and adjuster adoption.

‍

Nothing here argues that building is a mistake. Several of the strongest claims AI programmes in the market are internal. The question worth answering is narrower: what does your organisation have, and what will it have to acquire?

‍

What "build" means in a claims file context

‍

The word covers more ground than it appears to. A claims AI capability is not one system.

‍

The model is the smallest component

‍

Most build programmes start from a foundation model that somebody else trained. The build is everything around it.

‍

That includes:

‍

  • document intake;
  • optical character recognition for scanned and faxed records;
  • classification into claim document types;
  • deduplication across overlapping productions;
  • page-level citation;
  • an evaluation harness;
  • a review interface adjusters will actually open.

Each of those is a product in its own right. Teams that scoped the project as "connect an LLM to the claim file" discover this in month three.

‍

The evaluation harness is the part teams skip

‍

An evaluation harness is the test suite that tells you whether an answer is right. Without one, you cannot tell whether a change improved the system or broke it.

‍

Building this requires labelled claim files, which requires adjuster time, which competes with claim handling. This is the most commonly underestimated line in an internal build. It does not appear on a vendor invoice, because the vendor has already paid it.

‍

Cost: where the money actually goes

‍

Compute is rarely the dominant cost. People are.

‍

Staffing is the line that moves

‍

The US Bureau of Labor Statistics reports a median annual wage for data scientists of $120,230 in May 2025. It projects employment to grow 35 percent from 2025 to 2035, "much faster than the average for all occupations," with about 24,800 openings a year.

‍

That is the market you are hiring into, and you are competing with technology firms for the same people. A claims AI build typically needs more than data scientists: machine learning engineering, data engineering, platform operations, and a product owner who understands bodily injury.

‍

Model this as a standing team, not a project team. The system needs the same people in year three that it needed in year one.

‍

The cost that does not stop

‍

Vendor pricing is visible and negotiated. Internal cost is distributed and often uncounted.

‍

Three internal costs recur. Retraining or re-prompting when document formats change. Re-evaluation when a foundation model version is deprecated. And the adjuster hours spent labelling, reviewing and correcting.

‍

A fair comparison prices these on both sides. A vendor contract includes them; an internal build pays them out of departmental budgets where they are harder to see.

‍

Time to go live

‍

Both paths reach a working demo quickly. Neither reaches production quickly.

‍

The pilot is not the hard part

‍

A capable internal team can produce a convincing prototype on a sample of claim files in weeks. So can a vendor.

‍

The distance between that prototype and a system adjusters use on live files is where programmes stall. That distance is made of integration, security review, evaluation, change management and exception handling.

‍

Integration work is similar either way

‍

Buying does not remove the integration project. The claim system, the document management system, the identity provider and the audit trail all still need work.

‍

What buying removes is the build of the AI components themselves. Vendors typically expose this through an API or SDK. Some also offer a zero-integration claim file workspace as an interim step before systems work begins.

‍

Estimate the integration separately from the AI work. Teams that fold them into one number are usually comparing a vendor's AI timeline against their own combined timeline.

‍

Upkeep: the obligation that stays with you

‍

This is the section most build-versus-buy comparisons get wrong. Buying transfers work. It does not transfer accountability.

‍

Model drift is an explicit regulatory expectation

‍

One clause in the NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers matters more here than the rest. Among the expectations it sets, adopted on 4 December 2023, is oversight of third parties acting for the insurer.

‍

Read that clause carefully. Overseeing a vendor is the insurer's job, not the vendor's.

‍

The bulletin also asks for a written AI systems programme, plus testing and validation. And it "advises insurers of documentation that a state Department of Insurance may request during an investigation or examination."

‍

The governance framework is the same on both paths

‍

The NIST AI Risk Management Framework 1.0 organises AI risk work into four functions. "The Core is composed of four functions: govern, map, measure, and manage."

‍

Govern is not one stage among four. It is the oversight that runs across the whole programme. In the framework's words, it "is a cross-cutting function that is infused throughout AI risk management and enables the other functions of the process."

‍

A carrier that buys still performs govern, map, measure and manage. It performs them on a system it did not build, which is harder in some respects and easier in others.

‍

Security obligations do not move

‍

Where claim files contain protected health information, the HIPAA Security Rule applies to the carrier directly. Under 45 CFR 164.308(a)(1)(ii)(A), a covered entity must assess its security risks. The rule requires "an accurate and thorough assessment of the potential risks and vulnerabilities to the confidentiality, integrity, and availability of electronic protected health information."

‍

The same section governs the vendor relationship. Subsection 164.308(b)(1) allows a business associate to handle that information "only if the covered entity obtains satisfactory assurances" that it will be appropriately safeguarded.

‍

Building means the risk analysis covers a system your organisation controls end to end. Buying means it covers a system you assess through contracts, attestations and testing. Both are real work. Neither is nothing.

‍

Adoption: the dimension that decides it

‍

A technically excellent system that adjusters route around has failed. This is where most programmes are actually won or lost, and it is the least discussed line in a business case.

‍

Adjusters reject what they cannot verify

‍

An adjuster who cannot check an answer against the source page has two options. Verify it manually, which removes the time saving. Or accept it unchecked, which creates the exposure.

‍

Page-level citation is therefore an adoption feature before it is a compliance feature. Any evaluation, internal or external, should test it on a file the reviewer already knows.

‍

Internal builds have a genuine advantage here

‍

A build team sits inside the organisation. It can watch adjusters work, change the interface next week, and prioritise the two workflows that matter most to your book.

‍

That feedback loop is real and hard for a vendor to match. Carriers that built successfully almost always cite it. Carriers that built unsuccessfully usually had the team report into technology with no standing claims sponsor.

‍

Where building genuinely wins

‍

Five conditions make an internal build the stronger choice.

‍

  • You already run a machine learning organisation with production systems, not a data warehouse team.
  • Your claim mix is unusual enough that general claims tooling misses most of the value.
  • Data residency, hosting or contractual constraints rule out the vendors that fit otherwise.
  • The capability is strategically central and you intend to fund it permanently.
  • You have adjusters who will give the build team sustained time, with a claims executive accountable for adoption.

If four or five of those hold, build. The advantage compounds.

‍

Where buying genuinely wins

‍

Four conditions point the other way.

‍

  • The organisation has no standing ML engineering capacity and no plan to create one.
  • The claim types are common enough that a trained general system covers most of the volume.
  • The programme needs a result inside a budget cycle rather than a strategic horizon.
  • The document mix is the usual one, meaning scanned records, faxes, mixed-quality PDFs and duplicates, which is exactly the problem commercial systems have already solved.

The honest version of the buy case is not that it is cheaper. It is that someone else has already paid for the unglamorous parts.

‍

The hybrid worth considering on purpose

‍

A common pattern is neither pure path.

‍

Carriers buy the document and extraction layer, because it is expensive to build and undifferentiated. They build the layer above it: their own scoring, their own workflow rules, their own reserve logic, their own reporting.

‍

This is what an SDK or API is for. Several vendors package the document and extraction work as modular components built for this pattern. Evaluate the route explicitly rather than arriving at it by accident. It changes the selection criteria: you are buying an interface as much as an output.

‍

The comparison at a glance

‍

  • Dominant cost. Build: Standing engineering and data team. Buy: Licence plus integration.
  • Cost visibility. Build: Distributed, often uncounted. Buy: Explicit, negotiated.
  • Time to prototype. Build: Weeks. Buy: Weeks.
  • Time to production. Build: Governed by evaluation, integration and change management. Buy: Governed by vendor evaluation, integration and change management.
  • Document layer. Build: You build intake, OCR, classification, deduplication, citation. Buy: Already built, assessed by testing.
  • Upkeep. Build: Your team, your backlog. Buy: Vendor roadmap plus your oversight.
  • Regulatory accountability. Build: Yours. Buy: Yours.
  • HIPAA risk analysis. Build: Covers a system you control. Buy: Covers a system you assess.
  • Adoption feedback loop. Build: Fast, internal. Buy: Slower, contractual.
  • Main failure mode. Build: Underscoped support components and no standing sponsor. Buy: Poor fit, weak citation, slow roadmap.

Questions to settle before deciding

‍

Answer these before the business case, not after.

‍

  • Who is the accountable claims executive, and will they still hold the role in two years?
  • What is the evaluation harness, and who produces the labelled files it needs?
  • What happens when the foundation model version you built on is deprecated?
  • Which adjusters will use this daily, and what do they do today instead?
  • If the programme stalls at 40 percent of the intended scope, is that still worth the spend?

The last question separates the two paths more reliably than any cost model. A partial vendor deployment still works on the claims it covers. A partial internal build often leaves a team maintaining something nobody uses.

‍

That is not an argument against building. It is an argument for being honest about the failure mode you are choosing.

‍

Key takeaways

‍

  • The model is the smallest part of a build. Document intake, OCR, classification, deduplication, page-level citation, an evaluation harness and a usable review interface are each a product in their own right.
  • Cost is dominated by people, not compute. BLS reports a median data scientist wage of $120,230 in May 2025, with employment projected to grow 35 percent from 2025 to 2035.
  • Buying transfers work but not accountability. The NAIC bulletin makes oversight of third parties the insurer's obligation, and the HIPAA risk analysis stays with the covered entity either way.
  • Adoption decides most programmes. Adjusters who cannot verify an answer against the source page either check it manually, removing the time saving, or accept it unchecked, creating the exposure.
  • Build when you already run a production ML organisation, your claim mix is unusual, or constraints rule out vendors. Buy when there is no standing ML capacity and the document mix is the ordinary one.
  • The sharpest question: if the programme stalls at 40 percent of scope, is that still worth the spend? A partial vendor deployment still works. A partial build often does not.

‍