Every AI rollout story sounds the same in the pitch deck: train a model, deploy it, watch the metrics improve. The reality inside Continental Health Partners' claims workflow was slower, more political, and ultimately more interesting than that.
The starting point
Claims were reviewed by a team of 14 people working through a shared queue. Average processing time sat at four hours per claim, and — this was the part that actually got leadership's attention — two reviewers looking at the same claim would reach different decisions roughly a third of the time.
Why this mattered more than speed
The inconsistency, not the slowness, was what drove the original request. Slow is a cost problem. Inconsistent is a trust and compliance problem.
Shadow deployment, not a cutover
We trained the triage model on two years of labeled claims, but the model didn't touch a single live decision for the first three weeks. It ran silently alongside every human reviewer, and we compared its recommendation to theirs without either party knowing the other's answer.
Agreement between the model and senior reviewers climbed steadily as we retrained on the shadow period's near-misses. We didn't go live until agreement crossed 85% — a threshold the clinical team set, not us.
What actually shipped
The model doesn't make final decisions. It surfaces a recommendation and a confidence score, and claims below a confidence threshold route straight to a human. That single design choice is most of why adoption worked: reviewers experienced it as a second opinion, not a replacement.
Where this leaves reviewers today
Reviewers now handle roughly a third of the claim volume they used to, spending that reclaimed time on the genuinely ambiguous cases the model flags — the ones that actually need a human judgment call.
Processing time dropped 62% and decision variance between reviewers fell from 34% to 9%. Neither number was the original ask — leadership wanted speed. The variance number is the one that changed how they thought about the project.
