Two-thirds of clinicians now use AI in their work, according to the American Medical Association, and the FDA has cleared more than 1,450 AI-enabled medical devices, 295 of them in 2025 alone. Adoption is no longer the question.
The uncomfortable counterweight: fewer than 2% of those cleared devices were supported by randomized clinical trials, and more than 10,000 AI-related safety incidents have been reported in healthcare settings since mid-2024. That gap between how fast ai in healthcare is being deployed and how thinly it's been validated is the single most important thing for a buyer to understand. Here's an honest read on where it works, what actually goes wrong, and how to judge the return.
Where AI Is Genuinely Working

The strongest results cluster in administrative and operational workflows rather than in clinical decision-making, which is both less exciting and considerably safer. Documentation, scheduling, prior authorization, claims processing, and record digitisation all have clear before-and-after metrics, bounded scope, and a human reviewing anything ambiguous.
Diagnostic support is the more contested category. AI-assisted imaging genuinely improves detection rates in narrow, well-validated applications, and AI agents in healthcare covers the broader operational and clinical capabilities in more depth. The distinction that matters commercially is this: administrative AI failing means wasted money, while clinical AI failing means patient harm, and those two risk profiles justify very different levels of scrutiny before deployment.
The Three Failure Modes Behind Real Incidents
The 10,000-plus reported safety incidents concentrate into three recurring patterns, each with well-documented real-world examples.
|
Failure mode |
What happens |
Documented example |
|---|---|---|
|
Algorithmic bias |
Model performs unevenly across patient populations, reinforcing existing disparities |
A widely used follow-up care algorithm found to be systematically racially biased |
|
Data drift |
Accuracy degrades over time as real-world conditions diverge from training data |
Sepsis prediction models losing reliability as case mix and documentation practices shift |
|
Integration failure |
The tool disrupts clinical workflow instead of improving it |
Google Health's diabetic retinopathy screening slowed clinical workflows in field deployment |
The most instructive single case is the Epic Sepsis Model, which missed roughly two-thirds of actual sepsis cases at its validated thresholds despite wide deployment. IBM Watson for Oncology produced treatment recommendations judged unsafe. These weren't fringe pilots, they were flagship products at major institutions, which is precisely why vendor reputation is a poor substitute for independent validation on your own patient population.
The Evidence Gap Buyers Should Know About
A scoping review of 692 FDA-approved AI/ML medical devices found the documentation supporting them contains substantial blind spots. Only 3.6% of approvals reported the race or ethnicity of study subjects. 81.6% did not report subject age. 99.1% provided no socioeconomic data at all.
This matters practically rather than academically: without knowing which populations a device was validated on, a buyer cannot assess whether it will perform on theirs. A model validated largely on one demographic can underperform meaningfully on another, and the approval paperwork frequently won't tell you either way. Regulatory clearance answers whether a device met a bar, not whether it will work for your patients.
The direction of travel adds to this. A proposed HHS rule released in January 2026, known as HTI-5, would scale back certification criteria for AI used in decision support interventions, loosening a transparency requirement introduced only a year earlier. Buyers should expect to do more of this diligence themselves, not less.
Hallucination and Automation Bias in Clinical Settings
Studies of large language models used for clinical decision support estimate hallucination rates between 8% and 20%. In a text-based business workflow, a confidently wrong output gets caught at review. In a clinical setting, it interacts with a well-documented human tendency called automation bias, where a recommendation carries more weight simply because a system produced it.
The design principle that addresses this is algorithmic deferral: the system actively escalates to a human when confidence is low or when the situation falls outside its validated scope, rather than generating an answer the clinician may accept without scrutiny. It's considered a foundational safety feature, and many tools currently on the market lack it. It's worth asking any vendor directly how their system behaves at low confidence, and treating "it always returns an answer" as a warning rather than a feature.
How to Evaluate ROI Honestly

Healthcare AI ROI splits along the same line as the risk profile. Administrative automation produces returns that are straightforward to measure: hours returned, turnaround time reduced, error rates lowered, all against a baseline you can capture before deployment. Clinical AI returns are real but harder to attribute and slower to prove, and they carry a liability tail that administrative tools don't.
Three practical rules make the evaluation trustworthy. Capture the baseline before anything launches, since a retrospective comparison against remembered performance is not a measurement. Validate on your own patient population rather than on the vendor's published figures. And run it as a scoped pilot with a pre-agreed success threshold, building an AI PoC covers how to structure that so it produces evidence rather than a demo.
Questions to Ask Before You Buy or Build
What population was this validated on, and how does it compare to ours?
If the vendor can't answer with specifics, that's the FDA reporting gap showing up in your procurement process. Ask for the demographic breakdown of the validation cohort directly.
What happens when the model's confidence is low?
You're testing for algorithmic deferral. A system that always returns a confident answer, regardless of whether the case falls inside its validated scope, is the configuration most likely to produce automation-bias harm.
How will we detect drift, and who owns that monitoring?
Model accuracy degrades as case mix, documentation practice, and patient populations shift. Post-deployment monitoring needs a named owner and a defined cadence before go-live, not after the first incident. choosing an AI development partner covers the wider vendor-evaluation questions worth layering on top of these healthcare-specific ones.
