More legally defensible clinical AI may come with an unexpected tradeoff: the systems that performed best against medical malpractice standards also recommended substantially more tests, procedures, and referrals.
A peer-reviewed study published in Communications Medicine evaluated 15 large language models using 198 U.S. medical malpractice cases in which courts had already established what appropriate care should have included. Researchers then compared how often each model recommended the court-endorsed action with the number and estimated cost of procedures it proposed.
The result raises a potentially important question for medical professional liability insurers and healthcare organizations: as AI tools become better at avoiding omissions that could later appear indefensible, could they also encourage more intensive—and potentially more defensive—patterns of care?
What Did Researchers Find?
Across 3,072 simulated consultations, the models varied substantially in both legal defensibility and resource use. The highest-scoring model addressed the court-endorsed action in roughly 70% of consultations and recommended an average of 9.3 procedures per case. A lower-performing model reached the relevant standard in roughly one-third of consultations while recommending only 1.3 procedures.
Estimated Medicare procedure costs differed by more than fivefold between those models.
The relationship was remarkably consistent: models that scored higher for legal defensibility tended to recommend more medical interventions. The researchers found the same basic pattern when testing the models against a separate set of U.K. court cases.
More Defensible Does Not Mean Proven Safer
The study did not evaluate physicians actually using these systems with patients, nor did it show that additional testing prevented malpractice claims or improved clinical outcomes.
That limitation matters. The researchers themselves identify several possible explanations for the pattern, including better clinical reasoning and a tendency by newer models to avoid omitting potentially relevant actions. The study therefore identifies a relationship worth watching rather than proving that AI will cause physicians to practice defensive medicine.
Even so, the liability implications are difficult to ignore. If clinicians increasingly rely on AI recommendations, questions may eventually arise not only when a physician ignores an appropriate recommendation, but also when following AI routinely generates tests, consultations, or procedures that may not otherwise have been ordered.
GET THE SUMMIT
Sign up for news and stuff all about the stuff you wanna know about in your sector twice a month.




