The FDA Has Cleared 1,357 AI Medical Devices. Three Were Tested on Whether Patients Got Better
A PLOS Digital Health analysis found that of 1,357 FDA-authorised AI medical devices, 34 appeared in registered clinical trials and only 3 measured outcomes like mortality or hospitalisation.
Researchers led by Rawan Abulibdeh at the University of Toronto went through every AI-based medical device the US Food and Drug Administration has authorised for patient care and asked a simple question: was it ever tested on whether patients actually did better? Of 1,357 devices, 34 were linked to a registered clinical trial. Twelve had results published and peer reviewed. Three assessed patient-centred outcomes such as death, stroke, hospitalisation or quality of life. The analysis was published in PLOS Digital Health on 19 August.
The gap is not a scandal so much as a design feature, and understanding why is the useful part. Most AI medical devices reach the US market through a route that requires the maker to show “substantial equivalence” to a device already authorised. That means demonstrating the new thing works comparably to the old thing, typically by matching or beating it at a technical task: spotting a nodule on a scan, flagging an irregular heart rhythm, measuring a structure. Nobody is required to show that using it makes patients healthier. Clearance is a statement about equivalence, not about benefit.
Those two things come apart more often than you would hope. A tool can be measurably better at detecting something and still not improve outcomes, because it flags findings that would never have caused harm, or because the flag arrives at a clinician who is already overloaded, or because the follow-up test carries its own risk. This is a well-documented pattern in screening medicine generally, long before AI showed up. The authors also flag a second problem: the evidence that does exist comes overwhelmingly from well-resourced settings, and systematically excludes pregnant women, adults over 75 and people who do not speak English. Those are not edge cases. They are large slices of the people these tools will be pointed at.
None of this means the 1,357 devices are useless. Plenty of them are doing exactly what they should. The claim in the paper is narrower and harder to argue with: we do not know, for almost all of them, because the question was never asked in a way that would produce an answer. The authors argue for revised authorisation requirements covering effectiveness and equity, not for pulling anything off the market.
What this means for you: if you are a patient, the practical version is that “FDA cleared” on an AI health product means it passed a comparison test, not that a trial showed it helps. That is worth knowing before you read a marketing page, and it is a fair question to put to a clinician: what does this tool change about my care? If you work anywhere near health technology, the finding is a useful mirror to hold up to your own claims. And for everyone else, there is a general lesson that travels well beyond medicine: a system can be excellent at the thing it is measured on and still not move the thing you actually care about. That gap is where most disappointment with AI tools lives, and it is almost always visible in advance if you ask what was measured.
Sources
Source: https://medicalxpress.com/news/2026-08-ai-medical-devices-patient-outcomes.html
An Oxford Spinout Sold Anthropic 250 Million Dollars of Chips That Do Not Exist Yet
Fractile is reportedly raising 600 million dollars at a 6.5 billion valuation, six times its May mark, after an initial deal to supply Anthropic with inference chips. The chips are not expected until 2027.