Sharan Initiatives
๐Ÿง 
๐Ÿง AI & Medical Imaging

Radiology AI: How to Read the Accuracy Claims, and What Is Actually Deployed

Over a thousand AI devices have been authorised for medical imaging, most of them on substantial equivalence rather than outcome evidence. A guide to the evidence ladder, what the landmark studies really showed, and the five things deployed radiology AI actually does.

By Taresh Sharan ยท PhD, IIT BHUโ€ขMarch 4, 2026โ€ข11 min read

Radiology is the one medical specialty where AI has genuinely arrived at scale. Not in the sense the coverage implies โ€” nobody's job has disappeared โ€” but in the sense that a large amount of regulated software is in clinical use and a meaningful fraction of imaging departments in wealthy health systems have at least one such tool running.

The FDA's public list of AI-enabled medical devices now runs to well over a thousand entries, and roughly three-quarters of them are radiology devices. The annual authorisation rate has gone from a trickle in the 2010s to hundreds per year. Almost all of them arrived through the 510(k) route, which means they were cleared on the basis of substantial equivalence to an existing device rather than on fresh clinical outcome evidence.

That last sentence is the single most useful thing to know when reading any claim about radiology AI, and I want to unpack why before getting to what the tools actually do.

The evidence ladder, and where most claims sit

Performance claims in this field come from four quite different kinds of study, and they are routinely quoted as if they were interchangeable.

Retrospective testing on a held-out dataset. The model is evaluated on images set aside from the same collection it was trained on. This is the weakest evidence and the most commonly quoted. It tells you the model learned something about that dataset. It tells you very little about a different hospital, and as I have written elsewhere, it can look excellent while the model is keying on scanner artefacts rather than pathology.

External retrospective validation. Same idea, but on data from institutions that contributed nothing to training. Much better. Still retrospective, still curated, still using a population someone selected.

Reader studies. Radiologists read a set of cases with and without the model, and you compare. This measures the thing that actually matters โ€” does the combination beat the human โ€” but the cases are usually enriched for pathology, the readers know they are being studied, and there is no time pressure or fatigue from a real shift.

Prospective trials in a live screening or clinical pathway. Rare, expensive, and the only design that tells you what happens to patients. There are a handful in imaging and they are worth far more than the thousand retrospective papers beneath them.

Most of the numbers you will encounter come from the first rung. When someone says a system "matches radiologist performance," ask which rung.

What the landmark results actually showed

Two results get cited constantly, usually inaccurately, so it is worth stating them precisely.

CheXNet was a 2017 convolutional model for pneumonia detection on chest radiographs, and it reported exceeding the average performance of practising radiologists on that task. The caveats are substantial and were in the paper. The comparison was against four radiologists, on an F1 metric. The labels came from automated text mining of radiology reports, which is a noisy reference standard. And the human readers were given the frontal image alone โ€” no lateral view, no clinical history, no priors โ€” which is not how anyone reads a chest film. CheXNet was an important piece of work and a genuine advance in what was achievable. It was not evidence that a model could replace a radiologist, and reading the paper makes that clear.

The 2020 Nature paper on an AI system for breast cancer screening was more rigorous: large UK and US mammography datasets, a reported absolute reduction in both false positives and false negatives relative to the historical reading, and an independent reader study in which the model outperformed all six participating radiologists. It also simulated using the model as the second reader in the UK double-reading protocol and estimated a large reduction in second-reader workload.

It was still retrospective. It drew criticism, in the same journal, for not releasing enough code and data for independent groups to reproduce it โ€” a fair criticism and a recurring one in this literature.

The result that changed my mind about screening AI

What retrospective work cannot answer is whether a real programme, running on unselected women, with real radiologists who know the AI is there, does better. A Swedish randomised trial set out to answer exactly that, randomising a population screening cohort to AI-supported reading versus standard double reading.

The interim safety analysis, published in 2023, reported a roughly 44 percent reduction in screen-reading workload with no loss in cancer detection. Later analyses from the same trial reported a substantial increase in the cancer detection rate without an increase in false positives, and a lower rate of interval cancers.

This is the strongest evidence in the field, and it is strong for a reason worth noting: the workload finding is arguably more important than the detection finding. In a double-reading programme, the binding constraint is radiologist hours. A tool that lets a programme maintain detection performance while reading fewer studies twice is solving the actual problem. Detection improvements are a bonus; capacity is the crisis.

It also illustrates why generalisation still needs care. That trial ran in a national screening programme with standardised equipment, a trained reading workforce and a consistent protocol. A health system without those things cannot assume the same result.

What is actually deployed, in practice

Stripping away the categories that exist mainly in vendor materials, deployed radiology AI does about five things.

Worklist triage. The model scans incoming studies and reorders the queue, pushing suspected critical findings โ€” intracranial haemorrhage, large vessel occlusion, pneumothorax, pulmonary embolism โ€” to the top. This is the most defensible deployment pattern in the field. A false positive costs a radiologist thirty seconds of attention. A true positive found twenty minutes earlier can change an outcome. The failure mode is benign and the benefit is real.

Detection and marking. The model annotates candidate findings โ€” nodules, fractures, microcalcifications โ€” for the reader. Useful, and the deployment where automation bias is the biggest concern.

Quantification. Automatic measurement and comparison to prior: nodule volumetry, cardiac chamber volumes, brain volumetrics, lesion tracking across studies. This is quietly the highest-value category, because it replaces a manual task that is tedious, slow and genuinely variable between readers, and where being consistent matters more than being clever.

Image reconstruction and acquisition. Learned reconstruction that allows shorter MRI acquisitions or lower-dose CT at comparable image quality. Most patients and most radiologists never think of this as AI, and it may be the place where these methods have done the most measurable good.

Reporting assistance. Draft report generation, structured field population, protocol selection. Fast-moving, and the category where the gap between demo and dependable is currently widest.

Notice what is not on the list: autonomous interpretation. With the exception of narrow screening tasks such as diabetic retinopathy, deployed radiology AI produces an input to a radiologist's decision, and the radiologist signs the report.

What makes deployments fail

Having watched several of these go in, the failure patterns are consistent and almost never about the model.

The tool sits outside the reading workflow. If checking the AI output means switching applications, it will not be checked. Integration with the PACS and the reporting system is not an implementation detail, it is the product.

The operating point was never localised. Vendors ship a threshold chosen on their validation population. Prevalence at the deploying site is different. The result is a false-positive rate the department did not sign up for, and a reputation the tool never recovers from.

Nobody owns monitoring. A scanner gets replaced, a protocol changes, and performance drifts with no one watching. The only deployments I have seen age well are the ones where a named person reviews performance on a schedule and has authority to switch the thing off.

The department was not consulted. Radiologists who were told about a tool rather than asked about it will find reasons it does not work, and often they will be right, because they know things about their workflow that were never in the requirements.

Where this actually goes

The persistent framing is AI versus radiologists, and it continues to be the wrong axis. Imaging volume has been growing faster than the radiology workforce for two decades in most countries. The problem is throughput and turnaround, not a shortage of interpretive skill.

Against that problem, the realistic contribution is: read the normals faster, find the urgent cases sooner, do the measuring automatically, and reconstruct better images from less scanner time. None of that is a replacement narrative. All of it is worth having.

The specialty will change, and I would not tell a trainee otherwise. But the change looks less like being automated away and more like what happened when PACS replaced film โ€” the work is reorganised around a new tool, the tool absorbs the repetitive part, and the judgement stays where it was.

Tags

AIRadiologyHealthcareMedical ImagingClinical ApplicationsTechnology

About the Author

S

Taresh Sharan

PhD ยท IIT BHU

Research Scientist ยท Bangalore, India

PhD in Biomedical Engineering from IIT (BHU) Varanasi. Research Scientist based in Bangalore. Author of 200+ articles across AI, finance, photography, technical writing, careers, literature, and corporate ethics. Builder of the free Money and Health apps on this site.

Medical AITechnical WritingPhotographyPersonal FinanceLiterature
Full profile
Radiology AI: How to Read the Accuracy Claims, and What Is Actually Deployed | Sharan Initiatives | Sharan Initiatives