Radiology is where medical AI has gone furthest, and it is also where the distance between the literature and the reading room is easiest to measure. I work on the building side of this, and the thing I would most like to convey is how different those two worlds look.
The research picture is one of rapid, broad capability. The deployment picture is narrower, slower, and dominated by concerns โ integration, reimbursement, monitoring, liability โ that almost never appear in a paper.
Both pictures are accurate. Understanding the field means holding them at once.
The Regulatory Baseline
The most reliable public map of what is actually cleared for clinical use is the FDA's list of AI-enabled medical devices. It runs to over a thousand entries now, and radiology accounts for the large majority of them โ more than every other specialty combined by a wide margin.
That number sounds like saturation and is not. Read the entries and a pattern emerges: almost all of them are narrow. One finding, one modality, one anatomical region, cleared through the 510(k) pathway on the basis of substantial equivalence to a predicate device. A very small number are authorised to operate without a clinician reviewing the images.
The landmark exception is worth knowing precisely because it is exceptional. In 2018 the FDA authorised an autonomous diabetic retinopathy screening system for use in primary care โ the first AI permitted to deliver a diagnostic result without physician interpretation. The pivotal trial reported sensitivity of 87.2 percent and specificity of 90.7 percent against a reading-centre reference standard.
Those figures repay attention, because they are what a genuinely autonomous cleared system looks like: high eighties, low nineties. Not 99. And the indication is deliberately narrow โ screening a defined population for one condition, in a setting where the alternative is very often no screening at all. That last clause is the whole argument. Autonomy was justified not by the AI being better than an ophthalmologist, but by the comparator in that setting being nobody.
The Three Tiers, and Where Each Really Sits
Detection is mature, within limits. Flagging intracranial haemorrhage on a head CT, a pneumothorax on a chest radiograph, a large vessel occlusion on a CT angiogram, a pulmonary nodule on a screening CT. These are deployed at scale, they work, and the clinical value comes mostly from speed rather than from accuracy. A finding detected at three in the morning as the scan reconstructs, rather than when a radiologist opens the study, changes what happens to the patient.
Where I would push back on the standard framing is the accuracy comparison tables. Head-to-head accuracy figures against radiologists are usually measured under conditions that disadvantage the human badly: no clinical history, no prior imaging, no ability to ask a question, a forced binary decision on a single image. Radiologists do not work that way, and a comparison that removes everything that makes a radiologist useful is not measuring what it claims to.
Characterisation is genuinely advancing and less settled. Estimating malignancy probability rather than just flagging a mass, grading aggressiveness, measuring response across serial studies. Automated volumetric measurement is the underrated win here โ humans are inconsistent at estimating volume from cross-sectional images, and consistency across time points is what actually matters for tracking a lesion.
Inferring molecular or genetic characteristics from imaging features โ imaging genomics, radiomics โ is the part I would treat with the most caution. The published results are real and the reproducibility record is poor. Radiomic features are notoriously sensitive to acquisition parameters, reconstruction kernel and scanner vendor, which means a model can be learning the scanner as much as the biology. Some of it will survive. Much of it will not replicate.
Prediction is research. Forecasting cardiovascular events, future cancer, cognitive decline from current imaging is an active and legitimate field with encouraging retrospective results. Almost none of it has cleared prospective validation, and the failure mode is subtle: a model that predicts who gets diagnosed may be learning who gets screened, who has access to care, who returns for follow-up. Those correlations are strong, they are not causal, and they do not transfer to a population with different access patterns.
The Evidence That Actually Moved Me
Most AI radiology studies are retrospective. A randomised trial in a live screening programme is a different order of evidence, and there have now been a few.
The one worth knowing is a large Swedish randomised trial of AI-supported mammography screening, which published interim results in 2024. Women were randomised to standard double reading or to AI-supported reading. The AI-supported arm detected meaningfully more cancers without an increase in false positives, and cut the screen-reading workload substantially.
That is the strongest evidence in this field that I am aware of, and note its shape. It is not "AI outperformed radiologists". It is "a specific AI system, used in a specific way, inside an existing screening programme with radiologists still in the loop, improved outcomes and reduced workload". Every clause in that sentence is load-bearing. Change the population, the screening protocol, or the system and you are outside what was tested.
The operational counterpart is stroke triage, where AI notification systems have been repeatedly reported to shorten time from scan to intervention. Whether that translates into better neurological outcomes at ninety days is a harder question, and the evidence there is thinner than the process-measure evidence.
Treatment Planning: the Quiet Success
Everyone writes about diagnosis. The place where AI has changed daily practice most concretely is radiotherapy planning.
Contouring โ outlining the tumour and every organ at risk that must be spared โ traditionally took a clinician hours per plan, and the variability between two clinicians contouring the same scan was substantial. Automated segmentation now produces those contours in minutes, and the clinician reviews and edits rather than draws from scratch.
This is genuinely deployed, in routine use, and the reason it worked where flashier applications stalled is instructive. The output is a geometric object a human can inspect at a glance and correct directly. There is no black box problem, because the contour either follows the organ boundary or it does not. Errors are visible, correctable, and caught before they matter.
If you want a heuristic for which medical AI applications will succeed, that is a good one: those whose output a clinician can verify quickly and fix locally.
Surgical planning from imaging โ three-dimensional reconstruction, resection margin visualisation, implant sizing โ is real and expanding, though the AI contribution is often segmentation with conventional planning software doing the rest. Claims about predicted surgical outcomes or procedure-level recommendations should be read as decision support at best.
What Nobody Puts in the Brochure
Monitoring after deployment is the largest unsolved operational problem. A model's performance drifts as scanners are replaced, protocols change, and the patient population shifts. Unlike a drug, a model can degrade silently โ there is no adverse event report for gradually worse calibration. Very few institutions have real monitoring in place. The FDA's predetermined change control framework addresses the update path and puts a substantial ongoing burden on manufacturers, which is correct and expensive.
Automation bias is measurable and uncomfortable. Studies in which radiologists are shown deliberately incorrect AI output have found their own performance degrades โ including among experienced readers. This is the cost side of the assistance ledger and it is real. A tool that improves average performance while making the residual errors more correlated and harder to catch is not obviously a good trade, and it is not how these tools are evaluated.
Integration determines adoption more than accuracy does. If a result does not appear inside the PACS viewer at the moment the study is opened, it will be ignored. I have watched good models fail on this and mediocre ones succeed because they fit the workflow. It is not a technical detail; it is the product.
Reimbursement is the quiet bottleneck. A handful of AI services have established payment pathways. Most have not, which means the hospital pays for the software and the benefit accrues as faster throughput or avoided harm โ real value, hard to put on an invoice. This, more than any technical limitation, explains why deployment lags capability.
The Radiologist Question
The honest answer is that nothing currently deployed threatens the profession, and the near-term pressure runs the other way: imaging volume per radiologist has been rising for years, and AI that shaves time off routine work is addressing a staffing problem, not creating one.
But "AI will not replace radiologists" is too comfortable a formulation. What changes is the composition of the work. If routine detection and measurement is increasingly machine-assisted, the human residual is the hard cases, the clinical correlation, the conversations, and the oversight of the machines. That is a more cognitively demanding job, not a smaller one, and it requires skills โ understanding what a model's confidence means, recognising when a case is outside its distribution โ that were not part of anyone's training a decade ago.
The trainees I would worry about are the ones who learn on AI-assisted reads without ever developing independent pattern recognition. Deskilling is a genuine risk and there is no good answer to it yet.
Reading the Claims
Three questions handle most of what you will encounter.
Retrospective or prospective? Almost everything published is retrospective, and performance drops on prospective evaluation with enough regularity that you should assume it.
Internal or external validation? A model evaluated on held-out data from its training institution tells you it learned something. A model evaluated at a different hospital, on different scanners, in a different population tells you whether it learned the right thing.
What exactly was cleared, and for what? "FDA cleared" is not a quality statement. It is a statement that a specific device, for a specific indication, was found substantially equivalent to a predicate. Read the indication for use. It is often much narrower than the marketing around it.
The technology is real, the best of it is genuinely valuable, and the gap between what it can do in a paper and what it does in a hospital is where all the interesting work still is.
Tags
Taresh Sharan
support@sharaninitiatives.com