AI in Radiology: What These Systems Actually Do
A working ML researcher's view of medical imaging AI — what is genuinely deployed, why retrospective accuracy numbers mislead, and what breaks in real departments.
Artificial Intelligence is revolutionizing medical imaging by enabling faster, more accurate diagnoses. From detecting tumors to analyzing X-rays, AI is transforming how we approach healthcare diagnostics.
Explore our collection of articles in AI & Medical Imaging
A working ML researcher's view of medical imaging AI — what is genuinely deployed, why retrospective accuracy numbers mislead, and what breaks in real departments.
Benchmark charts will not tell you which model to use. What the real differences between GPT-4o, Claude, Gemini and open weights are, and how to test them on your own work.
Why on-device models became practical, what they are genuinely better at, where they fall down, and why "local" is not the same thing as "compliant".
Showing a model a screenshot beats describing it — but fluent image descriptions are not accurate ones, and in medical imaging the two have come apart almost entirely.
Reasoning plus reliable tool calling turns a text generator into something that acts. That is a real shift — and compounding errors make long-horizon autonomy much harder than the demos suggest.
Finding molecules was never the bottleneck. What AI genuinely changed, what the clinical results actually show, and the three questions that separate substance from marketing.
Voice and behavioural markers are real science. But models are trained to predict questionnaire scores, and base rate arithmetic means most people a screening tool flags do not have the condition.
The compounding-error arithmetic, why coding agents worked first, why multi-agent architectures usually make things worse, and why prompt injection has no clean fix.
What is actually cleared, what the randomised evidence shows, why radiotherapy contouring succeeded where flashier applications stalled, and the three questions to ask of any claim.
The detail everyone omits from the famous pathology AI benchmark, why stain variation makes generalisation harder here than in radiology, and what is genuinely cleared for clinical use.
Retinal screening went from research to regulatory authorisation to real clinics — and the clinics taught us more than the benchmarks did. What the evidence supports, and where it does not.
Saliency maps, LIME and SHAP answer a narrower question than clinicians are asking. What these methods actually measure, where they fail, and what to show instead.
Before a model can read a slide, the slide has to be a file. Why scanning, storage and validation determine whether computational pathology happens at all — and what it buys before any AI.
Headline accuracy numbers say almost nothing about clinical risk. A working ML engineer's account of the four failure modes that matter in deployed medical imaging models, and what they imply for how you deploy them.
Federated learning lets hospitals train a shared model without moving patient data. It is genuinely useful and routinely oversold — here is the real threat model, the real benefit, and the parts of the project that actually take the time.
Machine learning extended clinical decision support into images and free text, but it did not change why these systems fail. Alert fatigue, automation bias, and the gap between retrospective validation and live deployment.
Turning a tumour into a few hundred numbers is a good idea with a difficult history. What radiomic features actually measure, why so many published signatures fail to replicate, and where the field stands now.
Portals now deliver imaging reports straight to patients, but the reports are written for other doctors. What the sections are for, which words sound worse than they are, and why incidental findings are so common.
Diabetic retinopathy screening was the first task the FDA let AI perform without a clinician in the loop. What the pivotal trial numbers really mean, what a field deployment in Thailand revealed, and why this success story does not generalise as easily as people claim.
Over a thousand AI devices have been authorised for medical imaging, most of them on substantial equivalence rather than outcome evidence. A guide to the evidence ladder, what the landmark studies really showed, and the five things deployed radiology AI actually does.
Better classifiers address one slice of diagnostic error. The bigger determinant of whether a clinical model helps or harms is a decision made before any training starts: what you chose to predict.
Models can predict depression from language with a real, measurable effect. Whether that becomes useful screening depends on what the ground truth actually was, what a positive result costs, and whether anyone can see the people it finds.
Gigapixel slides, slide-level labels, and stain that differs between laboratories make pathology a distinct machine learning problem. But the real bottleneck is that most laboratories still read glass.
Software probably looked at your last scan before a radiologist did. A plain-language account of what these models actually compute, the five jobs they do in hospitals, and how to read the famous studies.
Cleared is not approved, the intended use statement is the foundational document, and the evidence package is mostly about process. What the regulatory regime for imaging AI actually asks of a development team.