Diabetic retinopathy is the textbook case for screening. The disease is common, it is silent until it is advanced, the damage it does is irreversible, and there is effective treatment if you catch it in time. The National Eye Institute puts the reduction in risk of severe vision loss from early detection and timely treatment at up to 95 percent.
It is also the textbook case for what goes wrong with screening programmes. Detection requires a retinal image and someone qualified to grade it. The people who need screening are numerous โ roughly a third of people with diabetes have some degree of retinopathy โ and the people qualified to grade are not. Guidelines say annual screening; adherence in most health systems falls well short of that, and falls shortest exactly where ophthalmologists are scarcest.
This mismatch is why diabetic retinopathy became the first place autonomous diagnostic AI was allowed to operate without a doctor in the loop, and why it remains the most informative case study the field has.
The regulatory first
In April 2018 the FDA authorised, through the De Novo pathway, an AI system for detecting more-than-mild diabetic retinopathy from fundus photographs in adults with diabetes. What made it notable was not the accuracy. It was the intended use: the system returns a screening result directly, in a primary care setting, without a clinician interpreting the images.
The pivotal trial was prospective, ran at ten US primary care sites, and enrolled 900 patients, with images acquired by existing staff rather than by ophthalmic photographers. Against a reading-centre reference standard it reported sensitivity of 87.2 percent, specificity of 90.7 percent, and โ the number that gets ignored and should not โ an imageability rate of 96.1 percent, meaning about one in twenty-five patients could not be assessed at all.
Those figures are worth sitting with, because they are far below what you see quoted in coverage of retinal AI. Research papers on this task routinely report AUCs above 0.99. The gap is not that the deployed system is worse technology. It is that the pivotal trial measured something much harder: a prospective, unselected population, images taken by whoever was in the clinic that day, on a specified camera, in a real workflow, against a rigorous reference. That is what the operating performance of a screening tool actually is.
I bring this up whenever someone shows me a retinal model with a spectacular retrospective number. The relevant question is not how it scores on a curated dataset. It is what happens when the person holding the camera has had twenty minutes of training and the patient will not stop blinking.
The deployment study everyone should read
The most useful piece of work on medical AI deployment I know of is not a modelling paper. It is a 2020 human-factors study of a retinal screening model rolled out across eleven clinics in Thailand.
The researchers interviewed and observed the nurses actually using the system, and what they found was a collision between the model's quality requirements and the conditions of a resource-constrained clinic. The model had been built to refuse images below a quality threshold โ a sensible safety design. In the field, a substantial fraction of images were rejected, often because the room could not be darkened enough for adequate pupil dilation, or because the camera and the lighting were what the clinic already had rather than what the model was developed on.
The downstream effects were the interesting part. A rejected image meant the patient had to be re-photographed or referred anyway, which removed the time saving the system was supposed to deliver. Nurses who had been told the system was accurate experienced it as unreliable. In clinics with poor connectivity, upload times added minutes per patient. Some nurses started working around the system.
Nothing in that story is about model accuracy. All of it determines whether the deployment works.
The same programme, once those issues were addressed, went on to report prospective real-time results from a multi-site national screening context โ which is the right arc. Build, deploy, discover the model was the easy part, fix the environment, then measure again.
The fairness problem is specific, not vague
Retinal AI has a well-documented performance gap across skin pigmentation, which affects fundus appearance through choroidal and retinal pigmentation.
One study quantified it directly: a baseline retinal diagnostic system showed roughly a 12.5 percentage point accuracy gap between lighter-skinned and darker-skinned individuals โ about 73 percent versus 60.5 percent. That is not a subtle disparity; it is the difference between a useful screening tool and one that is barely better than chance for part of the population it is deployed on.
The same work showed the gap was largely a data problem rather than an intrinsic one. Augmenting training with synthetic fundus images covering the underrepresented group narrowed the difference to under a percentage point, with the majority group's performance essentially preserved.
Two lessons follow. First, this class of disparity is fixable, and the fix is upstream โ get the training distribution right, or repair it deliberately. Second, and more importantly, you only find it if you stratify your evaluation. Aggregate accuracy on a test set that shares the training set's demographic skew will report a good number and hide the gap completely. Subgroup reporting should be the default in any validation of a screening model, and it still very often is not.
What I would want to know before deploying one of these
The technology is genuinely ready for this task in a way it is not ready for most tasks. That makes the remaining questions operational.
Which camera, and does the clinic have it? Model performance is tied to the acquisition device it was validated on. A system authorised with one camera is not automatically valid with another.
What is the ungradable rate here, not in the trial? This drives the real referral burden and the real staff time, and it varies enormously with lighting, operator training, cataract prevalence and whether dilation is used.
What happens to a positive result? An autonomous screening tool that flags patients into a referral pathway with a nine-month waiting list has not helped anybody. The ophthalmology capacity to absorb the referrals is the binding constraint, and screening capacity without treatment capacity just moves the queue.
How is performance monitored after go-live, by subgroup? Cameras get replaced, staff turn over, the patient mix drifts.
Who is accountable for a missed case? Autonomous means the software returns a result without clinician review. The liability and governance arrangements around that are not the same as for an assistive tool, and they need to be settled before the first patient, not after the first miss.
The realistic picture
Autonomous retinal screening is one of the few places where I would say AI has clearly moved from promising to useful. The task is narrow, the reference standard is well defined, the clinical pathway after a positive result is established, and there is a genuine access gap the technology addresses rather than a workflow inefficiency it optimises.
But the reason it works is not that the models are extraordinary. Retinal lesion detection is, in machine learning terms, a relatively tractable problem. It works because the surrounding conditions happen to be favourable, and the deployments that have succeeded got there by treating the camera, the room, the nurse and the referral pathway as part of the system rather than as someone else's problem.
Most medical AI does not have conditions that favourable. That is the part of the diabetic retinopathy success story that does not generalise, and it is the part most often left out when it gets cited as a template.
Tags
Taresh Sharan
support@sharaninitiatives.com