If you want to know whether medical AI can work, ophthalmology is the case to study. Not because it is the most sophisticated โ it is not โ but because it is the one place where the technology has been through the full cycle: research result, regulatory authorisation, real-world deployment, and the discovery of what goes wrong when you actually install it in a clinic.
The reason the retina got there first is mostly anatomical luck. It is the only place in the body where you can see blood vessels and neural tissue directly, non-invasively, with a camera that costs a fraction of a CT scanner. The image is two-dimensional, standardised, and quick to acquire. Diabetic retinopathy has an established grading scale, an enormous screenable population, and a clear referral threshold. If you were designing a first problem for medical computer vision, you would design something close to this.
What the Evidence Actually Says
Two numbers are worth carrying around.
The first is the aggregate. A 2025 systematic review and meta-analysis in the International Journal of Retina and Vitreous pooled AI performance across diabetic retinopathy screening studies and found sensitivity of 87.7 percent and specificity of 90.6 percent. The review found AI performing better than general ophthalmologists and comparably to retina specialists.
Hold on to that pair: high eighties, low nineties. That is what this technology does across the literature as a whole, and it is genuinely good for a screening test.
The second number comes from deployment rather than a benchmark, and it is the more interesting one. Google's diabetic retinopathy algorithm was evaluated across 45 sites in southern India, on several thousand fundus photographs, in the clinical workflow rather than on a curated dataset. For severe non-proliferative or proliferative retinopathy, sensitivity was 97.0 percent and specificity 96.4 percent. For sight-threatening disease overall, 95.9 and 94.9. The clinically important miss rate for severe disease was zero โ some patients were graded moderate rather than severe, but nobody who needed referral was sent home.
That last sentence is what a well-designed screening system looks like. The model made grading errors. It did not make the error that matters, because the system was built so that the consequential threshold had margin on the safe side.
The Deployment Lesson Nobody Quotes
The most instructive piece of work in this field is not a performance paper at all. It is a human-centred field study of a deep learning retinopathy system deployed in clinics in Thailand, which looked at what actually happened when nurses used it in practice.
The findings were sobering. A substantial fraction of images were rejected by the model's quality filter โ images that human graders would have read without difficulty, but which fell below the threshold the model had been given. Lighting in real clinics is not lighting in a study protocol. Some sites had internet connections too slow to upload images reliably, so results that were supposed to be instant were not. Nurses who had been told the system would speed things up found it adding steps, and some worked around it.
None of this is a criticism of the model, which was excellent. It is a demonstration that a model's accuracy and a system's usefulness are different quantities, and that the gap between them is made of image quality thresholds, network bandwidth, workflow fit and the patience of the person operating it.
I think about that study constantly when building anything intended for deployment. The clinical validation is the easy part.
Autonomy, and Why It Was Granted Here First
Diabetic retinopathy screening carries the distinction of being the first application where a regulator permitted AI to deliver a diagnostic result without a clinician reviewing the images. The FDA authorised such a system in 2018 for use in primary care.
The pivotal trial behind that authorisation reported sensitivity in the high eighties and specificity around ninety. Lower than several later systems, and the authorisation was granted anyway. The reason is the comparator. The relevant question was never whether the AI beats an ophthalmologist. It was whether AI screening in a primary care clinic beats what actually happens to those patients now, which is frequently no screening at all, because getting a person with diabetes to a separate specialist appointment has a failure rate that dwarfs any algorithm's.
That is the correct way to evaluate a screening intervention, and it generalises. The right baseline for medical AI is usually current practice in the setting where it will be used, not expert performance in a setting the patient cannot reach.
Where It Works Less Well
Glaucoma is much harder than the marketing suggests. The difficulty is not the imaging; it is that glaucoma has no clean reference standard. Diagnosis integrates optic disc appearance, retinal nerve fibre layer thickness, visual field testing, intraocular pressure and change over time. Experts disagree with each other on the same case. A model trained against one set of experts' labels inherits their disagreements, and a cup-to-disc ratio measured from a photograph is a weak proxy for a diagnosis that is fundamentally about progressive functional loss. Models that detect advanced glaucomatous damage work reasonably. Models claiming to catch early glaucoma should be examined closely for what they were validated against.
AMD is intermediate. OCT-based triage has strong published results โ a widely cited 2018 study at a London eye hospital showed referral decisions from OCT volumes performing on par with experts in retrospective evaluation. Detecting fluid and quantifying it for treatment monitoring is genuinely useful and increasingly used. Predicting which patients with early disease will progress is much less settled.
Oculomics โ inferring systemic disease from the retina โ is the most exciting and the least mature. Models can predict age, sex, smoking status and cardiovascular risk factors from fundus photographs with accuracy that surprised everyone when first reported, including the researchers. The signal is real. What it means clinically is not established, and whether these predictions add anything beyond information already available from a blood pressure cuff and a questionnaire is an open question in most cases. Claims that retinal imaging detects Alzheimer's pathology are research-stage; there is interesting work on retinal changes associated with neurodegeneration, and it is nowhere near a clinical test.
The Bias Problem Is Concrete Here
Retinal pigmentation varies with ethnicity, and fundus images vary accordingly. Camera hardware varies. Disease prevalence and presentation vary between populations. A model trained predominantly on one population can perform measurably worse on another, and in ophthalmology this has been documented rather than merely hypothesised.
This matters more than usual because the strongest argument for AI screening is extending care to underserved populations โ precisely the groups least represented in training data. A system that works well in the tertiary centre where it was built and less well in the rural clinic where it is most needed inverts its own justification. Demanding validation data from the deployment population is not a nicety; it is the difference between the technology delivering on its premise and undermining it.
What This Means in Practice
If you run a screening programme, the realistic offer is this: AI grading can handle the large majority of images, referring the abnormal and the uncertain to human graders, at a throughput and cost that makes screening feasible for populations that currently go unscreened. That is a substantial public health gain and it is achievable with technology that exists and is authorised today.
What you have to build alongside it is less exciting: a quality gate that fails gracefully rather than rejecting a fifth of your images, a referral pathway with capacity to absorb the positives, a monitoring system that would notice if performance drifted after a camera was replaced, and honest communication to patients about what a negative screen does and does not mean.
If you are a clinician, the part of your job under pressure is routine grading of normal images, and that is the part worth automating. The part that is not going anywhere is everything that happens after a positive result: the examination, the judgement about an atypical presentation, the treatment decision, the surgery, the conversation with a frightened patient about their sight.
The World Health Organization estimates that a very large number of people worldwide have vision impairment that could have been prevented or has gone unaddressed. Most of that gap is not caused by a shortage of diagnostic accuracy. It is caused by a shortage of access. AI happens to be well suited to the access problem specifically, and that โ rather than outperforming specialists โ is the honest case for it.
Tags
Taresh Sharan
support@sharaninitiatives.com