The case for screening mental health conditions earlier is not difficult to make. NAMI puts the average delay between the onset of mental illness symptoms and treatment at around 11 years. Anxiety and depression are common, treatable, and get harder to treat the longer they run untreated.
So the pitch writes itself. Depression and anxiety leave traces in behaviour — what people write, how they speak, when they sleep, how much they move, how they use their phones. Machine learning is good at finding patterns in high-dimensional behavioural data. Point one at the other and you get earlier detection at a scale no clinical workforce could match.
I build machine learning systems in medicine. This is the domain where I am least comfortable with how the pitch is usually presented, and the discomfort is not about the technology. It is about what these systems are actually being trained to predict, and what happens after they fire.
The reference standard problem, which is not a small one
Every supervised model needs a ground truth. In imaging, the ground truth is imperfect but tangible: a biopsy result, a confirmed outcome, an expert consensus read of the same pixels.
In psychiatry there is no equivalent. There is no blood test for depression and no imaging finding that establishes it. The reference standards available are a structured clinical interview, which is expensive and has its own inter-rater variability, or a self-report instrument like the PHQ-9, which is cheap and is what nearly everyone uses.
This matters more than it sounds. A model trained against PHQ-9 scores does not learn to detect depression. It learns to predict how someone will answer nine questions. Those two things overlap substantially, and they are not the same, and the places they diverge are systematic rather than random — they track culture, language, how comfortable someone is disclosing distress, and whether the presentation is somatic or emotional.
When a paper reports that a model detects depression with some accuracy, the precise claim is almost always that it agrees with a screening questionnaire. That is a worthwhile thing to be able to do. It is a weaker claim than the one usually made on the model's behalf.
What the evidence on language actually supports
The most studied signal is text. There is a substantial literature on predicting depression from social media language, and it has now been synthesised.
A 2025 systematic review and meta-analysis of 36 studies on text-based depression prediction from social media reported a large overall effect, with a pooled correlation of about 0.63. Demographic features, activity patterns, language features and temporal features all contributed. Platform type and modelling approach were significant moderators — which is to say results varied considerably depending on where the data came from and how it was analysed.
A real effect, then, and not a small one. But the label question comes back immediately. In most of these studies, the "depressed" group is defined by users who publicly disclosed a diagnosis, or by membership of particular communities, or by responses to a survey distributed to volunteers. People who publicly announce a diagnosis on a social platform are not a random sample of people with depression. They are people comfortable disclosing, in a particular language, on a particular platform, often while actively unwell.
A model that separates that group from a general-population comparison group is solving an easier task than screening. Some of what it learns is the disorder. Some of it is the act of disclosure, the vocabulary of a support community, and the demographics of who uses which platform.
This is not a reason to dismiss the work. It is a reason to be precise about what transfers to a clinical setting, where the person has not self-selected and is not writing for an audience.
The base rate problem, stated honestly
There is an arithmetic issue that the field has known about for years and that promotional material consistently omits.
Take a screening model with 80 percent sensitivity and 80 percent specificity, applied to a population where 10 percent have the condition. Out of a thousand people, it correctly flags 80 of the 100 affected — and also flags 180 of the 900 unaffected. Of 260 positives, 180 are wrong. Fewer than a third of flagged people have the condition.
Make the outcome rarer and it gets much worse. This is why suicide risk prediction, despite enormous effort and genuine clinical urgency, has been so difficult: systematic reviews of these models have repeatedly found positive predictive values in the low single digits, because the outcome is rare enough that even excellent discrimination produces overwhelming false positives. A model can have a respectable AUC and still be nearly useless as a trigger for action.
For a screening tool, a low positive predictive value is not automatically disqualifying — screening is supposed to over-refer into a cheap confirmatory step. It becomes disqualifying when the confirmatory step does not exist, or when being flagged carries consequences.
Which brings up the part that worries me most
Detection without treatment capacity is not a benefit. In most health systems, mental health services are the bottleneck, with waiting lists measured in months. A screening programme that identifies three times as many people with probable depression, in a service that cannot see the people it already knows about, has not helped anyone. It has generated distress and a longer queue.
I would want to see the referral pathway costed and staffed before the screening tool is switched on, and I would treat any proposal that does not address this as incomplete.
Then there is the consent and context question, which is sharper here than in imaging. A chest radiograph is taken because a clinician ordered it for a reason the patient knows about. Passive behavioural monitoring is different in kind. Phone usage patterns, typing dynamics, sleep inferred from accelerometer data, language from messages — this is data that people generate without thinking of it as clinical, and inferring a psychiatric state from it is not something most people would anticipate when agreeing to terms of service.
The settings where this gets proposed make it worse. Employer wellness programmes, insurers, universities, schools. In each case the party doing the screening has interests that are not purely the subject's, and a mental health inference is exactly the kind of information that can be used against someone in employment, insurance, custody or immigration contexts. A clinical screening result sits inside a framework of confidentiality and professional obligation. The same inference produced by an app does not automatically inherit that protection.
And there is the differential validity problem. The language of distress is culturally specific. Somatic presentation is more common in some populations than others; emotional vocabulary differs across languages and communities; what counts as a normal level of expressed negativity varies enormously. A model developed predominantly on one population and evaluated in aggregate will underperform on others and report a good overall number while doing so. Subgroup evaluation is the only way to see this, and it is not standard practice.
Where I think this is genuinely useful
Having listed the objections, I do not think this is a dead end. A few applications seem to me both defensible and valuable.
Making standard screening instruments easier to administer and harder to skip. Most of the detection gap is not that the PHQ-9 is insufficiently accurate; it is that nobody administers it. Digital delivery, sensible prompting and integration into routine visits address the actual bottleneck without needing any inference at all.
Longitudinal monitoring for people already in treatment, with their explicit and specific consent. Tracking change in someone whose baseline you know is a much easier and much better-defined problem than detecting a condition in a stranger, and it maps onto a real clinical need: knowing whether a treatment is working between appointments.
Triage support within services that already exist, to help order a waiting list by likely severity rather than by date of referral.
Reducing the burden of clinical documentation so that clinicians spend more of a session with the patient.
What these have in common is that they support a clinical relationship rather than substituting for one, and that the person being assessed knows it is happening and agreed to it.
The framing I would resist is the one where AI solves the mental health crisis by finding everybody. The crisis is mostly a shortage of clinicians and a shortage of funding. Detection is the part of the pipeline we are already least bad at.
Tags
Taresh Sharan
support@sharaninitiatives.com