Do mental-health chatbots help university students?
A preregistered review found small short-term symptom advantages over waitlists or low-intensity materials, based on three depression and four anxiety comparisons. Two active conversational-control trials found no reliable added benefit, and certainty was low to very low.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The review screened 3,294 unique records and maintained a registry of 23 independent study clusters, but only three depression and four anxiety comparisons met the strict rules for pooling against non-active or low-intensity controls.
- 2Short-term pooled effects favoured chatbots for depression (Hedges g -0.413) and anxiety (-0.435), with low certainty. Two MYLO trials against an active conversational control showed no reliable incremental effect and had very-low-certainty evidence.
- 3The interventions ranged from rule-based to generative or hybrid systems and lasted about seven days to 16 weeks. Results do not establish crisis safety, clinical replacement, durability or benefit from every architecture.
Research topic
Whether participant-facing AI conversational interventions reduce depression and anxiety symptoms among university students, with effects separated by comparator intensity
The direct answer: possibly, compared with doing little
Mental-health chatbots may produce a small short-term reduction in university students' depression and anxiety symptoms when compared with a waitlist or low-intensity material. The pooled standardised effects were minus 0.413 for depression and minus 0.435 for anxiety, where negative values favoured the conversational intervention. Yet only three depression comparisons and four anxiety comparisons supplied compatible data, and the review graded that evidence as low certainty.
The answer changed when the comparator could converse. Two trials of the MYLO problem-solving system against the ELIZA conversational control found no reliable incremental effect for depression or anxiety; confidence intervals were wide and crossed zero. That evidence was graded very low certainty. The contrast suggests that attention, expectancy, monitoring and the experience of being answered may contribute to apparent benefit. It does not show that therapeutic content is useless, only that current active-control evidence is too sparse to isolate it.[1]
The Impact Brief · Free
Follow the evidence in health & life sciences.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
How the review narrowed thousands of records
The authors prospectively registered the review in PROSPERO as CRD420261323382 and followed PRISMA. They searched PubMed, Embase, PsycINFO, CENTRAL and Web of Science through 24 August 2026 without a retrieval-stage language restriction. From 3,786 exports, they removed 451 DOI or PMID duplicates and 41 confirmed title-and-author duplicates, leaving 3,294 unique records for two-reviewer title and abstract screening. Forty-six records entered the updated full-text or research-level assessment.
The working registry contained 23 independent study clusters covering rule-based, generative and hybrid systems, with cognitive behavioural therapy, mindfulness, psychoeducation, problem-solving and mixed frameworks. Intervention periods ranged from about seven days to 16 weeks. The review retained feasibility and non-randomised evidence for narrative context, but restricted causal pooling to randomised trials with attributable, compatible two-group post-intervention means, standard deviations and sample sizes. Multiple reports from one trial were linked rather than double-counted.[1]
What the pooled estimates do and do not mean
For depression versus non-active or low-intensity controls, three comparisons produced Hedges g of minus 0.413 with a 95% confidence interval from minus 0.557 to minus 0.268 and no detected statistical heterogeneity. Four anxiety comparisons produced g of minus 0.435, with an interval from minus 0.603 to minus 0.268 and I-squared of 2.1%. The largest included randomised study described in the review enrolled 995 distressed Israeli students, while other contributing trials were substantially smaller and differed in country, duration and system design.
A standardised effect does not tell a student the number of symptom points they will improve, whether they will notice the change, or how many will recover. Low I-squared with three or four comparisons does not prove interventions are equivalent; heterogeneity tests have little power with so few studies. The estimates are immediate post-intervention averages. They should not be converted into claims about long-term recovery, diagnosis, academic functioning or reduced demand for counselling services.[1]
The active-control result is the most useful caution
Against ELIZA, two MYLO trials produced a depression estimate of plus 0.165, with a 95% confidence interval from minus 0.580 to plus 0.911, and an anxiety estimate of plus 0.190, from minus 0.556 to plus 0.937. Positive values here do not favour MYLO, but the intervals are so imprecise that both benefit and harm remain plausible. The result is not proof of no difference; it is failure to establish a reliable added effect with the available data.
Waitlist and ebook comparisons answer a different question from an active conversation. A student receiving a chatbot may be prompted to reflect, feel monitored, expect improvement or spend more time on coping exercises. Those effects can be valuable, but they are not necessarily evidence that an AI system's therapeutic reasoning works. Future trials need credible active controls that match attention, contact frequency and interface quality while separating the intervention's specific clinical content.[1]
Why certainty remained low
The recurrent risks were inability to blind participants, self-reported symptom outcomes, attrition and incomplete or completer-only analyses. One anxiety contribution used week-eight completers and was rated high risk for missing outcome data. Across the field, systems, therapeutic models and durations differed enough that a single pooled 'chatbot dose' is difficult to define. Architecture-specific subgroup analysis was not credible because technical design was entangled with country, year, therapy and comparator.
The review used random-effects models with restricted maximum likelihood and Hartung-Knapp adjustment, appropriate safeguards when studies are few and heterogeneous. Statistical technique cannot create information the trials did not collect. Funnel plots and small-study tests were not used with fewer than 10 comparisons because they would be hard to interpret. Publication bias, selective availability and language access may therefore remain. The most accurate evidence rating is the authors': low certainty against low-intensity controls and very low certainty against active conversation.[1]
What universities and students should do now
A chatbot can be offered as an optional, low-threshold support tool, but not as a substitute for counselling, clinical assessment or crisis care. Students need a clear description of what the system can do, what evidence applies to that exact product, whether conversations are stored or used for training, and how to reach a person. Systems should detect neither risk nor diagnosis unless that function has been prospectively validated; every service still needs a visible human escalation route for self-harm, abuse, psychosis, severe impairment or urgent distress.
Universities should measure adverse events, dropout, privacy complaints, symptom deterioration, delayed help-seeking and inequitable access as carefully as engagement. A high conversation count is not a treatment outcome. Procurement should distinguish a rule-based exercise, a generative companion and a hybrid clinical programme rather than treating 'AI chatbot' as one intervention. Students must be able to decline automated support without losing access to human care, disability support or academic accommodation.[1]
Funding, disclosure and what would change the assessment
The review team at Shandong Management University reported receiving no financial support for the work or its publication, declared no commercial or financial conflicts, and said generative AI was not used to create the manuscript. Its transparent registration, study-level report linkage and comparator-stratified analysis strengthen the synthesis, but the result still depends on a small and uneven underlying trial base.
Confidence would increase with adequately powered, preregistered multicountry trials using active conversational controls, complete follow-up and prespecified missing-data handling. Trials should report exact architecture, model version, human involvement, safety procedures and data governance; include clinician-rated or behavioural outcomes alongside validated self-report; and test durability after the intervention ends. Head-to-head comparisons with accessible human, peer and non-conversational digital support would clarify where chatbots add value rather than merely attention. Until then, the evidence supports cautious supplementation, not clinical replacement.[1]
What this means for people
- Students may gain a low-threshold support option, but need an immediate path to a trained person when distress is severe or worsening.
- Universities should not use weak evidence or engagement counts to replace counsellors or narrow access to human care.
- Consent, data minimisation and clear retention rules are essential because mental-health conversations can reveal highly sensitive information.
Global context
The review spans trials from several countries, but campus services, languages, stigma, connectivity and clinical referral systems differ widely. A system tested in one student population cannot be assumed culturally safe or effective elsewhere. Institutions should validate local accessibility and escalation pathways while preserving human services.
What the evidence does not yet show
- Only three depression and four anxiety comparisons were compatible with the main low-intensity-control meta-analyses; active-control estimates relied on two trials.
- Participants could not be blinded and outcomes were mainly self-reported, leaving expectancy and measurement effects possible.
- Attrition, completer analyses and missing data reduced confidence in several underlying trials.
- Rule-based, generative and hybrid systems with different therapies and durations cannot be assumed to share one effect.
- The review addresses short-term symptoms among university students, not crisis care, diagnosis, long-term recovery or unsupervised clinical treatment.
What to watch next
- Large preregistered trials with matched active conversational controls and complete follow-up.
- Independent safety reporting, including deterioration, crisis escalation and delayed human help-seeking.
- Product- and architecture-specific results instead of an undifferentiated chatbot category.
- Longer-term clinician-rated, behavioural and academic-function outcomes.
- Privacy, security and consent audits for highly sensitive conversation data.
Living evidence record
Impact record IAI-0PIEIMD
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Frontiers in Psychiatry published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can AI improve GLP-1 and weight-loss treatment?
A new analysis of research from China and South Korea: predicting treatment response and designing a different drug candidate are promising ideas, but neither establishes that using AI improves weight-loss outcomes.
6 min · 4 sources
Health & Life Sciences
Can a chest X-ray flag osteoporosis?
A peer-reviewed Korean study externally validated an AI prescreener in three cohorts totalling 153,058 people. Its simulated workflow preserved most osteoporosis detections while halving DXA use, but it has not yet proved benefit in prospective care.
7 min · 2 sources
Health & Life Sciences
Does a digital human improve AI stress counselling?
Seventy-one Beijing university students rated an avatar-based CBT system easier to use than the same text chatbot after one session. The study measured experience—not symptom improvement, long-term safety or therapeutic effectiveness.
7 min · 1 source
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.