Can AI diagnose ADHD reliably yet?
A 54-study meta-analysis found pooled sensitivity of 87% and specificity of 91%, but heterogeneity exceeded 96%, publication bias affected sensitivity, and recurring design weaknesses make clinical portability uncertain.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1Across 54 studies, pooled sensitivity was 0.87 (95% confidence interval 0.83–0.91) and pooled specificity was 0.91 (0.88–0.93).
- 2Residual heterogeneity exceeded 96%, meaning the headline average combines studies that differed substantially in populations, data, modelling and validation.
- 3PROBAST classified 11 studies as high risk of bias, 11 as unclear and 32 as low risk; the analysis domain was the main recurring concern.
Research topic
How accurately published AI models distinguish ADHD and whether pooled performance is dependable enough to support clinical diagnosis

The direct answer: promising averages, not a clinical green light
AI systems in the published ADHD literature often separated diagnosed cases from comparison groups with high apparent accuracy, but the evidence is not consistent enough to support a general claim that AI can diagnose ADHD reliably in routine care. The review combined 54 studies and estimated sensitivity at 87% and specificity at 91%. In plain terms, the average model result suggests that about 13 of every 100 people with ADHD could be missed and about nine of every 100 people without ADHD could be classified as positive, before accounting for differences between studies and real-world prevalence.
Those figures are pooled research estimates, not the performance of one clinically deployable product. ADHD diagnosis involves developmental history, symptoms across settings, functional impairment and exclusion of alternative explanations. A model trained on a particular imaging dataset, questionnaire, eye-tracking task or behavioural signal does not replace that evaluation. The more defensible interpretation is that several data types contain learnable patterns associated with ADHD, while clinical reliability and transportability remain unproven.[1]
What the researchers reviewed
The authors conducted a PRISMA-guided systematic review and meta-analysis of 54 studies using artificial intelligence for data-driven ADHD diagnosis. The studies did not test a single intervention. They used different populations, reference standards, inputs and algorithms, including neuroimaging, electroencephalography and behavioural or clinical data. The meta-analysis pooled sensitivity and specificity and examined whether performance differed across modalities. It found no statistically significant modality differences, which should not be read as proof that the modalities are equivalent: subgroup estimates can be underpowered and can conceal differences in dataset construction and validation.
Risk of bias was assessed with PROBAST. Eleven studies, or 20.4%, were judged high risk, another 11 were unclear and 32 were low risk. The analysis domain was the most frequent problem, a category that can include overfitting, inadequate event counts, poor handling of missing data or evaluation choices that make performance look better than it would on genuinely new patients. That structured appraisal is important because diagnostic-AI literature is especially vulnerable to information leakage and optimistic internal validation.[1]
Why heterogeneity is the central finding
Residual heterogeneity exceeded 96%. This is the statistic readers should keep beside the attractive pooled accuracy figures. It means most observed variation was not explained by sampling error alone. Studies could be measuring different things: a tightly screened research cohort is not the same as a general psychiatric clinic; distinguishing ADHD from healthy controls is easier than separating it from anxiety, autism, sleep disorders or learning difficulties; and validation inside one dataset is weaker than prospective testing in another service.
The authors repeated the analysis after excluding the 11 high-risk studies. Sensitivity remained 0.87, with a 95% confidence interval of 0.82–0.91, and specificity became 0.92, with an interval of 0.88–0.94. That stability is encouraging, but heterogeneity remained 97.9% for sensitivity and 98.5% for specificity. Removing the studies with the clearest bias concerns did not make the evidence base coherent, so the pooled estimate still should not be treated as a performance guarantee for a new clinic or patient group.[1]
Publication bias and the missing evidence
The review found evidence of publication bias for sensitivity but not for specificity. Positive or unusually strong results are more likely to be published, while failed models may remain unseen. If weaker sensitivity studies are missing, the 87% average could be optimistic. Statistical tests for publication bias are themselves imperfect when studies are highly heterogeneous, yet the signal strengthens the case for prospective registrations and reporting of negative evaluations.
A meta-analysis can only pool what the underlying papers report. It cannot repair inconsistent reference diagnoses, restricted samples, poorly documented preprocessing or missing subgroup results. It also cannot show whether an AI-supported pathway improves waiting times, patient experience, clinician agreement or outcomes after diagnosis. Those questions require prospective comparative studies in real services, including people with co-occurring conditions and people who do not fit clean research categories.[1]
What this means for people seeking an assessment
The practical opportunity is assistance rather than autonomous diagnosis. AI might help organise questionnaires, identify records needing review or provide an additional signal when clinicians disagree. Used carefully, that could reduce administrative burden or make assessments more consistent. Used as a gatekeeper, however, a model could deny referral to someone it misses or label someone whose symptoms have another cause. Those harms matter because access to medication, educational support and workplace adjustments can follow from the diagnostic process.
Health systems considering such tools should demand performance in their intended population, not rely on the pooled average. They should examine false negatives and false positives separately, test across age, sex, ethnicity and co-occurring conditions, and keep a route to human reassessment. People should know what data are used and whether the model affects access, while clinicians need explanations that are useful enough to challenge a result rather than merely display a confidence score.[1]
Funding, independence and what would change the assessment
The authors reported support from the China Brain Science and Brain-like Intelligence Technology Major Project, the National Natural Science Foundation of China, the China Postdoctoral Science Foundation and central university funds. They declared no competing interests. The review is peer reviewed and methods based, but its strongest contribution is mapping uncertainty rather than certifying a diagnostic product.
The assessment would change with preregistered, prospective, multicentre comparisons against a robust clinical reference standard. Studies should include consecutive referrals, relevant alternative diagnoses, independent external validation and transparent thresholds chosen before the test set is opened. Reporting should include calibration, subgroup errors, indeterminate cases, clinician overrides and downstream decisions. Evidence that an AI-supported pathway improves access or consistency without increasing harmful misclassification would be more consequential than another high score from a retrospective convenience dataset.[1]
What this means for people
- A missed case could delay support, while a false positive could redirect care or expose someone to unnecessary treatment.
- AI may help clinicians organise evidence, but it should not replace developmental history, impairment assessment and differential diagnosis.
- People need notice, meaningful human review and a route to challenge decisions influenced by a model.
Global context
The included literature spans multiple data types and settings, but ADHD assessment pathways, diagnostic criteria, referral populations and access to specialists vary internationally. A pooled global-looking estimate can therefore obscure local differences. Each health system would need prospective validation in its own languages, demographics and referral pathway before using AI to influence who receives an assessment or diagnosis.
What the evidence does not yet show
- The 54 studies differed greatly in population, input modality, reference standard, algorithm and validation design.
- Residual heterogeneity exceeded 96% and remained above 97% after high-risk studies were excluded.
- Eleven studies were high risk of bias and 11 had unclear risk; the analysis domain was the main concern.
- Publication bias was detected for sensitivity, potentially inflating the pooled ability to identify ADHD cases.
- Pooled research accuracy does not establish clinical utility, fairness, cost-effectiveness or better patient outcomes.
What to watch next
- Prospective, preregistered external validation in ordinary assessment services.
- Comparisons that include common alternative and co-occurring diagnoses, not only healthy controls.
- Age-, sex-, ethnicity- and comorbidity-stratified false-negative and false-positive rates.
- Whether AI-supported pathways improve waiting times, agreement and outcomes without restricting access.
- Full reporting of failed models, prespecified thresholds and independent replication.
Living evidence record
Impact record IAI-0YKE90Q
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
8 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what npj Digital Medicine published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 8 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can AI discover new MRI signs for glioblastoma?
It generated eight candidate visual signs from 106 glioblastoma scans, but only one was externally tested. Two radiologists achieved moderate discrimination between 50 glioblastomas and 50 metastases, so this is a discovery signal—not a clinical diagnostic.
7 min · 1 source
Health & Life Sciences
Can a new fracture dataset make orthopaedic AI more reproducible?
A new open dataset contains 19,940 CT-derived images from 1,579 patients, with multi-view expert labels and patient-level splits. It can support reproducible femoral-neck-fracture research, but its benchmark results are not evidence that an AI system is ready to diagnose or choose treatment.
6 min · 1 source
Health & Life Sciences
Can AI spot liposarcoma on ultrasound?
A peer-reviewed Chinese study reports strong results in a 95-patient internal test, including 0.97 accuracy. But all 317 patients came from one hospital and there was no external validation, so this is a proof of concept—not a clinically cleared diagnostic system.
9 min · 1 source
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.