Can AI prioritise cognitive assessment in sleep apnoea?
A random-forest model separated concurrent mild cognitive impairment in an external Chinese cohort, but overestimated probabilities and was evaluated almost entirely in men. It is not a diagnostic or prognostic tool.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1Development used 280 Changchun patients, algorithm selection used another 120, and external validation used 192 patients from Shenzhen.
- 2External ROC-AUC was 0.957 and PR-AUC 0.809, but calibration showed systematic probability overestimation with an intercept of -0.616 and slope of 0.701.
- 3The model identifies concurrent MCI classification rather than future decline, and performance evidence came almost entirely from male patients.
Research topic
Whether an interpretable model using sleep, mood, lifestyle and symptom measures can prioritise Chinese adults with obstructive sleep apnoea for formal cognitive assessment
The answer: the model separated groups well, but its raw risk numbers did not travel cleanly
In an independent Shenzhen sample, the selected random-forest model achieved a receiver-operating-characteristic area under the curve of 0.957 and a precision-recall area of 0.809 for concurrent mild cognitive impairment among people with obstructive sleep apnoea. Those figures indicate strong ranking discrimination in the study population. They do not mean that a predicted probability can be used unchanged in another clinic.
External calibration showed the problem: an intercept of -0.616 and slope of 0.701 indicated systematic overestimation and predictions that were too extreme. The authors explicitly position the model as a way to prioritise formal cognitive assessment, not diagnose MCI. It did not predict who would later deteriorate, compare an AI-assisted pathway with ordinary care, or show that referral decisions, treatment or quality of life improved.[1]
The Impact Brief · Free
Follow the evidence in health & life sciences.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
Who was included and how MCI was defined
The multicentre cross-sectional study included 400 eligible patients with obstructive sleep apnoea in Changchun. Researchers allocated 280 to training and 120 to an internal hold-out used to select the algorithm. A separate group of 192 patients from Shenzhen provided external validation. MCI was classified in 93 of 280 training participants, 38 of 120 internal participants and 57 of 192 external participants—roughly 30% to 33% in each group.
The outcome combined reported cognitive decline, objective evidence, largely preserved independence, exclusion of dementia and assessment-team consensus with physician participation. That is stronger than applying a single questionnaire cutoff, but it also makes the label dependent on a clinical process that other sites may not reproduce exactly. The paper says the retrospective operational specification cannot fully reconstruct the original case-report form from scale scores alone.[1]
How the model was built and selected
The investigators began with 27 candidate predictors, removed redundancy to leave 24 and used agreement between LASSO and Boruta to retain nine. The final variables were coffee consumption, physical activity, apnoea-hypopnoea index, mean oxygen saturation, depression and anxiety scores, daytime sleepiness, perceived stress and sleep quality. Eight algorithms were tuned with repeated ten-fold cross-validation, using SMOTE only inside training folds and precision-recall area as the primary metric.
The selected random forest used a threshold of 0.3228 derived from out-of-fold training predictions and locked before validation. Feature selection, however, was performed once on the complete training set rather than nested within every resampling fold. That choice may make the cross-validation stage more optimistic. The separate Shenzhen cohort is therefore important, though it validates one Chinese setting against another and does not establish transportability to different languages, health systems or patient mixes.[1]
The headline metrics need calibration and denominator context
In the 120-person internal hold-out used for algorithm selection, the random forest reached ROC-AUC 0.953, PR-AUC 0.936 and a Brier score of 0.060. Because that hold-out helped choose the algorithm, the paper appropriately warns that these are selection-conditioned results rather than an unbiased estimate of generalisation. In Shenzhen, ROC-AUC was 0.957, PR-AUC 0.809 and Brier score 0.051.
A high AUC means the model usually ranked a participant classified with MCI above one without it; it does not choose the right referral threshold or guarantee an accurate individual probability. Calibration matters when clinicians explain risk and allocate appointments. A model that overstates probability may send too many people for assessment, increase anxiety and divert capacity, even while its ranking metric remains impressive. Local calibration and capacity-aware threshold testing are necessary before any prospective use.[1]
What this could mean for sleep clinics and patients
Sleep clinics already collect several of the model's inputs, so a carefully evaluated score could help identify people who merit a formal cognitive review rather than waiting for concerns to become obvious. The likely benefit would be prioritisation, not automated diagnosis. Clinicians would still need to investigate medication, mood, education, language, hearing, vascular risk, sleep treatment and other explanations for performance on cognitive tests.
The model should not be applied to women on the strength of this paper because performance evidence came almost entirely from men. Nor should patients read coffee, exercise or mood variables as causal instructions: the cross-sectional model associates current characteristics with a current classification. It cannot show that changing one input prevents cognitive decline. Any clinical pathway should disclose uncertainty, avoid stigmatising labels and preserve access to a full assessment when symptoms conflict with the score.[1]
Limits, funding and evidence that would change the assessment
The study is cross-sectional, uses 592 participants across two Chinese centres and does not follow future cognitive outcomes. The external sample is independent but geographically and clinically related to development. Female representation was too small to support use in women. The once-only feature-selection step was not nested within resampling, and the selected predictors include self-reported lifestyle and symptom measures that can vary by culture, translation and time.
Funding came from Chinese national, provincial, university and education-development programmes. The paper reports ethics approvals at both participating centres. Confidence would rise with prospective validation in balanced-sex, multilingual and international cohorts; prespecified recalibration; comparison with a simpler clinical score; and decision-curve analysis tied to real appointment capacity. A trial showing earlier appropriate assessment without excessive false referrals, anxiety or unequal access would be needed to demonstrate patient benefit.[1]
What this means for people
- Some patients with sleep apnoea could be prioritised for formal cognitive assessment using information already collected in clinic.
- Overestimated probabilities could create anxiety and extra referrals if raw scores are used without local calibration.
- Women should not be assessed with this model until adequate performance evidence exists for them.
Global context
Obstructive sleep apnoea and cognitive concerns affect services worldwide, but questionnaires, referral thresholds and specialist capacity vary. Strong discrimination in two Chinese centres does not establish performance in another health system. Global usefulness depends on reproducible outcome assessment, local calibration, sex-balanced validation and evidence that prioritisation improves access rather than redistributing scarce appointments unfairly.
What the evidence does not yet show
- This cross-sectional model identifies concurrent MCI classification and cannot predict future cognitive decline or dementia.
- The two cohorts came from Chinese sleep-medicine settings; international and language transportability is unknown.
- Performance evidence came almost entirely from male patients, so validity in women is unestablished.
- External calibration showed systematic probability overestimation despite strong discrimination.
- Feature selection was performed once before repeated cross-validation rather than nested inside every resampling fold.
- No AI-assisted referral pathway, clinician behaviour, treatment decision or patient outcome was tested.
What to watch next
- Prospective validation with adequate numbers of women and transparent subgroup performance.
- External testing across languages, countries, primary care and sleep-clinic workflows.
- Local recalibration and comparison with simpler prespecified clinical scores.
- Trials measuring appropriate assessment, false referrals, capacity, anxiety, treatment and longer-term outcomes.
Living evidence record
Impact record IAI-1HQRH8P
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
11 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what BMC Psychiatry published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 11 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can AI predict risks before thymic-tumour surgery?
Models trained on 726 Chinese patients estimated complications, longer stays and severe pain before surgery. External AUCs of 0.660–0.776 make this promising risk support—not evidence for automated care decisions.
8 min · 1 source
Health & Life Sciences
Can AI estimate survival after a Parkinson’s diagnosis?
A peer-reviewed Chinese registry study compared four survival models in 3,148 people with Parkinson’s disease. A transparent Cox model matched the machine-learning alternatives, but validation stayed within the same registry and no clinical-impact study was performed.
9 min · 2 sources
Health & Life Sciences
When should AI autism-screening results reach families?
Twenty-six US participants favoured earlier support but warned against opaque EHR predictions arriving before useful next steps. The study maps implementation concerns; it does not validate a screening model.
7 min · 1 source
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.