Can a sleep-study ECG predict future heart risk?
A US study of 38,195 sleep-clinic patients found that a neural-network score added useful risk information for atrial fibrillation, heart failure and death. The gains were smaller for stroke and heart attack, and the retrospective study did not test whether using the score improves care.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The retrospective analysis covered 38,195 patients across Massachusetts General Hospital, Emory University Hospital and Beth Israel Deaconess Medical Center.
- 2Adding the neural-network score to clinical variables improved discrimination most consistently for atrial fibrillation, heart failure and death; gains for stroke and myocardial infarction were small or absent.
- 3Outcomes came from diagnosis codes, the population had already been referred for sleep testing, and the study did not show that acting on the score prevents illness or improves survival.
Research topic
External validation of a neural-network score derived from overnight sleep-study electrocardiography for long-term cardiovascular outcomes
The answer: promising for some outcomes, but not yet a care decision
A neural network applied to the single-lead electrocardiogram recorded during an overnight sleep study identified patients at higher long-term risk of atrial fibrillation, heart failure and death in two external US hospital cohorts. Its extra value was less convincing for stroke and myocardial infarction. The result suggests that an ECG already collected during polysomnography may contain useful risk information beyond routine clinical variables and standard sleep measurements.
It does not show that the system diagnoses hidden disease, selects a treatment or improves patient outcomes. The study looked backwards through health records, and the researchers did not put the score into a live pathway. For a clinician, the model is therefore a candidate for further validation and prospective risk-stratification trials, not a reason to start medication or label a patient with cardiovascular disease.[1]
The Impact Brief · Free
Follow the evidence in health & life sciences.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
The denominator: three sleep-lab cohorts and 38,195 patients
The investigators used three datasets assembled through the Human Sleep Project. Massachusetts General Hospital supplied the development data. The paper reports 17,328 MGH subjects with eligible ECG recordings; after outcome eligibility and splitting, 12,381 patients were used for training and 3,428 for internal validation, giving the 15,809-person MGH outcome cohort described in the abstract. External testing used 9,810 patients from Emory University Hospital and 12,576 from Beth Israel Deaconess Medical Center. Across the outcome analyses, the denominator was 38,195 patients.
Event counts were substantial but varied by endpoint. In the MGH training group there were 1,208 incident atrial-fibrillation events, 987 strokes, 341 myocardial infarctions, 1,099 heart-failure events and 827 deaths. The MGH validation group recorded 368, 312, 82, 313 and 237 respectively. Emory recorded 605 atrial-fibrillation events, 781 strokes, 464 myocardial infarctions, 832 heart-failure events and 498 deaths; Beth Israel recorded 409, 180, 130 and 520 events for the four cardiovascular outcomes.
These were sleep-clinic populations, not a random sample of the public. Referral for polysomnography often reflects suspected sleep apnoea, insomnia, unusual sleepiness or other clinical concerns. That makes the study relevant to sleep services but limits any claim that the same risk relationships or calibration will hold in primary care, wearable-device users or people without a reason for specialist sleep testing.[1]
What the model saw and what it was compared with
The team trained a deep residual neural network with an attention mechanism on the overnight single-lead ECG and expert-labelled sleep stages. The network had previously been trained for arrhythmia detection and was then fine-tuned to predict future outcomes. Follow-up extended to as much as ten years at Emory and seven years at Beth Israel. People with a prior record of the outcome were excluded, and atrial-fibrillation diagnoses made within one month of the sleep study were excluded from that incident analysis.
The comparison was not simply model versus nothing. Cox models included age, sex, body-mass index, smoking, hypertension, diabetes and sleep-study measures including the apnoea-hypopnoea index, arousal index, limb movements, sleep stages and sleep efficiency. The researchers then tested whether adding the neural-network output improved prediction beyond those variables. That is a more meaningful question than reporting the network in isolation.
Outcomes were identified from electronic records using ICD-9 and ICD-10 diagnosis codes. That makes a study of this size practical, but coding is an imperfect proxy for clinical confirmation and may miss disease treated elsewhere. The model also may be learning signals correlated with existing but undocumented illness rather than detecting a modifiable pathway years before disease begins.[1]
The strongest gains were for atrial fibrillation, heart failure and death
For each standard-deviation increase in the neural-network score, adjusted hazard ratios at Emory and Beth Israel were 2.03 and 2.72 for atrial fibrillation, 1.69 and 2.33 for heart failure, and 1.82 and 1.65 for all-cause mortality. Associations were smaller for stroke, at 1.19 and 1.40. For myocardial infarction they were 1.38 at Emory and 1.10 at Beth Israel; the Beth Israel confidence interval of 0.86 to 1.40 included no association.
Discrimination tells the same qualified story. At Emory, adding the neural score changed the concordance index from 0.70 to 0.75 for atrial fibrillation, 0.72 to 0.75 for heart failure and 0.74 to 0.78 for death. At Beth Israel the changes were 0.77 to 0.81, 0.74 to 0.78 and 0.80 to 0.81. For stroke, the changes were 0.71 to 0.71 and 0.75 to 0.76; for myocardial infarction, 0.71 to 0.72 and 0.79 to 0.79.
A hazard ratio is not an individual probability, and a higher concordance index does not show that treatment based on the score helps anyone. The modest or absent incremental gains for stroke and myocardial infarction are also important. A single headline saying the model predicts cardiovascular disease would hide meaningful variation between endpoints.[1]
What this could change for patients and clinicians
If prospective validation succeeds, a sleep service might use the score to identify patients who merit a closer cardiovascular history, a conventional risk assessment or targeted rhythm monitoring. That could be especially relevant for atrial fibrillation, which can be intermittent and silent. But an alert would need a defined follow-up route; otherwise it may create anxiety and additional appointments without improving care.
A safe pathway would separate screening from diagnosis. Clinicians would review symptoms, existing conditions and ordinary risk factors, then decide whether a standard ECG, ambulatory rhythm monitor or another established test is appropriate. Services would need to track false positives, missed cases, waiting time, workload and whether alerts improve outcomes across age, sex and racial or ethnic groups. The publication does not provide evidence for those operational choices.
Patients should also know when a recording collected for sleep assessment is reused to estimate other risks. Consent, data retention and the possibility of incidental findings need clear governance. A low score must not override symptoms or block ordinary assessment, and a high score should not become a diagnosis in the medical record without confirmation.[1]
Evidence limits, funding and interests
The retrospective design cannot show that the model causes better care. Diagnosis codes were not clinically adjudicated, cause of death was unavailable, and the records lacked biomarkers that might explain or improve predictions. The authors could not pair the original sleep-study ECG with an ECG at later diagnosis. Undiagnosed atrial fibrillation at baseline remains possible; only five 30-second segments were manually reviewed in recordings linked to future atrial fibrillation.
The National Heart, Lung, and Blood Institute supported the work through grant R01 HL161253. The authors reported no financial conflicts. They disclosed that one author, B.M.W., was a co-founder, adviser and consultant to Beacon Biosignals and held equity in the company. Code was made public, while patient-level data remained restricted. Independent testing would help determine whether the reported performance depends on local equipment, coding or referral patterns.
The assessment would strengthen with a prospectively registered, multicentre study that freezes the model before testing, reports calibration and decision-curve results, and compares it with simpler ECG and clinical baselines. Most importantly, a trial should test whether score-guided follow-up finds treatable disease earlier or improves outcomes without disproportionate investigations. It would weaken if calibration drifts, simpler models perform as well, or extra testing produces harm without measurable benefit.[1]
What this means for people
- Sleep-clinic patients could receive earlier cardiovascular review, but a risk score is not a diagnosis.
- Clinicians need a defined confirmation pathway before acting on any automated alert.
- Services must measure extra tests, false positives and missed cases rather than assuming better discrimination equals benefit.
- Patients should be told when sleep-study ECG data are reused for longer-term risk prediction.
Global context
All three cohorts came from US hospitals and from people already referred for polysomnography. Sleep-lab access, cardiovascular prevalence, coding, equipment and follow-up pathways differ between health systems. International or community use would require local validation and calibration rather than importing the reported score unchanged.
What the evidence does not yet show
- The retrospective cohorts consisted of patients referred for sleep studies, limiting generalisability to the wider population.
- Outcomes were derived from ICD diagnosis codes rather than complete clinical adjudication, and cause of death was unavailable.
- The study reported associations and discrimination, not prospective clinical benefit or treatment effects.
- Incremental prediction gains were small or absent for stroke and myocardial infarction.
- Possible undiagnosed baseline disease, missing biomarkers and site-specific referral or coding patterns could influence performance.
What to watch next
- Prospective validation with a frozen model and prespecified referral thresholds.
- Calibration and subgroup performance across different hospitals, devices and patient groups.
- Comparison with simpler ECG scores and established cardiovascular-risk pathways.
- Whether score-guided monitoring improves diagnosis, treatment or patient outcomes.
- False-positive burden, clinician workload and patient consent for secondary use of sleep recordings.
Living evidence record
Impact record IAI-02MGX3L
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
10 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what SLEEP published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 10 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can AI flag depression symptoms in adults with low haemoglobin?
A Chinese prediction study spanning 3,517 adults reported similar discrimination in two survey cohorts and a 400-person hospital cohort. It predicts a screening score, not clinical depression, and its proposed risk bands have not been prospectively tested in care.
8 min · 1 source
Health & Life Sciences
Did clinicians prefer AI discharge summaries after long hospital stays?
In a retrospective 60-case comparison, 12 physicians usually preferred GPT-5.2 summaries and annotated fewer omissions. Reviewers knew which summary was AI-written, one hospital supplied the records, and no patient outcome or time saving was tested.
7 min · 3 sources
Health & Life Sciences
Can clinical AI recognise when the patient record does not support an answer?
A new clinical-agent benchmark raises a practical question for health systems: can an assistant explain what the record cannot establish? Our analysis examines evidence, local testing and the burden on staff.
6 min · 2 sources
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.