Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisGlobalLatin AmericaBrazil

Is neonatal respiratory AI ready for routine care?

A peer-reviewed scoping review mapped 35 studies using AI and digital tools to predict, image or monitor newborn breathing problems. The evidence spans promising prototypes, but heterogeneity, small samples, limited external validation and few clinical-impact tests keep routine use unproven.

By The Impact of AI Health & Life Sciences DeskReleased 3 October 2026 at 05:00 BST9 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesNeonatal careRespiratory medicineClinical AIMedical imagingPredictive modelsDigital monitoring

Research topic

The maturity of AI and digital technologies for neonatal respiratory prediction, diagnosis and continuous monitoring

The Impact of AI research cover asking whether neonatal respiratory AI is ready for routine care, with conceptual lung, monitoring and imaging signals narrowing through an evidence funnel toward a protected cradle.
AI-generated editorial illustration. The cradle, lungs, signals and evidence funnel are conceptual; they do not depict an infant, medical device, patient record, scan or measured result.

At a glance

  • 1The registered scoping review screened 199 records, assessed 126 full texts and included 35 studies published from 2019 to 2026.
  • 2Twenty-one studies used machine learning, ten deep learning, three hybrid approaches and one an intelligent closed-loop system; electronic records and physiological signals were the most common inputs.
  • 3Because methods, populations and outcomes varied and external validation and clinical-impact evidence were limited, the review supports further multicentre prospective testing—not routine autonomous use.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-0BONK60

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

3 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

The review asked what has moved beyond a promising prototype

Newborn respiratory care demands rapid decisions from signals that change minute by minute. Premature infants can need support for respiratory distress, bronchopulmonary dysplasia, apnoea, infection or failed extubation, yet clinical examination, blood gases and conventional imaging each capture only part of that picture. AI research promises earlier warning, more consistent image reading and continuous monitoring.

The new Journal of Perinatology paper does not test one device. It maps the field. Nine Brazil-based authors conducted a scoping review under the PRISMA extension for scoping reviews and registered a protocol on the Open Science Framework. Their question was where digital technologies and AI have been applied to neonatal respiratory assessment, what inputs and outcomes they use, and which evidence gaps still block clinical integration.

A scoping review is suited to a fragmented field because it can show breadth, designs and missing areas. It is not a pooled estimate of benefit. The review did not ask whether one algorithm improves survival or safely changes treatment across hospitals; it organised heterogeneous studies that often measured different technical outcomes.[1][2]

Thirty-five studies survived screening

The authors searched studies published from January 2015 to March 2026 using controlled vocabulary and free-text terms for newborns, respiratory assessment, digital health and AI. The article reports database searches and says two reviewers independently screened titles, abstracts and full texts, with a third reviewer resolving disagreements. Data were extracted using a predefined form covering country, design, population, technology, respiratory parameters, clinical objective, setting, findings and limitations.

The flow began with 199 records. Ten duplicates were removed, leaving 189 for screening. Sixty-three were excluded at title or abstract stage and 126 full texts were assessed. Ninety-one full texts were excluded, leaving 35 studies—18.5% of the initially identified records—published between 2019 and 2026.

Eligibility required a neonatal population, a digital or AI technology and assessment of respiratory parameters. Editorials, non-original studies, non-respiratory applications and studies with fewer than ten patients were excluded. The paper contains a reporting inconsistency worth noting: its search-strategy paragraph names PubMed/MEDLINE, Embase, Scopus, Web of Science and IEEE Xplore, while the staged selection description names PubMed, Cochrane Library, LILACS, Web of Science, SciELO and ScienceDirect. The supplementary strategy may clarify the executed set, but the main text is not internally aligned.[1]

Most systems predicted outcomes or interpreted signals

Machine-learning approaches appeared in 21 of the 35 studies, or 60%. Ten used deep learning, three used a hybrid of machine and deep learning, and one used an intelligent closed-loop control system. Electronic health records were the most common data source at 34.3%, followed by physiological signals at 20%. Chest radiography, digital auscultation and ventilator data each accounted for 8.6%, with ultrasound, spectroscopy, video and other sources filling the remainder.

The review grouped applications into three broad domains. Predictive models estimated outcomes such as bronchopulmonary dysplasia, respiratory distress, mechanical ventilation or extubation failure. Imaging systems classified chest radiographs or lung ultrasound. Monitoring systems analysed sounds, movement, oxygen, ventilator interaction or other physiological streams, sometimes without contact.

Those categories matter to families and clinicians in different ways. A risk score may prompt observation or preparation; image support may reduce disagreement; continuous monitoring may flag deterioration between bedside checks. But a model that predicts a label accurately in a retrospective dataset is not the same as a system that improves a treatment decision, avoids an unnecessary intervention or works reliably during routine care.[1]

Promising performance was difficult to compare

The included studies often reported strong predictive or classification performance, but the review found substantial variation in datasets, outcomes, predictors, preprocessing, algorithms and validation strategies. That prevents a single meaningful performance number. An area-under-the-curve result for bronchopulmonary dysplasia cannot simply be compared with image classification, apnoea detection or closed-loop oxygen control.

Imaging studies had somewhat more standardised inputs, mainly radiographs and lung ultrasound, yet still lacked broad external validation. Across the field, many studies were retrospective, small or confined to one institution. Performance can fall when patient mix, gestational age, equipment, clinical definitions or data capture differ. A tool tuned on one neonatal intensive-care unit may encounter a different baseline risk and different artefacts elsewhere.

The review was descriptive and did not pool effect sizes. It also did not report a formal risk-of-bias rating for each included study in the methods available on the journal page. That is consistent with some scoping-review practice, but it means inclusion should not be mistaken for equal evidential strength. A prospective crossover of automated oxygen control and a retrospective prediction model answer very different questions.[1]

External validation and clinical impact are the main gaps

The review’s most consequential finding is not that neonatal AI performs badly. It is that the evidence rarely completes the path from technical performance to clinical value. External validation was limited, methods and reporting were not standardised, and prospective multicentre studies were scarce. Bias, equity and safety were also poorly explored.

For a prediction tool, useful validation means more than reproducing discrimination in another spreadsheet. Researchers need calibration across hospitals, gestational ages and populations; pre-specified thresholds; comparison with existing clinical scores; and evidence that acting on the output improves care without causing avoidable interventions. For monitors, alarm burden, sensor failure and staff response matter. For imaging, reader behaviour and missed disease matter as much as test-set accuracy.

The paper therefore stops short of routine implementation. Its conclusion calls for prospective, multicentre and standardised studies, transparent model reporting and integration with actual workflows. The authors describe AI as a complementary tool, not a replacement for neonatologists, nurses, respiratory therapists or radiologists.[1]

What the review means for families and clinical teams

For families, earlier warning sounds valuable, but false alarms and false reassurance both carry costs. A prediction can lead to extra tests, delayed extubation or escalation of respiratory support; missing a deteriorating infant can be catastrophic. None of those trade-offs is resolved by a catalogue of model-development studies, and the review does not establish that any included AI system improves survival, chronic lung disease, neurodevelopment or family experience.

For clinical teams, the map is useful for procurement and research planning. Hospitals should ask whether a tool was externally validated in comparable infants and equipment, whether calibration is local, how it handles missing data and whether clinicians can audit an alert. They should also measure workload and alarm burden rather than assuming automation saves time.

Equity deserves particular attention. Neonatal datasets can underrepresent hospitals with fewer digital records, different equipment or populations with the highest mortality. A system validated only where intensive monitoring and specialist staffing are already strong could widen gaps if adopted elsewhere without adaptation, infrastructure or a clear fallback when the model fails.[1]

Funding was public and competing interests were not declared

The article-processing charge was funded by Brazil’s Coordenação de Aperfeiçoamento de Pessoal de Nível Superior, or CAPES. Authors were affiliated with Faculdade Ciências Médicas de Minas Gerais and Universidade Federal do Rio Grande do Norte. They declared no competing interests. The paper was received in May, revised in August, accepted on 27 August and published as the version of record on 2 October 2026.

The review’s search stopped in March 2026, so it cannot represent studies released in the following six months. It also excluded samples below ten participants, but that threshold alone does not guarantee adequate power or generalisability. Search terminology focused on AI, digital technology and respiratory assessment; the authors acknowledge that extra prediction-model terms might have found more work.

Confidence would change with prospective multicentre trials that pre-register outcomes and compare AI-supported care with existing practice. The strongest studies would measure not only discrimination or agreement but treatment decisions, ventilation duration, bronchopulmonary dysplasia, adverse events, alarm burden, clinician workload, family-centred outcomes and costs. Until those arrive, neonatal respiratory AI remains a research and supervised decision-support field rather than a proven routine standard.[1]

What this means for people

  • Infants and families could benefit from earlier warning only if technical accuracy translates into safer decisions and better outcomes.
  • Clinicians may gain decision support, but poorly calibrated alerts can increase workload, unnecessary intervention or false reassurance.
  • Hospitals with fewer digital resources risk being underrepresented in development data and facing higher implementation burdens.

Global context

The review was led by researchers in Brazil and included studies from multiple countries and technical settings, but the evidence base was too heterogeneous to support a global performance estimate. Neonatal mortality, specialist availability, monitoring equipment, data quality and definitions of respiratory outcomes vary widely. Validation in Asia, the Middle East, Africa, Latin America, Europe and North America must be reported locally rather than inferred from a pooled narrative.

What the evidence does not yet show

  • The scoping review maps 35 heterogeneous studies and does not provide a pooled estimate of clinical benefit or safety.
  • Searches ended in March 2026, six months before publication, and the main text gives two inconsistent lists of databases searched.
  • Included studies differed in populations, outcomes, data, algorithms and validation, preventing direct performance comparison.
  • Formal study-level risk-of-bias ratings were not reported in the methods on the journal page, so inclusion does not imply equal evidence quality.
  • External validation, prospective multicentre testing, equity analysis and patient-relevant outcomes were limited across the evidence base.

What to watch next

  • Prospective multicentre validation across gestational ages, regions, equipment and levels of neonatal care.
  • Trials that pre-specify thresholds and measure treatment changes, ventilation duration, adverse events and family-centred outcomes.
  • Reporting of calibration, missing-data behaviour, subgroup performance, alarm burden and model drift.
  • Implementation studies comparing total workload, cost, infrastructure needs and safe fallback procedures with current practice.

Evidence trail

Sources used for this report

Links checked 3 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can AI support breast-ultrasound decisions across countries?

A peer-reviewed South Korean-led study validated an interpretable retrieval-augmented system across 8,311 images from 11 cohorts in seven countries and tested assistance with four readers. The retrospective evidence is encouraging, but it is not a prospective screening trial or proof of better patient outcomes.

8 min · 3 sources

Health & Life Sciences

Can this MRI model draw brain-tumour boundaries reliably?

A peer-reviewed multimodal segmentation model was developed on 2,422 public MRI volumes and externally tested on 125 cases. Accuracy remained useful but fell outside the development data; no prospective clinical workflow, radiologist comparison or patient-outcome test was performed.

8 min · 2 sources

Health & Life Sciences

Can a high-AUC diabetes model still be unsafe?

A peer-reviewed audit of 12 machine-learning approaches on two public diabetes datasets finds that similarly ranked models can differ materially in calibration, uncertainty and safe deferral. The work is a benchmark study—not a clinical trial, diagnostic approval or evidence of improved patient outcomes.

7 min · 3 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.