Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisUnited StatesSouth KoreaEast AsiaGlobal critical care

Can AI flag a longer ICU stay?

A peer-reviewed model separated longer-stay risk across 5,281 retrospective sepsis admissions in one US development cohort and two Korean external cohorts. Its AUROC fell from 0.848 internally to 0.781 at the smaller Korean site, and no prospective study tested whether alerts improve care or capacity.

By The Impact of AI Editorial DeskReleased 9 October 2026 at 15:01 BST7 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The retrospective study analysed 4,391 sepsis-related ICU admissions from MIMIC-IV, 728 from Chungbuk National University Hospital and 162 from Chungnam National University Hospital: 5,281 admissions in total.
  • 2Using 17 first-day variables, the hybrid model reported AUROCs of 0.848 internally, 0.823 at Chungbuk and 0.781 at Chungnam; 95% intervals were estimated with 2,000 stratified bootstrap resamples.
  • 3The study predicts stay length, not deterioration or benefit from an intervention. It excludes people who died, uses a retrospective four-day threshold and has not been tested prospectively in ICU workflow.
Key themesMedical AISepsisIntensive careLength of stayExternal validationClinical prediction

Research topic

External validation of first-day prediction for ICU stays longer than four days in adults with sepsis

The Impact of AI research cover asking whether AI can flag a longer ICU stay, with an empty conceptual ICU bed and three generic connected hospitals; it states 5,281 retrospective admissions and testing in two Korean hospitals.
AI-generated editorial illustration. The empty bed, hospitals and data connections are conceptual; they are not patient records, provider buildings, clinical alerts or evidence that the model improved an outcome.

The direct answer: it ranked risk across three cohorts, but did not improve care

A model using information available in the first 24 hours of intensive care identified adults with sepsis who were more likely to remain in the ICU for longer than four days. Discrimination was strongest in the US development cohort and lower, but still above chance, at two South Korean hospitals. The reported area under the receiver operating characteristic curve was 0.848 in internal validation, 0.823 at Chungbuk National University Hospital and 0.781 at Chungnam National University Hospital.

That is evidence of retrospective transport across institutions, not evidence that the model should allocate beds or change treatment. AUROC measures ranking over possible thresholds. It does not reveal how many additional admissions would be flagged at the operating point a hospital chooses, whether the probabilities are calibrated locally, or whether acting on a flag shortens a stay. The paper evaluates no live alert, clinician response, patient outcome, staffing change or cost.[1]

The Impact Brief · Free

Follow the evidence in health & life sciences.

Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.

Choose your topics (optional)

One concise, source-linked briefing. Unsubscribe at any time.

The denominator was 5,281 surviving ICU admissions

The investigators developed and internally validated the model on 4,391 sepsis-related ICU admissions from the public MIMIC-IV database. External validation used 728 admissions at Chungbuk and 162 at Chungnam, giving 5,281 analysed admissions. Longer stays were not equally common: 1,783 MIMIC admissions, or 40.6%, crossed the four-day threshold; 390 at Chungbuk, or 53.6%, did so; and 62 at Chungnam, or 38.3%, did so.

The flow chart matters because the analysed population is narrower than all adults arriving with sepsis. Admissions shorter than 24 hours, stays longer than 60 days, transfers or discharges from intensive care, records with extensive missingness and deaths were excluded under the study rules. Excluding people who died makes the target easier to interpret as resource use among survivors, but it also means the tool cannot be read as a general sepsis prognosis score. Death is a competing event that can produce a short stay for the worst possible reason.[1][2]

Seventeen routine variables fed a nested validation design

The model used 17 prespecified variables from the first ICU day: age, sex, body mass index, blood-pressure measures, lactate, atrial fibrillation and SOFA-related signals including oxygenation, Glasgow Coma Scale, vasopressor use and dose, mechanical ventilation, creatinine, platelets, bilirubin and urine output. The team says it applied no data-driven feature selection and used the same variables in every cohort. MIMIC missingness was under 2% and handled with median imputation for continuous variables; the Korean cohorts had less than 1% missingness after eligibility filtering and were not imputed.

Internal evaluation used five-fold stratified outer cross-validation with a separate five-fold inner loop for hyperparameter selection. Scaling parameters were learned inside each outer training fold, and external cohorts were standardised using parameters derived only from MIMIC. A ten per cent stratified subset of each training fold supported early stopping. These choices reduce leakage and make the internal estimate more credible than one random train-test split, although all model selection still begins from one US database.[1]

The hybrid beat its comparators, with wider uncertainty at the smallest site

The FT-TabNet architecture combines a feature transformer with TabNet's sequential attention and learns a gate that weights the two representations for each admission. It was compared with a meta-ensemble, transformer baselines, TabNet, a multilayer perceptron, XGBoost and a model based on the SOFA score. The hybrid produced the highest reported AUROC in the internal and both external cohorts. The paper also reports decision-curve results and SHAP explanations for which inputs influenced an individual prediction.

Uncertainty is largest where the denominator is smallest. The paper reports an internal AUROC confidence interval of 0.832 to 0.856, while the Chungnam estimate comes from only 162 admissions and has a much wider interval around the reported 0.781. Confidence intervals were generated with 2,000 stratified bootstrap resamples, and paired AUROCs were compared with DeLong tests. The authors state that gains over SOFA did not reach statistical significance, so the result should not be translated into a claim that deep learning has already surpassed standard clinical scoring in practice.[1]

What clinicians and patients could gain—and what they could lose

An early, well-calibrated estimate could help a critical-care team anticipate pressure on step-down beds, rehabilitation, staffing and family communication. Because the inputs are routinely collected, the model does not depend on a new scan or specialised assay. Its intended value is planning, not deciding that one patient deserves a bed more than another. A forecast should support clinicians who understand the trajectory, not become an automated discharge target.

Length of stay is partly shaped by hospital organisation, availability of ward beds, discharge customs, insurance and local thresholds for intensive care. A system can therefore learn institutional delay as well as physiology. If managers reward shorter predicted stays or use risk scores to ration care, errors could concentrate on people whose recovery is less predictable. Prospective evaluation needs to count false alerts, delayed transfers, readmissions and effects on groups defined by age, sex, disability, socioeconomic status and care pathway.[1]

Funding, availability and the evidence still missing

The article acknowledges support from a Korea Health Technology R&D Project through the Korea Health Industry Development Institute, funded by the Ministry of Health and Welfare, and names grant RS-2026-25540297. A separate funding statement says the authors received no specific funding for the work. That internal inconsistency should be clarified by the journal or authors. The authors declare no competing interests. MIMIC-IV is public under credentialled access; the Korean patient-level data are restricted for institutional and privacy reasons.

The authors call for broader validation, and the paper's limits support that caution. Both external sites are Korean tertiary hospitals operating in one national health system. Confidence would rise with a preregistered prospective study using frozen software and thresholds across health systems, direct calibration reporting, subgroup results and comparison with ordinary capacity planning. Most importantly, it should test whether showing the estimate changes care, resource use or patient outcomes without increasing premature transfer or inequity.[1][2]

What this means for people

  • A longer-stay score should support planning, not ration an ICU bed or compel discharge.
  • Patients who die were excluded, so a low predicted stay is not automatically reassuring.
  • Hospitals need to measure whether forecasts reduce delays without shifting risk to patients with less predictable recovery.

Global context

The development cohort comes from a de-identified US critical-care database and external cohorts from two South Korean tertiary hospitals. ICU admission rules, ward capacity, discharge practice, financing and sepsis definitions vary internationally, so length-of-stay models can transport less well than a physiological label suggests. Local recalibration and prospective governance are prerequisites for use elsewhere.

What the evidence does not yet show

  • Retrospective prediction cannot show that an alert improves care, capacity or outcomes.
  • The analysis excludes patients who died, so it does not represent all sepsis-related ICU admissions.
  • Both external cohorts are Korean tertiary hospitals within one health system, and the smaller cohort contains 162 admissions.
  • A four-day threshold is a study definition tied to the development distribution, not a universal clinical boundary.
  • The paper acknowledges a Korean government grant while separately stating that the authors received no specific funding; the discrepancy needs clarification.

What to watch next

  • Prospective, preregistered evaluation with frozen thresholds.
  • Calibration and predictive values at each intended hospital.
  • Subgroup performance and effects on transfer, readmission and mortality.
  • Comparison with existing clinical and operational planning.
  • Clarification of the funding statement and independent reproduction.

Living evidence record

Impact record IAI-05XWRG8

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

9 October 2026

Source trail

2 direct sources across 2 source types.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Evidence trail

Sources used for this report

Links checked 9 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can AI flag sepsis risk early in trauma intensive care?

Models using the first 24 hours of records moderately separated trauma patients who developed sepsis over the next 48 hours in one US intensive-care database. The best AUROC was 0.734, with sensitivity below 69%; this supports further surveillance research, not autonomous alerts or treatment.

7 min · 1 source

Health & Life Sciences

Can a chest X-ray flag osteoporosis?

A peer-reviewed Korean study externally validated an AI prescreener in three cohorts totalling 153,058 people. Its simulated workflow preserved most osteoporosis detections while halving DXA use, but it has not yet proved benefit in prospective care.

7 min · 2 sources

Health & Life Sciences

Can AI estimate survival after a Parkinson’s diagnosis?

A peer-reviewed Chinese registry study compared four survival models in 3,148 people with Parkinson’s disease. A transparent Cox model matched the machine-learning alternatives, but validation stayed within the same registry and no clinical-impact study was performed.

9 min · 2 sources

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.