Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisNorth AmericaUnited States

Did clinicians prefer AI discharge summaries after long hospital stays?

In a retrospective 60-case comparison, 12 physicians usually preferred GPT-5.2 summaries and annotated fewer omissions. Reviewers knew which summary was AI-written, one hospital supplied the records, and no patient outcome or time saving was tested.

By The Impact of AI Research DeskReleased 2 October 2026 at 23:15 BST7 min read3 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesClinical AIDischarge communicationLarge language modelsPatient safetyHuman oversight

Research topic

Clinician evaluation of LLM-generated discharge summaries for hospital stays lasting seven to 21 days

The Impact of AI research cover asking whether clinicians preferred AI discharge summaries, with abstract blank documents flowing through two comparison paths.
AI-generated editorial illustration. The documents and comparison pathways are conceptual and do not depict a patient record, clinician, hospital system or measured chart.

At a glance

  • 1Twelve attending physicians reviewed 60 historical encounters; one hospitalist and one primary-care physician assessed each case, and the AI summary was preferred in 57 of 60 cases.
  • 2LLM summaries received higher mean ratings for quality, readability, factuality and completeness, and reviewers annotated fewer omissions, but the comparison was unblinded and lacked a consensus reference standard.
  • 3The single-centre retrospective study did not measure readmissions, medication errors, adverse events, consultation time or patient comprehension, so it does not establish that automated summaries improve care.

Living evidence record

Impact record IAI-14SHKYE

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

2 October 2026

Source trail

3 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

The question was communication after a long admission

Discharge summaries compress days of tests, treatment changes, diagnoses and follow-up plans into a document for hospital teams, primary-care clinicians, patients and caregivers. The task becomes harder after a long admission, when relevant facts are scattered through many notes. This study asked whether a large language model could assemble a better clinician-facing summary for stays lasting seven to 21 days than the historical summary written during routine care.

This was not a deployment trial. Researchers retrospectively extracted 400 internal-medicine encounters from Stony Brook University Hospital between 1 January 2023 and 31 December 2024. After exclusions, 385 met the criteria. Patients were aged 19 to 101, with a mean age of 72; 162 of 385 were female. Sixty encounters were selected for intensive physician review because reviewing hundreds of long admissions was not feasible.[1][2]

The pipeline included expert curation before generation

The 60-case sample was stratified to resemble the length-of-stay distribution of the eligible dataset, while stays longer than 14 days were deliberately oversampled. Encounters without a historical discharge summary or with fewer than 25 notes were excluded. One physician informaticist chose 370 of 654 note types judged clinically relevant, reducing average text volume by about 55%. That may remove noise, but it also puts expert curation inside the pipeline rather than testing an automatic system against an untouched record.

The model was GPT-5.2, run through a private HIPAA-compliant Microsoft Azure OpenAI service. Notes were ordered chronologically and material after the final progress note was excluded to prevent information unavailable at discharge leaking into output. Long records were divided into 50,000-token segments. The model iteratively refined a chronological draft, separately extracted possible incidental radiology findings, and consolidated both into a final narrative.[2]

Twelve attending physicians compared the summaries

Six hospitalists and six primary-care physicians took part. Each reviewed ten encounters, and every case was independently evaluated by one clinician from each group, creating 120 sets of ratings but only two clinical perspectives per encounter. Reviewers used the same preprocessed notes supplied to the model through a custom application behind the hospital firewall.

Hospitalists inspected inpatient notes before rating both summaries. Primary-care physicians first read the summaries and then checked claims against the notes, reflecting how outpatient clinicians often meet a discharge document. All scored quality, conciseness and readability, factuality and completeness on five-point scales. Primary-care reviewers also judged ease of understanding, explaining the admission to patients or caregivers, and arranging follow-up.[2]

Clinicians preferred the model's summaries

Reviewers preferred the LLM summary in 57 of 60 encounters, or 95%. They preferred the human-authored version in three and never selected no preference. Across 120 ratings, model summaries had higher mean scores for quality, 4.71 versus 3.16; readability, 4.66 versus 3.39; factuality, 4.62 versus 3.73; and completeness, 4.69 versus 3.42. Reported p-values ranged from 4.28 × 10⁻⁶ to 1.93 × 10⁻⁴.

Primary-care physicians also rated the AI summaries easier to understand, communicate and use for follow-up. Mean scores were 4.75 versus 3.20 for understanding, 4.75 versus 3.27 for explaining the admission, and 4.68 versus 3.20 for follow-up. Those findings suggest that explicit structure may help readers who were not present during the admission. They do not tell us how quickly clinicians read each version or whether patients understood it.[2]

Fewer omissions did not mean error-free output

Reviewers marked 138 errors across the 60 encounters: 47 inaccuracies and 91 omissions. The mean error count was lower for model summaries, 0.42 versus 0.73 per case. That difference was driven by omissions, 0.26 versus 0.50, with p=0.004. Inaccuracies were not significantly different: 0.16 versus 0.23, p=0.45. Mean potential-harm and likelihood-of-harm scores also did not differ significantly.

These counts need caution. Evaluators identified errors without a separately adjudicated gold-standard summary or repeated ratings within each specialty. Primary-care physicians marked substantially more omissions than hospitalists and assigned higher likelihood-of-harm scores, showing that professional role changed what counted as a problem. A lower average count is encouraging, but it does not support accepting generated summaries without clinical verification.

The pipeline identified 31 radiology incidental findings. Reviewers judged 93.5% factually correct and 87.1% appropriate to include. Historical summaries were not evaluated systematically on that task, so this is not a direct model-versus-human comparison. The denominator is also small enough that a few changed judgments would materially shift the percentages.[2]

Unblinded review and one hospital limit the result

Reviewers knew which document came from the model. The authors retained each summary's original format because the versions looked substantially different and reformatting would reduce clinical realism. That makes the exercise practical, but allows expectations about AI, formatting or polish to influence preference and scores. Each case also came from one hospital, one EHR and one internal-medicine service, while a single model, prompt pipeline and curation strategy generated the new summaries.

No patient received care based on these outputs. The study did not measure readmission, adverse events, medication discrepancies, missed follow-up, clinician time, workload, patient comprehension or trust. For clinicians, the immediate value may be a draft that reduces synthesis work while leaving accountability with the treating team. For patients, the risk is that a fluent summary appears authoritative even when it omits context or states an unsupported relationship.[2]

What would change the assessment

The authors declared no competing interests and made analysis code and de-identified statistical results public. The clinical data contain protected health information and are not public. The publication does not report a separate funding statement. Future comparisons should include multiple locked model versions because performance, cost and failure modes can change when a hosted model is updated.

Confidence would rise with blinded review in a common format, multiple raters per specialty, adjudicated reference summaries and external testing across hospitals, services and languages. The decisive next step is a prospective trial in which clinicians review model drafts under defined oversight, measuring documentation time, medication and follow-up errors, patient understanding, readmissions and adverse events. Until then, the paper supports supervised drafting and further validation—not autonomous discharge documentation.[2][3]

What this means for people

  • A supervised draft could reduce the burden of assembling long records, but the study did not measure time saved.
  • Patients could benefit from clearer follow-up information only if factual claims and omissions are reliably checked.
  • Clinicians remain responsible for validating the document; fluent generation is not evidence of safe autonomous use.

Global context

The evidence comes from one US academic hospital and English-language records. Discharge formats, primary-care access, EHR systems, privacy law and responsibility for follow-up vary widely. International adoption would require local validation, including lower-resource settings where exhaustive human review or private model hosting may be harder to sustain.

What the evidence does not yet show

  • The comparison was retrospective, single-centre and limited to 60 long internal-medicine admissions selected from 385 eligible encounters.
  • Reviewers were not blinded, and each encounter had only one hospitalist and one primary-care reviewer.
  • Error annotations had no independent consensus reference standard and differed substantially by reviewer role.
  • The study measured preferences and annotations, not clinical outcomes, time savings, patient understanding or real-world workflow performance.

What to watch next

  • Blinded multi-centre replication using a common format and adjudicated reference summaries.
  • Prospective trials measuring clinician time, follow-up reliability, medication errors, readmissions and adverse events.
  • Performance across specialties, languages, hospitals and model versions without bespoke expert note curation.
  • Operational rules for clinician review, source traceability and responsibility for corrected output.

Evidence trail

Sources used for this report

Links checked 2 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can an AI digital twin shorten an ICU stay before it reaches a clinical trial?

ARPA-H has awarded the University of Vermont up to $38 million to model individual patients' immune responses. The five-year, milestone-based project targets a 25% reduction in ICU stays, but it has not yet demonstrated that result in patients and publishes no planned trial denominator.

6 min · 3 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.