Did clinicians prefer AI discharge summaries after long hospital stays?
In a retrospective 60-case comparison, 12 physicians usually preferred GPT-5.2 summaries and annotated fewer omissions. Reviewers knew which summary was AI-written, one hospital supplied the records, and no patient outcome or time saving was tested.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Clinician evaluation of LLM-generated discharge summaries for hospital stays lasting seven to 21 days

At a glance
- 1Twelve attending physicians reviewed 60 historical encounters; one hospitalist and one primary-care physician assessed each case, and the AI summary was preferred in 57 of 60 cases.
- 2LLM summaries received higher mean ratings for quality, readability, factuality and completeness, and reviewers annotated fewer omissions, but the comparison was unblinded and lacked a consensus reference standard.
- 3The single-centre retrospective study did not measure readmissions, medication errors, adverse events, consultation time or patient comprehension, so it does not establish that automated summaries improve care.
Living evidence record
Impact record IAI-14SHKYE
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
2 October 2026
Source trail
3 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
The question was communication after a long admission
Discharge summaries compress days of tests, treatment changes, diagnoses and follow-up plans into a document for hospital teams, primary-care clinicians, patients and caregivers. The task becomes harder after a long admission, when relevant facts are scattered through many notes. This study asked whether a large language model could assemble a better clinician-facing summary for stays lasting seven to 21 days than the historical summary written during routine care.
This was not a deployment trial. Researchers retrospectively extracted 400 internal-medicine encounters from Stony Brook University Hospital between 1 January 2023 and 31 December 2024. After exclusions, 385 met the criteria. Patients were aged 19 to 101, with a mean age of 72; 162 of 385 were female. Sixty encounters were selected for intensive physician review because reviewing hundreds of long admissions was not feasible.[1][2]
The pipeline included expert curation before generation
The 60-case sample was stratified to resemble the length-of-stay distribution of the eligible dataset, while stays longer than 14 days were deliberately oversampled. Encounters without a historical discharge summary or with fewer than 25 notes were excluded. One physician informaticist chose 370 of 654 note types judged clinically relevant, reducing average text volume by about 55%. That may remove noise, but it also puts expert curation inside the pipeline rather than testing an automatic system against an untouched record.
The model was GPT-5.2, run through a private HIPAA-compliant Microsoft Azure OpenAI service. Notes were ordered chronologically and material after the final progress note was excluded to prevent information unavailable at discharge leaking into output. Long records were divided into 50,000-token segments. The model iteratively refined a chronological draft, separately extracted possible incidental radiology findings, and consolidated both into a final narrative.[2]
Twelve attending physicians compared the summaries
Six hospitalists and six primary-care physicians took part. Each reviewed ten encounters, and every case was independently evaluated by one clinician from each group, creating 120 sets of ratings but only two clinical perspectives per encounter. Reviewers used the same preprocessed notes supplied to the model through a custom application behind the hospital firewall.
Hospitalists inspected inpatient notes before rating both summaries. Primary-care physicians first read the summaries and then checked claims against the notes, reflecting how outpatient clinicians often meet a discharge document. All scored quality, conciseness and readability, factuality and completeness on five-point scales. Primary-care reviewers also judged ease of understanding, explaining the admission to patients or caregivers, and arranging follow-up.[2]
Clinicians preferred the model's summaries
Reviewers preferred the LLM summary in 57 of 60 encounters, or 95%. They preferred the human-authored version in three and never selected no preference. Across 120 ratings, model summaries had higher mean scores for quality, 4.71 versus 3.16; readability, 4.66 versus 3.39; factuality, 4.62 versus 3.73; and completeness, 4.69 versus 3.42. Reported p-values ranged from 4.28 × 10⁻⁶ to 1.93 × 10⁻⁴.
Primary-care physicians also rated the AI summaries easier to understand, communicate and use for follow-up. Mean scores were 4.75 versus 3.20 for understanding, 4.75 versus 3.27 for explaining the admission, and 4.68 versus 3.20 for follow-up. Those findings suggest that explicit structure may help readers who were not present during the admission. They do not tell us how quickly clinicians read each version or whether patients understood it.[2]
Fewer omissions did not mean error-free output
Reviewers marked 138 errors across the 60 encounters: 47 inaccuracies and 91 omissions. The mean error count was lower for model summaries, 0.42 versus 0.73 per case. That difference was driven by omissions, 0.26 versus 0.50, with p=0.004. Inaccuracies were not significantly different: 0.16 versus 0.23, p=0.45. Mean potential-harm and likelihood-of-harm scores also did not differ significantly.
These counts need caution. Evaluators identified errors without a separately adjudicated gold-standard summary or repeated ratings within each specialty. Primary-care physicians marked substantially more omissions than hospitalists and assigned higher likelihood-of-harm scores, showing that professional role changed what counted as a problem. A lower average count is encouraging, but it does not support accepting generated summaries without clinical verification.
The pipeline identified 31 radiology incidental findings. Reviewers judged 93.5% factually correct and 87.1% appropriate to include. Historical summaries were not evaluated systematically on that task, so this is not a direct model-versus-human comparison. The denominator is also small enough that a few changed judgments would materially shift the percentages.[2]
Unblinded review and one hospital limit the result
Reviewers knew which document came from the model. The authors retained each summary's original format because the versions looked substantially different and reformatting would reduce clinical realism. That makes the exercise practical, but allows expectations about AI, formatting or polish to influence preference and scores. Each case also came from one hospital, one EHR and one internal-medicine service, while a single model, prompt pipeline and curation strategy generated the new summaries.
No patient received care based on these outputs. The study did not measure readmission, adverse events, medication discrepancies, missed follow-up, clinician time, workload, patient comprehension or trust. For clinicians, the immediate value may be a draft that reduces synthesis work while leaving accountability with the treating team. For patients, the risk is that a fluent summary appears authoritative even when it omits context or states an unsupported relationship.[2]
What would change the assessment
The authors declared no competing interests and made analysis code and de-identified statistical results public. The clinical data contain protected health information and are not public. The publication does not report a separate funding statement. Future comparisons should include multiple locked model versions because performance, cost and failure modes can change when a hosted model is updated.
Confidence would rise with blinded review in a common format, multiple raters per specialty, adjudicated reference summaries and external testing across hospitals, services and languages. The decisive next step is a prospective trial in which clinicians review model drafts under defined oversight, measuring documentation time, medication and follow-up errors, patient understanding, readmissions and adverse events. Until then, the paper supports supervised drafting and further validation—not autonomous discharge documentation.[2][3]
What this means for people
- A supervised draft could reduce the burden of assembling long records, but the study did not measure time saved.
- Patients could benefit from clearer follow-up information only if factual claims and omissions are reliably checked.
- Clinicians remain responsible for validating the document; fluent generation is not evidence of safe autonomous use.
Global context
The evidence comes from one US academic hospital and English-language records. Discharge formats, primary-care access, EHR systems, privacy law and responsibility for follow-up vary widely. International adoption would require local validation, including lower-resource settings where exhaustive human review or private model hosting may be harder to sustain.
What the evidence does not yet show
- The comparison was retrospective, single-centre and limited to 60 long internal-medicine admissions selected from 385 eligible encounters.
- Reviewers were not blinded, and each encounter had only one hospitalist and one primary-care reviewer.
- Error annotations had no independent consensus reference standard and differed substantially by reviewer role.
- The study measured preferences and annotations, not clinical outcomes, time savings, patient understanding or real-world workflow performance.
What to watch next
- Blinded multi-centre replication using a common format and adjudicated reference summaries.
- Prospective trials measuring clinician time, follow-up reliability, medication errors, readmissions and adverse events.
- Performance across specialties, languages, hospitals and model versions without bespoke expert note curation.
- Operational rules for clinician review, source traceability and responsibility for corrected output.
Evidence trail
Sources used for this report
Links checked 2 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can clinical AI recognise when the patient record does not support an answer?
A new clinical-agent benchmark raises a practical question for health systems: can an assistant explain what the record cannot establish? Our analysis examines evidence, local testing and the burden on staff.
6 min · 2 sources
Health & Life Sciences
Did machine learning beat logistic regression on surgical sepsis?
A 328,292-case US study found no significant predictive advantage for LASSO, XGBoost or random forest. Sepsis was rare, the best risk group still had low absolute incidence, and no model has external validation.
8 min · 2 sources
Health & Life Sciences
Can an AI digital twin shorten an ICU stay before it reaches a clinical trial?
ARPA-H has awarded the University of Vermont up to $38 million to model individual patients' immune responses. The five-year, milestone-based project targets a 25% reduction in ICU stays, but it has not yet demonstrated that result in patients and publishes no planned trial denominator.
6 min · 3 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.