Did the AI write a better clinic note?
A prospective single-centre study of 62 adults with abdominal pain found that AI-generated notes were more complete than clinicians' notes but scored worse on factuality. The assistant also produced only 50% diagnostic accuracy in the prospective cohort, so the result supports supervised documentation research—not autonomous clinical use.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The prospective cohort enrolled 62 adults attending four consultants at one Wuhan hospital with abdominal pain as their chief complaint.
- 2AI-generated records scored higher for integrity or completeness, but lower for factuality than human-generated records; the model's prospective diagnostic accuracy was 50%.
- 3The study did not randomise care, measure time saved, test downstream treatment or report patient outcomes, and its authors said the system required clean, turn-taking Mandarin speech.
Research topic
Whether a locked, clinician-supervised documentation assistant improves record quality and workload without adding unsupported information or harming diagnosis across complaints, languages, hospitals and patient groups

The direct answer: more complete notes, but a serious factuality trade-off
The AI assistant produced outpatient records that clinicians rated as more complete than the records written without it, but those AI-generated records were also less factual. In the prospective cohort, the mean integrity score was 3.50 for AI-generated notes and 3.21 for human-generated notes. The mean factuality score moved in the opposite direction: 3.34 for AI-generated notes versus 4.01 for human-generated notes. Both differences were statistically significant. For a clinical record, an unsupported addition can matter more than a missing stylistic detail, so completeness cannot be treated as the sole definition of quality.
The system, called SMART-ASSISTANT, combined speech transcription, conversion of dialogue into a structured electronic medical record, prompts about missing information and a list of possible diagnoses. Its prospective diagnostic accuracy was 50%, lower than in the retrospective test. The authors attributed part of the drop to the consultants' cohort containing more rare and complicated disease. Whatever the reason, the finding shows why a good extraction score does not make a system a reliable diagnostic authority.
The appropriate conclusion is therefore narrow. This study shows that a fine-tuned assistant can capture and organise parts of a Mandarin outpatient conversation under supervision. It does not show that clinicians should accept generated notes without checking them, that the software reduces workload, or that it improves diagnosis or care. The factuality result instead makes human verification the central safety requirement.[1][2]
What was tested in the 62-person prospective cohort
Four consultants at Renmin Hospital of Wuhan University recruited consecutive eligible outpatients between 22 July and 15 September 2024. Participants were adults attending for a first visit with abdominal pain, without a previous definitive diagnosis, and able to consent. Sixty-two people entered the prospective cohort. This was an observational, single-centre evaluation rather than a randomised comparison of two care pathways.
The assistant transcribed consultation audio and converted the dialogue into structured records. Mean semantic textual similarity for transcription was 0.9253 and the character error rate was 0.2058. For recognising predefined quality-control points, accuracy was 94.40%. The authors also calculated a theoretical completion rate: they assumed a missing item would be completed whenever the system correctly prompted the clinician to ask about it. That rate rose from a human record completion measure of 59.14% to a theoretical 97.85% with prompts.
The word theoretical is important. The calculation did not itself prove that every prompt was followed correctly, that the answer was recorded accurately or that the added information changed a decision. A real implementation study would need to log accepted, edited and rejected suggestions, identify unsupported additions and compare the final signed record with the conversation and other source data.[2]
How the assistant was built and compared
Development used 1,249 anonymised records, 119 paired free-text and structured-record examples, and a dictionary of medical terms. The researchers evaluated three base large language models before and after task-specific fine-tuning and compared them with GPT-4 and DeepSeek on extraction of quality-control points. After fine-tuning, HuatuoGPT-II achieved the strongest reported quality-control accuracy and became the model inside SMART-ASSISTANT.
The workflow also included a multi-reader, multi-case study. Five attending physicians used the assistant's prompts to revise 125 records. Ten physician readers were divided into two reading orders and, after a three-week washout, rated original and AI-assisted records for completeness and diagnostic correlation. Ratings were higher for the assisted set. This design helps reduce simple order effects, but it still evaluates records produced from one institution's data and a carefully constructed research interface.
Audio processing was separately tested on 30 simulated conversations. The selected transcription pipeline used FunASR with entity correction and was then taken into the prospective cohort. The pipeline depended on relatively controlled speech: the authors said it required little background noise, no overlapping speakers, limited speaker numbers and no dialect. Those conditions may be difficult to maintain in busy clinics and reduce confidence that the same performance will transfer to other languages, accents and settings.[2]
Why the factuality score matters more than fluency
Readers found no statistically significant difference between AI- and human-generated records for normativity, readability or logicality. That apparent parity could be reassuring until factuality is examined. A note can be clear, well structured and clinically plausible while still containing a detail that the patient never said. Because later clinicians may treat the record as evidence, an invented symptom, negative finding or medication history can propagate through handovers and future decisions.
The authors explicitly acknowledged that the model could generate implausible information and illusory outputs when it had only the doctor-patient dialogue. Their evaluation penalised values inserted into fields when the source record contained no supporting information. That is a stronger approach than scoring surface quality alone, but the prospective denominator remains small and the paper does not report downstream correction rates or whether any unsupported item reached a signed record.
Safe use would require provenance and review at field level: users should be able to see which words came from the transcript, which were reformatted and which were generated or inferred. High-risk elements such as allergies, medication, examination findings and diagnostic claims need explicit confirmation. A system should also abstain when audio quality is poor or when its confidence cannot be calibrated, rather than filling a template because the interface expects a value.[2]
What the study means for patients, clinicians and hospitals
For clinicians, the strongest potential benefit is a structured prompt that makes omissions visible during a consultation. That may help a busy team remember important questions. But the study did not measure documentation time, after-hours work, interruption, alert fatigue or the time required to verify a generated note. A more complete draft that takes longer to check—or that encourages automation bias—may not reduce workload.
For patients, this evidence does not establish better diagnosis, faster treatment or safer follow-up. It covers adults with one chief complaint at one hospital and excludes many people who could complicate a first deployment. Patients should be told when audio is processed, how long it is retained, who can access it and how an inaccurate record can be corrected. Consent to a study is not the same as consent to routine recording across a health system.
For hospitals, the deployment question is not whether a model can draft a note. It is whether the full process produces a verified record with fewer clinically meaningful errors at acceptable cost and workload. Procurement should require local prospective evaluation, audit logs, security controls, subgroup analysis, an incident route and clear accountability for the signed record. The assistant should support the clinician, not obscure who made a clinical claim.[1][2]
What would change the assessment
Confidence would increase with a preregistered, multicentre trial that randomises comparable consultations to usual documentation or supervised AI assistance. The primary outcome should count clinically meaningful unsupported statements and omissions in final signed notes, assessed by independent reviewers against recordings and source data. Secondary outcomes should include time, after-hours work, edits, rejected suggestions, diagnostic and treatment changes, patient experience and safety incidents.
The model should be frozen before evaluation and tested across complaints, languages, accents, ages and noisy real clinics. Hospitals need to know how often it abstains, how performance changes when speakers overlap and whether clinicians catch errors under time pressure. Analysis should separate transcription errors from generation errors because the remedy is different: better audio cannot fix a model that invents facts, and better generation cannot recover words the microphone missed.
The research was supported by the Hubei Province Key Research and Development Program and a Wuhan University college-enterprise reform project; the detailed manuscript said the funder had no role and reported no competing interests. Independent replication would help test those findings. Until it exists, the balanced reading is that SMART-ASSISTANT improved one measure of completeness while exposing the more consequential weakness that a fluent note can still be less factual.[1][2]
What this means for people
- Patients could benefit from fewer omitted history details, but unsupported additions could contaminate future care if they are not caught.
- Clinicians may gain a useful prompt and draft while also taking on hidden verification work and responsibility for every accepted field.
- Hospitals need consent, security, correction and accountability processes before routine audio-based documentation.
Global context
This prospective evidence comes from one Mandarin-speaking outpatient service in Wuhan and one tightly defined complaint. Documentation rules, consent, language, staffing and electronic-record systems differ internationally. The underlying question is global—whether AI can reduce clerical burden without degrading the clinical record—but each health system needs local validation and governance rather than importing the reported scores.
What the evidence does not yet show
- The prospective study included 62 adults with abdominal pain at one hospital and was observational rather than randomised.
- The development data and prospective cohort were Chinese-language records from the same institution, limiting transfer to other languages and workflows.
- The theoretical completion rate assumes clinicians correctly follow every useful prompt; it is not a direct measure of the final signed record.
- AI-generated notes scored lower for factuality and the prospective diagnostic accuracy was 50%.
- The study did not measure consultation time, verification workload, treatment changes, adverse events or patient outcomes.
- The transcription workflow required relatively clean, turn-taking speech without dialect or overlapping speakers.
What to watch next
- Multicentre randomised evaluation of final signed notes against recordings and source data.
- Field-level provenance, abstention and audit logs for generated or inferred information.
- Time saved after verification, rejected suggestions, alert fatigue and after-hours documentation.
- Results across languages, accents, complaints, noisy clinics and demographic groups.
- Patient outcomes and incident reporting rather than readability or completeness alone.
Living evidence record
Impact record IAI-0S3U5R1
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
2 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can multi-agent AI improve clinical notes without inventing facts?
In a blinded test of 30 resident-written hyperthyroidism notes, the specialist system raised median documentation scores from 39 to 42, cut positive factual hallucinations to 6.67% versus 26.67% for a general model and saved 112.9 seconds of physician editing per note. The scope is narrow and human review remains essential.
6 min · 1 source
Health & Life Sciences
Can an eye-care LLM improve referrals?
A peer-reviewed Shanghai study randomly assigned 180 community screening participants to standard reports with or without EyeSeek. Among 170 people analysed, referral-related action was higher with the tool, but the single-centre, Chinese-language study was small and short.
8 min · 2 sources
Health & Life Sciences
Can an AI agent take a better eye history than a resident?
A randomized trial at a specialist hospital in Guangzhou found that an LLM agent scored higher than ophthalmology residents on structured pre-consultation histories for 172 non-emergency patients. Interviews took much longer, examination findings narrowed the diagnostic gap, and the study did not test patient outcomes or emergency care.
8 min · 2 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.