Can paediatricians resist wrong AI advice?
A Greek study of 98 paediatricians found that doctors usually rejected deliberately incorrect advice labelled as AI-generated, but residents and private-practice clinicians were more likely than hospital attendings to change towards a wrong answer. The experiment used fabricated advice and fixed-order vignettes, not a live clinical system.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1Ninety-eight Greek paediatricians answered 16 multiple-choice questions across four vignettes, creating 1,568 clinician-question observations before modelling; the labelled AI advice was fabricated and standardized by the researchers.
- 2Doctors were more likely to end in agreement with the labelled AI when it was correct than when it was wrong, but residents and private-practice clinicians had substantially higher odds than hospital attendings of aligning with wrong advice.
- 3The fixed sequence always moved from the easiest, fully correct case to harder cases containing more wrong advice, so correctness, case difficulty and order cannot be cleanly separated.
Research topic
Whether paediatricians change their clinical answers towards advice labelled as AI-generated when that advice is correct or deliberately wrong

The experiment tested deference, not an actual language model
The study asked a narrow human-factors question: when a paediatrician has already chosen an answer, will a second answer presented as coming from ChatGPT or artificial intelligence pull the clinician towards it? Researchers recruited 98 paediatricians in Greece between February and July 2025. Forty-one worked in private practice and 57 in hospitals; the hospital group comprised 38 residents and 19 attendings. The sample was 89% female, the average age was 41.4 years and the average reported clinical experience was 13.4 years.
Each participant completed four paediatric case scenarios with four multiple-choice questions per case. That produced 16 decisions per clinician, or 1,568 clinician-question observations before the mixed-effects analysis. The cases were written by the research team and reviewed by three senior academic hospital paediatricians. Participants first answered independently, then saw a response described as AI-generated and could keep or change their answer. Crucially, no model generated the advice: the researchers constructed the responses so every participant encountered the same pattern of right and wrong suggestions.[1]
The advice became less reliable as the cases became harder
The four cases were delivered in a fixed order. The labelled AI was correct on all four questions in case one, three of four in case two, two of four in case three and one of four in case four. Across the full exercise, each participant therefore saw ten correct and six incorrect AI-labelled answers. This design creates a useful within-person contrast, but it also means the study cannot fully tell whether doctors reacted to answer correctness, increasing case difficulty, fatigue or a trust expectation established by the initially perfect advice.
The outcomes were also behavioural proxies rather than clinical endpoints. One measure asked whether the participant's final answer aligned with the labelled AI; another examined whether a clinician who initially disagreed switched towards it. Multiple-choice vignettes remove the uncertainty, team communication, examination findings and time pressure of real paediatric care. They can reveal susceptibility to a displayed recommendation, but not whether an AI assistant improves diagnosis, how often clinicians would consult one, or what happens to patients.[1]
Doctors distinguished right from wrong advice, but not perfectly
Final answers were much more likely to match the labelled AI when its suggestion was correct: the adjusted odds were 2.92 times those for an incorrect suggestion, with a 95% confidence interval from 2.35 to 3.63. Among decisions where the clinician initially disagreed, the odds of changing towards the advice were 1.97 times higher when it was correct than when it was wrong, with a 95% confidence interval from 1.40 to 2.77. These results show discrimination rather than blind obedience.
That discrimination was incomplete. An initially correct clinician was less likely to end up aligned with a wrong suggestion, with an odds ratio of 0.234 compared with someone whose initial answer was not correct. Yet some correct answers were still abandoned. Odds ratios describe relative odds after adjustment, not the percentage of doctors who would make a harmful error in ordinary practice. The article does not establish a patient-harm rate, and the repeated questions mean observations from the same clinician are not independent; the authors addressed that clustering with participant-level random intercepts.[1]
Residents and private-practice clinicians were more susceptible in this sample
Across all advice, residents had 2.25 times the adjusted odds of changing towards the labelled AI compared with hospital attendings, while private-practice paediatricians had 2.41 times the odds. The difference widened when the displayed suggestion was wrong: residents had 4.95 times the odds of final alignment with incorrect advice, and private-practice clinicians had 5.85 times the odds, again compared with attendings. The corresponding 95% confidence intervals were wide—1.28 to 19.1 and 1.70 to 20.2—reflecting small subgroups and uncertainty about the true size of the difference.
It would be a mistake to treat professional position as a fixed personal weakness. The attending comparator included only 19 people, private-practice clinicians may see a different mix of cases, and the vignettes leaned towards hospital-style acute care. Access to colleagues, recent examination training, familiarity with multiple-choice tasks and differences in case exposure could all contribute. The data identify groups for further study and tailored training; they do not prove that residents or community clinicians are inherently unsafe users of AI.[1]
The result points to safeguards around the decision, not only the model
Health organisations often focus AI assurance on model accuracy, but this experiment shows why the interface and the user's working environment matter as well. Even a system with a respectable average score will sometimes be wrong, and a confident recommendation can redirect a clinician who had the better initial answer. Deployment controls should make uncertainty visible, preserve the clinician's independent assessment, link claims to evidence, and make disagreement easy rather than framing acceptance as the default action.
Training can use deliberately mixed-quality examples like these, but it should be evaluated against behaviour rather than attendance. Teams can measure whether clinicians verify high-risk suggestions, how often they reverse a correct first judgment, whether error patterns differ by seniority or isolation, and whether escalation routes are actually used. Junior and private-practice clinicians may particularly benefit from rapid access to a second human opinion. The practical goal is calibrated reliance: accept helpful advice for defensible reasons and resist it when the patient's evidence points elsewhere.[1]
The study has important design and reporting limits
This was a convenience sample from university-affiliated clinics, regional hospitals and private practices in Greece, not a nationally representative survey. The authors did not measure confidence, reasoning steps or the time spent reconsidering each answer, so the mechanism behind a switch remains uncertain. The fixed case order and declining advice accuracy confound trust, difficulty and sequence. Because the prompts were described as AI-generated but manually fabricated, the experiment also cannot compare particular models, prompting strategies or versions of ChatGPT.
The paper reports no external funding and the authors declare no competing interests. Data and code are available from the corresponding author on reasonable request rather than in an open repository. One sentence in the limitations section says seven people declined and then parenthetically refers to a final sample of 105, while the abstract, methods, results and tables consistently report 98 participants. That internal inconsistency should be corrected by the journal or authors; this analysis uses the consistent 98-person denominator and does not infer an unreported group.[1]
What would change the assessment
Confidence would increase with a randomized design that varies advice accuracy and case order independently, uses outputs from named model versions, and includes enough attendings, residents and community clinicians for precise subgroup estimates. Multicountry studies should include outpatient as well as acute-care cases and report absolute switching rates alongside odds ratios. Capturing confidence, justification and source-checking would help distinguish thoughtful correction from automation bias.
The decisive evidence would come from prospective clinical workflow studies with safe oversight: real or realistically simulated records, representative case prevalence, documented model abstention, and patient-relevant error measures. Researchers should test whether interface changes, mandatory independent answers, evidence links or human escalation reduce harmful reversals without blocking useful corrections. Until then, the study supports a credible warning about reliance on labelled AI advice, especially for less senior or more isolated clinicians, but not a claim that current paediatric AI improves or worsens patient outcomes.[1]
What this means for people
- A clinician can be pulled away from a correct initial judgement when a confident AI-labelled recommendation is wrong.
- Junior and professionally isolated clinicians may need stronger access to verification and human escalation, not simply a generic warning to be careful.
- Patients are not represented directly in this vignette study, so it cannot quantify diagnosis errors, treatment changes or harm.
Global context
The evidence comes from 98 clinicians in Greece and a deliberately controlled vignette exercise. Training structures, access to senior advice, specialty practice and legal responsibility differ internationally. Hospitals and health systems elsewhere should test clinician–AI interaction locally and report absolute error patterns before treating these subgroup results as universal.
What the evidence does not yet show
- The convenience sample included 98 Greek paediatricians and only 19 hospital attendings, so subgroup estimates are imprecise and not nationally representative.
- Researchers fabricated the standardized advice; the study did not test a named language model, live ChatGPT output or model-to-model differences.
- Case difficulty increased while advice accuracy declined in a fixed order, confounding correctness with sequence, fatigue, priming and difficulty.
- Multiple-choice vignettes do not reproduce clinical examinations, team discussion, live records, time pressure or patient outcomes.
- The paper did not measure confidence, reasoning processes or verification behaviour, so it cannot explain why a clinician changed an answer.
- An internal sentence refers inconsistently to a 105-person final sample; the abstract, methods, results and tables consistently report 98, which is the denominator used here.
What to watch next
- Randomized studies that separate advice accuracy, case difficulty and presentation order.
- Prospective workflow tests using named model versions and reporting absolute harmful and helpful reversal rates.
- Whether evidence links, uncertainty displays and a required independent first answer reduce automation bias.
- Replication with larger attending, resident and community-clinician groups across countries and care settings.
Living evidence record
Impact record IAI-0083U45
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
5 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 5 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can a medical AI support a clinical board without becoming the decision-maker?
A peer-reviewed German study put OpenEvidence into discussions of 100 real inflammatory-disease cases. Specialists found the answers useful, but complete agreement with the board was limited, prompt tuning was inconclusive and outputs changed over time. The study tested workflow feasibility—not patient benefit or safety.
9 min · 1 source
Health & Life Sciences
Can an AI agent take a better eye history than a resident?
A randomized trial at a specialist hospital in Guangzhou found that an LLM agent scored higher than ophthalmology residents on structured pre-consultation histories for 172 non-emergency patients. Interviews took much longer, examination findings narrowed the diagnostic gap, and the study did not test patient outcomes or emergency care.
8 min · 2 sources
Health & Life Sciences
Can a health chatbot pass an equity audit?
A peer-reviewed audit generated 1,350 mental-wellness responses from three open models. It found complex reading levels, uneven crisis language and almost no explicit clinician-referral wording—but three repeated prompts and no community validation sharply limit the claim.
8 min · 2 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.