Wrong AI suggestions shifted five radiologists' bone-age estimates in a small Taiwan study
Six radiologists read 200 radiographs with accurate or deliberately shuffled AI advice. Five showed greater error with sham suggestions, but this controlled study cannot measure patient harm or a population-wide effect.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Would radiologists across multiple hospitals identify realistic AI errors under normal workload, and would review safeguards improve patient decisions?
At a glance
- 1A 28 September peer-reviewed exploratory crossover study used six radiologists and 200 bone-age radiographs, comparing accurate with shuffled AI suggestions.
- 2Five readers' accuracy was significantly affected by whether the AI advice was accurate; a strong unaided reader was less affected in this small sample.
- 3The experiment measured reading error, not changes in diagnosis, treatment or outcomes for children.
Living evidence record
Impact record IAI-0GN2NCJ
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
28 September 2026
Source trail
2 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
What the new study tested
A peer-reviewed Scientific Reports article published on 28 September describes a small randomized crossover test of how radiologists use an AI estimate when judging a child's bone age from a hand radiograph. Bone age is a measure of skeletal maturity used in assessments of growth. An AI suggestion could make readings more consistent when it is accurate, but it can also anchor a reader to an erroneous answer. The authors, based in Taiwan, deliberately compared both possibilities instead of testing only the tool at its best.
Six radiologists at three seniority levels assessed a set of 200 radiographs. In one condition they saw the model's accurate prediction; in the other, the researchers shuffled the model's predictions to create sham advice. The readers indicated whether they trusted the displayed result. The investigators compared each estimate with the study's reference bone age using mean absolute difference, and examined repeatability with an intraclass correlation coefficient. The earlier ClinicalTrials.gov record documents the intended crossover comparison, but its estimated enrollment and lack of posted results should not be mistaken for the final paper's reported sample.[1][2]
Where the AI suggestion changed a reading
Five of the six radiologists showed a statistically significant difference in error between the accurate-AI and sham-AI conditions in the paper's analysis. The reader with the strongest unaided reliability and lowest unaided error did not show a significant difference between those conditions. One junior reader with weaker baseline performance continued to express high trust even in the shuffled suggestions. When the incorrect suggestion differed from the accurate estimate by more than six months, error increased particularly for junior readers. These are results for named readers in a small experiment, not a population estimate for every radiologist.
The study also found that senior readers disagreed with AI suggestions more often, including some correct ones. That matters because rejecting an algorithm is not automatically evidence of better judgment; the useful distinction is whether the reader can identify when the suggestion is wrong while still taking advantage of accurate support. The paper's design makes the misleading advice visible in a controlled comparison. It does not show how a real hospital's workflow, time pressure, image mix or clinical review would change the same behaviour.[1]
What it means for clinical AI
For a child being assessed for growth or an endocrine condition, an erroneous skeletal-age estimate can lead to unnecessary follow-up or delay a needed investigation. This study did not track any child through diagnosis or treatment, so it cannot quantify those harms. It does show why a product's average accuracy is only one part of safety: the human and model interact, and an incorrect number displayed confidently can alter a clinical reader's answer. Interfaces and training need to make disagreement and uncertainty manageable, with a clear path to independent review.
The sample is the main limit. Six radiologists are too few to infer a general effect by seniority, and repeated readings of 200 radiographs are not equivalent to 1,200 independent clinicians or patients. The sham predictions were produced by randomizing real model outputs, which tests susceptibility to deliberately wrong advice but may not match the pattern of errors made by a deployed system. The paper reports funding from Taiwan's National Science and Technology Council and Cheng Hsin General Hospital and declares no competing interests. A useful next study would enroll more readers across hospitals, test realistic error distributions and examine downstream decisions rather than image-reading error alone.[1][2]
What this means for people
- Children receiving growth assessments need a clinician able to challenge a wrong AI estimate rather than treating the displayed number as an answer.
- Hospitals evaluating AI aids should test how staff use them under realistic error conditions, not rely solely on the model's standalone accuracy.
Global context
This is a Taiwan-based exploratory study with six readers. It cannot establish how radiologists elsewhere, different products or healthcare systems would respond.
What the evidence does not yet show
- Six radiologists and a controlled 200-image set limit generalization; repeated readings do not create a large independent human sample.
- Randomly shuffled AI outputs are a deliberate stress test and may differ from a deployed model's typical mistakes.
- No patient outcome or real clinical workflow was measured; the study registry has no posted results beyond the linked paper.
What to watch next
- Larger, multisite reader studies with realistic error patterns and predefined safety measures.
- Whether interface design, uncertainty display and second review reduce harmful reliance on incorrect advice.
Evidence trail
Sources used for this report
Links checked 28 September 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Pathology agent learns where clinicians look on cancer slides; clinical value remains unproven
A published research system uses pathologists' viewing behaviour to guide slide analysis and reports stronger performance on a lymph-node metastasis task. It has not shown improved patient outcomes.
5 min · 1 source
Health & Life Sciences
Healthcare review treats model safety and hospital security as one connected problem
A Nature review assesses the safety and security of large language models themselves and the risks created when they are connected to hospital data, software and human workflows.
4 min · 1 source
Health & Life Sciences
AI-guided formulation keeps experimental mRNA vaccines active after heat storage
A peer-reviewed MIT-led study reports solid-state mRNA–lipid nanoparticle formulations that retained laboratory bioactivity after more than two months at 37°C. The result is preclinical and does not yet establish a human vaccine shelf life.
5 min · 3 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.