Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisTaiwanEast Asia

Wrong AI suggestions shifted five radiologists' bone-age estimates in a small Taiwan study

Six radiologists read 200 radiographs with accurate or deliberately shuffled AI advice. Five showed greater error with sham suggestions, but this controlled study cannot measure patient harm or a population-wide effect.

By The Impact of AI Editorial DeskReleased 28 September 2026 at 14:28 BST4 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesClinical AIAutomation biasRadiologyPatient safety

Research topic

Would radiologists across multiple hospitals identify realistic AI errors under normal workload, and would review safeguards improve patient decisions?

At a glance

  • 1A 28 September peer-reviewed exploratory crossover study used six radiologists and 200 bone-age radiographs, comparing accurate with shuffled AI suggestions.
  • 2Five readers' accuracy was significantly affected by whether the AI advice was accurate; a strong unaided reader was less affected in this small sample.
  • 3The experiment measured reading error, not changes in diagnosis, treatment or outcomes for children.

Living evidence record

Impact record IAI-0GN2NCJ

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

28 September 2026

Source trail

2 direct sources across 2 source types.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

What the new study tested

A peer-reviewed Scientific Reports article published on 28 September describes a small randomized crossover test of how radiologists use an AI estimate when judging a child's bone age from a hand radiograph. Bone age is a measure of skeletal maturity used in assessments of growth. An AI suggestion could make readings more consistent when it is accurate, but it can also anchor a reader to an erroneous answer. The authors, based in Taiwan, deliberately compared both possibilities instead of testing only the tool at its best.

Six radiologists at three seniority levels assessed a set of 200 radiographs. In one condition they saw the model's accurate prediction; in the other, the researchers shuffled the model's predictions to create sham advice. The readers indicated whether they trusted the displayed result. The investigators compared each estimate with the study's reference bone age using mean absolute difference, and examined repeatability with an intraclass correlation coefficient. The earlier ClinicalTrials.gov record documents the intended crossover comparison, but its estimated enrollment and lack of posted results should not be mistaken for the final paper's reported sample.[1][2]

Where the AI suggestion changed a reading

Five of the six radiologists showed a statistically significant difference in error between the accurate-AI and sham-AI conditions in the paper's analysis. The reader with the strongest unaided reliability and lowest unaided error did not show a significant difference between those conditions. One junior reader with weaker baseline performance continued to express high trust even in the shuffled suggestions. When the incorrect suggestion differed from the accurate estimate by more than six months, error increased particularly for junior readers. These are results for named readers in a small experiment, not a population estimate for every radiologist.

The study also found that senior readers disagreed with AI suggestions more often, including some correct ones. That matters because rejecting an algorithm is not automatically evidence of better judgment; the useful distinction is whether the reader can identify when the suggestion is wrong while still taking advantage of accurate support. The paper's design makes the misleading advice visible in a controlled comparison. It does not show how a real hospital's workflow, time pressure, image mix or clinical review would change the same behaviour.[1]

What it means for clinical AI

For a child being assessed for growth or an endocrine condition, an erroneous skeletal-age estimate can lead to unnecessary follow-up or delay a needed investigation. This study did not track any child through diagnosis or treatment, so it cannot quantify those harms. It does show why a product's average accuracy is only one part of safety: the human and model interact, and an incorrect number displayed confidently can alter a clinical reader's answer. Interfaces and training need to make disagreement and uncertainty manageable, with a clear path to independent review.

The sample is the main limit. Six radiologists are too few to infer a general effect by seniority, and repeated readings of 200 radiographs are not equivalent to 1,200 independent clinicians or patients. The sham predictions were produced by randomizing real model outputs, which tests susceptibility to deliberately wrong advice but may not match the pattern of errors made by a deployed system. The paper reports funding from Taiwan's National Science and Technology Council and Cheng Hsin General Hospital and declares no competing interests. A useful next study would enroll more readers across hospitals, test realistic error distributions and examine downstream decisions rather than image-reading error alone.[1][2]

What this means for people

  • Children receiving growth assessments need a clinician able to challenge a wrong AI estimate rather than treating the displayed number as an answer.
  • Hospitals evaluating AI aids should test how staff use them under realistic error conditions, not rely solely on the model's standalone accuracy.

Global context

This is a Taiwan-based exploratory study with six readers. It cannot establish how radiologists elsewhere, different products or healthcare systems would respond.

What the evidence does not yet show

  • Six radiologists and a controlled 200-image set limit generalization; repeated readings do not create a large independent human sample.
  • Randomly shuffled AI outputs are a deliberate stress test and may differ from a deployed model's typical mistakes.
  • No patient outcome or real clinical workflow was measured; the study registry has no posted results beyond the linked paper.

What to watch next

  • Larger, multisite reader studies with realistic error patterns and predefined safety measures.
  • Whether interface design, uncertainty display and second review reduce harmful reliance on incorrect advice.

Evidence trail

Sources used for this report

Links checked 28 September 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.