Can AI spot liposarcoma on ultrasound?
A peer-reviewed Chinese study reports strong results in a 95-patient internal test, including 0.97 accuracy. But all 317 patients came from one hospital and there was no external validation, so this is a proof of concept—not a clinically cleared diagnostic system.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1Researchers retrospectively analysed 317 patients from Beijing Jishuitan Hospital: 126 had liposarcoma and 191 had benign masses. The patient-level split used 222 people for development and 95 for internal testing.
- 2In the internal test, the model reported an AUC of 0.97, accuracy of 0.97, sensitivity of 0.96 and specificity of 0.97. It outperformed two resident radiologists on accuracy after multiple-comparison adjustment but did not significantly outperform two specialists.
- 3The evidence is preliminary: every patient came from one hospital, the test set was small, 2,964 correlated frames represented only 317 independent people, and there was no external, prospective or workflow validation.
Research topic
Whether a deep-learning model can distinguish liposarcoma from benign soft-tissue masses on ultrasound images

The finding is promising, but it is not ready to diagnose patients
The direct answer is that the model distinguished liposarcoma from benign soft-tissue masses very well in one hospital's internal test, but the study cannot show that it will work safely in another clinic. The peer-reviewed paper reports 0.97 accuracy and an area under the receiver-operating-characteristic curve, or AUC, of 0.97 in 95 held-out patients. Those are strong discrimination results. They are still estimates from a small internal sample drawn from the same institution and historical period as the development data.
That distinction matters because liposarcoma is a rare malignant soft-tissue tumour that can resemble benign lesions on imaging. Ultrasound is accessible and comparatively inexpensive, but interpretation depends on the operator and the lesion. A useful assistant could help less experienced readers decide which masses need specialist imaging, biopsy or referral. A wrong benign classification could also delay cancer care, while a false alarm could expose a patient to anxiety, cost and unnecessary procedures.
The authors appropriately describe the system as having preliminary potential and say external validation is required before clinical use. The study therefore supports further testing, not deployment, regulatory clearance or replacement of radiologists and histopathology.[1]
What the researchers did
The team retrospectively assembled B-mode ultrasound examinations performed at Beijing Jishuitan Hospital between March 2015 and March 2022. The paper says 504 patients were initially enrolled. It excluded 185 people with incomplete baseline information and two with atypical lipomatous tumours, a borderline entity that can overlap benign lesions and well-differentiated liposarcoma. The final cohort contained 317 patients: 126 with liposarcoma and 191 with benign masses. Mean age was 51.6 years; 178 were male and 139 female.
The system used two stages. A U-Net first segmented the tumour region in an ultrasound image, and a ResNet50 classifier then predicted whether the mass was liposarcoma or benign. The researchers kept the split at patient level, assigning 222 people to model development and 95 to an internal test. That is an important safeguard because assigning frames from the same person to both sets would leak patient-specific information and exaggerate performance.
The data nevertheless contained 2,964 images from only 317 independent patients. Multiple frames from one examination are correlated, so the apparent image count is not the true denominator for clinical generalisation. The authors used patient-level five-fold cross-validation within the development cohort, then estimated final performance once on the 95-person internal test. A few changed classifications could therefore move the reported rates materially.[1]
What the numbers show
In the internal test, the model's AUC was 0.97 with a 95% confidence interval from 0.92 to 1.00. Accuracy was 0.97, also with a 95% confidence interval from 0.93 to 1.00. Sensitivity—the proportion of liposarcomas identified as such—was 0.96, with a confidence interval from 0.86 to 1.00. Specificity—the proportion of benign masses correctly classified—was 0.97, with a confidence interval from 0.93 to 1.00.
The study also compared the AI with four radiologists on the same internal test. After Holm adjustment for multiple comparisons, the model had significantly higher accuracy than each of two residents, with adjusted P values of 0.0156 and 0.0352. Differences between the AI and either of two specialist radiologists were not statistically significant; both adjusted P values were 1.0000. That result is more informative than a generic claim that the model ‘beat doctors’: it suggests possible assistance for less experienced readers while providing no evidence of superiority over specialists.
The comparison still does not tell us whether AI assistance improves radiologists' decisions. The paper compared readers with the model; it did not randomise clinicians to read cases with and without assistance. It also did not measure whether using the system changes referrals, biopsy decisions, diagnostic delay, patient outcomes or workload. Standalone discrimination is an early technical endpoint, not a clinical-benefit endpoint.[1]
Why an internal test can look better than real-world use
All patients came from a single specialist hospital. The same institution's scanners, image-acquisition habits, referral patterns, patient mix and annotation practices shaped both development and testing. A model can learn signals linked to those local conditions as well as features of the disease. Performance can fall when scanners, operators, tumour prevalence or competing diagnoses change.
The retrospective design creates additional uncertainty. The included cases had already reached a diagnosis and had usable baseline information, while 185 of the 504 initially enrolled patients were excluded for incomplete data. The paper does not show how the model performs on every consecutive person who arrives with an uncertain soft-tissue mass. Excluding atypical lipomatous tumours also removes an especially difficult clinical boundary from the binary task.
The authors highlight the uneven number of frames between benign and liposarcoma groups, which may encode acquisition behaviour. They also acknowledge that one 95-patient test partition remains vulnerable to overfitting and partition-specific variation. Resampling within the same source population cannot replace testing on new hospitals, vendors, technicians and patients.[1]
What this could mean for people
If later studies reproduce the result, an ultrasound assistant could be most useful where experienced musculoskeletal radiologists are scarce. It might flag a suspicious mass for specialist review or help standardise an initial assessment. Because ultrasound is widely available and does not require ionising radiation, that pathway could be valuable in settings where MRI access is limited or delayed.
The tool should not be framed as a definitive cancer test. Liposarcoma diagnosis still depends on the wider clinical picture and, when indicated, cross-sectional imaging and pathology. A safe workflow would present the model as one input, preserve human escalation, record disagreement and prevent a low-risk output from closing a case that remains clinically concerning.
Benefits and harms may also differ across groups. The paper reports overall age and sex counts, but this single-centre sample cannot establish performance across countries, ethnicities, body types, tumour subtypes, anatomical sites or scanning equipment. Prospective studies should pre-specify subgroup analyses and report calibration as well as discrimination, because a probability estimate used to guide action must mean the same thing across settings.[1]
Funding, interests and the next evidence needed
The authors report support from China's National Natural Science Funds, grant 82572229. They declare no commercial or financial relationships that could be construed as a conflict of interest and say generative AI was not used to create the manuscript. The data are available from the corresponding author on reasonable request rather than through an open repository, which can make independent replication slower.
The decisive next step is a locked-model external validation across several hospitals, ultrasound systems and operators, using consecutive patients and a pre-specified protocol. That study should include the difficult borderline and alternative diagnoses seen in practice, publish calibration and subgroup results, and keep all frames from each patient together. Confidence would rise further with a prospective reader-assistance trial comparing unaided care with AI-supported care.
A clinically meaningful evaluation would measure more than AUC: missed cancers, unnecessary referrals and biopsies, time to definitive diagnosis, reader workload, patient anxiety, costs and safety incidents all matter. Until those results exist, the fairest conclusion is narrow. The model performed strongly in an internal 95-patient test and matched specialist accuracy statistically, but it remains a single-centre proof of concept rather than evidence of safe clinical benefit.[1]
What this means for people
- Patients could benefit from faster escalation of suspicious masses, but an incorrect benign result could delay cancer diagnosis.
- Less experienced radiologists may gain decision support; specialists were not significantly outperformed in this study.
- Health systems need prospective workflow and cost evidence before treating the model as a safe clinical tool.
Global context
The study comes from a major orthopaedic centre in Beijing and addresses a globally relevant diagnostic problem. Its results cannot yet be assumed to transfer to community hospitals, other countries, different patient populations or other ultrasound equipment. Multicentre external validation is the minimum evidence needed to test that transfer.
What the evidence does not yet show
- This was a retrospective study from one specialist hospital; there was no external, prospective or multicentre validation.
- The final test contained only 95 independent patients, so a small number of different classifications could materially change the metrics.
- The 2,964 images represented 317 patients, and frames from the same examination were correlated; the image count is not the clinical sample size.
- Exclusions included 185 patients with incomplete baseline information and two with atypical lipomatous tumours, limiting applicability to an all-comers diagnostic pathway.
- The reader study compared standalone AI with four radiologists; it did not test whether AI assistance improves clinician decisions or patient outcomes.
What to watch next
- External validation using a locked model at multiple hospitals and with multiple ultrasound vendors.
- Prospective studies enrolling consecutive patients, including borderline and difficult alternative diagnoses.
- Randomised or crossover reader studies comparing radiologists with and without AI assistance.
- Calibration, subgroup performance, workflow safety, missed cancers, unnecessary biopsies and time to diagnosis.
- Independent replication with accessible protocols and data or a well-governed validation dataset.
Living evidence record
Impact record IAI-08MWR71
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
7 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Frontiers in Medicine published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 7 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Does 99% MRI accuracy mean an Alzheimer’s diagnostic is ready?
No. A peer-reviewed model classified images from one 6,400-image Kaggle collection with reported 99% test accuracy, but it has no external clinical validation and the published test counts do not reconcile cleanly.
8 min · 2 sources
Health & Life Sciences
Can AI discover new MRI signs for glioblastoma?
It generated eight candidate visual signs from 106 glioblastoma scans, but only one was externally tested. Two radiologists achieved moderate discrimination between 50 glioblastomas and 50 metastases, so this is a discovery signal—not a clinical diagnostic.
7 min · 1 source
Health & Life Sciences
Can mouse MRI predict glioblastoma treatment response?
A deep-learning model separated cure from relapse across 298 MRI examinations, but the effective test cohort was only 10 treated mice from one laboratory. It is preclinical evidence, not a patient-ready predictor.
7 min · 1 source
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.