Can a high-AUC diabetes model still be unsafe?
A peer-reviewed audit of 12 machine-learning approaches on two public diabetes datasets finds that similarly ranked models can differ materially in calibration, uncertainty and safe deferral. The work is a benchmark study—not a clinical trial, diagnostic approval or evidence of improved patient outcomes.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Whether predictive ranking, probability calibration, uncertainty estimates and selective deferral tell different safety stories when machine-learning models are benchmarked for diabetes detection

At a glance
- 1The study compared 12 models using a balanced 70,692-record derivative of the 2015 US BRFSS survey and a separate 520-person early-stage symptom dataset.
- 2On the larger dataset, the three leading PR-AUC values were tightly clustered at 0.804, 0.803 and 0.800, while calibration and uncertainty measures produced different model rankings.
- 3The tests are retrospective and dataset-bound: no model made a clinical decision, improved an outcome or underwent external, prospective or multi-centre validation.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-19REDRR
Evidence stage
Studied
Confidence
Corroborated
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
4 October 2026
Source trail
3 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
The study asks a broader question than which model scores highest
A diabetes classifier can separate higher-risk from lower-risk records reasonably well and still give probabilities that are too confident, unstable or difficult to act on. The new Scientific Reports paper treats that distinction as a safety problem. Rather than selecting a winner from one accuracy table, the authors compare discrimination, probability calibration, uncertainty, selective prediction and conformal coverage across 12 classical, ensemble and deep tabular models.
That framing matters in health care because a probability is often used downstream: to set a referral threshold, communicate risk or decide which cases need human review. A high area under a curve does not show that a quoted 70% risk occurs in roughly seven of ten comparable people. It also does not show that the model recognises unfamiliar cases. The paper therefore provides an audit framework, not a claim that any tested model is ready to diagnose diabetes.[1][2]
Two public datasets create very different tests
The larger dataset contains 70,692 records and 21 predictors derived from the 2015 US Behavioral Risk Factor Surveillance System. It is balanced between diabetes and non-diabetes classes, unlike natural prevalence, and relies heavily on self-reported survey information. Age, body-mass index, blood pressure, general health and related indicators can be predictive without being a clinical work-up. The balancing makes model comparison convenient but changes the probability environment in which calibration would be used.
The second dataset contains only 520 records with 16 symptom and demographic features. It has much clearer class separation and yields near-ceiling results, including PR-AUC values around 0.99 for several models. That can reveal how uncertainty methods behave in a small, easy benchmark, but it is also a warning: 520 observations leave wide scope for sampling instability, and symptoms collected for a diabetes dataset may not resemble unselected primary-care populations.[1][2][3]
Twelve models were compared inside a leakage-controlled pipeline
The authors evaluated logistic regression, support-vector and probabilistic baselines, tree ensembles including random forest, XGBoost and CatBoost, and deep tabular approaches including multilayer perceptrons, TabNet and FTTransformer. They used a stratified five-fold outer cross-validation design. Preprocessing was fitted inside training data, the held-out fold was reserved for testing, and a separate calibration subset was used for temperature scaling. Those choices reduce a common source of optimistic leakage.
On the BRFSS derivative, FTTransformer had the highest reported mean PR-AUC at 0.804, followed by XGBoost at 0.803 and CatBoost at 0.800. The differences are too small to support a sweeping claim that one model family dominates. The more useful finding is that models with nearly indistinguishable discrimination had different expected calibration errors, uncertainty scores and composite safety rankings. Optimising only the leaderboard metric would conceal those operational differences.[1][2]
Deferral improved error rates, but coverage is a design choice
The researchers estimated epistemic uncertainty with repeated predictions and tested a selective strategy that retained 90% of cases while deferring the most uncertain tenth. On the larger dataset, mean error fell by roughly 7%. On the small early-stage dataset, the largest relative reduction was 43.7% for the multilayer perceptron, although several confidence intervals for that dataset included zero. A large relative gain from a small initial error rate is not the same as a large clinical benefit.
Deferral does not make difficult cases disappear. It transfers them to another process—usually a clinician, laboratory test or follow-up pathway. A real service would need to measure how many patients are deferred, whether uncertainty is concentrated in particular demographic groups, how quickly review occurs and what errors remain among confidently automated cases. The paper reports age- and sex-based checks, but those limited attributes cannot establish equity across ethnicity, disability, income, geography or access to care.[1][2]
Conformal coverage is conditional, not a safety guarantee
Split conformal prediction produced empirical coverage close to nominal 90% and 95% targets. This means the generated prediction sets contained the true class at approximately the promised marginal rate in held-out data under the study's exchangeability assumptions. Some uncertain cases receive both possible labels, making ambiguity visible rather than forcing a single answer. That is useful behaviour for triage systems, where an honest 'both remain plausible' can trigger review.
The guarantee does not survive every distribution shift. If deployment populations differ in prevalence, measurement, access, survey response or disease definition, past residuals may no longer represent future cases. Marginal coverage can also hide undercoverage in smaller subgroups. The datasets were not collected as a prospective clinical pathway, and no threshold was tied to treatment benefits or harms. Conformal validity here is a statistical benchmark result, not evidence that patients would be safer.[1][2]
What would change the assessment
The next step is not another random split of the same tables. Stronger evidence would freeze a model and evaluate it on later patients from multiple health systems, with clinically verified outcomes and prespecified calibration, subgroup, deferral and decision-curve measures. Natural disease prevalence must be retained or corrected explicitly. The evaluation should compare the AI-assisted pathway with usual care and report missed cases, unnecessary tests, time to diagnosis, workload and patient-relevant outcomes.
Until then, the study supports a procurement question rather than a purchasing conclusion: show the calibration plot, uncertainty behaviour and deferred-case workflow, not only AUC. Its most consequential result is that similar headline rankings can mask different safety properties. It does not establish that any of the 12 models should screen, diagnose or prioritise a person outside the two public datasets.[1][2][3]
What this means for people
- Patients could benefit if uncertainty routes difficult cases to appropriate review, but a poorly calibrated probability could still mislead referrals or reassurance.
- Clinicians and screening services need to know how many cases a model defers and whether the review workload is safe and sustainable.
- Developers and buyers should not treat a strong AUC as evidence of reliable probabilities, equitable performance or clinical benefit.
Global context
The larger dataset is derived from a US telephone survey, the small symptom dataset originated in Bangladesh, and the study includes researchers and funding from Saudi Arabia. That geographic spread does not create global validation. Diabetes prevalence, coding, clinical access and predictor distributions vary across countries, so local temporal and multi-centre testing remains essential.
What the evidence does not yet show
- Both datasets are retrospective public benchmarks; neither represents a prospective clinical deployment or a completed diagnostic pathway.
- The 70,692-record BRFSS derivative is artificially balanced and based largely on self-reported US survey variables, so its probabilities do not directly represent natural prevalence.
- The 520-record symptom dataset is small and highly separable, making near-ceiling results and uncertainty gains vulnerable to sampling variation.
- Fairness checks were limited to age and sex, while conformal coverage was marginal and depends on exchangeability.
- The composite safety score depends on chosen metrics and weights; it is an analytic device, not a clinical regulatory standard.
What to watch next
- Locked external validation across later time periods, health systems and naturally occurring prevalence.
- Prospective comparisons reporting downstream tests, clinician workload, missed diagnoses, subgroup calibration and patient outcomes.
- Transparent thresholds and review capacity for deferred cases rather than uncertainty scores without an operational pathway.
Evidence trail
Sources used for this report
Links checked 4 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Did clinicians prefer AI discharge summaries after long hospital stays?
In a retrospective 60-case comparison, 12 physicians usually preferred GPT-5.2 summaries and annotated fewer omissions. Reviewers knew which summary was AI-written, one hospital supplied the records, and no patient outcome or time saving was tested.
7 min · 3 sources
Health & Life Sciences
Can this MRI model draw brain-tumour boundaries reliably?
A peer-reviewed multimodal segmentation model was developed on 2,422 public MRI volumes and externally tested on 125 cases. Accuracy remained useful but fell outside the development data; no prospective clinical workflow, radiologist comparison or patient-outcome test was performed.
8 min · 2 sources
Health & Life Sciences
Wrong AI suggestions shifted five radiologists' bone-age estimates in a small Taiwan study
Six radiologists read 200 radiographs with accurate or deliberately shuffled AI advice. Five showed greater error with sham suggestions, but this controlled study cannot measure patient harm or a population-wide effect.
4 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.