Can genomic AI reduce missed antibiotic resistance?
A peer-reviewed model trained on 5,942 Klebsiella pneumoniae genomes made fewer false-susceptible predictions than two established genomic tools in retrospective global, African and Thai benchmarks. It has not yet shown that it improves prescribing or patient outcomes in live care.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1PredKlebAMR used 38 Kleborate-derived genomic predictors and 5,942 Klebsiella pneumoniae genomes, with a stratified 80:20 internal split and five-fold cross-validation.
- 2Its reported sensitivity was 95.45% in a global external cohort, 95.56% in an African cohort and 95.62% in a Thai cohort; only the African sensitivity comparison was reported as statistically significant.
- 3The evidence is retrospective and genomic: it does not show faster results, better antibiotic choices, fewer deaths or safe deployment in a clinical laboratory.
Research topic
Does a locked genome-based model reduce false-susceptible meropenem reports without creating harmful false-resistance calls or delaying effective treatment when tested prospectively across laboratories and bacterial lineages?

The direct answer: fewer dangerous misses in retrospective benchmarks, not yet better care
PredKlebAMR made fewer false-susceptible predictions than AMRFinderPlus and ResFinder in three external genomic benchmarks, including cohorts from Africa and Thailand. That is an important technical result because a false-susceptible call can make a resistant Klebsiella pneumoniae isolate look treatable. But the study did not place the model in a live laboratory, compare prescribing decisions or follow patients. It therefore supports further prospective evaluation, not a claim that the system already improves treatment.
The model was designed to prioritise sensitivity: finding as many resistant isolates as possible. On the authors' held-out global test set it reached 91.92% accuracy, 94.67% sensitivity, 89.47% specificity and a Cohen's kappa of 0.839. Those figures describe agreement with the study's resistance labels under its retrospective protocol. They do not measure the probability that a particular medicine will work for a particular patient, and they should not replace phenotypic antimicrobial-susceptibility testing where that testing is clinically required.[1]
What the model learned from 5,942 genomes
The researchers retrieved 6,000 Klebsiella pneumoniae genomes from the Bacterial and Viral Bioinformatics Resource Center. After preprocessing, 5,942 genomes were used for model development. Rather than feeding raw sequence directly into a black box, the team used 38 predictors generated through Kleborate, a genomic surveillance tool. The features included known resistance determinants and bacterial lineage information. A random-forest classifier was trained with a stratified 80:20 split, while five-fold cross-validation was used to select hyperparameters.
Reported influential predictors included fluoroquinolone-resistance mutations, high-risk sequence types, the KPC-2 and KPC-3 carbapenemase genes, and changes in the OmpK35 and OmpK36 porins. That makes the model more biologically legible than a purely opaque sequence classifier: its decisions draw on features already associated with resistance. It also creates a boundary. A model built from known, engineered signals can underperform when new mechanisms, unfamiliar lineages or combinations absent from its training data appear.
The paper concerns resistance profiling for K. pneumoniae, a pathogen associated with hospital and community infections and with globally important antimicrobial resistance. It is not a universal bacterial-resistance model. Performance for this species, these data sources and the studied drugs cannot be transferred automatically to other organisms, laboratories or resistance phenotypes.[1]
How it compared with established genomic tools
The authors compared PredKlebAMR with AMRFinderPlus version 4.2.7 and ResFinder version 4.1 on independent cohorts described as global, African and Thai. In the global cohort, PredKlebAMR's sensitivity was 95.45%, compared with 90.91% for AMRFinderPlus and 86.36% for ResFinder. The model produced one very major error, versus two and three respectively. In the African cohort, sensitivity was 95.56%, compared with 91.11% and 82.2%, with two very major errors rather than four or eight.
The African sensitivity comparison was the one reported as statistically significant, using McNemar's test with p=0.041. In the Thai cohort, sensitivities were 95.62% for PredKlebAMR, 92.5% for AMRFinderPlus and 63.12% for ResFinder; the reported very major error counts were seven, 12 and 59. A very major error here means a resistant isolate was predicted to be susceptible. It is a high-consequence analytical error category, but it is not the same as a documented adverse event in a patient.
The accessible source record does not provide the exact number of isolates or drug-isolate comparisons behind each external percentage. That missing denominator limits interpretation: one percentage point can represent very different evidence in a small cohort and a large one. It also prevents a reader from independently reconstructing the uncertainty around every comparison from the public summary. The full evaluation should be read with cohort-specific sensitivity, specificity, confidence intervals and denominators side by side, not sensitivity alone.[1]
Why the African and Thai tests matter—and what they cannot prove
External evaluation matters because bacterial lineages, resistance mechanisms, sequencing practices and clinical breakpoints vary across place and time. A model can look strong when tested on data resembling its training set and then fail where the mix of strains changes. Testing separate global, African and Thai cohorts is therefore more informative than reporting only an internal split, and the African collaboration is particularly valuable in a literature that often under-represents genomic data from the continent.
Geographic labels are not substitutes for representative sampling. An 'African cohort' does not capture every country, laboratory, patient group or circulating lineage across Africa, just as one Thai cohort cannot represent South-East Asia. Public genomic repositories can over-represent well-resourced laboratories, outbreak investigations and unusual resistant isolates. They may also combine data produced under different sequencing, curation and susceptibility-testing protocols. Those composition effects can make performance look better or worse than it would in routine local practice.
The relevant transfer test is prospective and local: freeze the model, run it on consecutive eligible samples in several independent laboratories, and compare its calls with a predefined phenotypic reference under current breakpoints. Results should be reported by drug, lineage, specimen, hospital and demographic group where appropriate. New or uncertain mechanisms need an abstention or escalation path rather than a confident answer.[1]
What the result could mean for patients and laboratory teams
If prospective studies reproduce the sensitivity advantage, genomic prediction could help laboratories flag likely resistance earlier in investigations where whole-genome sequencing is already available. Fewer false-susceptible calls could reduce the chance that a stewardship team is falsely reassured by genomic evidence. The model's dashboard may also help researchers compare resistance profiles and prioritise cases for review. Those are plausible workflow benefits, not outcomes measured by this paper.
Real clinical value depends on the full pathway. A laboratory needs a suitable isolate or sample, sequencing capacity, quality control, a reproducible bioinformatics pipeline, current interpretive rules and a secure route to the clinician. Turnaround time and cost must be compared with existing culture and susceptibility testing. A result that arrives after treatment is chosen, or that cannot be explained and confirmed, may add little. A higher-sensitivity system can also generate more false-resistance calls if specificity falls, potentially steering clinicians away from useful narrower-spectrum antibiotics.
For patients, the evidence does not support changing medication on the basis of this model alone. Antibiotic decisions depend on infection site, severity, dose, allergies, kidney function, local guidance and laboratory confirmation, not just the presence of genomic features. The safe near-term role is decision support inside a governed microbiology and antimicrobial-stewardship process, with qualified staff accountable for interpretation and a route to resolve disagreement between genomic and phenotypic results.[1]
Evidence quality, interests and the next decisive study
Scientific Reports published the work as a peer-reviewed open-access article on 9 October 2026. The publisher describes the available item as accepted research that may still receive editorial changes before the final version of record. The authors declared no competing interests and said that the research received no specific funding. They acknowledged National Institutes of Health Office of Data Science Strategy support for an April 2026 codeathon while stating that it did not provide financial support for the article.
This is one research group's retrospective study. Independent teams have not yet shown that a locked version reproduces the results prospectively or that using it changes care. The next decisive evaluation should pre-register primary endpoints, publish exact cohort denominators and confidence intervals, and compare the model with current laboratory practice as well as competing software. It should measure false susceptibility, false resistance, abstentions, turnaround, staff workload, antibiotic changes and disagreement resolution.
A clinically persuasive trial would then ask whether adding the genomic model to usual care shortens time to effective treatment without increasing unnecessary broad-spectrum use. Patient outcomes could include treatment escalation, intensive-care admission, length of stay and mortality, with appropriate adjustment and safety oversight. Until that evidence exists, PredKlebAMR is a promising sensitivity-focused research tool whose strongest contribution is a better-tested hypothesis: genomic machine learning may reduce missed resistance across diverse datasets, but whether that benefit survives the bedside pathway remains unproven.[1]
What this means for people
- Fewer false-susceptible predictions could help prevent false reassurance about a resistant infection, but no patient benefit was measured.
- False-resistance predictions could also deny useful medicines or encourage broader-spectrum treatment, so sensitivity cannot be judged alone.
- Patients should expect qualified laboratory and clinical review; this study does not support autonomous prescribing or replacing indicated phenotypic tests.
Global context
Antimicrobial resistance crosses borders while genomic surveillance capacity remains uneven. This multi-country work links researchers in South Africa, Saudi Arabia, Ghana, Nigeria, Burkina Faso and the United States and tests data described as global, African and Thai. That breadth is valuable, but it does not make the model universally validated. Countries need representative local data, interoperable laboratories, sustainable sequencing capacity and governance that connects genomic evidence to stewardship rather than treating software accuracy as the whole solution.
What the evidence does not yet show
- The study is retrospective and does not test the model prospectively in a clinical laboratory or treatment pathway.
- The exact denominator for each external benchmark was not visible in the accessible source record, limiting independent interpretation of percentages and error counts.
- Only the African sensitivity comparison was reported as statistically significant; numerical advantages elsewhere should not be treated as confirmed superiority.
- Repository genomes may not represent consecutive patients, all circulating lineages or routine sequencing quality in each named region.
- Prioritising sensitivity can increase false-resistance predictions; cohort-specific specificity, calibration and predictive values remain important for stewardship.
- No study outcome covered prescribing, turnaround, cost, adverse events, length of stay or mortality, and the work has not been independently replicated prospectively.
What to watch next
- Prospective, pre-registered testing of a locked model on consecutive samples across independent laboratories.
- Exact cohort and drug-isolate denominators, confidence intervals, specificity, calibration and abstention rates.
- Performance on new lineages and resistance mechanisms that were rare or absent in the training data.
- Turnaround time, cost and safe reconciliation with phenotypic susceptibility testing and clinical guidance.
- Evidence that model-assisted reporting improves time to effective treatment without increasing unnecessary broad-spectrum antibiotic use.
Living evidence record
Impact record IAI-03QXCZ8
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can AI sort four stages of cognitive decline?
In one Egyptian dataset of 140 people, LightGBM produced the strongest average four-class result. The study used internal cross-validation, not an independent cohort or a prospective clinical test.
7 min · 1 source
Health & Life Sciences
Can interpretable AI distinguish ovarian tumours before surgery?
A peer-reviewed retrospective study trained five classifiers on laboratory data from 349 patients at one Chinese hospital. Its best model reached 93.7% leave-one-out accuracy, but the small reused cohort and absence of external or prospective validation rule out clinical deployment.
7 min · 2 sources
Health & Life Sciences
Can AI narrow the search for a phage—and what would prove a patient benefit?
A new strain-level prediction study offers a way to rank laboratory candidates. We compare its evidence stage with clinical-trial requirements and the wider antimicrobial-resistance challenge.
6 min · 3 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.