Can researchers rerun clinical AI?
A peer-reviewed scoping review found accessible analytical code in 12.2% of 3,967 prediction-model papers that cited TRIPOD or TRIPOD+AI. Even among shared repositories, dependency versions, tests and reusable structure were often missing, so code availability alone did not establish reproducibility.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The review included 3,967 open-access prediction-model articles that cited TRIPOD or TRIPOD+AI; 12.2% reported accessible analytical code.
- 2Among retrieved repositories, 80.5% had a README, 37.6% specified dependencies, 21.6% constrained dependency versions and 42.4% were modular.
- 3The LLM-assisted pipeline was checked against 500 human-annotated papers, but subjective repository features were less reliable and the sampling frame may overestimate sharing in the wider literature.
Research topic
Whether clinical prediction-model research supplies enough accessible, documented and executable analytical code for independent researchers to reproduce and scrutinise published work

The direct answer: usually not from the paper alone
Most clinical prediction-model papers in this review did not provide accessible analytical code. Among 3,967 eligible articles, 12.2% included a code-sharing statement that the review counted as accessible. The proportion rose over time and reached 15.8% for papers published in 2025, but that still left the large majority without code that another researcher could inspect or rerun.
Availability was only the first test. Of the repositories the team examined, 80.5% contained a README, 37.6% specified software dependencies, 21.6% pinned dependency versions and 42.4% were organised modularly. The paper also found infrequent licences, tests, citation metadata and reproducibility aids. A public link can therefore preserve files without supplying the environment, instructions or structure needed to reproduce an analysis.
The appropriate conclusion is not that 87.8% of models are wrong, nor that every shared repository fails. The study did not rerun every model or audit clinical validity. It establishes a reporting and scrutiny gap: in a large, guideline-aware sample, independent readers usually lacked accessible code, and many available repositories omitted practical features that make reuse credible.[1][2]
How the 3,967-paper sample was assembled
The researchers identified PubMed-indexed articles that cited the TRIPOD or TRIPOD+AI reporting statements by 11 August 2025. Eligible primary studies developed, updated or validated a multivariable prediction model using statistical or machine-learning methods. To keep full-text retrieval reproducible and free of subscription barriers, the review included only articles available through the PubMed Central Open Access API.
That produced 3,967 eligible papers. The frame is deliberately policy-relevant because these authors had acknowledged reporting guidance, but it is not representative of every prediction model. Papers behind paywalls were excluded, as were studies that did not cite the statements. The authors argue that a guidance-aware cohort may be more likely to share code, so the 12.2% estimate could be higher than the rate across the broader field.
The review followed PRISMA guidance for scoping reviews and registered a protocol. Code counted as shared when the retrieval tool successfully downloaded a repository containing at least one non-empty source-code file. Statements pointing to supplementary material or unsupported repository providers were also counted even when the tool could not independently inspect their contents, a decision intended to avoid underestimating reported availability.[2]
What the AI-assisted review did—and how it was checked
A structured pipeline used GPT-5.2 to screen full-text articles, identify code-availability statements, retrieve repository links and record the first author's country. Two independent reviewers manually annotated a random set of 500 articles, resolving disagreements through discussion. Against those labels, the automated selection achieved a weighted F1 score of 0.97 and repository-link retrieval accuracy of 92.3%.
The team then converted repositories into structured text and used a model to score 14 predefined features covering documentation, dependencies, testing, reproducibility, citation and sample data. On a manually annotated subset of 29 usable repositories, objective features such as the presence of a README, licence or tests performed strongly. More interpretive judgments were weaker: the reported F1 was 0.71 for sufficient code documentation and 0.67 for listing hardware requirements or linking to the paper.
Those validation results make the pipeline auditable, but they do not remove error. Automated misclassification could move the point estimates, particularly for subjective features. The authors therefore frame the numbers as indicators of overall practice rather than exact counts. They also published the review code and intermediate datasets, letting others inspect the method that was used to inspect everyone else's code.[2][3]
Why a repository link is not the same as reproducibility
Analytical code depends on more than source files. Another team needs to know which software and package versions were used, how data were transformed, what command starts the analysis and which outputs should appear. Random seeds, tests and a small synthetic dataset can expose hidden assumptions. A licence determines whether reuse is legally clear; an archived release protects against a living repository changing after publication.
The review found examples of good practice, including step-by-step instructions, sample or synthetic data, fixed random seeds, typed functions and persistent archives. It also found hard-coded file paths and loosely referenced external utilities during manual review. Those details explain why binary code-availability policies can be satisfied while meaningful rerunning remains difficult.
Reproducibility is not identical to clinical validity. A perfectly rerunnable model can still be biased, poorly calibrated or useless in another hospital. Conversely, restricted patient data can prevent full public execution even when researchers document the model carefully. The value of better code sharing is that it enables faster error detection, method comparison and adaptation; it does not turn a prediction model into a safe clinical product by itself.[2]
What this means for patients, researchers and journals
For patients, the effect is indirect but important. Prediction models can influence who is investigated, treated or discharged. If external teams cannot inspect preprocessing, missing-data handling and implementation details, errors may survive longer and local validation becomes harder. Code sharing is therefore part of the evidence chain, not a guarantee that an individual will receive better care.
For researchers, the paper supports planning reproducibility at the start of a project. That means separating confidential data from shareable analysis code, documenting dependencies, providing a minimal runnable example and archiving the exact release associated with the paper. Funders and institutions also need to recognise maintenance work rather than treating a repository as a last-minute publication attachment.
For journals and reporting-guideline groups, requiring a URL is too weak. Policies can ask whether the link resolves, whether dependencies are versioned, whether documentation describes inputs and outputs, and how restricted data can be replaced with synthetic examples. The authors are developing a TRIPOD-Code extension; this review is intended to supply an empirical baseline for that work rather than pre-empt its final recommendations.[1][2]
What would change the assessment
A broader assessment would sample prediction-model studies regardless of whether they cite TRIPOD, include subscription articles and stratify results by statistical versus machine-learning methods, clinical specialty, country and journal policy. Independent reviewers should repeat the repository scoring and publish uncertainty around prevalence estimates, especially for subjective documentation features.
The stronger outcome is successful execution. Future work could draw a preregistered sample of repositories, recreate environments from scratch and measure how many analyses produce the reported outputs without author assistance. It should record failure causes—missing data, unavailable packages, undocumented preprocessing or non-deterministic results—rather than collapse them into one pass-or-fail number.
Nature Medicine published the peer-reviewed version on 9 October 2026; the earlier manuscript appeared on 16 March. The authors' own public code makes the central recommendation testable. Until larger execution studies are available, the evidence supports a focused claim: accessible analytical code remains uncommon even among reporting-guideline-aware clinical prediction studies, and a public repository often needs substantially more documentation before another team can rerun it.[1][2][3]
What this means for people
- Patients benefit when model errors can be found and local performance checked before a tool influences care.
- Researchers face extra documentation work but gain a more inspectable, reusable scientific record.
- Journals and funders can shift incentives from merely posting files to supplying code that another team can actually use.
Global context
Prediction models cross borders only when data, clinical practice and implementation assumptions are made explicit. This review spans an international literature but uses PubMed, English-language reporting infrastructure and open-access retrieval. Countries with different publishing systems or limited repository infrastructure may face additional barriers. Shared minimum standards should allow secure alternatives for patient data while keeping analytical steps open to scrutiny.
What the evidence does not yet show
- The sample was restricted to open-access papers that cited TRIPOD or TRIPOD+AI and may not represent the wider prediction-model literature.
- The study measured reported code availability and repository features; it did not rerun all analyses or establish clinical validity.
- Some stated code locations could not be independently retrieved but were counted to avoid underestimating reporting.
- The LLM-assisted pipeline performed well on article selection but was less reliable for subjective repository qualities.
- Repositories were inspected at the time of analysis and may have changed since publication.
What to watch next
- Independent attempts to execute a preregistered sample of clinical prediction-model repositories.
- The final TRIPOD-Code recommendations and whether journals adopt enforceable repository checks.
- Results outside reporting-guideline-aware and open-access samples.
- Versioned environments, tests, synthetic data and archived releases becoming routine requirements.
- Evidence that better code practice reduces errors or improves external validation efficiency.
Living evidence record
Impact record IAI-01T1VH8
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
3 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Science & Research
Did the arthritis AI travel between cohorts?
A peer-reviewed analysis of gut-microbiome data from 2,238 people reached a mean internal ROC-AUC of 0.834, but its genus-level model fell to 0.439 in 39 independently processed Shanghai samples. The study is a warning about cross-cohort transfer, not evidence for a rheumatoid-arthritis diagnostic test.
8 min · 1 source
Science & Research
Can AI reveal hidden disease patterns in tissue?
A peer-reviewed method used ensembles of neural networks and statistical testing to find disease-associated spatial structures across rheumatoid arthritis, ulcerative colitis and dementia datasets. It is a research-discovery tool, not a diagnostic test or evidence of improved patient care.
8 min · 2 sources
Science & Research
Can an automated map keep pace with AI-oncology research?
A peer-reviewed study used transparent rules to classify 20,766 open-access cancer-AI papers from 2019 to 2025. Audits show useful article-level accuracy, but the map is not a clinical-readiness score, a complete field census or a comparison of which algorithms work best.
9 min · 2 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.