Can AI choose the next drug-chemistry experiment?
A peer-reviewed Roche–ETH study used active learning to select 50 new substrates from a 22,253-molecule pool, then tested about 96 reaction conditions for each. A later graph model identified the observed reaction site for all ten prospectively tested substrates—but only within a narrow chemistry workflow, not drug discovery as a whole.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The active-learning system selected 50 previously unreported substrates from a Roche pool of 22,253 drug-like aromatic compounds across three prospective rounds; laboratory teams generated 4,821 new reactions.
- 2The resulting open C–H borylation resource contains 6,865 reaction outcomes across 568 substrates and atom-level reaction-site annotations covering 812 substrates.
- 3A graph model's top-ranked reaction site matched the observed result for all ten prospectively tested substrates, but ten examples are too few to establish reliability across other reactions, laboratories or drug programmes.
Research topic
Active-learning selection of laboratory experiments and geometric graph prediction of C–H borylation outcomes and reaction sites
The direct answer: it chose informative experiments in one chemistry domain
Yes, within a tightly defined medicinal-chemistry workflow. Researchers at Roche and ETH Zurich used an uncertainty-guided model to select which molecules to test next for C–H borylation, a reaction that gives chemists a useful handle for making related molecular variants. Across three prospective rounds, the system chose 50 previously unreported substrates from a pool of 22,253 drug-like aromatic compounds. Laboratory teams then tested roughly 96 reaction configurations for almost every selected substrate and added 4,821 measured reactions to the dataset.
A second stage used geometric graph neural networks to predict whether reactions would work and where on a molecule the borylation would occur. In a final prospective check on ten unseen substrates, the model's highest-ranked site matched the experimentally observed site in all ten. That is a concrete model-to-lab result. It is not evidence that an AI designed a medicine, selected a biological target, proved safety or shortened development time from discovery to patient treatment.[1]
The Impact Brief · Free
Follow the evidence in science & research.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
What the closed loop actually did
The study separates two jobs that are often blurred in claims about autonomous science. First, a relatively lean ensemble of XGBoost decision-tree models estimated whether candidate substrates were likely to react and where the model was uncertain. The team compared ten ways of choosing the next examples and selected Bayesian Active Learning by Disagreement, or BALD, because retrospective tests suggested it would explore uncertain and structurally diverse regions rather than repeatedly sampling familiar chemistry.
Second, human chemists ran the chosen experiments, characterised the products and returned the results to the dataset. The loop was repeated three times: 30 substrates in round one, then ten in each of rounds two and three. This was model-guided experimental design, not a laboratory operating without people. Trained personnel handled air- and moisture-sensitive reagents, heated reactions and performed scale-up and structural characterisation before labels could be trusted.
The final public resource contains 6,865 high-throughput reaction outcomes across 568 substrates for yield and binary success prediction. It also contains atom-level regioselectivity information for 812 substrates, including 106 negative examples, and 920 borylated species. The paper and repository release the data, while a separate Zenodo record provides code and model weights. That openness gives independent groups a route to check the reported splits and build comparisons rather than accepting a company-only benchmark.[1][2]
The experiment broadened the chemistry, but barely touched the full pool
The chosen molecules were not merely close copies of the starting set. Across the three rounds, the nearest-neighbour molecular-distance measure averaged about 0.69, and 42 previously unseen structural scaffolds were added. The number of represented scaffolds rose from 105 to 147, a 40% increase. In the first active-learning round, selections covered all six clusters used to divide the candidate chemical space.
The denominator tempers that result. The 22,253 candidates contained 6,547 unique scaffolds. Coverage increased from 0.87% to 1.34%, so almost all possible structural families remained untested. Only 24% of the new reactions met the study's positive threshold of at least 5% yield. Exploring failure is valuable because negative results teach a classifier where reactions do not work, but it also shows why this is a data-acquisition study rather than a claim of universally accurate synthesis planning.[1]
Graph models generalised better, not perfectly
After collecting more experiments, the researchers compared ten geometric graph neural networks with XGBoost baselines. Random splits ask a relatively easy question because training and test molecules can be structurally similar. Cluster and scaffold splits are harder because they hold out broader molecular families. On the final scaffold split for binary reaction success, the graph models produced mean Matthews correlation coefficients between 0.43 and 0.50, while the best condition-aware XGBoost model reached 0.35 with a standard deviation of 0.12.
Those values represent an improvement, not near-perfect classification. For reaction-site prediction, the best condition-aware graph model also outperformed the T5Chem language-model comparator on harder splits: one Ponita configuration reached an MCC of 0.59 with a 0.05 standard deviation on the scaffold split, compared with 0.38 with a 0.08 deviation for T5Chem. Adding self-supervised training tasks—masking atom identities and reconstructing perturbed coordinates—usually helped, especially on unfamiliar scaffolds.
No single graph architecture consistently won, and the paper found no clear general advantage for equivariant over invariant models. The authors' interpretation is important for research teams deciding where to spend effort: better experimental data and training strategy may matter more than choosing the most elaborate neural architecture once a baseline level of molecular geometry is represented.[1]
Ten successful prospective checks are promising but narrow
For the most visible test, the team retrained an EquiformerV2 model on the complete reaction-site dataset and asked it to rank possible sites on ten previously unreported substrates. Five had one dominant predicted site; five had several comparable sites. In every case, the highest-ranked site corresponded to the experimental result. For one molecule, the predicted aromatic position was confirmed in a product isolated at 44% yield. For others, multiple observed products were consistent with sites receiving similar scores.
The probabilities were not calibrated as reaction yields, and the sample of ten is much too small to estimate a dependable failure rate. All examples concerned variants of the same broad borylation problem, were tested inside the collaborating Roche workflow and were chosen after the earlier data-building campaign. The paper did repeat the modelling comparison on a distinct Minisci C–H alkylation dataset containing 13,490 reactions, 80 heterocyclic substrates and 59 radical sources; graph models again did better on the scaffold split. That supports a transferable modelling pattern, but it is still retrospective for the second reaction class.[1]
What this changes for scientists and drug programmes
For medicinal chemists, the practical gain is prioritisation. A system that points to informative experiments and plausible reaction sites can help a team spend scarce compounds, laboratory time and analytical capacity on tests that expand knowledge rather than repeating familiar examples. It can also support virtual enumeration of molecular variants after a viable reaction site is identified. The human role shifts toward defining the chemistry, checking uncertainty, executing safely, interpreting failures and deciding which molecular changes serve the biological programme.
For patients and health systems, the effect is remote and unproven. Faster analogue generation could eventually help discovery teams find useful candidates more efficiently, but this study measures neither drug quality nor time or money saved across a development programme. No molecule entered animal testing or a clinical trial because of the system. Claims about cheaper medicines, faster approvals or better outcomes would therefore outrun the evidence.[1][2]
Funding, interests and what would change the assessment
The Roche Access to Distinguished Scientists programme funded the research, and open-access publication support came from ETH Zurich. Ten named authors are Roche employees. One ETH author discloses co-founding two companies and consulting for the pharmaceutical industry. These interests do not invalidate the prospective experiments, but they increase the value of independent reproduction on another company's or an academic laboratory's compounds, equipment and reaction protocols.
Confidence would rise with a preregistered external test covering hundreds rather than ten unseen substrates, fixed decision rules, calibration of reaction-site probabilities, and direct measurement of failed suggestions, experiment count, material use, scientist time and cost. The more consequential test is whether the workflow improves a real drug programme against expert-selected experiments while maintaining safety and documentation. Until then, the strongest conclusion is that active learning helped build a broader reaction dataset and that geometry-aware models can use it to guide a narrow, experimentally verified synthesis task.[1][2]
What this means for people
- Medicinal chemists could use the system to prioritise scarce experiments, while remaining responsible for safe execution and scientific judgment.
- Research managers should measure saved experiments and laboratory time rather than equating model accuracy with productivity.
- Patients should not expect an immediate treatment benefit: the study did not produce or test a clinical drug candidate.
Global context
The study was conducted by ETH Zurich and Roche teams in Basel using a proprietary candidate pool, while releasing the resulting datasets, code and model weights through public repositories. Laboratories differ in compound collections, equipment, reaction protocols and analytical standards. Independent replication across institutions, including public-sector and lower-resource discovery settings, is needed before treating the workflow as a broadly transferable route to more efficient drug research.
What the evidence does not yet show
- The prospective reaction-site validation contains only ten substrates and one broad reaction family.
- The work was performed in a Roche–ETH workflow; independent external laboratory validation is absent.
- Only 1.34% of the candidate pool's unique scaffolds were represented after active learning.
- Reaction-site probabilities were not calibrated as yields, and model performance on hard scaffold splits remained imperfect.
- The study tests chemistry prediction, not a complete drug-discovery programme, candidate safety, clinical efficacy or patient benefit.
- Roche funded the work and employs ten authors; an ETH author reports pharmaceutical-industry interests.
What to watch next
- Independent prospective reproduction in another laboratory and chemical collection.
- Larger blinded tests on unseen scaffolds with fixed thresholds and calibrated probabilities.
- Head-to-head comparison with expert-selected experiments on material use, time, cost and failed reactions.
- Prospective validation across additional reaction families rather than retrospective transfer alone.
- Evidence that model-guided chemistry improves an actual discovery programme without weakening safety or documentation.
Living evidence record
Impact record IAI-1W129BJ
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
2 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Science & Research
Can AI-controlled droplets make colour tests more reliable?
In laboratory tests, a digital microfluidic platform moved nanolitre droplets, watched colour reactions and adapted its measurement strategy. The peer-reviewed system is technically promising, but it was not clinically validated.
9 min · 1 source
Science & Research
Can self-supervised AI help neutrino detectors learn from fewer labelled events?
A Nature Machine Intelligence paper reports strong low-label performance across simulated detector tasks. The million-event dataset is substantial, but it models a proposed detector and cannot substitute for validation on real collisions.
5 min · 2 sources
Science & Research
Robin links literature agents and laboratory data in a closed discovery loop
Researchers describe Robin, a multi-agent system that generates hypotheses, proposes experiments, analyses results and revises its ideas, including work on candidate therapies for dry age-related macular degeneration.
2 min · 1 source
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.