Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisUnited KingdomCanadaInternational

Can we trust AI gene-perturbation models yet?

A peer-reviewed benchmark across 14 single-cell datasets and 18 metrics finds that common scores can hide useful gene-perturbation signals. Several deep-learning models beat uninformative baselines after calibration, but the study predicts gene-expression responses—not medicines, safety or patient benefit.

By The Impact of AI Health & Life Sciences DeskReleased 4 October 2026 at 11:04 BST9 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesDrug discoverySingle-cell biologyGene perturbationBenchmarkingModel evaluationReproducibility

Research topic

Whether standard benchmark metrics underestimate useful signal in deep-learning models that predict gene-expression responses to genetic perturbations

The Impact of AI research cover asking whether AI gene-perturbation models can be trusted, with a conceptual cell-and-gene signal passing through two evaluation lenses.
AI-generated editorial illustration. The cell, gene signals and evaluation lenses are conceptual; they do not depict experimental data, a medicine, a patient result or a validated screening system.

At a glance

  • 1The researchers tested positive and negative controls across 14 public single-cell perturbation datasets and 18 metrics, finding that frequently used error and correlation scores were often poorly calibrated.
  • 2Under better-calibrated weighted, retrieval and differential-expression-focused metrics, models including scGPT, GEARS, PRESAGE, scLambda and CellFlow often beat an uninformative mean baseline.
  • 3The result repairs part of the evaluation problem; it does not show that a model can discover an effective drug, predict an unseen patient context, replace wet-lab experiments or improve health outcomes.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-0MI19KN

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

4 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Why a benchmark can make a useful model look useless

AI systems for genetic perturbation modelling try to predict how a cell's gene-expression profile will change after one or more genes are inhibited or activated. If those predictions were reliable, researchers could use them to prioritise laboratory experiments, explore biological pathways and narrow a large search space before spending time and money on wet-lab screening. Recent benchmark papers, however, produced a sobering result: sophisticated deep-learning systems often failed to beat a simple average of the training data.

The new Nature Biotechnology study asks whether some of that apparent failure belongs to the models and some to the ruler used to measure them. A mean prediction can look accurate when most genes barely change, because error scores average a small biological signal across thousands of mostly stable genes. Correlation scores can also be inflated or distorted when both predictions and observations are measured relative to a common control. A metric may therefore reward an uninformative answer or miss a narrow but real perturbation response.

That is not a technical footnote. Benchmark scores steer research funding, model selection and claims about whether virtual-cell systems are ready to support drug discovery. If the metric cannot distinguish a positive control from a negative one, a leaderboard position should not decide which approach is adopted.[1]

The study builds a positive control for model evaluation

The authors assembled 14 public single-cell perturbation datasets and evaluated 18 metrics. Their negative control was the familiar mean baseline: predict the average perturbed expression profile from the training set. For a positive control, they split cells exposed to the same perturbation into two groups and used one half to predict the other. Because both halves contain the same perturbation-specific biology, a useful metric should usually score this technical duplicate above the uninformative mean.

A plain duplicate can still be noisy when a perturbation affects only a few genes. The team therefore introduced an interpolated duplicate that blends the duplicate and mean predictions according to statistical evidence that each gene changed. They then defined a dynamic range fraction: where a model's score sits between the negative and positive controls. This calibration test is valuable because it first asks whether the metric has enough sensitivity before using that metric to judge a model.

The analysis covers direct error, control-referenced correlation, weighted, differentially expressed gene and retrieval-based metric families. The authors released code and relied on datasets already deposited in Zenodo and Figshare. That makes the computational analysis reproducible in principle, although independent teams still need to rerun it and test alternative control definitions.[1][2]

What changed when the metrics were calibrated

Common measures such as mean squared error and Pearson correlation on control-referenced changes were frequently poorly calibrated: in some datasets they did not reliably separate the positive control from the mean baseline. Metrics that concentrated weight on perturbation-sensitive genes or asked whether the correct perturbation could be retrieved from the prediction space produced a clearer usable range.

Under those better-calibrated measures, several deep-learning models recovered perturbation-specific signal that earlier scorecards had obscured. The benchmark included earlier systems such as scGPT and GEARS and newer architectures including PRESAGE, scLambda and CellFlow. Performance was not uniform. A model could lead on weighted error and differential-expression correlation yet remain only marginally above the mean baseline on another view of the prediction space. The authors explicitly caution against declaring any one metric definitive.

The downstream checks add practical meaning. Rankings on calibrated metrics correlated with recovery of known biological pathways, with average rank correlations of 0.76 and 0.84 in two highlighted datasets. Neighbourhood-structure recovery showed a similar but more variable relationship. These are retrospective consistency tests, not prospective discoveries, but they suggest the calibrated scores capture more than numerical convenience.[1]

Dataset coverage still sets a hard ceiling

The paper shows why the training landscape matters as much as the architecture. In the widely used Norman19 combination dataset, the training split covered only 31 of 4,950 possible gene pairs—0.63%—and about 96% of measured combination effects were additive. A baseline that simply sums single-gene effects therefore captured roughly 88% of the available performance gap on the best-calibrated metrics. Deep models had little non-additive structure to learn and little coverage from which to learn it.

In Wessels23, training covered 10.3% of possible combinations and perturbations were stronger. There, multiple systems surpassed the additive baseline, and PRESAGE exceeded it on 15 of 18 metrics. That comparison supports the narrower conclusion that deep learning can outperform a strong baseline when the data contain enough informative combinations. It does not establish a universal advantage across cell types, genes, doses or experimental platforms.

Two other datasets contained a median of one or zero differentially expressed genes per perturbation, making the unseen-perturbation task especially difficult. A low score in that setting can reflect weak experimental signal, metric insensitivity, genuine model failure or all three. The calibrated framework helps separate those causes but cannot create biological information that was never measured.[1]

What this means for drug discovery

For laboratory teams, the immediate implication is procedural: validate the benchmark before interpreting the model. Include explicit positive and negative controls, report several complementary metric families, show uncertainty across perturbations and test whether better scores recover pathways or neighbourhoods that matter for the intended scientific decision. A single average error number is not enough to justify replacing one model with another.

For biotechnology companies and funders, the finding is encouraging but modest. It restores evidence that some models learn perturbation-specific signals and that architectures using biological prior knowledge can perform well. It does not prove that a predicted expression shift identifies a safe target, that the result transfers to another donor or tissue, or that a model can anticipate toxicity, delivery, dose response or clinical efficacy. Each of those steps needs separate evidence.

People should therefore be wary of translating 'outperforms a baseline' into 'discovers drugs'. These systems can help rank experiments. Human scientists still choose the biological question, inspect failure modes, test interventions in cells and organisms, and decide whether evidence supports development. The benefit is a potentially better search tool, not an automated medicine pipeline.[1]

Commercial interests and remaining uncertainty

The paper was peer reviewed and reports no external funding. Five authors are employees of Shift Bioscience, a biotechnology company, and another is Chief AI Scientist at Xaira Therapeutics. Those affiliations create a clear commercial interest in whether perturbation models appear useful. The journal publishes the competing-interest statement and reviewer information, while the open code enables scrutiny. Neither peer review nor code availability substitutes for independent replication.

The authors leave the unseen-context task unresolved: predicting responses in cell types, donors or conditions absent from training. They also compare only a limited number of combinatorial datasets, and the positive control itself depends on statistical decisions about which genes changed. Dataset processing, cell counts, sequencing depth and highly variable gene selection can alter calibration. The paper studies public experimental data rather than a prospective decision made in a drug programme.

Confidence would rise if independent groups reproduced the results, pre-registered metric choices and showed that calibrated model rankings predict genuinely new wet-lab outcomes. Multicentre datasets with unseen donors and tissues, richer combination coverage, dose information and prospective target-selection tests would reveal whether the signal survives the leap from benchmark to biological discovery. Until then, this is an important correction to evaluation practice—not validation of AI-designed medicines.[1]

What this means for people

  • Researchers gain a more defensible way to decide whether a model captures real perturbation-specific biology.
  • Biotechnology teams may avoid discarding useful models—or overbuying weak ones—because of a poorly calibrated leaderboard.
  • Patients should not expect near-term treatment changes: this work evaluates research tools, not clinical outcomes or approved therapies.

Global context

The datasets and modelling community are international, while the authors are based mainly in the United Kingdom and Canada. Virtual-cell and perturbation-prediction programmes are expanding across academic and commercial laboratories worldwide. Shared public datasets and code make cross-border replication possible, but the biological representativeness of donors, tissues and experimental conditions remains a global equity issue: a metric can be calibrated within one dataset without proving that the model works across populations or laboratories.

What the evidence does not yet show

  • The analysis is retrospective and uses public single-cell perturbation datasets; it does not prospectively discover or validate a medicine.
  • The study does not address unseen contexts such as new cell types, donors or biological conditions.
  • Performance varies by dataset and metric, and no single calibrated measure captures every useful or unsafe behaviour.
  • Combination coverage ranges from 0.63% to 10.3% in highlighted datasets, limiting conclusions about non-additive genetic interactions.
  • Most authors work for Shift Bioscience and one works for Xaira Therapeutics; independent replication remains important despite peer review and open code.

What to watch next

  • Independent reproductions of the calibration framework and alternative definitions of positive and negative controls.
  • Prospective wet-lab tests in which model-ranked perturbations are evaluated before outcomes are known.
  • Benchmarks covering unseen donors, tissues, disease states, doses and substantially denser gene-combination spaces.
  • Reporting that connects benchmark gains to experimental yield, failed experiments, cost, safety and downstream drug-development decisions.

Evidence trail

Sources used for this report

Links checked 4 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can AI make tumour mitosis counts more consistent?

A preprint study paired 13 pathologists' unaided and AI-assisted reviews of 385 tumour slides from three European centres. Agreement rose and counting time fell, but the unreviewed study did not establish which counts were correct, whether diagnoses improved or whether patients benefited.

11 min · 1 source

Health & Life Sciences

CSL plans AI and cloud tools for drug research with AWS

The Australian biotechnology company says the collaboration will support target discovery and clinical-development paperwork. No faster trial, approved treatment or patient benefit has yet been measured.

4 min · 1 source

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.