Back to the news portal
Climate & EnergyResearch paperResearchSource analysisRepublic of KoreaGlobal food systems

Can AI shortlist sustainable protein sources?

A peer-reviewed Korean study compared 1,345,631 protein records across 50 grain, seaweed, legume and mushroom sources. AlterProtX can rank molecular similarity to animal proteins, but it has not shown that a shortlisted ingredient will taste, process or perform the same in food.

By The Impact of AI Editorial DeskReleased 8 October 2026 at 14:32 BST7 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1AlterProtX analysed 1,345,631 protein records from 50 non-animal sources: 15 grains, 14 seaweeds, 11 legumes and 10 mushrooms.
  • 2Its thermal-stability component was developed on a separate Meltome Atlas set of 24,572 proteins and combined sequence, structural and physicochemical features.
  • 3The framework compares whole-proteome distributions with five animal-derived reference categories and then adds predicted allergenicity and essential-amino-acid balance.
Key themesAlternative proteinsFood systemsProteomicsProtein language modelsAllergenicitySustainable agriculture

Research topic

Proteome-scale computational ranking of non-animal protein sources by molecular similarity, allergenic potential and nutritional adequacy

The Impact of AI research cover showing conceptual grains, legumes, seaweed and mushrooms flowing through an abstract comparison system toward a shortlist, with the evidence limit that no ingredient performance was tested.
AI-generated editorial illustration. The protein chains, comparison system and shortlist are conceptual and do not depict a measured food product, validated ingredient performance or commercial recommendation.

The direct answer: AI can narrow the search, but cannot yet choose a food ingredient

AlterProtX shows that a computational system can compare very large collections of proteins and produce an interpretable shortlist of non-animal sources whose molecular profiles resemble animal-derived food proteins. The framework moves beyond judging one protein at a time. It represents the full distribution of predicted thermal stability and six interaction-related attributes across each source, then compares those distributions with animal-protein reference groups. For researchers facing millions of possible sequences, that is a useful way to decide which candidates deserve laboratory time first.

The study does not show that any candidate can replace egg, milk, meat or fish in a finished food. Food functionality emerges after proteins are extracted, concentrated, heated, mixed with fats and carbohydrates, and processed under specific pH and salt conditions. The paper contains no ingredient-scale gelation, foaming or emulsification experiment, no sensory panel, no digestibility study and no life-cycle assessment. The defensible finding is therefore about research prioritisation, not a successful substitute or an environmental saving.[1][2]

The denominator is 1,345,631 proteins across 50 candidate sources

The main comparison covers 50 non-animal food-protein sources selected using criteria linked to the Korean Ministry of Food and Drug Safety: 15 grains, 14 seaweeds, 11 legumes and 10 mushrooms. The authors retrieved complete proteomes and associated three-dimensional structures from UniProt and AlphaFold, removed entries with missing annotation and excluded sequences shorter than 10 or longer than 1,000 amino acids. The resulting source-level dataset contained 1,345,631 proteins. Five animal-derived categories—egg white, egg yolk, milk, meat and fish—served as references.

That denominator is large in protein records but small in source diversity. Fifty candidate organisms cannot represent the full range of cultivars, growing conditions, harvest stages, fermentation states or industrial preparations used in food. A proteome is also not the same thing as an ingredient. Some proteins are barely present in an edible fraction, while storage proteins may dominate what manufacturers actually extract. The paper's equal-contribution assumption gives every protein a role in the source profile because standardised abundance data were unavailable.[1]

A separate 24,572-protein dataset trained the thermal-stability model

Thermal denaturation matters because heating changes how proteins unfold, aggregate and interact during processing. The researchers used experimentally measured denaturation temperatures from the Meltome Atlas, excluding incomplete entries and taking a median when one protein had multiple reported values. The final set contained 24,572 proteins from humans, other animals, plants, fungi and bacteria or archaea. An 80:20 stratified split preserved the denaturation-temperature distribution across training and test sets.

The selected graph-convolutional model combined complete primary amino-acid sequences with detailed atomic three-dimensional structural context and physicochemical descriptors. Those descriptors included hydrophobicity, van der Waals volume, charge, secondary structure, solvent accessibility and polarizability. The authors compared sequence and structure representations before selecting the best-performing combination. This is a prediction of one processing-relevant property, not a direct measurement for all 1.35 million source proteins, so errors in predicted structures and transfer across species propagate into the later rankings.[1]

Whole distributions matter more than one average protein

For each source, AlterProtX converts the predicted thermal values and six interaction attributes into kernel-density distributions. It then measures similarity between the shape and distance of source distributions using Wasserstein distance rather than comparing only their means. This matters because a food source contains a heterogeneous mixture: two sources can share the same average value while one has a narrow cluster and the other contains several very different protein populations.

The framework selects the top 10 candidates for each animal-derived reference and highlights which molecular attributes drive a match or mismatch. That interpretability is more useful than a single opaque score because a developer can see whether thermal response, charge or another interaction feature is doing the work. Even so, a mathematically close distribution is not a validated functional equivalent. The authors explicitly describe the scores as relative source-level priorities within the analysed search space, not a universal threshold for substitution.[1]

Allergenicity and nutrition change what a promising match means

AlterProtX adds predicted allergenic potential and essential-amino-acid balance after the proteome-similarity comparison. That multi-criteria step is important: a source can look attractive for processing but be unsuitable because of a well-established food allergy or a limiting amino acid. The paper uses quinoa as an example of a candidate combining high molecular similarity with a balanced essential-amino-acid profile and relatively limited clinical allergy reports, while wheat illustrates how functional similarity can coexist with established allergenicity and lysine limitation.

Those screens are not clinical clearance. Absence of a documented allergen does not establish safety for an unfamiliar protein source, and sequence-based predictions can miss how processing creates, destroys or exposes epitopes. Nutritional adequacy also depends on digestibility, bioavailability and how an ingredient is eaten, not amino-acid composition alone. Comprehensive allergen testing, protein digestibility measurement, exposure assessment and regulatory review would be required before a candidate entered the food supply.[1]

What this could change for people—and what evidence comes next

A reliable early-screening tool could reduce the number of expensive extraction and formulation experiments needed to find useful crop, fungal or seaweed proteins. That could broaden the raw materials available to food producers and help researchers investigate sources that are currently overlooked. The public interest lies in affordable, acceptable foods with lower land, water or emissions burdens—not in producing a sophisticated ranking that cannot survive a factory process or reach consumers at a viable price.

The next decisive evidence is a prospective benchmark. Researchers should lock the ranking, select candidates across high and low predicted similarity, manufacture comparable protein ingredients and measure denaturation, solubility, gelation, foaming and emulsification under blinded protocols. Results should be compared with simple baselines and established ingredients, followed by digestibility, allergenicity, sensory and life-cycle studies. Independent replication across cultivars and production methods would show whether AlterProtX is discovering transferable food properties or mainly organising public databases.[1]

What this means for people

  • Better shortlisting could make alternative-protein research faster and less expensive, widening the range of affordable ingredients considered.
  • Consumers gain only if candidates remain nutritious, safe, acceptable and functional after industrial processing.
  • People with allergies could face harm if a computational screen were mistaken for clinical safety evidence.

Global context

The study was developed at Kookmin University in Seoul with Korean public funding, but the protein databases and food-system question are global. Regions differ in staple crops, seaweed and mushroom use, allergy surveillance, processing capacity and dietary needs. Local validation should therefore begin with locally consumed sources and manufacturing conditions rather than assuming one global ranking.

What the evidence does not yet show

  • The 1,345,631 records are predicted source proteomes, not measured protein abundance in edible or industrial ingredients.
  • The 50 non-animal sources cover four broad groups but not cultivar, environment, harvest or processing variation.
  • The study reports no ingredient-scale functional, sensory, digestibility, safety, cost or life-cycle experiment.
  • Predicted structures, denaturation temperatures and allergenicity introduce uncertainty before candidates are ranked.
  • The peer-reviewed Article in Press is citable but may receive copy-editing before the final Version of Record.

What to watch next

  • A locked prospective test of high- and low-ranked candidates in real ingredient formulations.
  • Measured protein abundance and processing conditions rather than equal contribution from every proteome entry.
  • Independent allergenicity, digestibility and nutritional validation.
  • Life-cycle and cost comparisons showing whether a functional substitute is also environmentally and economically preferable.

Living evidence record

Impact record IAI-0KNPVDH

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

8 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Evidence trail

Sources used for this report

Links checked 8 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Climate & Energy

Can machine learning stabilise renewable microgrids in real time?

A peer-reviewed simulation study generated 3,000 solar, wind, temperature and load scenarios, then trained ensemble models to imitate an optimiser's controller settings. Gradient boosting reproduced those settings closely, but no physical microgrid or live disturbance was tested.

6 min · 2 sources

Climate & Energy

Can AI make local rainfall projections more useful?

A new peer-reviewed dataset uses deep learning to translate four global climate models into daily 0.1° precipitation fields for China’s Loess Plateau from 1950 to 2100. It adds spatial detail for research and planning, but cannot remove climate-model uncertainty or verify rainfall decades ahead.

7 min · 1 source

Climate & Energy

Can neurosymbolic AI make electricity grids safer?

A simulation study combined learned grid recommendations with hard topology checks and fuzzy control, cutting losses and outage energy on benchmark feeders—but no live network, operator or physical controller was tested.

7 min · 1 source

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.