Can a general-purpose LLM search crystal compositions?
In a closed computational benchmark, GPT-5.4 recovered 95.65% of 3,740 low-energy Elpasolite targets within 5,000 proposals. Iterative evaluator feedback drove the result; no new material was synthesised, and the approximate energy labels are not proof of stability or usefulness.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1Across three runs of 5,000 proposals, GPT-5.4 recovered 3,577 plus or minus 9 of the 3,740 benchmark targets, or 95.65% plus or minus 0.24%.
- 2Each proposed composition was scored and returned to the model, creating an active-learning-like feedback loop. The purpose-trained GAN, VAE and reinforcement-learning comparators were one-shot generators, so the architecture comparison is not like-for-like.
- 3The experiment searched compositions within fixed crystal prototypes using approximate precomputed energy labels. It did not predict full structures, synthesise a compound or establish stability, safety or practical performance.
Research topic
Whether a general-purpose language model can efficiently recover a known low-energy region in constrained inorganic-crystal composition spaces using iterative evaluator feedback
The direct answer: yes in this closed search task
A general-purpose language model searched this constrained computational benchmark unusually efficiently. Across three independent runs of 5,000 proposed compositions, GPT-5.4 recovered an average of 3,577 of the 3,740 Elpasolite compositions labelled below the target formation-energy threshold. That is 95.65%, with a spread of plus or minus 0.24 percentage points. The first repeated proposal occurred after 297 plus or minus 178 attempts, suggesting that the model initially explored broadly rather than immediately cycling through the same candidates.
This is a result about navigating a known mathematical search space, not discovering a usable material. The candidates were compositions conforming to a fixed ABC2D6 prototype and were evaluated against a pre-tabulated energy landscape. The researchers did not generate complete atomic structures, run an experimental synthesis, measure a physical property or establish that a candidate is stable under real conditions. 'Recovered a target label' is therefore the accurate claim; 'invented 3,577 materials' would be misleading.[1][2]
The Impact Brief · Free
Follow the evidence in science & research.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
The benchmark contained a known answer set
The Elpasolite benchmark covers almost two million possible compositions formed by assigning chemical elements to four crystallographic sites. Formation energies were available from a kernel-ridge-regression model trained on roughly 10,000 density-functional-theory calculations. The target was the 3,740 compositions with predicted formation energy below minus 2.26 electronvolts per atom—less than 0.2% of the full space. Because the entire landscape was known, the team could count recall, diversity and duplication rather than judging only a few apparently good outputs.
That completeness is a strength for benchmarking but also a boundary. The score is only as meaningful as the reference evaluator. The paper notes chemically suspicious features: 262 target compositions contain noble gases, and 688 composition pairs created by swapping two sites straddle the threshold. Those cases expose uncertainty in a learned energy model and the underlying representation. A target classification may be useful for comparing search behaviour without being reliable evidence that a chemist should attempt synthesis.[1]
Feedback, not a single prompt, was the central mechanism
At every step, the model proposed one composition in the required formula. A validator checked compositional consistency, retrieved the precomputed formation energy and returned that value along with the growing history of previous proposals. The model then used the enlarged context to make its next choice. No task-specific weights were trained. In effect, the prompt became an active-learning notebook in which the model could infer which regions of the periodic table led toward the target.
The feedback is crucial. When the model generated a large batch without frequent scores, performance deteriorated. Batches of 10 to 50 retained much of the sequential mode's performance with fewer model calls, while larger batches degraded progressively. Medium reasoning effort recovered about 96% within 5,000 proposals; low and minimal settings recovered about 88%, and removing reasoning produced a sharper fall. More reasoning also increased cost because output tokens are more expensive than input tokens, so recovery alone is not a complete efficiency measure.[1]
The comparison with specialist models is informative but asymmetric
Earlier task-specific generative adversarial, variational-autoencoder and reinforcement-learning models recovered about 40% to 46% of the same 3,740 targets in their first 5,000 proposals, rising to 75% to 94% after 250,000. GPT-5.4's 95.65% within 5,000 is therefore a large benchmark advantage, achieved without task-specific fine-tuning. The authors also report performance nearly four times that of a Bayesian-optimisation comparison in their supplementary analysis.
But the language model saw an evaluator score after every proposal, while the older generative baselines operated in a one-shot manner. The authors explicitly call this an architectural advantage and say it must be considered in comparison. A fair operational contest would include strong iterative optimisers, equal evaluator access, matched compute and monetary cost, and multiple modern foundation models. The experiment demonstrates that an LLM can implement a flexible feedback-driven strategy; it does not establish that language modelling is intrinsically superior to every materials-search algorithm.[1]
Tests against memorisation strengthen—but do not settle—the case
One concern is that a general model may have encountered Elpasolite data during pre-training. The researchers tested an abstracted version in which element identities were anonymised and the target unit removed; the model still adapted through feedback. They also applied the workflow to a structurally different AB2X4 spinel dataset whose target property values were released after the model's training cutoff. Those tests make simple recall of the complete target landscape less plausible.
They do not eliminate every pre-training advantage. Chemical regularities and token representations learned from public literature can still help a model reason about element combinations. Indeed, a realistic starting composition improved early performance compared with an anonymous formula. That is not misconduct—it is part of what a general model offers—but it complicates attribution. Repetition with independently held-out simulators, unpublished landscapes and controlled starting information would clarify how much comes from prior chemical knowledge versus feedback adaptation.[1]
What this could change for materials scientists
The promising use is as a configurable proposal engine around a trusted scientific evaluator. A researcher could specify a formula family, return computed scores, inspect diversity and redirect the search without building a new neural generator for each property. Code and outputs are publicly linked, which should support independent reproduction. For small teams, that flexibility may reduce setup time and make exploratory inverse-design workflows easier to prototype.
Human scientific control remains essential. Experts must define physically meaningful constraints, inspect why the evaluator favours a candidate, check charge balance and structural feasibility, and decide whether higher-fidelity simulation or laboratory work is justified. Safety review cannot be delegated to a model optimising a number. A system may exploit defects in the scorer just as readily as it learns chemistry. The practical gain should be measured as validated candidates per unit of compute, researcher time and experimental cost—not as the count of strings below an approximate threshold.[1][2]
Funding, disclosure and what would change the assessment
The authors are based at the Fritz Haber Institute of the Max Planck Society. They report using Max Planck Computing and Data Facility resources; open-access funding was organised through Projekt DEAL. The paper declares no competing interests. It provides a Zenodo dataset with model outputs and a public GitLab workflow, important foundations for checking sensitivity to prompts, model versions and evaluator settings.
The assessment would strengthen if independent groups reproduce the result with held-out property landscapes, modern iterative baselines and full accounting of tokens, compute and failed proposals. The scientific bar is higher still: selected candidates should survive higher-fidelity calculations, uncertainty analysis and prospective laboratory synthesis, then deliver a relevant property under realistic conditions. Until that chain is demonstrated, the paper supports a precise conclusion—general-purpose LLMs can be strong adaptive search components in constrained computational composition spaces—not autonomous materials discovery.[1][2]
What this means for people
- Materials researchers may gain a flexible search assistant, but remain responsible for physical constraints, validation and laboratory safety.
- Funders and laboratories should not treat benchmark hits as discovered products or use them to bypass staged scientific verification.
- Open outputs and code make it easier for independent teams to test whether the apparent efficiency survives different models and evaluators.
Global context
The study was conducted at Germany's Max Planck Society using a commercial general-purpose model and established public benchmarks. Access to high-end models, compute, scientific evaluators and wet-lab validation is uneven internationally. Open data and code help, but reproducibility also depends on continued access to the same model behaviour and transparent cost reporting.
What the evidence does not yet show
- The main task used a known, pre-tabulated composition space and an approximate learned formation-energy evaluator rather than prospective experiments.
- The model proposed compositions within fixed prototypes, not complete crystal structures or synthesis routes.
- No candidate was synthesised, and a low predicted formation energy does not prove stability, manufacturability, safety or useful performance.
- The LLM received iterative feedback while headline specialist baselines were one-shot generators, making the comparison architecturally asymmetric.
- Prompt history grows with the number of proposals, raising context-window and cost limits for larger or less structured spaces.
What to watch next
- Independent reproduction with equal-feedback Bayesian and evolutionary search baselines.
- Prospective benchmarks whose evaluator landscape was unavailable during model training.
- Higher-fidelity calculations and uncertainty estimates for selected candidates.
- Laboratory synthesis and property measurements rather than benchmark recovery alone.
- Full compute, token, carbon and researcher-time accounting across search methods.
Living evidence record
Impact record IAI-1U9O3UC
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
2 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Science & Research
Can AI choose the next drug-chemistry experiment?
A peer-reviewed Roche–ETH study used active learning to select 50 new substrates from a 22,253-molecule pool, then tested about 96 reaction conditions for each. A later graph model identified the observed reaction site for all ten prospectively tested substrates—but only within a narrow chemistry workflow, not drug discovery as a whole.
9 min · 2 sources
Science & Research
Can a failing AI policy reveal what a tumour model is missing?
A peer-reviewed simulation study used the different ways ten reinforcement-learning policies failed when moved into a richer tumour model to identify missing spatial dynamics. It is a method for interrogating models—not clinical evidence for an AI-designed cancer treatment.
8 min · 1 source
Science & Research
Can a polymer AI predict beyond its simulations?
AdaptDelivery predicts polymer size and solvent-accessible surface area through neural networks trained on 20,000 values generated from fitted molecular-dynamics curves. The low errors show the networks learnt those curves; they do not validate new polymer chemistry, wet-lab delivery or clinical performance.
8 min · 1 source
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.