Did AI just produce 372 families of new mathematics?
OpenAI has released 722 manuscripts grouped into 372 result families after posing about 4,000 problems to an internal model. The scale is extraordinary, but the repository says some unformalised results may contain errors and most have not yet received independent human review.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1OpenAI's repository contains 722 manuscripts grouped into 372 result families, selected after an internal model was posed approximately 4,000 problems.
- 2The company reports about three hours of ChatGPT Pro-equivalent thinking compute per result on average and publishes ten abridged reasoning summaries, but it has not released the model itself or a complete reasoning record for every result.
- 3Some proofs have Lean artifacts, but not all; the repository explicitly says some unformalised results could have issues. Independent checking and human understanding are now the decisive tests.
Research topic
Whether a large release of AI-generated mathematical manuscripts constitutes verified new mathematics, and what the published denominator, formal artifacts and independent release standards allow readers to conclude

The answer is a vast claim set, not 372 verified breakthroughs
OpenAI has published an unusually large body of AI-generated mathematical work: 722 manuscripts organised into 372 families. The catalogue includes claims that, if correct, would settle or materially advance problems across number theory, geometry, theoretical computer science, probability and mathematical physics. The release is consequential because it moves beyond benchmark answers into a public corpus that specialists can inspect, challenge and try to formalise.
It is not responsible to count the 372 families as established discoveries. OpenAI's own repository says the collection contains results at different stages of verification, that not every result has an accompanying Lean formalisation and that some unformalised results could have issues. Mathematics normally moves from a claimed proof to scrutiny, correction, expert understanding and eventual acceptance. Publication by a model developer starts that process; it does not complete it.[1][2]
The denominator is about 4,000 attempted problems
The repository gives a denominator that prevents a misleading success narrative. OpenAI says its internal model was posed approximately 4,000 problems. Requiring what the company judged an appropriate level of significance and aggregating related outputs produced 372 families and 722 manuscripts. A family can include a principal claim, companion argument, consequence or alternative proof, so manuscript count is not the same as the number of independent problems solved.
On average, each reported result used the equivalent of roughly three hours of ChatGPT Pro thinking compute with the unreleased model. That is a compute estimate, not a measurement of human research time saved. The release does not provide a random or representative sampling frame for all open mathematics, nor a blinded comparison with teams of mathematicians. Problem selection, significance screening and the grouping of related outputs remain important human and organisational choices.[1][2]
The corpus contains papers, formal artifacts and ten reasoning summaries
The public materials are more substantial than a list of claims. The GitHub repository includes manuscript PDFs and sources, an overview organised by mathematical discipline, citation instructions, a revision policy, Lean files for some proofs and comparator instructions for checking formal artifacts. It also includes ten abridged summaries of the model's reasoning, chosen from families spanning subjects such as multiplicative functions, the irrationality exponent of pi, optimisation hardness and mathematical physics.
Those disclosures help researchers reproduce builds, map related papers and distinguish a machine-checked artifact from an ordinary natural-language proof. But ten reasoning summaries cover only a small fraction of 372 families, and a formal proof checks what has been encoded against a specified logical environment; it does not automatically establish that a formal statement captures the intended informal theorem, uses the right assumptions or has been interpreted correctly. Formalisation is powerful evidence, not a substitute for expert reading.[1][2]
Independent guidance sets a higher bar than repository publication
OpenAI says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study. The group's 29 September recommendations were informed by more than 600 community responses. They call for clear attribution to earlier ideas, conventional mathematical exposition, persistent repositories outside a lab's control, disclosure of model and prompts, reasoning summaries, time and cost, formalisation where possible, and the number of comparable failed attempts.
The group also states that it does not endorse testing advanced mathematical problems on inaccessible proprietary models. Its central concern is human responsibility: when nobody fully understands an AI-generated proof, the lab should fund community-led work to develop that understanding rather than treat volume as marketing. OpenAI's release addresses parts of the requested transparency through a repository, failed-attempt denominator, revision policy and planned support for workshops. Gaps remain, including proprietary model access, incomplete reasoning disclosure and reliance on a lab-controlled repository while other hosting options are explored.[1][3]
The first practical burden falls on mathematicians
For researchers, the immediate opportunity is also a workload problem. A potentially important proof can redirect a field, but hundreds of long manuscripts cannot be trusted or rejected by headline. Specialists will need to check novelty, citations, hidden assumptions, edge cases and whether a claimed theorem already follows from known work. When errors are found, version history and an explicit correction channel will matter as much as the initial release.
The people impact is therefore indirect but real. Early-career mathematicians may gain new conjectures, proof techniques or formalisation projects, while also facing pressure to review machine-produced material at a scale that was cheap for the model developer to generate. Universities, journals and funders will need norms for credit, authorship, triage and compensation. Broad access matters too: if only a well-funded lab can run the model that produced the corpus, researchers elsewhere can inspect outputs without being able to test the instrument on their own questions.[2][3]
What would change the assessment
Confidence would rise family by family, not through a single verdict on the repository. The strongest path is independent specialist review with public reports, successful reproduction of formal checks in clean environments, corrections that preserve earlier versions, seminars where people can explain the arguments and eventual peer-reviewed publication. For headline claims, multiple independent experts should verify both correctness and novelty before anyone describes a problem as solved.
Assessment of the model would also improve with a complete problem-selection protocol, per-problem prompts and compute, the status of every attempted problem, and evaluation by an independent body on a prospectively chosen set. Public access to the model—or at least controlled access for geographically diverse researchers—would test whether the capability is reproducible rather than a one-off curated release. Until then, the appropriate conclusion is precise: OpenAI has made a large, inspectable set of mathematical claims; the mathematical community has not yet converted that set into 372 accepted results.[2][3]
What this means for people
- Mathematicians may gain useful proofs and new research directions, but they also inherit a large verification and interpretation burden.
- Students and non-specialists should not treat a released manuscript or a formal-looking proof as settled mathematics without expert validation.
- Access to proprietary research models could deepen inequalities between wealthy institutions and the wider global mathematical community.
Global context
OpenAI is based in the United States, but mathematics is a global commons whose validation depends on researchers, seminars, journals and repositories across countries. The independent advisory group explicitly warns that unequal access to powerful models could create a multi-tier research system. Verification efforts should therefore include diverse institutions and avoid making acceptance depend on proximity to the model developer or access to its funding.
What the evidence does not yet show
- The 722 manuscripts are grouped into 372 related families, so neither number is a count of independently verified solved problems.
- The collection has not been peer reviewed as a whole, and OpenAI explicitly warns that some unformalised results may contain issues.
- Only some results have Lean formalisation, and a formal check does not by itself establish novelty, intended interpretation or adequate exposition.
- The generating model remains internal, limiting independent reproduction and equitable access.
- Problem selection and significance screening were not a random sample of mathematics, so the release does not establish a general success rate on open problems.
What to watch next
- Independent expert verification, correction or rejection of the most consequential theorem claims.
- A public per-family verification status linking human reviews, formal builds and preserved revisions.
- Community-hosted repositories with persistent identifiers and transparent commentary or review.
- Broader researcher access to the model and a prospective evaluation on independently chosen problems.
Living evidence record
Impact record IAI-09C5HBB
Evidence stage
Announced
Confidence
Corroborated
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
7 October 2026
Source trail
3 direct sources across 3 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 7 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Science & Research
Can an LLM spot what a clinical-trial paper leaves out?
A peer-reviewed benchmark found that a retrieval-plus-LLM pipeline reached F1 0.822 across 119 reporting questions on 40 held-out trial articles. Better evidence retrieval lifted F1 to 0.893, showing both the promise and the main bottleneck.
9 min · 2 sources
Science & Research
Can AI reveal hidden disease patterns in tissue?
A peer-reviewed method used ensembles of neural networks and statistical testing to find disease-associated spatial structures across rheumatoid arthritis, ulcerative colitis and dementia datasets. It is a research-discovery tool, not a diagnostic test or evidence of improved patient care.
8 min · 2 sources
Science & Research
Can an AI choose your next experiment before you spend the compute?
A new preprint tests training intuition. We compare its narrow task with engineering, replication and architecture-search benchmarks to explain what research teams should measure next.
6 min · 4 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.