OpenAI has published on GitHub a collection of 722 mathematical studies produced by an internal artificial intelligence model that the company has not yet released for public use. The collection covers 372 families of results, including attempts to address long-standing open problems, according to an evaluation by the independent Advisory Group on Mathematics and Artificial Intelligence (AGMAI).
However, publishing this work does not mean that it has been accepted as final mathematical achievements. OpenAI explains that the results are at different stages of verification, and that some informal studies may contain errors, while all of them require review by the mathematical community and confirmation through the usual academic channels.
From Around 4,000 Problems to 722 Studies
The model was tested on around 4,000 mathematical problems during the evaluation process. After filtering the results according to their level of importance, OpenAI selected the published collection, which consists of 372 families and 722 studies. The company says that producing a single result required, on average, computing power equivalent to around three hours of the reasoning capacity available to ChatGPT Pro users.
To improve inspectability, OpenAI published additional data about 10 problems, including brief summaries of the reasoning process, estimates of the computing power used, and the number of problems attempted. The examples span various fields, including a measure of the irrationality of the number π, Mahler conjectures, the Kaplansky conjecture concerning direct finiteness, and the three-dimensional relativistic Vlasov-Maxwell system.
Verification Using Lean Is Not Comprehensive
A significant portion of the published proofs underwent review using Lean, a formal verification system that enables computers to test the logical consistency of a proof. However, not all studies in the collection have formal verification through Lean so far.
OpenAI intends to update the GitHub repository as the formal verification processes are completed. The company says that most of the results were produced according to a standardized procedure, with some exceptions; one study concerning a zero-free region for the Riemann zeta function was obtained in a different way, and a human edited its text to improve readability.
Why Does This News Matter?
The collection’s main value lies not only in the number of studies, but also in testing a new method for producing mathematical results that can be reviewed. Making the files, computing-cost data, reasoning summaries, and partial automated verification available provides researchers with the elements they need to assess whether the model produces discoveries that can be reexamined or merely texts that appear convincing.
At the same time, clear limitations remain: verification is incomplete, human and academic review has not concluded, and the model’s results do not yet represent a system available to researchers or a public product. AGMAI calls on artificial intelligence companies to publish mathematical results through academic channels, disclose the models, costs, and methods used, and avoid turning discoveries into a purely marketing tool.
OpenAI also announced that it will fund workshops, conferences, and special programs to understand important results produced by artificial intelligence, in parallel with its work to make the internal model responsibly available. The current collection therefore represents a test step in how to document automatically generated scientific results, not a final judgment on the model’s mathematical capabilities.