What happened?
OpenAI has shared 722 mathematics papers produced by an unreleased internal frontier model in a public GitHub repository. The work was filtered by significance and organized into 372 result families.
The model worked on roughly 4,000 problems
According to OpenAI, the model was confronted with about 4,000 math problems during evaluation. The company says producing an average result took compute roughly equivalent to three hours of ChatGPT Pro thinking capacity.
Verification is still ongoing
The independent Advisory Group on Mathematics and Artificial Intelligence (AGMAI) says the collection contains results on hundreds of open problems that had long resisted solution. However, OpenAI itself notes the results are at different verification stages and that some not yet formalized papers may contain errors. A significant portion of the proofs was checked with Lean, a system that uses computers to verify the logical validity of mathematical proofs; not all papers have Lean verification yet.



