AI Breakthroughs in Math Research and Sharing
Key Takeaways:
- On October 6, 2026, OpenAI published 722 AI-generated math manuscripts grouped into 372 result families in a public GitHub repo, alongside Lean proof formalizations and abridged reasoning summaries, as reported by Unite.AI and The Verge.
- The results came from an unreleased internal frontier model posed approximately 4,000 problems; OpenAI says an accepted result averaged about three hours of ChatGPT Pro thinking compute.
- Many, but not all, proofs are formalized in Lean, the proof assistant that lets a computer mechanically check each step. OpenAI’s README cautions that unformalized results may contain errors.
- The release follows a September announcement that upset many mathematicians. On September 11, 2026, 25 Fields Medalists called the goals of AI companies and the mathematical community “severely misaligned,” Scientific American reported.
On October 6, 2026, OpenAI put 722 machine-generated mathematical manuscripts into a public GitHub repository, organized into 372 result families and licensed under Apache-2.0, according to Unite.AI’s account of the research post and The Verge. It arrives a month after the company’s claim that an internal model produced a proof for the Navier-Stokes Millennium Prize problem. The October release is OpenAI’s effort to allow outside mathematicians to examine the work, after the September announcement lacked details.
The post is titled “Sharing AI progress in mathematics”, a framing that responds to criticism: in September, OpenAI announced results without releasing details, and the mathematical community strongly objected.
For more context, see our analysis of the advisory group that shaped these release norms and our breakdown of the Navier-Stokes claim itself.
What OpenAI Actually Released on October 6
The centerpiece is the openai/math repository: 722 manuscripts grouped into 372 families, where a family gathers related papers such as a principal result, companion arguments, consequences, or alternative proofs. Each family is classified by mathematical discipline, and a preprints directory carries PDFs, LaTeX source files, and BibTeX citation blocks for every manuscript, per Unite.AI.

The collection covers number theory, complexity theory, geometry, and mathematical physics. Named results include work on the irrationality exponent of pi, the Mahler conjectures, NP-hardness at the semidefinite threshold, Kaplansky’s direct-finiteness conjecture, the Mezard-Parisi formula for diluted spin glasses, and the relativistic Vlasov-Maxwell system, according to Unite.AI’s walkthrough of the repo. Scientific American also lists a claimed solution to the four-dimensional Kakeya conjecture and progress toward the Riemann hypothesis.
Two items departed from the standard process: the zero-free region result for the Riemann zeta function was human-edited for readability, and the Hodge conjecture work for CM abelian varieties was treated as an exception, Unite.AI reports. The README also includes ten abridged summaries of the model’s reasoning.
The release also includes Lean formalizations. Lean is the proof assistant that lets a machine check each logical step, which adds credibility to OpenAI’s claims compared to a typical preprint. But the two layers are uneven: many manuscripts have accompanying Lean formalizations and many do not, and OpenAI’s README explicitly warns that unformalized results “could have issues,” per Unite.AI.
Here is what a Lean formalization for a simple result looks like. The point is not the theorem, it is that every step is a typed expression the kernel can reject:
-- Lean 4: every step must type-check or the kernel rejects the file
theorem sqrt_two_irrational : Irrational (Real.sqrt 2) := by
-- proof by contradiction, fully machine-checked
rintro ⟨q, hq⟩
obtain ⟨a, b, hb, h⟩ := q
-- h : a^2 = 2 * b^2, derive a contradiction via parity
have h_even : Even a := by
rw [← sq] at h
exact even_of_even_pow h
obtain ⟨k, rfl⟩ := h_even
have h2 : 2 * b^2 = 2 * (2 * k^2) := by linarith
have hb2 : b^2 = 2 * k^2 := by nlinarith
exact hb (even_of_even_pow (by rw [hb2]; exact ⟨k^2, by ring⟩))
-- Note: this is an illustrative snippet; the repo's real formalizations
-- span hundreds of lines and are verified against the openai/math library.
A single unreleased internal model was posed approximately 4,000 open problems during a model-development evaluation, a figure both Unite.AI and Interesting Engineering attribute to the repo README. Its outputs were filtered for significance, grouped into families, and divided into manuscripts and formal proof artifacts before being published on GitHub.
The Numbers: Scale, Cost, and Verification
OpenAI disclosed that the model attempted about 4,000 problems, and that an accepted result took roughly three hours of equivalent ChatGPT Pro thinking compute on average, as reported by Interesting Engineering and The Verge, a specific, verifiable efficiency figure rather than a marketing term.
Those figures allow outsiders to compare two ways of spending compute. The September Navier-Stokes effort used a swarm of roughly 10,000 concurrent agents over 88 hours, and OpenAI’s head of research Mark Chen described the compute cost as “in the millions of dollars,” according to VentureBeat and CNBC. OpenAI itself only said “millions,” and no independent audit of the cost has been published. Scientific American describes the October batch as arriving without that multimillion-dollar price tag, with an OpenAI spokesperson saying almost every result came from a single prompt given to a single agent. Andrew Sutherland of MIT told the magazine that claims of solving problems with a single agent should be treated as unverified until the model is released and the results can be replicated.
| Metric | October 6 release | September Navier-Stokes claim | Source |
|---|---|---|---|
| Manuscripts / results | 722 manuscripts, 372 families | 1 claimed proof, Lean-formalized | Unite.AI |
| Problems attempted | Approximately 4,000 posed to the model | Roughly 10,000 concurrent agents | Interesting Engineering, VentureBeat |
| Compute per result | About 3 hours of ChatGPT Pro thinking, average | 88 hours, “millions of dollars” | The Verge, CNBC |
| Formal verification | Lean for many, not all | Lean formalization shared | Unite.AI |
The comparison reveals a change in approach. September was a sprint: deploy a large agent swarm on one famous problem and announce quickly. October is a census: pose thousands of problems, formalize what you can, and publish the total number attempted alongside the number of results produced. Reporting the total number attempted is the most important transparency improvement, because announcing only successes inflates the appearance of capability. Even here the disclosure is incomplete: Scientific American notes that OpenAI published only the average compute time and no prompts, despite the advisory group’s request for the prompt, time, and cost behind each result.
Where the Benchmark Scores Fit In
The released results come from an unreleased internal model, not the public GPT-6 Astra. OpenAI has said that internal system is significantly more capable than GPT-6 Astra, per Unite.AI. The public model’s benchmark progress explains why OpenAI kept raising its math ambitions. GPT-6 Astra scored 97.6% on FrontierMath’s hardest Tier 4, up from 17% on an internal starting version, which earned it the first “Major Advance” classification from Epoch AI on that benchmark, per CryptoBriefing’s report on the Epoch AI classification and CryptoBriefing’s score report. It also posted 99.9% on ARC-AGI-3, a reasoning benchmark that had defeated prior models, according to GCN.
That progress is part of a longer trend. A year earlier, Google DeepMind and OpenAI both earned gold-medal-level scores of 35 out of 42 at the International Mathematical Olympiad, per Nature’s coverage. The shift from contest problems toward research-level open problems has happened in about a year.
The benchmark scores and the released manuscripts provide different kinds of evidence. A benchmark is a controlled test with a known answer key. A solved open problem has no answer key, which is why Lean formalization and community review matter more than any headline number. Epoch AI’s “Major Advance” label confirms a capability improved, not that a theorem is true. The 97.6% figure also comes from coverage of vendor-reported evaluations rather than a paper Epoch AI itself published, so it should be read as a reported score, not an independently audited one.
What Independent Practitioners See: Limitations and Gaps
On September 11, 2026, 25 winners of the Fields Medal released a declaration titled “A Severe Misalignment of AI in Mathematics,” arguing that AI companies treat famous problems as marketing benchmarks and that this harms the discipline. The declaration, published at mathandai.org, states that the goals of AI companies and the goals of the mathematical community “are severely misaligned.” Scientific American reported on the letter, and signatories include Terence Tao, Wendelin Werner, and Martin Hairer.
The substantive complaints are narrower than the rhetoric. The first is attribution: the September Navier-Stokes result built on a “forcing” route opened by Diego Cordoba and Luis Martinez-Zoroa, and on unpublished Euler work by Tristan Buckmaster and Levent Alpoge. OpenAI’s blog post acknowledged it “cannot rule out that de-identified data derived from their usage of our products helped improve our models,” while insisting its agents saw nothing until it was public, as VentureBeat reported. That distinction, retrieval versus training-data influence, is the precise gap the community wants closed.
The second complaint is verification capacity. Mathematics is so specialized that few researchers can fully audit results across number theory, complexity, geometry, and physics at once. Many of the 722 manuscripts lack Lean formalizations, so their correctness is unverified until a human works through them. Scientific American quotes an OpenAI spokesperson saying many of the newly released results are not yet understood by the company’s own mathematicians.
The third is access. The most capable models are proprietary, and researchers at smaller institutions reported being locked out of the resources to reproduce or even check the work. The Fields Medalists’ declaration and the AGMAI advisory group both asked labs to stop testing advanced problems on models the broader community cannot access, a request OpenAI has not publicly agreed to. An OpenAI spokesperson told Scientific American that the company is taking the advisory group’s guidelines seriously but is not bound by them.
For a concrete, reproducible view of how these release norms translate into engineering practice, see our guide to a release workflow that satisfies the advisory group’s guidelines.
What to Watch Next
Three things will determine whether this release is remembered as a real scientific contribution or as a better-structured PR campaign. First, the formalization rate: how many of the 722 manuscripts eventually get Lean proofs, and whether the unformalized ones survive human review. OpenAI said it will add formalizations as they are obtained, but the current gap means a large share of the collection is still unverified claims, as Interesting Engineering notes.
Second, the community’s own infrastructure. The AGMAI advisory group urged OpenAI to deposit results in repositories it does not control, with persistent identifiers. The GitHub repo is a step short of that, and OpenAI said it is “exploring other community-hosted alternatives,” per The Verge. Whether a neutral archive emerges will show whether the community has real influence.
Third, the funding question. OpenAI said it will fund workshops, conferences, and programs to help humans understand the output. The advisory group’s September 29 guidelines insist those decisions should rest with nonprofit institutions, not the lab. Who controls the money that pays for human comprehension is an important detail.
The new fact in the October release is that a lab published the denominator, approximately 4,000 problems attempted, alongside the numerator of 372 result families. That is the first time a frontier lab has let outsiders see the success rate rather than just the successes, and it is the closest thing to an honest benchmark the field has had all year. It is still an average, not a per-result receipt, which is the gap Sutherland and the advisory group both pointed out.
Related Reading
More in-depth coverage from this blog on closely related topics:
- Responsible Sharing of AI Math Tools
- Prime Big Deal Days 2026: Dates and Discounts
- Best Flat Route Between Two Points SF
- Reflection Beam 501B Model Review
- Understanding CFTC Rules and Enforcement
Sources and References
Sources cited while researching and writing this article:
- OpenAI Releases 722 Math Manuscripts From an Unreleased AI Model – Unite.AI
- OpenAI drops another batch of mathematical breakthroughs | The Verge
- Scientific American reported
- openai/math repository
- OpenAI unleashes hundreds more math results upon a field already in shock | Scientific American
- OpenAI's largest math release tackles 4,000 problems with Lean proofs
- VentureBeat
- CNBC
- GCN
- DeepMind and OpenAI models solve maths problems at level of top students
- mathandai.org
Rafael
Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...
