Skip to content
D-CSIL

AI News · 2026-10-06 · 9:00 PM CT

OpenAI's model proved 722 theorems — 162 are already machine-checked

TL;DR

OpenAI released 722 mathematical manuscripts from an internal frontier model into a public GitHub repo — 162 already machine-checked in Lean, including the quasi-Riemann hypothesis. The honest story is a verification ladder: claimed, manuscript, machine-checked. The bottleneck in mathematics just moved from finding proofs to verifying them.

The equation L(s, χ) ≠ 0 for Re(s) > 7/8 — the quasi-Riemann hypothesis
Photo: D-CSIL AI Briefing

What OpenAI released

On October 6, OpenAI put 722 mathematical manuscripts — organized into 372 families — into a public GitHub repo called openai/math, all produced by an unreleased internal frontier model. The announcement says they consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to release it.

The headline claims are enormous: a proof of the quasi-Riemann hypothesis, the rational Hodge conjecture for CM abelian varieties, the Birch–Swinnerton-Dyer formula in low Selmer corank, counterexamples disproving Kaplansky's conjectures, the Artin K(pi,1) conjecture, and Yau's uniformization conjecture — spanning 17 mathematical disciplines.

How it was produced

OpenAI's README describes the procedure plainly. They were evaluating the model on open research problems after it saturated their existing math benchmarks. They posed roughly 4,000 problems. Each result took, on average, three hours of ChatGPT Pro-level thinking compute. Then they aggregated outputs into families and applied a significance bar.

That is an 18% hit rate on open research problems. This is not a model thinking really hard once — it is a research pipeline, AI as a scientific instrument at a scale no human mathematical community could match in a year. It found proofs the way a particle collider finds particles: aim the instrument at enough targets and record what comes out.

The verification ladder

The repo is explicit that results sit at different stages of verification. Tier one — machine-checked: 162 manuscripts have Lean formalizations, every logical step verified by computer. That includes the quasi-Riemann hypothesis (every Dirichlet L-function zero-free in Re(s) > 7/8), the Kaplansky disproofs, the Artin K(pi,1) conjecture, and the symmetric Mahler conjecture.

Tier two — about 560 manuscripts awaiting full verification. OpenAI is blunt: some of the unformalized results could have issues. This tier holds the Hodge, BSD, and Yau results. And note what is not claimed: nobody proved the full Hodge conjecture or the full Riemann hypothesis. These are enormous special cases — not the Millennium prizes themselves. Independent verification by unaffiliated mathematicians has not happened yet.

What it means

The bottleneck in mathematics just moved from finding proofs to verifying them. The 162 Lean-checked results are the honest number for actually solved, verified — everything else is claimed, pending. And the pattern generalizes: generate at superhuman scale, verify rigorously, publish with the verification status attached. Mathematics is the ideal first field because proof is checkable.

The practical takeaway: the era where AI assists mathematics is over. The era where AI produces mathematics, and humans plus proof checkers verify it, started today. The open question is not whether the machine can do research — it is how fast the rest of us can check its work.