OpenAI dropped 722 math papers on GitHub overnight, including claimed progress on three famous open problems. The mathematicians asked to vet them say they didn't.
OpenAI released 722 math manuscripts on GitHub overnight, spanning roughly 17 research directions and including claimed progress on three of the field's most famous open problems: results adjacent to the Riemann hypothesis, the Hodge conjecture, and the Birch and Swinnerton-Dyer (BSD) conjecture.
The release was generated by an OpenAI model the company has not yet shipped. During evaluation OpenAI posed the system about 4,000 problems. The average result cost roughly three hours of ChatGPT Pro compute. Most of the 372 result families came out of the same fixed pipeline, a search process rather than a confirmed proof corpus.
The rest of the catalogue ranges from solid number theory to more speculative items. OpenAI lists a Hilbert 10th result over the rationals (negative), a separate proof pushing the Riemann zero-free region further, partial results on Siegel zeros, a claimed irrationality exponent of 2 for pi, an irrationality result for Catalan's constant, and a claimed proof of the Unique Games conjecture, an open problem in theoretical computer science with implications for the difficulty of approximating many optimisation problems. None of these sits in the headline three, and none has full Lean, the standard machine-checked proof language, formalisation attached.
The three Millennium-adjacent claims sit inside the same caveat. One extends a zero-free region for the Riemann zeta function, which describes where the function's complex values avoid zero. Another handles a specific case of the Hodge conjecture for CM abelian varieties, a class of complex shapes that generalise elliptic curves. A third extends the BSD formula, which connects the rank of an elliptic curve to an associated L-function, to curves whose Selmer group, an arithmetic object attached to the curve, has corank zero or one, and adds a density result covering most quadratic twists. OpenAI labels the Riemann and BSD items as "partial."
OpenAI convened a group of mathematicians in August to discuss the release, including members of a newly formed Advisory Group on Mathematics and AI at the Institute for Advanced Study. The group includes Fields medalists Timothy Gowers and Martin Hairer, plus theoretical physicist Edward Witten. Their statement is blunt: advising the project is not endorsement of the results or the process.
A second group, AGMAI, issued a separate statement that also declines to endorse. It objects to testing advanced mathematical problems on proprietary models and calls for releases that require follow-on human understanding to remain community-led.
Francesco Maggi, a mathematician at the University of Texas at Austin, raised the velocity concern publicly: a prior fluid-mechanics result from the same project is still being analysed, and several hundred more have arrived at once. The bottleneck has moved.
Producing a result used to be the hard part. A working mathematician might spend a year on a single conjecture. Now a single model run can produce several hundred drafts in a night, and the same fixed pipeline that writes them can rewrite them. Reading them is the bottleneck now. The IAS advisory statement names the missing infrastructure directly: formal verification, distributed human-AI reading groups, and AI-assisted review.
Lean, the proof-checking language that has matured over the last decade, gives a path to a verifiable manuscript on submission. Reading groups, where small groups work through a single paper over weeks, scale better when AI summarises background and humans handle the proof structure. AI-assisted review, where tools flag gaps, sketch counterexamples, and check consistency with cited results, is a near-term engineering problem rather than a research one.
OpenAI says the openai/math repository will keep accepting community annotations. The 722 manuscripts sit there now, indexed and timestamped. The next few months of mathematics will look like whichever part of the field builds the verification stack fastest.