OpenAI Releases 722 Mathematical Manuscripts While Verification Questions Mount
The company published hundreds of claimed solutions to open problems on GitHub, but scepticism remains high after earlier controversies and incomplete disclosure of methods.
A Flood of Claims, a Scarcity of Details
OpenAI has published 722 manuscripts on GitHub detailing claimed progress on 372 unsolved mathematical problems, the company disclosed on 7 October. The results originate from an unreleased model referred to internally as "ChatGPT pioneer," and OpenAI states that each solution required an average of approximately three hours of ChatGPT Pro compute time. Yet the company declined to provide per-problem compute figures or the specific prompts used, omissions that have drawn immediate criticism from researchers who say such transparency is essential for independent verification.
The release follows OpenAI's announcement last month that its models had resolved "more than 100 long-standing open problems across most areas of mathematics." Among the headline claims are a purported solution to the four-dimensional Kakeya conjecture, improvements to unspecified computer algorithms, and steps toward the Riemann hypothesis, one of the most famous unsolved problems in number theory. Nearly every manuscript was generated in response to a single prompt given to a single agent, OpenAI told Scientific American.
Advisory Group Guidelines, Selectively Applied
OpenAI formed an independent Advisory Group on Mathematics and Artificial Intelligence (AGMAI) to guide the release process. The group recommended prompt publication through traditional academic channels, disclosure of model names, prompts, and compute costs, and protocols for revisions and citations. OpenAI followed some of these guidelines: the manuscripts are public on GitHub, the model name is disclosed, and the company has outlined a protocol for paper revisions and citations. The firm also stated it is "continuing to explore other community-hosted alternatives for this release which meet the committee's guidelines."
But OpenAI diverged from the advisory group's recommendations on two critical points. It did not release specific compute times for individual problems, nor did it disclose the prompts used. These omissions matter. Compute time per problem signals the difficulty and cost of reproducing a result; prompts reveal the degree of human scaffolding behind each claim. Without them, independent researchers cannot assess whether the breakthroughs are the product of general reasoning or highly engineered inputs.
The Navier-Stokes Shadow
The scepticism surrounding this release is not abstract. Earlier this year, OpenAI claimed its models had made progress on the Navier-Stokes existence and smoothness problem, one of seven Millennium Prize Problems carrying a one-million-dollar reward. That claim ignited a fierce debate in the mathematical community, with several prominent researchers questioning the validity of the work and the lack of reproducibility. The controversy has not been resolved, and it now casts a long shadow over the current release.
Andrew Sutherland, a mathematician at the Massachusetts Institute of Technology, told Scientific American that until OpenAI releases the model itself and allows independent replication, "you should treat any claims about one-shotting problems with a single agent as unverified." That position reflects a broader concern: without access to the model, the mathematical community is being asked to trust rather than verify.
What Verification Actually Requires
Mathematical proof is not a matter of opinion. A valid proof can be checked by any competent mathematician, and the history of the discipline is built on this principle. When a human mathematician claims a result, the proof is published in full, referees examine it line by line, and the community debates edge cases and assumptions. The process is slow, but it is also robust.
AI-generated proofs pose new challenges. The reasoning traces OpenAI has published may run to dozens of pages, and they are not written in the concise, human-readable style that mathematicians expect. Checking them is labour-intensive, and without the prompts and model weights, it is impossible to know whether a given result is reproducible or an artefact of a specific prompt-engineering strategy.
At Opentechwire, we've tracked the growing use of AI in formal mathematics, from DeepMind's work on theorem proving to Meta's efforts in automated reasoning. In each case, the most credible advances have been those where the model, data, and methods were fully disclosed. OpenAI's partial transparency here falls short of that standard.
The Economics of Compute and the Cost of Opacity
OpenAI's statement that the "average result" required around three hours of ChatGPT Pro use is revealing, but only to a point. ChatGPT Pro, priced at two hundred dollars per month for consumers, offers priority access to the company's most capable models. If we assume a rough compute cost in line with that subscription tier, three hours per problem suggests non-trivial inference expense, but not the multi-day runs associated with reinforcement learning or large-scale search.
Yet averages obscure variance. Some problems may have taken minutes; others, days. Without per-problem breakdowns, it is difficult to assess which results represent genuine advances in automated reasoning and which may be the product of brute-force search over vast solution spaces. That distinction matters for the broader AI research agenda. If the breakthroughs are compute-bound rather than algorithm-bound, they may not generalise beyond well-resourced labs.
Regional Implications for AI Research Infrastructure
The release also highlights a structural asymmetry in global AI research. OpenAI's compute infrastructure, likely spanning tens of thousands of high-end GPUs, is beyond the reach of most academic institutions and nearly all researchers in Asia outside the largest corporate labs. If breakthroughs in formal mathematics increasingly require this level of compute, then the ability to verify, reproduce, and extend such work will be concentrated in a small number of US-based firms.
This has implications for research communities in Seoul, Singapore, Bengaluru, and other centres where AI talent is abundant but access to frontier-scale compute is constrained by cost, export controls, or infrastructure. The GitHub release is open, but the model is not. That gap limits the ability of researchers outside OpenAI to engage meaningfully with the work.
What Happens Next
The mathematical community now faces the task of assessing 722 manuscripts, many of which address problems that have resisted solution for decades. Some results may be correct and significant; others may contain subtle errors that only emerge under close scrutiny. The verification process will take months, and it will require coordinated effort from researchers who are already stretched thin.
OpenAI has said it is committed to "further improving the quality of the papers via the citations, mathematical exposition, and presentation of the results for better understanding." That commitment will be tested. If the company releases the model weights, prompts, and per-problem compute data, it will substantially strengthen the credibility of these claims. If it does not, the results will remain in a kind of limbo: published, but not fully reproducible, and therefore not fully trusted.
The tension here is not new. It is the same tension that has accompanied every major AI capability release over the past three years: the gap between what companies claim their models can do and what independent researchers can verify. Until that gap closes, the mathematics community will remain sceptical, and rightly so.



