OpenAI published 722 mathematical manuscripts on October 6, 2026, crediting the work to an internal frontier model it has not named and will not release. The drop landed on GitHub, not in a press conference, and it left mathematicians, AI researchers, and rival labs parsing a strange question: what exactly is proof of progress when the thing that produced it stays locked away?

What OpenAI Actually Published on October 6

The company said it was “releasing a broad range of new mathematical results produced by an internal frontier model.” That is a careful sentence, and it matters that OpenAI chose those words rather than announcing a new model launch. The release itself landed in the public repository openai/math, and coverage of the drop was confirmed across outlets including Unite.AI, The New York Times, The Indian Express, and Business Standard within hours of the GitHub push going live.

The headline figures are 722 manuscripts, organized into 372 result families. A result family, in this context, groups related theorems, lemmas, or problem variants that share a common thread rather than treating each write-up as a standalone contribution. That distinction matters for anyone trying to size up the release: 722 sounds like a flood of new mathematics, but 372 families is a more honest measure of how many genuinely distinct lines of reasoning the model actually pursued.

Inside the openai/math Repository on GitHub

Beyond the raw manuscripts, the repository includes supporting proof artifacts, Lean formalizations for many of the results, and ten abridged summaries describing how the model reasoned its way to a solution. Anyone can clone the repository and start reading through the material today. A basic entry point looks like this:

git clone https://github.com/openai/math.git
cd math
ls

The inclusion of Lean formalizations is the detail worth sitting with. Lean, the formal proof assistant maintained by the Lean community, forces a proof into a form a computer can check step by step, with no room for the hand-waving that sometimes slips into human-written mathematics. Pairing natural-language manuscripts with machine-checkable Lean code gives outside reviewers a way to test claims instead of simply trusting them, which is precisely the kind of verification a release like this needs if it wants to be taken seriously by working mathematicians rather than just AI watchers.

From Problem to Proof: Roughly 4,000 Attempts

According to reports on the release, the unnamed model was presented with approximately 4,000 mathematical problems. Run the arithmetic against the headline numbers and the picture sharpens: out of roughly 4,000 attempts, the model produced material judged worth publishing in under 400 result families. That is not a success rate anyone should read too precisely, since OpenAI has not published a breakdown of how many attempts failed outright versus how many produced partial or redundant results. But it does frame the release as a curated selection rather than a raw dump of everything the model touched.

That framing lines up with how OpenAI’s math review board, formed earlier this year to evaluate claims tied to the company’s models, has approached similar announcements: results get filtered through internal review before anything reaches the public, and the company has leaned on that process repeatedly as scrutiny of AI-generated math claims has grown.

The Compute Bill Behind Each Result

OpenAI disclosed one figure that puts real weight behind the release: the average result used computing power equivalent to roughly three hours of ChatGPT Pro thinking time. That is a meaningful amount of inference-time compute per problem, and it is also a figure that doubles as a cost signal. ChatGPT Pro’s extended reasoning mode is already the most compute-intensive consumer tier OpenAI sells, and burning three hours of it, on average, across 372 result families implies a serious internal compute budget dedicated purely to mathematical exploration.

It also raises a practical question nobody outside OpenAI can answer yet: how much of that three-hour average was spent on the roughly 370 families that made the cut, versus the much larger pool of attempts that did not produce anything worth formalizing. Three hours per published result is one number. Three hours per attempt, across 4,000 attempts, is a very different and far larger number.

What’s Confirmed Versus What’s Still Unverified

It helps to separate what OpenAI has actually stated from what has spread around the release. Confirmed: the 722-manuscript count, the 372 result families, the GitHub repository, the Lean formalizations, the ten abridged reasoning summaries, the roughly 4,000 problems attempted, and the three-hour average compute figure. Also confirmed, bluntly: OpenAI has not released the model itself, and has given no public identifier for it.

Not confirmed: whether every one of the 722 manuscripts contains a correct, independently validated proof. Reports on the release note that some of the proofs have not been formally checked, which is a notable gap given how much the release leans on its Lean formalizations as a credibility signal. Not every manuscript appears to have gone through that machine-checking step, and OpenAI has not published a count of how many have.

A Separate, Smaller Figure Floating Around

One complication: at least one report put the collection at 377 new mathematical results, a figure noticeably different from the 372 result families OpenAI itself cited. The gap is small in absolute terms but it is a reminder that secondary coverage of a dense technical release can drift from the source numbers fast, and readers chasing the “real” count should treat OpenAI’s own repository and announcement as the baseline rather than any single aggregator’s restatement of it.

The Quasi-Riemann and Navier-Stokes Claims OpenAI Has Not Made

Within hours of the repository going public, claims began circulating that the model had made progress on the quasi-Riemann Hypothesis or had meaningfully advanced the Navier-Stokes existence and smoothness problem, two of the most famous open questions in mathematics. Those claims did not come from OpenAI’s announcement or from the repository material itself. Nothing in the cited release establishes either claim, and conflating a batch of 722 manuscripts with a solved Millennium Prize problem is the kind of leap that tends to happen in the first 48 hours after any AI-and-math story breaks, before anyone outside the lab has had time to actually read the proofs.

Readers should treat any headline claiming OpenAI “solved” either problem with real skepticism until a named mathematician outside the company has reviewed the specific manuscript in question and said so on record. That is exactly the kind of independent check the Lean formalizations exist to make possible, and it is also exactly the check that has not happened yet for most of the 722 files.

Why OpenAI Won’t Release the Model Itself

OpenAI said it is “working to responsibly release the model,” language that echoes how the company has talked about other internal systems it has held back or shelved before shipping, including the GPT-6.1 Astra model it shelved in favor of Sol earlier this year. Releasing research outputs while withholding the underlying model is a pattern OpenAI has used before, and it buys the company time to run safety evaluations, cost down inference, or simply decide whether a math-specialized frontier system has any commercial product shape at all.

It is also a lower-risk way to claim a research milestone. A paper dump invites scrutiny of the math. A model release invites scrutiny of the math, the safety profile, the misuse potential, and the cost structure all at once. Given how much turnover and internal friction OpenAI’s safety organization has seen this year, following the departure of its safety lead after a string of model launches and the exit of several safety researchers amid leak disputes, a narrower, lower-stakes release looks like the more defensible move right now.

Market Impact: A Research Flex, Not a Product Launch

Markets did not move on this the way they move on a new consumer model launch, and that is the point. The release functions as a capability signal aimed at researchers, competing labs, and enterprise customers evaluating which lab’s reasoning systems are worth building on, not at retail ChatGPT subscribers. Math benchmarks have become one of the clearest, hardest-to-fake proxies for raw reasoning capability in frontier AI, which is why labs keep publishing them even when there is no consumer product attached.

For OpenAI specifically, the timing lands in the middle of a crowded few months of model releases and shelvings, from Sol’s discounted launch to the ongoing scrutiny from the FTC’s inquiry into AI agent behavior at both OpenAI and Anthropic. A clean research win, with no agent misbehaving and no model escaping a sandbox, is good press in a year that has had plenty of the opposite.

Historical Context: AI and the Hardest Problems in Math

AI systems tackling serious mathematics is not a new story, but it has accelerated fast. Google DeepMind’s AlphaProof and AlphaGeometry systems reached silver-medal-equivalent performance at the 2024 International Mathematical Olympiad, according to DeepMind’s own announcement at the time, which was widely treated as a watershed moment for AI reasoning on competition-level problems. OpenAI has been racing to match and exceed that bar ever since, and its own math review board, formed to evaluate claims attached to its models after earlier disputes over what its systems had actually proven, is part of the same arms race to be taken seriously by the mathematics community rather than dismissed as a hype machine.

What makes the October 6 release different from an Olympiad benchmark run is scale and format. Olympiad problems are well-defined, time-boxed, and scored by a known rubric. Research-level mathematics is open-ended, and judging whether 372 result families represent genuine novel contributions requires the kind of slow, specialist peer review that competition math never needed. That is a much higher bar, and it is one OpenAI has only partially cleared by shipping Lean formalizations alongside some, not all, of the manuscripts.

Competitive Comparison: Where the Major Labs Stand on AI Mathematics

The table below lays out how OpenAI’s October 6 release compares with the public posture of other major labs on AI-driven mathematics, based on each company’s own disclosures.

LabLatest public math milestoneModel released?Verification method disclosed
OpenAI722 manuscripts, 372 result families (Oct. 6, 2026)No, model withheldLean formalizations for many, not all, results
Google DeepMindAlphaProof / AlphaGeometry, silver-medal-standard IMO 2024 performanceResearch systems, not shipped as consumer productsFormal proof checking tied to Olympiad scoring
AnthropicNo equivalent large-scale math manuscript release disclosedN/AN/A
xAINo equivalent large-scale math manuscript release disclosedN/AN/A

The comparison is lopsided because the categories themselves are new. No other lab has published a manuscript count or a result-family figure at this scale, which is exactly why OpenAI’s framing of the release as a pure research artifact, detached from any model launch, reads as an attempt to claim first-mover status in a metric nobody else is reporting yet.

By the Numbers: The October 6 Release

Here is every verified figure from the release in one place.

MetricFigureStatus
Total manuscripts published722Confirmed by OpenAI
Result families372Confirmed by OpenAI
Alternate result count reported elsewhere377Reported by a separate outlet, not matching OpenAI’s own figure
Abridged reasoning summaries included10Confirmed by OpenAI
Problems the model attempted~4,000Reported
Average compute per result~3 hours of ChatGPT Pro thinkingConfirmed by OpenAI
Model itself releasedNoConfirmed by OpenAI

Five Predictions for What Happens Next

First, expect a slow trickle of independent verification over the coming weeks rather than one clean verdict. Mathematicians working through 372 result families, especially ones without Lean formalizations attached, will take months, not days, to confirm or challenge the strongest claims.

Second, expect at least one of the 722 manuscripts to be publicly disputed or walked back. Releases at this scale, produced without full external peer review before publication, almost always contain at least a handful of errors, redundant results, or claims that need qualification once specialists dig in.

Third, expect Google DeepMind or another major lab to respond with its own math-focused disclosure within the next quarter. This is a benchmark race now, and labs rarely let a rival’s headline-grabbing number sit unanswered for long.

Fourth, expect continued pressure on OpenAI to either name the model or explain, in more specific terms than “working to responsibly release it,” what a release would actually require. Vague release language tends to draw more scrutiny over time, not less, especially once competitors start asking the same question publicly.

Fifth, expect this release to get cited in enterprise sales conversations long before it gets cited in a peer-reviewed journal. A manuscript count and a GitHub link are useful marketing collateral for selling reasoning-heavy AI tools to technical customers, regardless of how the math community eventually grades the work.

Why Mathematicians Are Both Intrigued and Cautious

The mathematics community has been burned before by AI claims that outran their evidence, which is part of why the reaction to this release has split cleanly between genuine interest in the Lean-formalized results and open irritation at the viral claims about Navier-Stokes and the quasi-Riemann Hypothesis that OpenAI never actually made. The company’s decision to publish supporting proof artifacts rather than just final answers is a step in the right direction for anyone trying to build trust with a skeptical field. Shipping it without full formal verification on every manuscript, and without naming the model that produced any of it, leaves plenty of room for that trust to erode just as quickly if early spot-checks turn up errors.

That tension, impressive scale paired with incomplete verification, is likely to define how this release is remembered. It also mirrors a broader pattern across the industry this year, visible in how benchmark comparisons between rival frontier models keep getting contested almost as soon as they are published, and in how Anthropic’s own public filings have leaned heavily on caveats and risk disclosures rather than unqualified victory laps. Big numbers travel fast. Independent verification travels slower, and it is verification, not the manuscript count, that will decide whether this release holds up.

What to Watch in the Coming Weeks

Three things are worth tracking directly. First, whether any named mathematician outside OpenAI publishes a review of specific manuscripts, since that is the signal that will actually move this story from “interesting GitHub drop” to “verified mathematical contribution.” Second, whether OpenAI adds Lean formalizations to the manuscripts that currently lack them, which would meaningfully close the verification gap critics have flagged. Third, whether the company says anything more concrete about naming or releasing the model itself, given how much attention a fully unnamed frontier system has already drawn just from its research output alone.

Frequently Asked Questions

What did OpenAI actually release on October 6, 2026?

OpenAI published 722 mathematical manuscripts, organized into 372 result families, in the public GitHub repository openai/math. The release includes supporting proof artifacts, Lean formalizations for many results, and ten abridged summaries of the model’s reasoning.

Did OpenAI release the model that produced these results?

No. OpenAI has not released the model and has not given it a public name. The company said it is working to responsibly release it, without a specified timeline.

Did the model solve the Riemann Hypothesis or the Navier-Stokes problem?

No. Claims linking the release to the quasi-Riemann Hypothesis or the Navier-Stokes existence and smoothness problem were not made in OpenAI’s own announcement or in the cited repository material, and have not been established.

How many problems was the model given?

Reports on the release put the figure at approximately 4,000 mathematical problems presented to the model.

How much computing power did each result take?

OpenAI said the average result used computing power equivalent to roughly three hours of ChatGPT Pro thinking time.

Why is there a discrepancy between 372 result families and 377 results?

OpenAI’s own figure is 372 result families. At least one outlet separately reported 377 new mathematical results. The two numbers do not match, and OpenAI’s own repository and announcement should be treated as the primary source.

Are all 722 manuscripts verified as correct?

Not all of them. Reports note that some proofs in the release have not been formally checked, even though Lean formalizations accompany many of the results.

Where can I read the manuscripts myself?

The full release is public at the openai/math GitHub repository, free to clone or browse without an account.