OpenAI's Navier-Stokes Proof Controversy with Mathematicians: The Truth Is Far More Complex Than 'Cherry-Picking'

The OpenAI–Buckmaster dispute reveals deeper questions about information value and attribution in AI-era mathematics.
OpenAI announced that ~10,000 AI agents spent 88 hours finding a Navier-Stokes-related proof, prompting mathematician Buckmaster to accuse the company of "cherry-picking" his team's work. But the real issues run deeper: beyond disputed communications on September 6th, the case raises unresolved questions about whether merely knowing a breakthrough has occurred provides competitive value, how academic contributions should be tracked across AI-assisted research, and why the traditional priority system may be poorly equipped for this new landscape.
A Rumor That Ignited the Academic World
OpenAI recently announced that approximately ten thousand AI agents spent around 88 hours finding a proof related to a difficult problem involving the Navier–Stokes equations. No sooner had the news broken than mathematician Buckmaster publicly stated that during collaboration negotiations, the company had excluded his collaborator and even pressured him regarding his career prospects. Online commentary quickly reached for the term "cherry-picking" to describe the situation — a resource-rich AI company swooping in to claim the fruits of traditional mathematicians' labor.
But if we stop at that dramatic opening, we may have already missed the most important questions worth asking. The publicly available materials are not sufficient to conclude that OpenAI stole mathematicians' drafts. And a deeper question than "who stole from whom" is this: Even without access to any draft, does the mere knowledge that someone has achieved a breakthrough carry value in itself?
The Simplified Story: The Real Context of Academic Collaboration
First, let's clarify a fact that public discourse has glossed over: Buckmaster has long studied fluid equations, while his collaborator Albritton has a background in number theory and also works at Anthropic. According to Buckmaster's own statement, this was a personal academic collaboration — and both researchers were themselves using Claude and OpenAI's tools.
Framing them as "a group of AI-resistant traditional mathematicians vs. a tech company" simply doesn't hold up from the start. More importantly, their research built on prior work by Córdoba and Martínez-Zoroa, and another paper on porous media models lists a third author, Coiculescu, on its front page.
When coverage gets compressed into "two people vs. one company," specific academic contributions are already being discarded in the process. This is precisely the central hazard of this case — the record of contributions is being systematically erased in transmission.
What Makes the Navier-Stokes Millennium Problem So Hard
To judge "what each side proved," we first need to understand the problem.
What Is "Blow-Up" in Fluid Equations
Imagine a swirling mass of water. What the Navier–Stokes equations do is take the current state of that flow and predict how every point will move at the next moment. The Millennium Prize problem asks: can a well-behaved three-dimensional flow, starting from perfectly smooth initial conditions, develop a point of unbounded velocity in finite time? This is called "blow-up."
It's worth emphasizing that blow-up doesn't mean the water literally explodes — real fluids don't move infinitely fast. The problem lies with the mathematical model: even when total energy is finite, that doesn't guarantee gentle behavior everywhere. Energy can concentrate into smaller and smaller regions, with spikes growing ever taller while the total remains constant.

Viscosity, External Forces, and the Boundaries of the Problem
Comparing the two sides' results requires examining two key variables: viscosity (internal friction) and external forcing (pushing and pulling from outside). Remove viscosity, and the discussion shifts to the Euler equations.
To lay it out clearly: the academic team's result concerned the Euler equations with smooth external forcing; OpenAI claims to have produced a result for the Euler equations without external forcing, as well as a result for the Navier–Stokes equations with smooth external forcing.
Having external forcing doesn't automatically mean "dodging the problem" — the Clay Institute's official problem statement itself permits forces that meet certain requirements. The true difficulty lies in satisfying three conditions simultaneously: keeping the external force smooth, keeping total energy finite, and still allowing local velocity to blow up. All three must hold at once. So before anyone talks about "who stole from whom," we need to be clear about what each side actually proved — and all of these new proofs still await verification by the mathematical community.
The Heart of the Dispute: The September 6th Communication and the Data Gap
The controversy centers on a communication that took place on September 6th. Buckmaster alleges that the collaboration proposal excluded Albritton, and that the conversation touched on his career prospects. OpenAI's Bubeck acknowledged that his wording was inappropriate and apologized, but denied demanding that Albritton be removed from the paper; his explanation was that they had discussed rewriting OpenAI's results and talked about how internal models could be made accessible to employees of a competing company.
There is clear responsibility for the wording involved, as well as facts that the two sides have yet to reconcile. We have no verified recording of the call, so "an apology" cannot be expanded into "an admission of all allegations."

The data question also needs to be read in full. OpenAI denies that it failed to solve the problem independently, denies having accessed the other team's unpublished work or specific user data; but it also stated that it could not rule out that "de-identified product usage data" may have helped improve its models. This leaves a gap that needs to be examined: did any draft enter the training process? When? Did it affect the outcome? The available public materials do not provide complete answers.
OpenAI's announcement states that after hearing rumors of a major breakthrough on September 1st, the company launched attempts on multiple hard problems; its system first obtained its own Euler result, after which the company directed more resources toward Navier–Stokes. Finding the proof took approximately 88 hours; formalizing it afterward took another 17 hours. These timelines and resource figures are the company's own account.
Why the Information That "a Breakthrough Exists" Is Valuable in Itself
This is the most intriguing aspect of the case: Even if we fully accept OpenAI's account, the message that "research has progressed" may itself be enormously valuable.
Suppose you have been working on a problem for months, with no idea whether to keep going. Then someone tells you the problem has been cracked — they haven't handed you the proof, and you still have to solve it yourself, but your assessment of the probability of success has already changed. When another team can simultaneously marshal enormous computational resources, that shift in assessment can quickly translate into tangible investment.
This is why "we didn't see their draft" still doesn't fully answer the question of contribution. Someone may have provided a critical signal of progress without providing the final derivation; in this case, exactly how much that signal helped cannot be quantified. Nor does a rumor equate to revealing a specific solution path.

What genuinely deserves to be documented is: which pieces of information influenced the choice of problem? Which results changed resource allocation? And what work was independently completed afterward?
Why the "Cherry-Picking" Metaphor Misleads: Academic Contribution Can't Be Judged by the Final Step Alone
The term "cherry-picking" invites the impression that the last step was just a matter of reaching out a hand. But the final obstacle in mathematics may be precisely the hardest one.
Whoever spent more time doesn't automatically own the entire problem; whoever worked faster doesn't automatically deserve less credit. But looking only at "who submitted the final paper first" also misses the problem framing, the key constructions, the foundational tools, and the verification work. A more accurate approach is to break the work apart: who posed the problem, who found the structure, who closed the final gap, who made the proof checkable and comprehensible.
The AMS's ethics guidelines have long called for proper acknowledgment of unpublished materials and announced results, while opposing premature appropriation of results and blocking others' research. Communicating a piece of progress may deserve acknowledgment, but saying "I'm working on this" cannot permanently fence off a problem. The guidelines also allow independently obtained coincident results, and suggest considering joint authorship in appropriate circumstances — but merely proposing joint authorship is not, by itself, evidence of misappropriation.
Research on Academic Priority: First Arrivals Don't Take All, and Reputation Amplifies the Gap
Researchers Hill and Stein analyzed over 1,600 races in protein structure research, estimating that papers that were scooped received roughly 21% fewer citations. That is certainly a loss — but far from zero.

That study focused on structural biology and cannot be directly applied to this mathematics dispute. What is especially glaring in the present case is that a team with lower name recognition may receive less attention even if they finish first. Timing determines only part of the reward — whose name carries more weight, and who is more easily seen, also shapes the history of discovery.
OpenAI has stated it does not intend to claim the Millennium Prize for this result, but "not claiming the prize" does not automatically resolve questions of citation, authorship, and how contributions should be distributed. If researchers come to expect that their early exploratory work will simply help better-resourced players get there first, they may reduce open communication — that is a consequence worth guarding against, not a universal fact already proven by this case.
Academic Attribution in the Age of AI: We Need Two Sets of Records
Some practical approaches already exist. When Gowers launched the Polymath massively collaborative mathematics project in 2009, he considered using public discussion logs to preserve fragmented contributions: who first proposed an idea, who identified a flaw, who patched it — all retrievable by later readers. PLOS's complementary research policy preserves space for novelty assessment of recent simultaneous work — the fact that someone published a step earlier doesn't have to strip another piece of work of all its value.
These practices don't guarantee fairness, and they can't directly adjudicate this case, but they at least turn "contributions should be acknowledged" into something concrete and operational.
For AI-driven mathematics research, an additional layer of documentation is needed. Formalization tools like Lean can verify derivations but won't tell you which drafts the system was exposed to. Mathematical conclusions require definitions, assumptions, and proofs; the discovery process requires draft timestamps, information access logs, and run records. The two sets of materials answer different questions.
If we want researchers to continue communicating openly, early ideas, independent methods, and verification work must be visible, citable, and assessable. Otherwise, what academia may learn is a cold, self-protective lesson: Until you're done, say as little as possible.
Related articles

LangChain + MCP: From Core Concepts to Agent Tool Calling in Practice
Learn how LangChain and MCP work together — covering LLM tool calling, Agent architecture, and conversation history management to build real-world AI applications.

Probabilistic Machine Learning: Why It's the Cornerstone to Unlocking the ML Black Box
Without probability theory, ML is always a black box. This article explores why probabilistic foundations are essential for understanding machine learning algorithms, Bayes' theorem, MLE, and more.

Optimization Pitfalls in Self-Evolving LLM Agents: Value Concentration and Budget-Splitting Problems
HARNESSEVO research reveals 3 key LLM agent harness optimization findings: value concentrates in reflection/control slots, uniform budget splitting is harmful, and credit assignment must precede structured evolution.