AI Is Conquering World-Class Math Conjectures — Will Mathematicians Become "Curators"?

AI is solving world-class math conjectures, pushing mathematicians to become curators of AI-generated proofs.
AI is cracking world-class mathematical conjectures at an unprecedented scale — from IMO gold medals to the Langlands Program. Fields Medalist Terence Tao says math has entered an era of "proof abundance," but AI has a fatal blind spot: it cannot judge research significance. The mathematician's core edge is shifting from "can you prove it?" to "can you judge what's worth proving?"
From "Proof Scarcity" to "Proof Abundance": A Historic Reversal
AI is now cracking world-class mathematical conjectures at scale — from olympiad problems to the Langlands Program. A field once considered the highest threshold of human intellect is being upended in ways never seen before.
A recent in-depth report in New Scientist vividly captured this disruption. A journalist with no math background typed a single prompt and watched GPT-5.5 Pro produce a logically coherent, complete proof of Erdős Problem #710 in just 22 minutes. The journalist briefly thought they'd made a major breakthrough — until a professional scholar reviewed it and discovered that the AI had simply restated a conclusion proven decades ago.
Understanding the deeper significance of this mishap requires knowing the special status of the Erdős problem canon. Paul Erdős (1913–1996) was one of the most prolific mathematicians of the 20th century, publishing over 1,500 papers and collaborating with more than 500 scholars across combinatorics, number theory, graph theory, and virtually every corner of discrete mathematics. He left behind hundreds of prize problems, with bounties ranging from $25 to $10,000, forming a unique "Erdős prize problem" ecosystem. These problems vary enormously in difficulty — some appear simple yet have gone unsolved for decades. Erdős #710 and #1196 both fall into this category, involving combinatorial properties of integer sets and density estimates of primitive sets, respectively. It is precisely this "simple surface, deep core" quality that makes AI-generated "proofs" nearly impossible for non-experts to evaluate.
The incident seems almost comical, yet it cuts to the heart of the deepest anxiety now gripping the mathematics community. A decade ago, a rigorous proof of a significant theorem might take top scholars years to develop, and the field operated in a state of chronic "proof scarcity." In just two years, the situation has completely reversed.

Google DeepMind's AlphaProof has already earned an IMO gold medal. Its architecture is elegant: it combines a large language model with a formal proof verifier (the Lean proof assistant), first generating natural-language reasoning with the language model, then translating it into formal logic that a computer can rigorously verify. This "generate-and-verify" feedback loop mechanically validates every step, fundamentally eliminating logical gaps. In the 2024 IMO test, AlphaProof solved 4 of 6 problems — gold medal level. Problems that once only the most elite students could crack are now within reach for ordinary people using AI. Thomas Bloom, a scholar at the University of Manchester who has long catalogued Erdős's hundreds of unsolved problems, reports that complete solutions are now flooding in from amateurs — even second-year Cambridge undergraduates can use AI to independently crack several classical conjectures.
Fields Medal winner Terence Tao offered a precise summary. Tao received the Fields Medal in 2006 at the age of 31 — the prize is awarded only once every four years and exclusively to mathematicians under 40 — and he himself is an active practitioner of AI-assisted mathematical research, having publicly documented experiments using large language models to accelerate literature review. His assessment carries particular weight: mathematics has moved from "proof scarcity" to a new era of "proof abundance." Where mathematicians once competed to be first to complete a proof, the competition of the future will be about who can make sense of the vast torrent of derivations generated by AI.
AI's Breakthroughs: Escaping Human Mental Ruts
At their core, large language models are extraordinarily powerful pattern-recognition engines. Having ingested a century of mathematical literature, they can reassemble scattered theoretical fragments and occasionally break free of the rigid mental frameworks humans get stuck in, charting entirely new solution paths.
The clearest example is Erdős Conjecture #1196 on primitive sets, which had stumped the field for a full 60 years. Human researchers had long been constrained by a fixed analytical framework, but AI independently introduced the von Mangoldt function — a tool originally developed for the study of prime numbers — and neatly sidestepped every technical obstacle.
The von Mangoldt function (Λ(n)) is a classic tool in analytic number theory, originally designed to study the distribution of primes: it returns a logarithmic value when n is a prime power and zero otherwise, serving as a key bridge between the Riemann zeta function and the prime number theorem. It had long been regarded as a "primes-only tool." AI's application of it to the primitive set conjecture amounted to discovering a deep analogy between two seemingly unrelated mathematical structures — primitive sets and the set of primes share similar extremal properties in density control. Cross-domain tool migration of this kind is not unheard of in mathematical history (e.g., Fourier analysis entering number theory), but it typically requires years of intuition-building by top scholars. By training on vast bodies of literature, AI has in some sense "compressed" this discovery process, sparking deep debate in the mathematical community about whether mathematical intuition can be computationalized.
Stanford researchers described these as "the first AI-generated proofs with lasting academic value," and noted they have given rise to genuinely new general research methods.

What makes this case significant is that it demonstrates AI is not merely a tool for accelerating existing human approaches — it may also drive genuine methodological innovation by transplanting mature tools from one mathematical domain into another that appears entirely unrelated, breaking through long-standing deadlocks. This capacity for cross-domain "knowledge connection" is a distinctive advantage that emerges from AI's exposure to enormous bodies of literature.
The Fatal Blind Spot: AI Cannot Judge "Research Significance"
Yet AI has a fundamental blind spot: it cannot judge the significance of research. It cannot distinguish between conjectures at the cutting edge and conclusions that are redundant or long obsolete. As the journalist's mishap at the opening illustrates, AI produces logically coherent derivations with no ability whatsoever to assess whether a result has genuine academic value.

A scholar at the University of Toronto put it plainly: models can complete derivations, but they can never perform the human work of understanding, judging, and choosing. This statement defines the current boundary between AI and human mathematicians — machines handle computation and logical deduction; humans handle value assessment and directional choices.
Put differently, "the correctness of a proof" and "the importance of a proof" are two entirely distinct dimensions. AI can now efficiently handle the former, but the latter remains firmly in human hands. A logically airtight proof might do nothing more than restate the obvious; what actually advances a field is often the key result that points toward the unknown and connects disparate branches of mathematics.
The Mathematician's New Role: From Creator to Curator
At a closed-door academic gathering in San Francisco this past April, mathematicians from around the world openly shared their collective confusion. Someone posed the question: if future mathematical research requires nothing more than typing a prompt and waiting for AI to output a proof, would you still want to be a mathematician? Only half of the scholars in the room raised their hands.
Everyone was forced to accept a fundamentally new professional identity: "mathematical curator."

The standard research workflow of the future may be completely restructured: AI handles unlimited trial and error, exhaustively exploring all derivation paths and batch-generating vast quantities of theorems and proofs; human mathematicians handle screening and evaluation, judging which lines of reasoning have extensibility, which conclusions can bridge different branches of mathematics, and selecting from among them the core results that truly advance the discipline.
Mathematics is no longer "the art of creating proofs" — it is becoming "a systematic engineering of selecting value." This is a profound role reversal: the core competency of the mathematician is shifting from "can you prove it?" to "can you judge what's worth proving?"
The Ultimate Open Question: Where Are Human Mathematicians Headed?
Today, AI has already solved the 80-year-old "unit distance problem in the plane" and is beginning to make inroads into the core components of the grand unifying theory of mathematics — the Langlands Program.
The Langlands Program, proposed by Canadian mathematician Robert Langlands in 1967, is often described as mathematics' "grand unified theory." Its ambition is to reveal deep correspondences among number theory, algebraic geometry, representation theory, and harmonic analysis. The Taniyama–Shimura conjecture — the central tool Andrew Wiles used to prove Fermat's Last Theorem — is a special case of the Langlands Program. The program is enormous in scope, encompassing dozens of interrelated conjectures; mathematicians have estimated that a complete proof would require the efforts of several generations. AI beginning to engage with its core components means that machine reasoning has extended beyond "closed problems" like competition questions and into genuine research frontiers — a development that has shaken the mathematical community far more profoundly than the IMO gold medal itself. And the machine's derivation capabilities continue to grow at an accelerating pace.
But the question that truly merits reflection is a more extreme vision of the future: if there comes a day when AI can both independently propose frontier conjectures and produce beautiful, profound, complete proofs, how will those we today call "mathematicians" define themselves?
If AI eventually breaches that last fortress — the judgment of value — will the human role in mathematics disappear entirely? No one can give a definitive answer today. What we can say is that the intrinsic aesthetics, curiosity, and drive to explore that define mathematics as a human intellectual endeavor may be precisely what AI can never replicate. At least for the foreseeable future, the division of labor and collaboration between humans and AI in mathematics will be the most thrilling transformation this ancient discipline has ever seen.
Key Takeaways
Related articles

The Truth Behind Codex 'Build a Website in 5 Minutes': AI Isn't Creating Sites—It's Helping You Copy Them
Exposing the truth behind viral Codex 5-minute website videos: creators aren't building original sites with AI—they're copying shared prompts or scraping others' work. Learn AI coding tools' real limits.

Getting Started with AI Agent Development: A Complete Guide from Concept to Practice
A comprehensive guide to AI Agent architecture and development, covering automated marketing, intelligent customer service, and investment analysis scenarios with single and multi-agent collaboration.

The Truth Behind Codex 'Build a Website in 5 Minutes': AI Isn't Creating Sites — It's Helping You Copy Them
Exposing the truth behind viral Codex 5-minute website videos: creators aren't building original sites with AI — they're copying shared prompts or scraping others' work.