Crypto

Mathematicians warn AI is solving too many problems too fast

Terence Tao calls for 'analysis-required' standard as AI models solve decades-old math problems within days

A rapid acceleration in artificial intelligence capabilities at major technology laboratories has touched off a debate within the global mathematical community over how research is conducted, verified, and valued. In a warning published on the mathematics-focused social platform Mathstodon, Tao argued that the indiscriminate deployment of frontier AI models risks depleting the finite ecosystem of “good” open problems—unsolved questions that serve as catalyst frameworks for broader mathematical understanding—before human researchers can develop new theoretical fields around them.

Mathematically valid open questions range widely in significance. While arbitrary calculations—such as computing the googol-th digit of pi—remain technically open, they offer little structural insight into the discipline. By contrast, fertile open problems act as benchmarks that drive the development of new techniques, proofs, and conceptual frameworks across subfields.

Led by University of California, Los Angeles mathematician Terence Tao, researchers are calling for structural changes in how mathematical contributions are evaluated as automated systems begin resolving long-standing open problems through computational power. “In short, the indiscriminate use of powerful solution-extraction tools can achieve the immediate short-term goal of solving problems at hand, but at the cost of sustaining the ecosystem for the next wave of progress, or in understanding the progress already obtained,” wrote Tao, who won the Fields Medal in 2006 and was the youngest full professor appointed in UCLA’s history.

The friction comes as artificial intelligence companies allocate compute resources toward pure science and higher mathematics. In May, an OpenAI model disproved the Erdős unit-distance conjecture, an 80-year-old problem first introduced by Hungarian mathematician Paul Erdős in 1946. The conjecture concerns the maximum number of pairs among $n$ points in a two-dimensional plane that can sit at a distance of exactly one unit from each other. The model’s result was verified by external mathematicians, including 1998 Fields Medalist Timothy Gowers of the University of Cambridge.

Historically, mathematical progress has relied on an established “difficulty landscape” that categorizes problems into trivial calculations, solvable challenges requiring novel methods, and intractable questions beyond current techniques. While technological advances have routinely flattened parts of that landscape, they traditionally revealed new frontiers behind them. Tao contends that modern reasoning models disrupt this dynamic because their capabilities operate unpredictably, making it difficult to discern where automated computational search ends and genuine theoretical insight begins.

Within days of OpenAI’s result, Anthropic researcher Levent Alpöge tested the same conjecture using Claude Mythos, Anthropic’s unreleased flagship model. Operating offline to prevent the model from accessing OpenAI’s published materials, Mythos independently derived a solution. Anthropic engineer Sholto Douglas characterized the output as a “cute, simple proof” that was shorter than OpenAI’s version, while mathematician Daniel Litt evaluated it as “a bit worse” than OpenAI’s proof, though Mythos also successfully derived OpenAI’s original solution.

Tao expressed concern that competition between AI laboratories is altering academic research norms. “We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential,” Tao wrote, warning that the dynamic risks reversing centuries of open scientific collaboration.

The pace of automated problem-solving has accelerated across other classical domains. Anthropic recently formalized a proof of Fermat’s Last Theorem—a conjecture originally noted by Pierre de Fermat in 1637 and famously proved by Andrew Wiles in 1994—translating the logical framework into machine-checked code. Days later, OpenAI resolved a 90-year-old mathematical problem hours after a human researcher published a separate proof, resulting in a co-authored paper alongside an Anthropic researcher.

To counter the trend, Tao proposed establishing a formal “analysis-required” designation for select mathematical problems. Under this standard, a raw numerical answer or bare logical output produced by an AI model would hold minimal academic standing unless accompanied by explanatory reasoning that provides insight into adjacent mathematical questions. He compared the proposal to food banks that establish strict quality standards, rejecting donations that are merely edible if they do not meet broader nutritional and institutional criteria.

The integration of computational tools into pure mathematics has a documented history of structural debate. In 1976, Kenneth Appel and Wolfgang Haken of the University of Illinois used computer algorithms to assist in proving the Four Color Theorem, marking the first major mathematical proof that could not be manually verified entirely by hand. In 1998, Thomas Hales used extensive computer calculations to prove Kepler’s 1611 conjecture on sphere packing, leading to the Flyspeck project, which spent more than a decade formally verifying the proof using interactive theorem provers such as Lean, Isabelle, and HOL Light.

More recently, corporate research institutions have expanded from computer-assisted verification to direct AI-driven discovery. Google DeepMind developed systems such as FunSearch, which paired large language models with evaluators to find new bounds for combinatorial problems, and AlphaGeometry, which solved high-level geometry problems from the International Mathematical Olympiad. Tao noted that attempting to ban artificial intelligence from mathematical research is “technically infeasible.” However, his proposed “analysis-required” classification has not yet been formally adopted by peer-reviewed academic journals, research grant agencies, or institutional bodies, leaving the mathematical community to navigate how automated tools interact with traditional theoretical research.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *