Paper: The Gold Rush in AI4Math: Where Are We Now?
Cecile G. Tamura's post:
The Gold Rush in AI4Math: Where Are We Now?
Jiashun Jin, Zheng Tracy Ke, Bingcheng Sui
Carnegie Mellon University, Harvard University 2026
Is substantive AI use in math research increasing? Which subfields are most affected? What problems have been solved? Who is writing them—and which AI models are they using? An empirical study of 32,944 math arXiv submissions from Mar. 1–Aug. 20, 2026.
The AI Math "Gold Rush": A Reality Check
Lately, the mathematics world has been buzzing with a mix of excitement and anxiety over artificial intelligence. Some see AI as a powerful new tool for groundbreaking discoveries, while others worry it could undermine traditional research. But until now, this debate has been driven more by speculation than by hard data.
To find out what’s *actually* happening, researchers took a deep dive into the data, analyzing nearly 33,000 mathematics papers posted to the preprint server arXiv over a six-month period in 2026. They were looking for papers where authors explicitly admitted to using AI in meaningful, substantive ways.
What they found paints a picture of a genuine "gold rush" that is accelerating rapidly, but remains highly concentrated:
* A Meteoric Rise: AI adoption is exploding. In March, only about 1.4% of math papers disclosed meaningful AI use. By mid-August, that number had skyrocketed to over 14%.
* Winning the Lottery on Hard Problems: AI isn’t just being used for basic formatting or simple calculations. Researchers are actively applying it to genuine, unsolved mathematical mysteries. Strikingly, authors reported that they were able to fully resolve 71% of the open problems they tackled with AI, mostly by successfully proving new conjectures.
* A Concentrated Effort: This AI-driven research isn’t spread evenly. It is heavily clustered in specific branches of mathematics (like Combinatorics and Metric Geometry). Geographically, the United States and China dominate the field, accounting for roughly two-thirds of this research.
* Big Tech’s Footprint: The tools powering this shift are also concentrated, with OpenAI’s systems being the most frequently used, followed by Anthropic.
The integration of AI into mathematical research is no longer just a hypothetical future—it is happening right now, and at a breathtaking pace. However, the "gold rush" is still in its early, uneven stages. As these tools become more powerful and widespread, the mathematical community will need to navigate both the thrilling new opportunities for discovery and the valid concerns about how this technology will reshape the future of the field.
1. Donald E. Knuth. Claude’s cycles, 2026. Stanford Computer Science Department.
AI Overview
Donald E. Knuth published a short paper titled "Claude's Cycles" in early 2026 after Anthropic's Claude Opus 4.6 AI model successfully solved an open graph theory and combinatorics problem he had been working on for weeks. [1, 2]
What is the Paper About?
- The Problem: Knuth was wrestling with a directed Hamiltonian cycle decomposition challenge for a future volume of The Art of Computer Programming involving a m × m × m grid graph where vertices must be split into three separate cycles. [1, 2, 3]
- The AI's Role: Collaborator Philip Stappers guided Claude Opus 4.6 through 31 systematic explorations over the course of about an hour until the AI produced a working concrete construction and Python program for odd values of m. [1]
Significance
- The Reaction: Knuth famously opened his paper with the words "Shock! Shock!" and noted that the experience forced him to revise his skepticism toward generative AI.
Paper: Claude’s Cycles
---
2. Noga Alon, Thomas F. Bloom, W. T. Gowers, Daniel Litt, Will Sawin, Arul Shankar, Jacob Tsimerman,
Victor Wang, and Melanie Matchett Wood. Remarks on the disproof of the unit distance conjecture, 2026.
AI Overview
The paper titled "Remarks on the disproof of the unit distance conjecture" is an influential 19-page mathematical expository note published on arXiv on May 20, 2026. [1]
Co-authored by a prestigious panel of nine mathematicians—Noga Alon, Thomas F. Bloom, W. T. Gowers, Daniel Litt, Will Sawin, Arul Shankar, Jacob Tsimerman, Victor Wang, and Melanie Matchett Wood—it serves as the definitive human-verified translation and commentary on OpenAI's autonomous disproof of Paul Erdős's 80-year-old unit distance conjecture. [1, 2]
Core Purpose of the Paper
- Human Verification: The authors thoroughly reviewed and verified a 125-page computational counterexample generated by an internal OpenAI reasoning model (utilizing contributions from OpenAI researchers like Boris Alexeev and Lijie Chen). [1, 2]
- Simplification & Digestion: The 19-page document provides a "human-digested," simplified, and generalized version of the highly technical AI proof. [1, 2]
- Contextual Analysis: It maps the AI's unexpected techniques to existing human mathematical theories. [1]
The Mathematical Breakthrough
The paper confirms that the AI successfully disproved this by proving:
- Polynomial Bound: There exists an absolute constant \(\delta > 0\) such that the number of unit distances is at least \(n^{1+\delta }\) for infinitely many \(n\). [1, 2]
- Shift in Field: Instead of traditional geometric or graph-theoretic approaches, the proof uniquely utilizes algebraic number theory. It constructs configurations using infinite class field towers, Minkowski lattices, and Golod–Shafarevich theory. [1, 2, 3]
- Historical Precedents: The paper highlights that the underlying logic draws on deep arithmetic concepts that can retrospectively be attributed to Ellenberg–Venkatesh, Golod–Shafarevich, and Hajir–Maire–Ramakrishna. [1]
---
3. Mohammed Abouzaid, Nikhil Srivastava, Rachel Ward, and Lauren Williams. First proof second batch, 2026.
AI Overview
First Proof Second Batch (arXiv:2606.18119) is a June 2026 research report and project release by mathematicians Mohammed Abouzaid, Nikhil Srivastava, Rachel Ward, and Lauren Williams that tests frontier artificial intelligence models on original, unpublished research-level mathematics problems. [1, 2, 3]
Project Overview
- The Editors: Mohammed Abouzaid (Stanford), Nikhil Srivastava (UC Berkeley), Rachel Ward (UT Austin), and Lauren Williams (Harvard). [1, 2]
- The Goal: Assess whether leading AI systems can perform genuine mathematical reasoning rather than just pattern-matching against previously solved problems on the internet. [1, 2, 3]
Paper: First Proof Second Batch
---
4. Levent Alpöge and Claude Fable 5. A counterexample to the jacobian conjecture in dimension three, 2026.
Announced July 2026.
AI Overview
What is the Jacobian Conjecture?
- First proposed in 1939 by mathematician Ott-Heinrich Keller.
- It stated that if a polynomial function has a constant, non-zero Jacobian determinant, it must have a reverse function (be invertible).
The Counterexample
- Levent Alpöge found a function in three dimensions (\(C^{3}\)).
- The function has a constant Jacobian determinant of \(-2\).
- It fails to be one-to-one (injective) because it maps three different input points to the same output point.
Why This Matters
- It disproves the conjecture for every dimension greater than two.
- The original two-dimensional case remains an open problem.
- Other mathematicians quickly verified the short formula.
---
5. OpenAI. Ten advances in mathematics and theoretical computer science, 2026.
AI Overview
On August 1, 2026, OpenAI published ten major breakthroughs in mathematics and theoretical computer science achieved by an internal AI prototype named Astra. [1, 2, 3]
Key Details of the Project
- The AI Model: An unreleased version of OpenAI's model codenamed Astra did the core reasoning.
- The Cost: The compute time cost roughly $2,000 at standard API rates.
- The Output: A massive document containing ten distinct chapters or papers spanning nearly 250 pages.
Fields and Problems Solved
- Group Theory: Construction of non-sofic groups.
- Operator Algebras: A disproof of Connes's rigidity conjecture.
- Extremal Combinatorics: Solutions tackling specific Erdős problems (146, 180, and 183).
Current Status
- Peer Review: The formal human peer review process is still underway.
- Attribution: OpenAI openly stated that the AI generated the math arguments, while humans helped prepare manuscripts and check the framing.
- Public Access: You can view the full documents and technical details through the OpenAI Research portal. [1, 2, 3]
---
6. Claude. More than two thirds of the zeros of the riemann zeta function are simple and on the critical line,
2026. Anthropic.
AI Overview
In August 2026, an unreleased research version of Claude developed by Anthropic made a major breakthrough in number theory by discovering and writing a paper titled "More than two thirds of the zeros of the Riemann zeta function are simple and on the critical line". [1, 2]
The paper was published on arXiv (authored by Claude, with verification and responsibility taken by Anthropic mathematicians Levent Alpöge and Ralph Furman). [1]
Key Takeaways of the Breakthrough
- The Record Leap: Claude unconditionally proved that at least 67.25% (more than two thirds) of the non-trivial zeros of the Riemann zeta function are simple and lie on the critical line. [1, 2]
- Previous Human Record: This shattered a mathematical bottleneck that had stood since 2020, raising the known lower bound from 41.6% to 67.25%. [1, 2]
- Distinct Zeros: The proof also established that at least 83.62% (five sixths) of all non-trivial zeros are distinct. [1]
Paper: More than two thirds of the zeros of the Riemann zeta function are simple and on the critical line
-----
How many LLM were participating in IMO 2026?
AI Overview
An independent evaluation conducted by former Google engineer Deedy Das tested 7 state-of-the-art LLMs on the IMO 2026 problems, while other independent sets or corporate disclosures highlighted varying counts of participating or evaluated models. [1, 2]
Key Details on AI Evaluation at IMO 2026
- Independent Runs: Deedy Das ran an evaluation featuring 7 frontier models in his GitHub benchmark, with multiple models achieving a perfect 42/42 score. [1, 2]
- Top Performers: Models such as Claude Fable 5, GPT-5.6 Sol, and Kimi K3 achieved perfect scores in independent or self-administered harnesses. [1, 2, 3]
Article: GitHub benchmark
-----
Không có nhận xét nào:
Đăng nhận xét