AI solves Fermat's Last Theorem and a string of famous maths problems in a remarkable 2026 stretch
From a 78-year-old geometry puzzle to a proof nobody managed in 358 years, AI models are cracking open problems that stymied the world's best mathematicians. Here is what happened, and why researchers are also urging caution.

Key points
- In September 2026, Anthropic produced the first complete computer-checked proof of Fermat's Last Theorem using Claude, in a Lean file 13 million lines long.
- Between May and September 2026, AI models from OpenAI and Anthropic contributed to at least ten significant mathematical results, spanning number theory, geometry and probability.
- Several results, including OpenAI's claimed progress on the twin prime conjecture, have not yet been independently verified by outside mathematicians.
- An anonymous contributor using GPT-5.6 Sol solved a long-standing prime-gap problem and fully formalised the proof in Lean, a computer language used to check mathematical arguments line by line.
- Researchers and mathematicians warn that some AI-assisted proofs have been published on corporate websites or social media without peer review.
For most of human history, proving a hard mathematical theorem took years of lonely work. Starting in mid-2026, AI models began doing it in days.
The headline result came in September 2026, when Anthropic announced that its AI assistant Claude had helped produce a complete, computer-checked proof of Fermat's Last Theorem. Fermat's Last Theorem is one of the most famous problems in all of mathematics: the claim, first written in 1637, that no three whole numbers can satisfy a particular simple-looking equation involving cubes or higher powers. Andrew Wiles finally proved it in 1994 after seven years of work, but his proof was never fully checked by computer. Anthropic's version, produced in 11 days and written in a proof-checking language called Lean, is 13 million lines long, making it the largest formally verified proof ever written.
That was not an isolated event. It capped roughly four months of striking results.
What else happened between May and September 2026?
The run of results was wide-ranging. Each involved AI models doing work that previously required expert mathematicians working for months or years.
| Date | Result | Model involved |
|---|---|---|
| May 2026 | Disproof of the unit distance problem | OpenAI internal model |
| July 2026 | Counterexample to the Jacobian conjecture (dimension 3) | Claude Fable 5 |
| August 2026 | Riemann zeta zeros on critical line raised from 41.6% to 67.2% | Anthropic unreleased Claude |
| August 2026 | Twin prime gap narrowed from 246 to 186 (unverified) | OpenAI internal model |
| September 2026 | Full formal proof of Fermat's Last Theorem | Claude |
Mathematician Levent Alpöge worked with Claude Fable 5 in July to find a counterexample to the Jacobian conjecture, a problem in algebra that had stood for 87 years. In August, Alpöge separately claimed Claude had helped him prove that a certain six-dimensional mathematical sphere admits a complex structure, which would resolve the Hopf problem, an open question since 1948. That result has not yet been checked by outside experts.
An anonymous contributor, posting under the name DottedCalculator and crediting authorship largely to GPT-5.6 Sol, uploaded a paper to GitHub that improves on a known result about gaps between prime numbers. Mathematician Ben Green noted the new approach, while correct, uses ideas elementary enough that they could have been discovered in the 1960s.
Should readers trust these results?
Some results are on stronger ground than others. Proofs formalised in Lean, a proof-checking program that verifies every logical step automatically, carry more certainty than results published only as written papers or blog posts.
Fermat's Last Theorem and the dying percolation conjecture both come with Lean formalisations. The twin prime conjecture progress and the Hopf problem claim do not yet have independent verification.
Critics point to one cautionary case. OpenAI announced ten mathematical advances in August 2026, including a claimed disproof of Connes's rigidity conjecture. Mathematicians later showed that result was wrong. Neither the original claim nor the refutation has cleared peer review.
A broader concern, noted by researchers studying AI-generated proofs, is that companies and individuals often publish results on corporate websites or social media without sharing the prompts they used, the exact model version, or the workflow that produced the proof. That makes it hard for other mathematicians to check the work or reproduce it.
Common questions
Does this mean AI has replaced mathematicians?
No. Every result listed here involved human mathematicians directing the work, checking the output, or building on what the AI produced. The tools are powerful, but a person still has to ask the right question.
What is Lean, and why does it matter?
Lean is a programming language that checks mathematical proofs step by step, the way a spell-checker checks words. A proof that has been formalised in Lean is as close to definitely correct as mathematics can get, which is why it matters when Anthropic's Fermat proof carries that stamp.
How can I follow which results have been verified?
The arxiv.org preprint server is where most formal papers land before peer review. Results published only on company blogs or GitHub, without an accompanying arxiv paper or Lean file, should be treated as unverified claims until independent mathematicians weigh in.



