Two articles from late April and early May 2026 about AI in mathematics are best read together. Tim Gowers, a Cambridge mathematician and Fields medalist, wrote a detailed blog post describing how ChatGPT 5.5 Pro produced what he called “PhD-level research” in about an hour, with “no serious mathematical input” from him. Terence Tao, a UCLA mathematician and Fields medalist, said in a Nature interview called “The job description is changing” that mathematics now has to reconsider basic questions: what counts as a proof, what counts as a paper, and what the profession is for. As Tao said, “if we don’t ask these questions ourselves, then they will get answered for us by a technology company or decided by financial incentives.” What interests me is how much expert judgment surrounds the model’s hour of work, and who will supply that judgment if the volume of results keeps growing.
A case study, not a sample
Gowers explained his methods, asked Isaac Rajagopal—the MIT student whose earlier paper the model built on—to check the result, and was explicit that he considered the output non-trivial. The model took a result from Rajagopal’s recent paper and pushed it further. Rajagopal described the main technique as “completely original” and said it was the kind of idea he would have been proud to come up with after a week or two of thinking. Gowers’s own view is cautious: he calls the result “a perfectly reasonable chapter in a combinatorics PhD,” not a major breakthrough, but “definitely a non-trivial extension.”
There isn’t a matching blog post called “ChatGPT 5.5 Pro spent an hour producing a confident, plausible, but subtly wrong proof of a small open problem and I almost believed it.” Cases like that almost certainly exist, but we don’t see them because people rarely write up failures with the same care. This is the same issue with every viral “AI did X” story: we’re looking at the right tail of a distribution whose shape we’ve barely begun to measure.
We don’t know how many similar problems the model would fail on, or how much the success depended on how the problem was presented. One example can’t show us how often fluent proofs would fall apart under close review, or whether the problem or its main technique is similar to anything in the training data.
The jagged frontier
FrontierMath is a benchmark of original research-level problems, put together with input from Fields medalists — Gowers among them. When it launched in late 2024, the best models solved under two percent. By early May 2026, several frontier models, GPT-5.5 Pro included, score above fifty percent on its main problem set. The numbers come from BenchLM’s leaderboard, a secondary tracker but a useful snapshot.
DeepMind’s math agent Aletheia got six of ten curated research-level problems right. On seven hundred open problems from Thomas Bloom’s online database of Erdős conjectures, it solved four on its own. That’s about sixty percent on the curated set, but less than one percent on the open problems.
Dell’Acqua and colleagues found a similarly uneven pattern in a recent Organization Science study with 758 Boston Consulting Group consultants. AI help improved results on tasks within its strengths, but on tasks just outside its abilities, consultants using AI were nineteen percent less likely to get the right answer.
The people behind the result
In Gowers’s case, a person posed the problem, the model built on earlier published work with the original author involved, the AI generated arguments, and people checked and evaluated the outcome. Calling this “AI-produced”—as one popular headline said, “with zero human help”—oversimplifies what actually happened. A disclosure should identify who posed the problem, what the model supplied, and who checked the result.
Gowers suggests that maybe nobody needs credit in the usual sense for an AI-assisted result. I think that’s too hasty. Credit records whose prior work contributed to the answer. And if a proof is wrong or a claim of novelty is overstated, readers need a named person who will answer for it.
Checking proofs takes expert time, and that doesn’t scale as quickly as AI can generate new work. If thousands of AI-assisted papers start showing up on arXiv, the main challenge will shift from creating proofs to checking them.
Equity
Gowers says he was “fortunate to have been given access” to ChatGPT 5.5 Pro before it was widely available. Now, the model is only available through ChatGPT’s Pro, Business, and Enterprise plans, or a separate paid API. Other AI labs also have internal tools that only some researchers can use. If those tools become necessary for research, their price will affect who can produce publishable work, and therefore who gets admitted and hired. One early commenter on Gowers’s post brought this up directly. Gowers replied that this was “potentially a very bad aspect” of the current situation and suggested some ways to address it.
Learning to do the work
In the Nature interview, Tao says graduate students who avoid using AI may be at a disadvantage. Gowers adds that people who have solved hard problems themselves are usually better at using AI for them, “just as very good coders are better at vibe coding than not such good coders.” If graduate students use AI to skip the slow process of making mistakes and learning from them, the field might see a short-term boost in productivity but lose the deep expertise needed to guide and check AI’s work. Graduate programs will have to decide which work students should learn to do before delegating it to a model.
In his 1994 essay “On Proof and Progress in Mathematics,” William Thurston said the real purpose of mathematics is human understanding, not just proving things. For Thurston, the goal is to know why something is true, in a way that others can also understand. Francis Su, in Mathematics for Human Flourishing, argues that doing math builds patience, focus, and persistence.
If a graduate student uses AI to skip a problem they would have spent a month struggling with, they get the answer but lose a month of the work that helps make them a mathematician.
Whose institutions?
Tao puts the main point clearly: the rules for AI-assisted math are being set right now. DeepMind, when talking about Aletheia, has already suggested a five-level system for classifying AI-assisted math results, from “negligible novelty” to “landmark breakthrough,” combined with three levels of AI involvement. DeepMind says the system came from “extensive discussions with the mathematical community.” It remains the company’s proposal. Mathematical societies and journals should decide whether to adopt it.
Journals need to say how authors must disclose AI use and who answers for a proof. Departments need to decide how students will learn to check the work, and how researchers without expensive accounts will get access. Those decisions require mathematicians’ time and judgment, too.
This is the kind of data-driven justice work I do in my book Unlocking Justice, now available from Princeton University Press.