In an hour of back and forth, Daniel Litt got OpenAI's Codex to turn five recent conjectures in algebraic geometry into three papers that were bad but correct.
Every new AI proof is written up as another step toward machine mathematicians. Litt, who works in algebraic geometry at the University of Toronto, said the models are extremely good at applying techniques that already exist and still weak at the parts of the job he cares about most: building a theory, and working out which question is worth asking at all.
"The goal of mathematics is not to produce mathematics papers. It's to produce some kind of understanding. Maybe some of that understanding resides in model weights. To me, that's pretty unsatisfying."
Litt has published a paper whose lemmas were improved by working with a frontier model, runs his own harness on top of Codex, and has been public enough about his changing view of AI in mathematics that the host said checking in two weeks apart produces a different answer.
I listened to the full interview so you can skip it. 64 minutes of audio, 21 minutes of reading.
Here are the 15 takeaways that matter.
👤 Guest: Daniel Litt, assistant professor of mathematics at the University of Toronto, who works in algebraic geometry and has one published paper in which AI models proved lemmas alongside him
🎙️ Host: Lisha Li, an infrastructure partner at a16z who says she no longer has the time to practice mathematics herself
📰 Published: 1 September 2026 on the show's own feed and on YouTube (a16z)
🔴 YouTube | 🟣 Apple Podcasts | 🔗 Episode page | ⏱️ 1 hr 4 min | ✅ Time saved: 43 min
Key Takeaways
The models are very strong at applying known techniques and weak at choosing what to prove Litt put their output in the same band as a good career spent applying everything already known
An hour with Codex produced 3 correct papers he has not bothered to send to anyone He told it to go online, find five recent conjectures in algebraic geometry, and prove them
Claude and ChatGPT are solving almost the same collection of problems
AI proofs are short because short is what anyone can currently check An 800-page AI-generated proof posted to arXiv is one he says cannot be right
The projects he has worked on for three to five years are where AI has helped least There, he said, it is mostly a faster substitute for a web search
Failing to grind through an ugly proof is what produced a better one
The conjectures he works on are probably true, so there is no counterexample to construct
Papers carrying the same proof of the same theorem are landing on arXiv days apart He reads three, four or five of them as the model collapsing onto one path of reasoning
The rank-30 elliptic curve is a fun construction rather than a major result
He teaches his 3-year-old daughter addition inside a general group
1. The Erdős Result Stands Out
Litt opened by dividing the recent results into three kinds: the ones produced autonomously, the ones produced semi-autonomously, and the ones where the AI's contribution is not clear at all. He said he can only comment on areas where he has expertise, and that another mathematician would give different answers.
His favorite fully autonomous result is the counterexample to the Erdős unit distance problem. "So my favorite, like, fully autonomous result by an AI so far is the solution to the Erdős unit distance problem, which I think was announced in mid-May."
It was creative in a way most results are not, he said. People working in the area thought the statement was true, and a counterexample was found
The techniques came from a different field. He described them as classical ideas from the 1960s, not deep or new in themselves, but new to the study of point configurations in the plane
The test he applies is what happens afterward, and here other mathematicians used the ideas. They found counterexamples to other open questions, among them the sum-product conjecture for the real numbers
The contrast is with what he called last-mile results: deep work done by a group of human mathematicians, with the AI taking the final step
2. The Proofs Look Human
Li asked where the model differences show up, and whether the impressive results are creative or just inhuman speed at logical implication.
Litt said reinforcement learning works well on mathematics because a proof can be verified relatively cheaply compared with other domains, but that symbol-pushing misses most of what is interesting about the subject
He found the released reasoning trace recognizable rather than alien. He said it looked like what his own chain of thought might look like on the same problem, while noting the public version is a summary and may have been cleaned up
Nothing in the output reads as a machine move. "If you just look at problem model output, it's not like there's some move 37 or whatever."
The non-human parts are stamina and breadth, not style. "So the ways they might be a little bit unhuman is, like, they don't get tired, they know a lot."
He added that the informal, natural-language route rather than formal proof code suggests to him the techniques will generalize to other domains
3. Claude And ChatGPT Are Level
On capability he put the two frontier families in the same place. "You asked a little bit about Claude versus ChatGPT. My sense is that they're pretty similar in terms of capabilities."
The pattern he sees is that when one lab publishes a solution, the other reports having it too. "It actually seems like they're solving a very similar collection of problems."
What they are good at is computation and synthesis across the literature — grounding out a long calculation, or pulling together technical ideas from many papers a single mathematician might not have read
Where they fall down is the part he spends his own time on. "But they're not, like, they seem like weaker in things like intuition or, like, having some big picture point of view."
He was careful not to make that dismissive. "There are mathematicians who have had great careers doing very high quality work of that flavor."
On theory building the models need to be led. "They definitely are not good at it autonomously, at least with, like, whatever scaffolding I've set up. But with some hints, you can kind of get them to do something interesting." His caveat is that hints make attribution hard, because you cannot tell which part came from the model If a model needs a hundred bits of hints today, he said, in six months it may need none
He uses ChatGPT more than Claude out of habit rather than judgment: ChatGPT was useful for research mathematics earlier, and he put the point at which Anthropic's models more or less caught up at Opus 4.5 or Opus 4.6
He finds one of the two clearer in exposition and better at judging what he already knows, and said both are poor at modeling what a reader understands
4. Problem Solver Or Theorist
Asked what he actually does, Litt separated the problem solvers from the theory builders and put himself with the problem solvers.
An open problem is a measuring instrument, not a goal. "At least for me, the point of an open problem is it's supposed to measure your failure to understand something." The method is to find the smallest case you cannot handle, fiddle with it, then turn what you used to win into a theory
The other mode is an analogy pursued for decades without a problem at the end. His own work is driven by an analogy between the cohomology of algebraic varieties and representations of fundamental groups, which he said has produced beautiful mathematics over 30 or 40 years, naming Carlos Simpson and Shinichi Mochizuki
Finding the question is itself hard, he said. A published list of open conjectures does not capture what the field actually does not know
His example of a conjecture found in data is Birch and Swinnerton-Dyer. The two computed statistics on elliptic curves in the 1960s in what he called one of the first pieces of computer-aided mathematics, graphed them, and noticed that the slope of a line matched an algebraic invariant they already knew
That process, he said, is closer to running an experiment and looking for an explanation than to deduction
5. Where AI Helps His Day Job
The work he has been doing longest is the work AI has helped least. "Actually, what I've found is that the projects that I have that kind of predate AI, like the projects I've been thinking about for three or four or five years, it's just not that useful." There he uses it as a faster substitute for a web search — discussing a related topic with the model instead of searching and reading a paper
The change is that he now takes on projects he would previously have skipped. He said he is bad at coding, so a question that needed code used to sit for six months; Li called the model a tireless PhD student
Where it earns its place is running many examples at once. Asked to work through a thousand examples, it splits the job up: "It would be 10 examples at a time, 10 different sub-agents."
He described those as different activities layered on top of what he was doing before, rather than a speed-up of it
6. Beauty Is Not The Goal
Li asked what drives the choice of what to study in a field with no external motivation for it.
Litt said he tries not to be guided by beauty, which he called a controversial position. "I don't really have aesthetic feelings." The failure mode he sees in young mathematicians is abandoning a proof because it feels ugly, when the proof may be right and the feeling wrong
Asked what ugly means, he offered only that it may involve a lot of grinding, or a calculation that does not illuminate anything
His rule is to finish. "But you should like win by any means necessary in my opinion."
The frame he substitutes for beauty is scientific rather than artistic. "Like I like to think of what I'm doing like doing kind of physics except with concepts." He asks instead what is fundamental and what will open up the most further understanding
He allowed that plenty of mathematicians think of themselves as closer to poets, and said progress in the field comes from many people following their own curiosity rather than a single agreed direction
Li brought in a point she attributed to Thurston in the 1970s: which structures mathematics ends up studying is a sociological phenomenon, driven by what enough people find interesting
7. Why His Problems Resist AI
The reason is not difficulty, it is the shape of the answer. "I think one reason the models might not be useful for some of these things is like the conjectures are true." The unit distance problem was believed true and turned out false, so there was a specific construction to find The conjectures he works on sit inside a broad framework that gives a lot of evidence they hold, which means there is nothing to construct
He said resolving them probably requires genuinely new techniques rather than very strong application of existing ones, while allowing that a clever construction might avoid the need for a big new idea
He was explicit that this is not a prediction of a ceiling. "I'm just being clear, like I'm not saying the models won't be able to do this."
On the trajectory he is not a skeptic. "So at this point, my expectation is just like that the trajectory will continue upwards. I'm not a skeptic of continued capabilities growth." He said his guess is that theory building is "probably totally doable and it just hasn't been done yet", perhaps needing a different reinforcement-learning environment
The hard part is the reward. A finished theorem can be rewarded at the end; the intermediate steps of building understanding of a poorly understood object are much harder to score His comparison is the human version: a PhD advisor saying an idea is good, which he described as an element of human taste
Two open problems he named in this stretch came through the captions garbled beyond recognition and are left out here
8. Failing To Grind Paid Off
Litt has one paper so far in which the models were useful, and the story runs against the idea that more grinding power is straightforwardly better.
The models proved a couple of lemmas in it, and he worked with Gemini Deep Think, which he said was on the frontier at the time and no longer is
None of the frontier models could prove the lemma he wanted. So he worked out a large number of examples himself, saw why it might be true, and found a better statement of it — which the models then proved very quickly
The failure was the useful part. "So like our inability to prove it led to an improvement in the result."
Fed the original lemma, he said, ChatGPT 5.6 Pro returns "the worst proof you've ever seen, like 10 pages of like just brutal calculation with no insight whatsoever." That proof would have been perfectly valid and would have produced no understanding
He noted he was being hypocritical about ugly proofs: he could see the horrible grinding proof would work and could not bring himself to do it, which is why he looked for another argument
Li's reading was that the drive is still aesthetic, or at least a wish to understand; Litt's counter is that compression is a weak measure of interest, though a fair anthropological account of why theory building happens
9. Playing The Slot Machine
Asked how the mathematics community should adapt, Litt went straight to incentives, and said the current ones point the wrong way.
The job market rewards volume, and volume is now cheap. "The best way to do that is, like, you want to produce a lot of papers, which maybe prove, old conjectures or whatever, and you can do that by playing the slot machine until, the model produces a hopefully correct proof of such a result." He pointed out you do not even have to pick the theorem in advance
He ran the experiment himself. "You can say, go online and find five recent conjectures in algebraic geometry and prove them." The result: "And, okay, I've run this experiment and with some back and forth, I was able to, in an hour, get like three, quite bad papers, but correct papers." They are on his hard drive, unsent, because emailing the relevant people is not worth his time
The duplicates are the tell. "Like sometimes, we've seen examples where like three or four or five papers with the exact same proof of the exact same theorem out within a couple days of each other" He and Li read that as mode collapse: the model consistently finding the same path of reasoning
The volume of postings has risen sharply, and while some of it is interesting, much of it shows no sign that a human engaged with the result
Li's addition was that this is a problem for the labs, not only for mathematics: if the proofs converge because the models draw on the same literature, it is unclear where a more diverse set of intuitions would come from, since today's come from practicing mathematicians
10. Why Humans Still Matter
Li pushed the case to its limit: suppose the models become robustly superhuman and add all the cognitive diversity themselves.
Litt's answer was about who decides, not about capability. "It's like, despite something being optimal doesn't mean you do it." Handing mathematics research to the models gives no guarantee they pursue the wide range of directions that turns out to be valuable; instructed to make life better, they may take the direct path
His stake is in broad, undirected research. "So, I don't know, if you believe that there is any value in sort of this broad-based, like, fundamental research, which I do, look, I think that's one of the most valuable things, humans or whatever, or the models can be doing."
The way to guarantee it happens, he said, is to keep a community with broad interests directing the models, on the assumption that humans keep control of where things go
The pipeline argument is the practical one. "Like you need an entire mathematical community to support a small group of people who are on the frontier." Frontier work depends on a large base of people learning to think mathematically, and that base has to be given a reason to invest the years it takes
He said the appeal of the models is real — they may answer questions that have kept him up at night, and he would get to learn the answers — but understanding that sits only in model weights does not satisfy the reason he does the work
11. Education And Human Capital
The classroom effect he hears about is a split, not an average. "You know, I've heard people start talking about a bimodal distribution in their classes where there's some people who are really like figuring out how to take advantage of new tools and other people who are just like letting them do their homework and then bombing everything else."
His general worry is a technology that is cheaper and slightly worse. "And so you get a lot of suddenly a lot of, like, low quality outputs that are displacing previous high quality outputs." He said the models let people do a lot more things more cheaply, but it is not clear they are improving the quality of what comes out
The remedy is institutional rather than technical. "But I think it's possible to use the tools in a way that actually, like, improves, the quality among, along every dimension. It just requires some thoughtfulness and some redesign of institutions."
Li's version of the risk is that it is easy to hand thinking over to a model that cannot yet think at the level being handed over, and that the harder path is using the tools to deepen understanding rather than replace it. She also noted that entry-level roles are the hardest to adapt
Litt said mathematics is a test case for every profession, not a special case. He thinks a great deal of computer-based work is already being done by models without a public reckoning, and that mathematics is reckoning publicly because model capability in mathematics is useful for the labs to talk about
Li said she had heard that primary schools in Budapest teach group theory, and argued that more accessible explanation is a reason to spread that further rather than less
He also refused to treat the flood of amateur work as a loss. "But to me, it seems like just the fact that there's lots of people excited about math is like kind of a positive." He noted that a good deal of the low-quality output comes from professionals, where the incentive to publish volume already existed "It's like suddenly there's a spotlight on it and I can nerd out about math more."
12. The Rank-30 Curve
Li raised the rank-30 elliptic curve announced the day before the recording.
Litt said no details had been released. His understanding was that it came from Claude, prompted by Levent Alpoge and a collaborator whose name he could not recall, and he later suggested Ava Howell as the likely third name
Without the method, he said, the significance cannot be judged. "What I would say is that if you want to understand kind of historically how such results have been proved, or sorry, have been understood by the community, they're like cool, but I wouldn't say they're like a big deal."
His placement of it was deliberately modest. "So like the typical place of result like this might go is like someone's website of records." Noam Elkies and Zev Klagsbrun have been pushing the record for years and recently found a rank-29 example
The general rule he applies to any model result is that "you cannot evaluate it except in retrospect" — and he said the same holds for human mathematics, where a problem thought to need deep new ideas sometimes does not
On disclosure: Alpoge posts his results on X and has been slowly releasing write-ups, which Litt read as someone having fun on the internet rather than a policy of secrecy
13. Short Proofs Are Checkable
Li said two OpenAI researchers had told her the most delightful thing about the recent proofs was their length.
Her framing was that the proofs are short rather than sprawling. "So I had Mark Sellke and Mehtaab Sawhney on from OpenAI recently, and they were saying how, what's kind of been the most charming or delightful is just that the proofs have been relatively short."
Litt's explanation was a constraint, not a preference. "I think my sense is that the reason they're not producing long, complicated proofs is that they cannot." The reason: "Just like the ability to check correctness is not yet there."
Ask a model whether the short proof it just produced is correct and it will often say no. He said they are much more reliable than six months ago and still sometimes produce things that are wrong and known to be wrong The danger with a long proof is that the model may not know it is wrong
He assumes the labs are sitting on results they cannot verify. A recent list of 10 problems released by OpenAI was formalized in Lean, which he called good evidence the proofs are true; others could not be formalized because the prerequisites are not in Mathlib yet, and longer ones are harder to check
What an unverifiable long proof looks like in public is already on arXiv. "So, for example, someone recently posted a claimed proof of resolution of singularities and positive characteristic, which was 800 AI-generated pages." He has not read it and said it cannot be right, because a correct version would be a major result and is beyond what current models can do
14. Harnesses Trade Off Rigor
Getting length out of a model requires a harness, and the harness costs reliability. "And when you make a harness whose goal is to elicit a proof, I think it often decreases reliability because you're just trying to produce output." He said the strongest models try hard not to assert things they cannot support, and pushing them to be creative means pushing them out of that rut
The length of a paper is therefore evidence about the harness. "So, I think any harness that can elicit a 250-page paper is probably not being very, careful about what is producing." He tied this to people trying to arbitrage the prestige mechanics of academic mathematics by eliciting proofs nobody checks
Humans cannot check a 250-page paper line by line either. What they do instead is understand the global structure of the argument and stress test it — does it prove something known to be false, does it work in a special case — and he said the models are not yet good at that kind of fuzzy testing
His standing test is a wrong paper the models still cannot break. A paper published a couple of years ago claimed a big result close to his own area; the specific error was very hard to find, and he and other experts knew immediately it was wrong because the argument proved something too strong to be true. "But it was also very clear from the structure of the argument that it couldn't work."
Li's parallel from software was a report from Cursor on a long-horizon test, reproducing SQLite in Rust, which she said needed a harness rather than raw models and showed how far off the equivalent is
He has built a harness and mostly leaves it alone. "So I mostly do not use the harness. I mostly try to use it to help me understand stuff." He said he does not enjoy autonomous mathematics
15. Teaching A 3-Year-Old Math
His daughter is three and has never used AI. She is starting to add: "She can add single digit numbers, like, by counting on her fingers." He said she counts to about 30 reliably and 50 semi-reliably
The teaching is not remedial arithmetic. He said that when he teaches her addition and subtraction, they do it in the context of a general group
He described waking her up to find her hiding under the blankets announcing that she was doing math, and thinks she is interested because she can tell he is
Li's daughter, a year younger, learned the word icosahedron at about one from a toy her grandparents gave her, and went on to learn the Platonic solids
His argument for teaching mathematics does not depend on what AI can do. "I think the reason to learn math has always been like to think clearly and like better understand the world." He expects the world to look very different in 20 years and thinks that reason survives it, and said the same case applies to reading widely and to the humanities
His closing suggestion was that mathematics graduate students could focus on training a much younger generation to use AI to get better at mathematics rather than to skip the understanding
Bonus Insights
Li framed the ladder of increasingly difficult conjectures as a curriculum that serves humans and models alike, and asked why humans are good at formulating structures in mathematics and physics at all
Litt said mathematicians do informal mathematics rather than writing out long formal strings of symbols precisely because the informal version is where non-rigorous understanding comes from
On the model comparison, he said both frontier systems are poor at judging what a reader already knows, and that one of them alternates between explaining something trivial and assuming he knows the rest
Litt's bottom line is that AI is now genuinely good at the part of mathematics that applies existing technique, and that the field's problem is an incentive structure which rewards exactly that output while the understanding it was supposed to measure goes unproduced.
Products, Companies & Tools Mentioned
ChatGPT and Claude (The two he compares directly; he says they are solving a very similar collection of problems, and he uses ChatGPT more out of habit)
OpenAI and Anthropic (The labs behind them; he puts the point at which Anthropic's models caught up on research mathematics at Opus 4.5 or Opus 4.6)
Codex (OpenAI's coding agent, which he pointed at five recent algebraic geometry conjectures and got three correct papers from in an hour; he also runs his own harness on it)
Gemini Deep Think (The model he worked with on the lemmas in his one published AI-assisted paper, which he says was on the frontier at the time)
Lean and Mathlib (The proof assistant and its mathematics library; OpenAI's recent 10 problems were formalized in Lean, and he says others could not be because the prerequisites are not in Mathlib yet)
arXiv (Where the uptick in machine-produced papers is showing up, including an 800-page claimed proof he says cannot be right)
Cursor (Li's software parallel: its published long-horizon test, reproducing SQLite in Rust, needed a harness rather than raw models)
Books & Resources Mentioned
The Birch and Swinnerton-Dyer conjecture (His example of a conjecture found in data — statistics on elliptic curves computed in the 1960s, then graphed)
Daniel Litt on X (The account the show's notes point to, and where his changing view of AI in mathematics has been posted)
Lisha Li on X (The host's account, also listed in the episode notes)
The Erdős unit distance problem (The counterexample he calls his favorite fully autonomous AI result, announced in mid-May)
The sum-product conjecture for the real numbers (One of the open questions other mathematicians then refuted using the same techniques)
If this was worth your time, send it to someone closer to the industry than you are.
Get the latest market chatter as it happens:

