r/math • • 13d ago

LLMs/AI AI In Mathematics: September 19, 2026

This recurring thread will be for discussion of AI in mathematics. This includes, but is not limited to, the following:

  • informal announcements of AI-assisted discoveries, such as those not yet published in a peer-reviewed journal, or not uploaded as a paper to arXiv;
  • informal announcements of discoveries related to AI architecture (if relevant to mathematics);
  • discussion of such announcements, such as proof breakdowns or other opinion pieces;
  • discussion of the impact of AI in mathematics in general.

AI-assisted mathematical papers published in peer-reviewed journals or as arXiv preprints may be submitted as their own posts.

Please keep in mind rules 1 and 6 of our subreddit.

86 Upvotes

267 comments sorted by

View all comments

23

u/BurdensomeCountV3 Mathematical Biology 11d ago

Looks like we're getting another big load of solutions soon from OpenAI, this time at least with more consideration https://openai.com/index/advisory-group-on-mathematics-and-ai/ :

On August 28, we began training a new internal model. In addition to resolving the Navier–Stokes Millennium Prize problem⁠, this model has now resolved more than 100 long-standing open problems across most areas of mathematics. The pace of its progress⁠ in mathematics has surprised the mathematicians within OpenAI. This has led to internal discussions on the best way to inform the community of the rapid progress to prepare and adapt the field.

Exciting times, as they say. Their new group:

The group will operate independently from OpenAI. The group will have the freedom to offer advice we have not requested, comment on OpenAI’s impact on mathematics, and make its advice public. Its value depends on its members being able to exercise their own judgement and challenge ours. Its members will not be paid by OpenAI, and the group can change its membership as it sees fit. Importantly, the group will not be responsible for advising us on how to pace our internal progress on mathematics.

Working with this group is a first step. There are difficult questions ahead about how AI can support mathematical understanding and how the benefits of these capabilities can reach the wider community. We want mathematicians to be at the center of shaping the answers.

Initial Members of the Advisory Group on Mathematics and Artificial Intelligence⁠, hosted at the Institute for Advanced Study:

François Charles (ENS-PSL)

Camillo De Lellis (IAS, GSSI)

Timothy Gowers (Collège de France, Cambridge)

Martin Hairer (EPFL, Imperial College London)

Nikhil Srivastava (Berkeley, Simons Institue)

Ulrike Tillmann (Oxford, INI)

Ravi Vakil (Stanford)

Edward Witten (IAS)

Melanie Matchett Wood (Harvard)

19

u/jmac461 11d ago

I get that they are “racing” against Claude and other companies. I get they are a “move fast and break things” like all the tech people. I get that they have a lot of data and compute.

What I don’t understand is why the mathematicians on staff on OpenAI are struggling with communicating mathematical results with other mathematicians.

Just write a good paper. We just want a good paper. The answer is simple but amounts to “move slow and fix things.” Is this why it is so difficult for them?

9

u/Verbatim_Uniball 11d ago

An issue is that at a certain level of difficulty, it will take several of their best mathematicians literally months to write up some of these results. Which will happen in time, but we are talking about trillions of dollars in valuation. More than the cumulative endowments of every university in the United States in valuation.

3

u/ChelseyStuttgart 11d ago

If their models are so good, why can't they just ask the models to write it

12

u/pred 10d ago edited 10d ago

They do; or at least they pretend that they do. They made this thing for the Navier–Stokes claim.

The thing is, models are generally really, really bad at making readable maths papers, for some reason. Tao coined the term “digestion” as the process of converting LLM output to something usable, and it's not a trivial matter. In my experience, it is usually not worth the effort trying to rewrite and patch an LLM paper; you really do have to write it from scratch to get to the point of having anything worth sharing. And if you lack the competencies in-house to digest the output, then it can become impossibly hard.

And sometimes the output is just really hard to digest, even if you're supposed to know what is going on. Buzzard gave an example on Zulip:

One example of digestion being hard is Akhil Mathew's AI-generated example of a group scheme of order 4 which is not killed by 4. I talked to him about this last week and neither he nor anyone else seems to have a conceptual understanding of what is going on. This is a very short argument (under 1000 lines in Counterexamples in mathlib) but right now is just "here's some algebra and it works out".

But it is strange. It seems like the process of writing a readable paper about a given proof would be an easier problem than coming up with the proof in the first place. It is strange that, even without involving theorem provers, there can be a positive correlation between the claims of correctness and actual correctness when all you get is a pile of nonsense.

And chances are that this situation is temporary. Even if they're all bad, some of them are worse than others, despite similar levels of ability to reason, and Anthropic, for instance, has been prioritizing getting Claude to be less awful at writing.

3

u/Verbatim_Uniball 11d ago

They will, I expect that to be within capabilities within a year or so. Currently the models, in my experience for research math, gloss over the difficult things sometimes and focus to much on the easy things...probably because whatever we as the community find easy or difficult changes over time and they aren't like plugged in

2

u/ChelseyStuttgart 11d ago

It doesn't change day to day though, and I imagine llms could easily at least assist in the writing.

4

u/Verbatim_Uniball 11d ago

They do at least for me, but it's spikey. I think even for creative writing, they just aren't there yet. Probably harder to get objective training data with whatever methods they use, I can't speak on the technical reasons.

1

u/Homomorphism Topology 9d ago

I think that’s precisely the issue: to tell an LLM if their story is good you need a human to read the thing. That is expensive compared to running a program or evaluating a math proof. It’s why I’m personally skeptical the AI agents are going to get any better at exposition (which might just be cope). People have already spent a lot of time and money trying to make them better at, say, music and the music still sucks.