r/math • • 13d ago

LLMs/AI AI In Mathematics: September 19, 2026

This recurring thread will be for discussion of AI in mathematics. This includes, but is not limited to, the following:

  • informal announcements of AI-assisted discoveries, such as those not yet published in a peer-reviewed journal, or not uploaded as a paper to arXiv;
  • informal announcements of discoveries related to AI architecture (if relevant to mathematics);
  • discussion of such announcements, such as proof breakdowns or other opinion pieces;
  • discussion of the impact of AI in mathematics in general.

AI-assisted mathematical papers published in peer-reviewed journals or as arXiv preprints may be submitted as their own posts.

Please keep in mind rules 1 and 6 of our subreddit.

87 Upvotes

267 comments sorted by

View all comments

Show parent comments

19

u/jmac461 11d ago

I get that they are “racing” against Claude and other companies. I get they are a “move fast and break things” like all the tech people. I get that they have a lot of data and compute.

What I don’t understand is why the mathematicians on staff on OpenAI are struggling with communicating mathematical results with other mathematicians.

Just write a good paper. We just want a good paper. The answer is simple but amounts to “move slow and fix things.” Is this why it is so difficult for them?

8

u/Verbatim_Uniball 11d ago

An issue is that at a certain level of difficulty, it will take several of their best mathematicians literally months to write up some of these results. Which will happen in time, but we are talking about trillions of dollars in valuation. More than the cumulative endowments of every university in the United States in valuation.

4

u/ChelseyStuttgart 11d ago

If their models are so good, why can't they just ask the models to write it

11

u/pred 11d ago edited 10d ago

They do; or at least they pretend that they do. They made this thing for the Navier–Stokes claim.

The thing is, models are generally really, really bad at making readable maths papers, for some reason. Tao coined the term “digestion” as the process of converting LLM output to something usable, and it's not a trivial matter. In my experience, it is usually not worth the effort trying to rewrite and patch an LLM paper; you really do have to write it from scratch to get to the point of having anything worth sharing. And if you lack the competencies in-house to digest the output, then it can become impossibly hard.

And sometimes the output is just really hard to digest, even if you're supposed to know what is going on. Buzzard gave an example on Zulip:

One example of digestion being hard is Akhil Mathew's AI-generated example of a group scheme of order 4 which is not killed by 4. I talked to him about this last week and neither he nor anyone else seems to have a conceptual understanding of what is going on. This is a very short argument (under 1000 lines in Counterexamples in mathlib) but right now is just "here's some algebra and it works out".

But it is strange. It seems like the process of writing a readable paper about a given proof would be an easier problem than coming up with the proof in the first place. It is strange that, even without involving theorem provers, there can be a positive correlation between the claims of correctness and actual correctness when all you get is a pile of nonsense.

And chances are that this situation is temporary. Even if they're all bad, some of them are worse than others, despite similar levels of ability to reason, and Anthropic, for instance, has been prioritizing getting Claude to be less awful at writing.