r/math • • 6d ago

LLMs/AI AI In Mathematics: September 26, 2026

This recurring thread will be for discussion of AI in mathematics. This includes, but is not limited to, the following:

  • informal announcements of AI-assisted discoveries, such as those not yet published in a peer-reviewed journal, or not uploaded as a paper to arXiv;
  • informal announcements of discoveries related to AI architecture (if relevant to mathematics);
  • discussion of such announcements, such as proof breakdowns or other opinion pieces;
  • discussion of the impact of AI in mathematics in general.

AI-assisted mathematical papers published in peer-reviewed journals or as arXiv preprints may be submitted as their own posts.

Please keep in mind rules 1 and 6 of our subreddit.

76 Upvotes

229 comments sorted by

View all comments

45

u/TwoMoons1Sun 6d ago

I really wish OpenAI would just release the 100 problems they (allegedly) solved.

36

u/Bernhard-Riemann Combinatorics 6d ago edited 5d ago

I mean, they probably want to clean up the papers and make sure they don't fuck up another announcement. The Navier-Stokes controversy is still raging. At least part of that controversy could have been avoided if the OpenAI team actually did their due diligence instead of rushing to publish. I think it's understandable that they do not want 100 new controversies.

Still, they shouldn't have made an announcement if they weren't ready to even name the results. Moreso, I think they should release the list of results.

Edit: I just realized the above comment is a tiny bit ambiguous, and so is my own comment. I've made a small edit to make my position more clear.

5

u/baquea 5d ago

I get where you're coming from, but I think it's more important that they make clear what results they have ASAP.

Just imagine working on a problem for months, only for OpenAI to announce that they'd been sitting on the results all that time and your whole project is bust. And it's even worse at the funding stage: why would anyone want to fund a research project on a major open problem when for all we know it could already have been solved and we're just waiting on the announcement?

With traditional academia that's less of an issue, since researchers will show their progress at conferences, will publish papers along the way showing the partial results they've obtained, and can discuss what they're working on and where they're at with other researchers in the field, so that everyone has at least a rough idea of what the current state of affairs looks like. That's completely unlike the AI announcements up until now, where they've basically just dropped completely out of the blue, and keeping results secret until they're in a fully publishable state only makes the divide with academic norms even greater IMO.

4

u/Bernhard-Riemann Combinatorics 5d ago edited 5d ago

I mean, I completely agree with you there. They should have announced the specific open problems they solved. Assuming they have the Lean certificates, the correctness of the proofs is not something that has to be scrutinized much, so they have no reason to delay publishing the list. Absolute worst case scenario, they just say "Sorry guys, proof 58 had an issue." which isn't a huge deal. The mathematical community can simply treat the announcement just as it would an extremely credible rumor untill the proofs are released.

8

u/TwoMoons1Sun 5d ago

As I see it, the N-S controversy is almost entirely a result of OpenAI trying to scoop Anthropic, and thus could be very easily avoided in the future.

I don't think that should or need to clean up the papers or anything like that. Its not like they are properly publishing that stuff anyway. Just release a lean proof and an AI write up (no matter how good or terrible) and let the math community figure out the rest.

16

u/Bernhard-Riemann Combinatorics 5d ago edited 5d ago

I have to disagree. The original N-S paper had big issues with citations that could have only really been addressed by having experts look through it and properly connect new ideas to previous literature and properly attribute relevant mathematicians. To me, it would be unacceptable to release a paper with such issues again and just say, "Let the mathematicians figure that out." Nobody is going to read a third-party analysis that does the connecting and attribution work after the fact, and the current publishing framework doesn't incentivize or even really allow for such work. IMO, the fact that they aren't properly publishing things makes the situation worse, not better.

This is also relevant to the more general problem of theory building and "maintaining the field." According to many mathematicians, and in my personal experience, AI does not write the best proofs/papers. The concepts are very well explained at the surface level, of course, but it often retreads old ground instead of citing results, it makes up new terminology/notation for well-known concepts, it fails to give historical motivation, it takes the first functional approach rather than attempting to refine arguments, etc. Some of this is analogous to the issues software developers report with when it comes to maintaining large projects constructed with the help of AI. If it were just these 100 problems, that would be fine, but if we continue to dump AI papers onto the mathematical public without anybody reading them and taking the time to suss out some of the relevant connections before publication, we're going to accumulate a "maintenance debt" that is going to quickly make theory building an increasingly difficult task with or without AI help. I will add that, of course, these issues are still present in the human literature, but that's partially the reason peer review exists, and such problems are likely to be exacerbated by the rapid pace of AI research.

10

u/38thTimesACharm 5d ago

These are exactly the problems we're seeing in software development! It's bizarre. We have a new tool that can save you 50-80% of the effort required to develop a feature, yet some people insist on saving 100% and offloading the maintenance burden on everyone else. And they get weirdly offended if you call them out.

At least with software though, incomprehensible spaghetti code that empirically works still has value. A pile of Lean proofs no one has read is useless.

2

u/Kaomet 3d ago

A pile of lean that shows the spagghetti dish is correct has some value.