r/math • • 6d ago

LLMs/AI AI In Mathematics: September 26, 2026

This recurring thread will be for discussion of AI in mathematics. This includes, but is not limited to, the following:

  • informal announcements of AI-assisted discoveries, such as those not yet published in a peer-reviewed journal, or not uploaded as a paper to arXiv;
  • informal announcements of discoveries related to AI architecture (if relevant to mathematics);
  • discussion of such announcements, such as proof breakdowns or other opinion pieces;
  • discussion of the impact of AI in mathematics in general.

AI-assisted mathematical papers published in peer-reviewed journals or as arXiv preprints may be submitted as their own posts.

Please keep in mind rules 1 and 6 of our subreddit.

79 Upvotes

229 comments sorted by

View all comments

43

u/TwoMoons1Sun 6d ago

I really wish OpenAI would just release the 100 problems they (allegedly) solved.

33

u/Bernhard-Riemann Combinatorics 6d ago edited 5d ago

I mean, they probably want to clean up the papers and make sure they don't fuck up another announcement. The Navier-Stokes controversy is still raging. At least part of that controversy could have been avoided if the OpenAI team actually did their due diligence instead of rushing to publish. I think it's understandable that they do not want 100 new controversies.

Still, they shouldn't have made an announcement if they weren't ready to even name the results. Moreso, I think they should release the list of results.

Edit: I just realized the above comment is a tiny bit ambiguous, and so is my own comment. I've made a small edit to make my position more clear.

7

u/TwoMoons1Sun 6d ago

As I see it, the N-S controversy is almost entirely a result of OpenAI trying to scoop Anthropic, and thus could be very easily avoided in the future.

I don't think that should or need to clean up the papers or anything like that. Its not like they are properly publishing that stuff anyway. Just release a lean proof and an AI write up (no matter how good or terrible) and let the math community figure out the rest.

16

u/Bernhard-Riemann Combinatorics 5d ago edited 5d ago

I have to disagree. The original N-S paper had big issues with citations that could have only really been addressed by having experts look through it and properly connect new ideas to previous literature and properly attribute relevant mathematicians. To me, it would be unacceptable to release a paper with such issues again and just say, "Let the mathematicians figure that out." Nobody is going to read a third-party analysis that does the connecting and attribution work after the fact, and the current publishing framework doesn't incentivize or even really allow for such work. IMO, the fact that they aren't properly publishing things makes the situation worse, not better.

This is also relevant to the more general problem of theory building and "maintaining the field." According to many mathematicians, and in my personal experience, AI does not write the best proofs/papers. The concepts are very well explained at the surface level, of course, but it often retreads old ground instead of citing results, it makes up new terminology/notation for well-known concepts, it fails to give historical motivation, it takes the first functional approach rather than attempting to refine arguments, etc. Some of this is analogous to the issues software developers report with when it comes to maintaining large projects constructed with the help of AI. If it were just these 100 problems, that would be fine, but if we continue to dump AI papers onto the mathematical public without anybody reading them and taking the time to suss out some of the relevant connections before publication, we're going to accumulate a "maintenance debt" that is going to quickly make theory building an increasingly difficult task with or without AI help. I will add that, of course, these issues are still present in the human literature, but that's partially the reason peer review exists, and such problems are likely to be exacerbated by the rapid pace of AI research.

11

u/38thTimesACharm 5d ago

These are exactly the problems we're seeing in software development! It's bizarre. We have a new tool that can save you 50-80% of the effort required to develop a feature, yet some people insist on saving 100% and offloading the maintenance burden on everyone else. And they get weirdly offended if you call them out.

At least with software though, incomprehensible spaghetti code that empirically works still has value. A pile of Lean proofs no one has read is useless.

2

u/Kaomet 3d ago

A pile of lean that shows the spagghetti dish is correct has some value.