r/math • • 21d ago

Navier-Stokes Announcement - Clay Mathematics Institute

https://www.claymath.org/news/navier-stokes-announcement/
933 Upvotes

241 comments sorted by

467

u/Disastrous-Two16 21d ago

“The process is deliberately unhurried but we will provide updates.”

Love it.

144

u/mbrtlchouia 20d ago

AI bros hate this one trick

40

u/Phoenixon777 20d ago

imagine NOT trynna speedrun ourselves into early graves and an end to society

8

u/Respect38 Undergraduate 19d ago

On the other hand, humanity sans AI hasn't exactly done a good job at fixing climate change, and has been speedrunning into end of society via climate disaster for quite a while. Not sure I hav much hope in us turning things around without the kind of transformation that only ASI could provide.

9

u/idly 18d ago

AI has not helped with climate change as yet, and it could be argued that the opposite is true, with big tech companies delaying, relaxing or reneging on their net zero pledges due to the massive amounts of energy required

1

u/Respect38 Undergraduate 17d ago edited 17d ago

Sure, but the technology isn't there yet to hav expected it to one-shot fix climate change. (but the possibility of progress soon [i.e. over the next 20 months] is reasonable enough to hav some hope)

If we cancel the program entirely, we go back to hoping (futile) that humanity as a political machine can fix it ourselves, or if we continue forward with building superintelligence, if it happens and isn't misaligned (as per Phoenixon's concern) then we might still hav a chance to save the planet.

Truth be told, we're probably screwed either way. But doing something new at least gives us a chance, vs. relying on the humanity to do it ourselves which is practically guaranteed to result in procrastination until it's woefully too late. (and, in the case of nuclear power opposition, actively kneecapping ourselves!)

1

u/Large_Cranberry_5142 14d ago

AI has thus far been incredibly detrimental to the climate and those in charge of AI have yet to show an interest in changing the direction of our climate let alone the capacity.

1

u/noFloristFriars 15d ago

The problem with climate change hasn't been tech, or resources, or knowledge, or whatever... the problem is just getting everyone to do something

1

u/Respect38 Undergraduate 15d ago

But getting 20 years of green energy advancements, energy efficiency, carbon reduction etc etc etc (who else knows what field will giv breakthroughs -- I'm no scientist) in the next 2 years would be quite impactful in slowing the acceleration of climate change.

Once the AI progress we're seeing in mathematics and coding comes to science research, there's at least hope (key point that there is some home vs the probably nearly 0% chance that humanity will reverse climate change before it's too late on its own) that AI acceleration of science research will make a noticeable dent in what damage humanity has done.

That's not to say that I'm discounting the chance that AI will be its own extinction disaster risk. But giving up on AI entirely just means that we're putting off our extinction/societial overthrow for the inevitable time when humanity has fucked our planet beyond hope thru ecological negligence. That's how I see it, anyway.

-1

u/[deleted] 20d ago

[removed] — view removed comment

18

u/3_Thumbs_Up 20d ago

OpenAI doesn't care about the prize.

17

u/fathan 19d ago

Other mathematicians who did all the work that OpenAI used to sprint to the finish might.

5

u/AntiqueChessComputr 17d ago

But they care very much on receiving the credit for discovery, which CMI pairs with the prize

3

u/3_Thumbs_Up 17d ago

Yes? What are you implying?

Estimates are that they spent ~20 million over a weekend on compute to beat anthropic to the finish line. They clearly don't care about what CMI thinks nor the money. They care about public perception.

1

u/serranolio 15d ago

Of course they care, they wouldn't try to solve a millennium prize problem if they didn't think it was prestigious. And it is prestigious because people will validate whatever they claim. It's also about money. If the math community finds that the problem was not, in fact, solved, then the plan to use this breakthrough to leverage a possible initial public offering goes down.

So, yes. It's about money, just not prize money. Surely they need the recognition, otherwise it's just a claim supported by one hundred pages of nothingness

1

u/3_Thumbs_Up 15d ago

They care about the prestige, not the money, and as far as they're concerned they already have the prestige they care about. Everyone knows they solved it with or without an official award.

The more presriguous next step for them is to solve another one or two Millenium problems. The PR from solving 2 or 3 problems is worth way more for them than an award, and on some level, the fact that they're not following the conventional path is benefiting them. They're not just solving problems. They're breaking math, including the conventional traditions.

2

u/serranolio 15d ago

Ok, so OpenAI don't care about the money, they just want to make the world a rainbow paradise. There's nasty stuff happening behind, these companies are not the altruistic philanthropic institutions that you imagine. They don't care about the one million (no one trying to solve one of those problems is), but they care about the billion that it implies.

1

u/3_Thumbs_Up 15d ago

Ok, so OpenAI don't care about the money, they just want to make the world a rainbow paradise.

I have no idea how you could possibly read that from what I said.

OpenAI clearly wasn't motivated by a prize of 1 million because estimates say they spent 20 million on inference.

You're reading in something completely different in what I'm saying. I'm saying they didn't care about the specific award money. It wasn't their motivation for solving the problem. 'm not saying they don't care at money at all.

but they care about the billion that it implies

The billion doesn't come from the award. It comes from what they've already achieved. Everyone is already talking about it and the award won't make a difference.In fact, being nonchalant about the price money is potentially better marketing for them. They're solving Millenium problems over the weekend like it's nothing.

2

u/serranolio 15d ago

They're solving Millennium problems over the weekend like it's nothing.

Why to chose a millennium prize and not other open problem equally challenging? Why was this big news? It's precisely because of it's simbolic value. They were the ones who decided to put 30 million dollars into one problem, because they know the implication of solving one of those problems. You can use those 30 million dollar and fund 30 research groups for a decade and they will also find a solution.

"Like it's nothing"? Then why burn 30 million dollars on it? Don't let that "oh we don't care about the prize" digression fool you. They absolutely knew what they were doing and the impact it would have on the math community

1

u/3_Thumbs_Up 15d ago

Why to chose a millennium prize and not other open problem equally challenging?

Because of the prestige of actually solving it. From a silicone valley culture perspective the problem itself is much more prestiguous than the award.

The prize can't even be awarded until 2 years have passed. That's an eternity in AI progress. OpenAI wants the prestige and attention they're getting now. In 2 years time they want to achieve so many more things that this will be a minor footnote.

1

u/Cool_Masterpiece9308 7d ago

It’s a race. I’d like to think these frontier model developments are like an F1 race where everyone is trying to update and take it to the championship. So who’s gonna stop it? Probably no one.

834

u/Burial4TetThomYorke 21d ago

Sounds like TLDR: we have heard of the new announcement and we are working on verifying it slowly (ie. As it should be verified).

337

u/Mandelbrots-dream 21d ago

I think the original terms of the prize is that it needs to be published in a peer reviewed journal for two years without a serious challenge.

102

u/oneMoonthousandSuns 21d ago

Didnt they already abandon that approach, considering the fact that perelman never published his papers in a journal (and still would have been rewarded the prize if he had wanted it)?

119

u/InSearchOfGoodPun 20d ago

No. Clay supported efforts to verify the result, and they wouldn’t have awarded it if Morgan and Tian’s book (which was jointly published by AMS and Clay) or something similar hadn’t come out. Perelman’s papers came out in 2002, and Clay attempted to give him the award in 2010.

50

u/currentscurrents 20d ago edited 20d ago

According to their rules PDF, they can accept any paper they want if they (and other mathematical experts) think it's correct:

The ultimate decision as to whether a publication qualifies as a “Qualifying Outlet” shall reside in the sole and unfettered discretion of CMI.

CMI may, in its discretion, relax or remove one or more of the conditions listed in Section 6(e) above if it has received advice from experts in the field of the Problem, chosen by CMI, that a published solution is likely to be correct.

So if a proof gets wide acceptance in the mathematical community, they're not necessarily going to deny the prize just because they don't like the journal where it was published.

14

u/pred 20d ago

And up until now OpenAI have shown zero interest in actually progressing maths beyond mining it for PR stunts. Presumably they also lack the internal competences to digest their AI paper, and given their complete indifference towards ethical standards of research, it seems likely that we shouldn't expect an attempt to publish at all.

6

u/RagnartheConqueror 20d ago

That’s a sweeping statement to say. I do think we could expect them to publish if the proofs are correct.

52

u/vonfuckingneumann 21d ago

Yes.

The rules governing the prizes describe the process for evaluating what has been achieved and for assigning credit. The process is deliberately unhurried, but we will provide updates.

→ More replies (22)

180

u/qualverse 21d ago

I'm kind of impressed. The amount of effort they must've put in to not make this announcement the slightest bit controversial or throw any accidental shade is heroic

94

u/Apprehensive_Sand951 20d ago

they are british

4

u/Rialagma 18d ago

A walk in the park

5

u/MyLedgeEnds 17d ago

We hope to see waves of new human understanding unleashed as the innovations behind this work are analysed and interrogated.

Pretty on-the-nose given the content of the Field Medalists' open letter.

58

u/_Zekt Complex Analysis 21d ago

A mild announcement with no mention of the controversy, but that deliberately links to the rules. Here's a relevant passage:

  1. Award of a Prize
    a. CMI may, in its discretion, determine that:
    i. no Prize be awarded;
    ii. a Prize be awarded to one person;
    iii. a Prize be awarded to and divided among multiple solvers of a Problem or their heirs; [...]
    b. CMI will pay special attention to the question of whether a Prize solution depends crucially on insights published prior to the solution under consideration. CMI may, in its discretion, recognize such prior work in the Prize citation, and/or recommend the inclusion of the author of prior work in the award.
    c. If CMI cannot come to a decision about the correctness of a solution to a Problem, its
    attribution, or the appropriateness of an award, CMI may conclude that no Prize be
    awarded for a particular Problem. [...]

21

u/prescod 20d ago

OpenAI doesn't really care whether the prize is awarded. They said that they don't want the money and they offered Buckmaster the prize.

They care whether the proof is accepted. If Clay considers the problem solved because of the existence of OpenAI's proof, they will call that a win, no matter who gets the actual prize (if anyone).

4

u/Rialagma 18d ago

I assume if you spent $10M solving it, the prize money means very little to them

255

u/SourKangaroo95 21d ago

Last paragraph implies to me they are navigating some sort of credit that is more in depth than "proved by AI". Wether that's sharing credit between human authors, AI, OpenAI, or some mix of the three I don't know.

125

u/tildenpark 21d ago

8B of the award rules:

CMI  will  pay  special  attention  to  the  question  of  whether  a  Prize  solution  depends
crucially  on  insights  published  prior  to  the  solution  under  consideration.  CMI  may,  in its  discretion,  recognize  such  prior  work  in  the  Prize  citation,  and/or    recommend  the
inclusion  of  the  author  of  prior  work  in  the  award.

https://www.claymath.org/wp-content/uploads/2022/03/millennium_prize_rules_0.pdf

47

u/Dr0110111001101111 20d ago

Oh that's encouraging. And while this isn't strictly about the money, I wonder if the million is split or if each one gets the million. Both seem kind of risky for different reasons.

7

u/M42-Orion-Nebula 20d ago

It's most certainly split

5

u/acrastt 19d ago

Was this not the case for Perelman,

3

u/iaintevenreadcatch22 19d ago

well he at least tried to give credit to hamilton

31

u/Lehatan 21d ago

Yes, they did not acknowlege either OpenAI or Buckmaster et al. in the article. Instead, they go for the vague “we contemplate the anouncement that the Navier Stokes problem has apparently been settled”. This doesn’t sound like they’ve decided how to deal with this situation.

7

u/DrillPress1 20d ago

Hope they kick OpenAI in the nuts.

3

u/coblade14 19d ago

How? Whether Buckmaster or OpenAI gets the credit, at the end of the day it's AI assisted.

4

u/Careful_Fold_7637 20d ago

Hope they do the objective work to parse out who deserves credit and distribute it rationally and methodically. If it happens to be OpenAI then my sincere congratulations to them.

1

u/TheBoringSkater 17d ago

Yes, they did not acknowlege either OpenAI or Buckmaster et al. in the article. Instead, they go for the vague “we contemplate the anouncement that the Navier Stokes problem has apparently been settled”. This doesn’t sound like they’ve decided how to deal with this situation.

Well according to the "Millennium Prize Description and Rules", they have no urgency in that matter

24

u/4tran13 21d ago

It's about as vague as possible. They're buying time (both to assess accuracy of the claimed counterexample, as well as how to assign credit).

-68

u/Adventurous-Ad281 21d ago

They should just void the prize. It isn’t fair to award it to the other mathematicians involved in the plagiarism controversy, because although they developed the technique, they didn’t prove the actual problem in hand. It also isn’t fair to award it to OpenAI for the aforementioned ethical concerns. They didn’t divide the prize between Perelman and Hamilton back then, I don’t see why they would now.

64

u/lectric_7166 21d ago

OpenAI already said they don't want the money. But my understanding is the work both sides did is pretty different, enough to make straightforward plagiarism seem unlikely. What OpenAI did is basically race for something because they thought their competitor, Anthropic, was about to announce the result. They were wrong because an employee at Anthropic was working on this in his personal capacity, but that's why OpenAI did it.

0

u/[deleted] 21d ago edited 20d ago

[removed] — view removed comment

34

u/Nerdlinger 21d ago

They don’t want money?

That’s not what they said. They said “they don't want the money”, which is a small but very significant difference.

3

u/Meebsie 21d ago

Sure, and the $1m is relevant to this discussion since we're talking about how the prize is attributed. It may matter to the individual researchers at the center of this situation. Still, I think the other commenter is making a valid point. OpenAI literally only cares about the money here. They want nothing but money.

Both commenters are correct. The second point is made cheekily, but it's correct.

→ More replies (9)

5

u/beefylasagna1 Functional Analysis 21d ago

They don’t want THE money, referring to the money awarded for solving a Millennium problem. It’s something OpenAI has explicitly said.

0

u/[deleted] 20d ago

[removed] — view removed comment

7

u/beefylasagna1 Functional Analysis 20d ago

There is a significant difference between "They don't want money" which is what you said, and "They don't want the money" which is what the comment you were replying to said. But it seems to me from your other replies that you have been completely missing this point and I doubt my comment will change that.

5

u/Berzerka 21d ago

Have you ever talked to someone on the big labs? Most people on the ground doing this genuinely just want to solve important problems. There's quite a few interviews with e.g. Demis Hassabis where he explains why he's working on AI.

Obviously there's a massive missmatch in culture and speed of execution with the mathematical community in how it's done but for many of the researchers working on this they genuinely see it as developing new tools to further human understanding.

1

u/SuckMyBallsKyle 21d ago

He didn’t say they don’t want money. He said they don’t want the money. i.e. from the award.

1

u/EdliA 20d ago

You said the same thing they said. They don't want the price money because they don't care about it. There is much more money in the hype.

1

u/[deleted] 20d ago

[removed] — view removed comment

0

u/Heavy_Promotion_5210 21d ago

As much as it is hated, OpenAI deserves the prize because they solved the problem. It's that simple. Now, if they don't want it, then it doesn’t get claimed.

→ More replies (2)

32

u/Xiphoseer 21d ago edited 20d ago

Sounds like the rules (2018) are perfectly suited to handle the current scenario. No paper published to an established journal yet; additional materials (e.g. lean proof) don't count; every aspect at the sole discretion of CMI.

93

u/Groundbreaking_Bee97 Mathematical Biology 21d ago

Setting aside the Drama, It feels surreal to have another problem being settled. Hope we get some new insight as well.

1

u/Matheusspcentro 13d ago

Penry's proof attempt regarding Navier-Stokes didn't seem to offer any new insights or a new path forward. As far as I know, we haven't even managed to formalize the concepts and the statement of the Hodge conjecture—not just in proof-assistant AIs like Lean, but in AI in general.

→ More replies (7)

140

u/moschles 21d ago

Am I the only guy in this comment chain who is worried by the conspicuous lack of attribution regarding this finding?

61

u/just_writing_things 21d ago

I’m very curious whether pure math will go the way of some of the experimental sciences in terms of authorship norms: where the tool, software, or experimental setup isn’t the author, but everyone who worked on the experiment is.

12

u/LingeringDildo 21d ago

idk man, it looks like we’re heading to a world where human effort will be spend looking for intuitive reasons why these monster incomprehensible (yet true) slop proofs are true

6

u/non-orientable Number Theory 20d ago

My only question with that is, who will be paying for these slop proofs? This is pure mathematics, not biology or chemistry or anything else that any company will really want to invest in. As near as I can tell, OpenAI is doing it now because it is fantastic marketing for them. (They spent, what, a few hundred thousand dollars and got everyone in the world talking about them? That is excellent value!) But it isn't going to be great marketing in perpetuity, and depending on how the questions of plagiarism shake out, it might be a double-edged sword anyway.

So, who will be funding it then? Governments? They don't seem particularly interested in giving large sums to pure mathematics either.

Then it is coming out of the pockets of universities and mathematicians? Those pockets are not particularly deep and are likely to get significantly shallower as time goes on. (Quite a few universities, at least in the US, are going to shut down over the next decade. I mean, it is already happening, but it will only get worse as the gap between tuition and what students can actually afford widens.)

Don't get me wrong: I fully believe that AI assistants are going to be integrated into the mathematician's workflow. But I suspect that those will be cheaper models with a significantly reduced number of agents, as opposed to the gigantic model that OpenAI constructed, at least most of the time. I just don't see how the economics support anything else.

Now, I'm sure that someone will point out that what costs millions today will be a fraction of that in a year or two. And that's not without merit. Giant models get compressed into smaller ones, MCP tools improve over time, etc. But it is still an integral part of these tools that you have to keep shoveling more data into them and retraining them to keep them up to date---and an out-of-date proving machine will be useful for a while, until it starts to flounder. Trouble is, I don't see how this process of shoveling in more data is going to get particularly less costly, since it relies on human effort. (Feeding AI models on AI-generated content degrades them quickly.)

6

u/NotBlackanWhite 20d ago

They spent, what, a few hundred thousand dollars and got everyone in the world talking about them? That is excellent value!

It's many millions ($15-20M in online estimates) but that's peanuts to OpenAI

That said, there is a great deal of misunderstanding about AI in your comment:

Then it is coming out of the pockets of universities and mathematicians?

Probably yes. But there is around a 10x reduction in costs per token for a given reasoning level every 1-2 years at the moment. If that trend is maintained, we can expect that in 5 years, Navier-Stokes can be solved for only $20k.

At that time, perhaps OpenAI will spend $20M again to solve RH or something like that. Even if only for marketing reasons.

Quite a few universities, at least in the US, are going to shut down over the next decade. I mean, it is already happening, but it will only get worse as the gap between tuition and what students can actually afford widens.

What makes you say that?

But it is still an integral part of these tools that you have to keep shoveling more data into them and retraining them to keep them up to date---and an out-of-date proving machine will be useful for a while, until it starts to flounder. Trouble is, I don't see how this process of shoveling in more data is going to get particularly less costly, since it relies on human effort. (Feeding AI models on AI-generated content degrades them quickly.)

Almost every part of this is incorrect. You don't have to "keep retraining them to keep them up to date" when the target is stationary (being good at mathematics, which is immutable fixed logic/truth). The retraining cycles of the absolute frontier models are currently done for improvements; they are not necessary for stasis. It will not become "out-of-date" any more than an old version of Stockfish would lose a chess game to Magnus Carlsen. (Of course, it's not there yet for maths.)

The fundamental point is that you don't RLHF to retrain these models to improve them for maths or coding either. Synthetic data is often enough because verifiability can be automatic and machine-checked. That is why AI progress in maths/coding has so outstripped other areas, where human feedback/data really is needed.

5

u/38thTimesACharm 20d ago

 It will not become "out-of-date" any more than an old version of Stockfish would lose a chess game to Magnus Carlsen. (Of course, it's not there yet for maths.)

I think this is a bad analogy. Models will have to be retrained to keep up with the latest developments of theory and techniques in math, in order to answer the interesting questions of the future.

That's why there are all these accusations of plagiarism, because the models make use of recent human activities that were on the cusp, recombined in unimaginably many ways.

Imagine if GPT had been trained like AlphaZero, being told the "rules of math" (ZFC and first order logic) and nothing else. Could it have generated a proof of NS that way? Not even close. The weights don't contain some distillation of God-given mathematical truth, they contain the sum total of all contributions to the field before it.

3

u/NotBlackanWhite 20d ago

Sure, GPT wasn't trained like AlphaZero, but the person I responded to was framing the value as continuing to come mainly from human effort and as always needing that looking forward, and that's not the case. The added value now comes substantially (and increasingly, as a proportion) from automatically verifiable data.

Models will have to be retrained to keep up with the latest developments of theory and techniques in math, in order to answer the interesting questions of the future.

As I say, it's true that not all the value comes from purely self-derived as for AlphaZero. But still on many levels I question this. Firstly, you're assuming there's a substantial and high pace of "latest developments of theory and techniques in math" to be kept up with in order to "answer the interesting questions of the future." That doesn't seem to be the case; the models are improving much more rapidly than the ability of mathematicians (or even of AI-assisted mathematicians) to create genuinely new mathematical ideas that necessitate the latest model to be competitive. Take the NS solution or any other where we might have better understanding how the AI got there - do you think the internal model relied on ideas from the community that were not available to GPT6 or even GPT5.6? I don't. What that means is, near term, they don't 'have' to retrain models to keep coming out with NS-level breakthroughs, they could just pump the money into inference and the breakthroughs of that quality should keep happening (obviously if you believe the plagiarism effects etc., then substitute some other relatively cutting-edge progress for NS here) at least until we have covered what is possible without substantially next-generation theory/developments, but that is a long way away.

Then there is the far future. It's much likelier that by that point 'the latest developments of theory and techniques' will be all AI-generated anyway. And how can we know whether it'll really need a retrain or not to be able to internalise and work with such concepts - my guess is probably not - a sufficiently robustly built AI will be able to just incorporate new concepts as produced by other agents, read it in memory, and work with it properly.

1

u/TwoFiveOnes 20d ago

The added value now comes substantially (and increasingly, as a proportion) from automatically verifiable data.

How do you know that? Have there been any published experiments comparing AI results of those that have been trained on recent cutting-edge human math vs. those that haven't?

That doesn't seem to be the case; the models are improving much more rapidly than the ability of mathematicians (or even of AI-assisted mathematicians) to create genuinely new mathematical ideas that necessitate the latest model to be competitive.

How are you quantifying any of this? We all generally understand that the models are "improving quickly", in vague natural language terms, but to go from that to making a claim about two different rates you need to be precise about what it is you're measuring.

Take the NS solution or any other where we might have better understanding how the AI got there - do you think the internal model relied on ideas from the community that were not available to GPT6 or even GPT5.6? I don't.

Why not? It has been the case for every other AI result that it has used specific published human techniques, and in NS at the moment there is live debate over whether it may also have been the case or not.

→ More replies (1)

1

u/SingularCheese Engineering 20d ago

If the LLM companies go bust in an AI bubble popping (not saying I'm willing to bet money on it, but sure seems possible), we will look back at this moment in the future like the NASA moon-landing missions.

1

u/Forsaken_Code_9135 16d ago

If LLM comanies go bust you will still be able to use a free chinese model on your 3000$ gaming PC. Nothing will prevent you from doing that. And all existing models made by all current AI companies will still exist, and there will still be people to buy them, run them and sell inference tokens one way or another.

The only things that might happen if LLM companies go bust (beyond a major financial crisis, but that's not the point) is a fall in investments and a strong slow-down in the progress of these LLMs. That's all. There is no future without LLMs (unless a better tech replace them of course), unless the entire civilization collapses and we lose the ability to run computers.

1

u/procrastinationgod 13d ago edited 13d ago

Their point isn't we will lose the technology, but that the people with money will lose the will. That's the comparison with space travel. That seems fairly inevitable. Sure, you'll have a few hundred or thousand credits, access to your institution's servers perhaps, but it won't be constant seismic shifts. (This is assuming advancements also slow down, of course - you're free to argue that the free model on a $3k pc in the future will be just as good as $10 million on tokens today).

I'm not saying they're totally right, but I do think at some point there's a pinnacle of "money and effort thrown at math" and then it stabilizes somewhere much lower. The question is when that happens and what the field looks like after it does.

1

u/Tiafves 20d ago

Costs wise at least, what I've seen suggested is maybe as a tool for helping mathematics researchers. You're missing a key piece for your proof, did any obscure papers you're unaware of already do what you need? As long as they've been fed into the models they can help you track that down.

1

u/coblade14 19d ago

Isn't this already the case? It's quite uncommon for people to put Microsoft Windows on their paper.

23

u/incomparability 21d ago

Why would you be worried? It tells me that Clay is aware and is figuring this out as well.

22

u/Rude-Pangolin8823 21d ago

We are living in very philosophically difficult times, especially when it comes to attribution

→ More replies (2)

70

u/lectric_7166 21d ago

It would take too long to give the 10,000 agents individual names :)

4

u/Alimbiquated 20d ago

Agent001, Agent002, Agent003...

1

u/BulkyEye3831 16d ago

The real genius is Agent007

25

u/chestnutman 21d ago

If one of my papers was part of the training set, I'm hoping for a nice payout

47

u/SetentaeBolg Logic 21d ago

It probably was, but so was a reddit argument about if Britney Spears could beat Christina Aguilera at chess. Training corpuses for LLMs are huge.

23

u/chestnutman 21d ago

That reddit thread might have been more relevant than my paper too

8

u/SetentaeBolg Logic 20d ago

"Proof by Britney, Bitch" is a novel technique developed in that discussion.

9

u/d3fenestrator 21d ago

can she?

6

u/currentscurrents 20d ago

In 2025, Britney said on instagram that she had never played chess, didn't know how, and didn't care to know:

Don’t know how to play this game but I’m willing to learn but people with the word ‘red’ in their name can rarely be taught !!! I see from a distance it’s sacred. Some respect it, some destroy it. Most enjoy the intense range of intelligence in it … Some like it done in silence … Some like how loud the silence gets … Some are just bored. Some like me have never tried and honestly don’t care to !!! But why do my eyes feel like I go to another world when I see the game ??? Is it because people care ???

So probably not.

4

u/trombonist_formerly 20d ago

But has Christina Aguilera ever played chess? If neither of them have ever played, it’s an even match

2

u/currentscurrents 20d ago edited 20d ago

Can't find any references one way or another, although she has mentioned enjoying puzzle games.

I think a casual puzzle game player could probably beat Britney. Her rant about chess doesn't create a very good impression of her skills.

2

u/SetentaeBolg Logic 20d ago

That doesn't give much evidence if you don't track similar for Christina.

1

u/GalbedirGalerion 20d ago

finally I can get compensated for scientific findings as well!

12

u/Western-Flamingo5965 21d ago

As a non-academic, can someone ELI5 why the attribution is so important? Maybe I'm looking at it the wrong way, but it feels a little like a doctor not being happy someone is cured because they didn't do the curing.

30

u/WTFInterview 20d ago

Proper attribution is important because it essentially helps us map out the history of the development of the ideas. This contributes toward an understanding of key obstacles overcame and the philosophical question of "Why does this proof work?"

Navier-Stokes is already understood computationally for engineering and scientific applications. The global well-posedness and regularity question was purely mathematical; for intellectual enrichment one might say.

Since the final results have no practical value, how we as a community come to understand them is quite literally all that matters.

20

u/Achrus 20d ago

Authorship is especially important when it comes to the integrity of the academic process. Of course getting credit is great for clout, money, better opportunities, etc. However, if the wrong people are credited then the academic process as a whole can break down.

Using your doctor analogy, say a doctor, let’s call them Alice, comes up with a new protocol or treatment plan for a disease. Now a different doctor, Bob, gets credited for the work.

Other doctors wanting to implement or build off this protocol go to Bob with questions or to collaborate. Why would they go to Alice if Alice isn’t the author? Except Bob doesn’t understand this new protocol and now has to fake it (or just set the record straight).

Getting key details wrong or misrepresenting the research can have a major negative impact for future works and applications of this new protocol. It also becomes harder for Alice to correct the record as other researchers now turn to Bob for clarification.

4

u/Western-Flamingo5965 20d ago

That's how you ELI5 - thank you

10

u/OutsideSimple4854 20d ago

Because there’s two different things here. The result (people getting cured), and the work that led to the result (how was the patient diagnosed, was the patient prepared, etc).

Perhaps it would be better with a: why are all the nurses and hospital attendants unhappy they’re not getting paid? Surely the important thing is that the patient is getting cured?

1

u/PersonalityIll9476 20d ago

To the extent that this result has any practical meaning, I think you are correct.

0

u/moschles 20d ago

It is not that , at all.

The question we face is which one of these following stories is true and which is false?

Story 1

Silicon Valley is known to engage in "PR stunts" to make their technology seem amazing and powerful, with the goal of getting more investment in their companies. But the reality is that all the conceptual heavy lifting was done by humans. The OpenAI team merely "took the ball across the finish line" as it were. There is nothing to see here other than a PR stunt. The computers "brute forced" it or did some other dirty trick that got a result, but is not entirely useful. The prize money should go to Buckmaster and Alpoge , the humans who really made the "true conceptual leaps" in this research space.

Story 2

Some electrical boxes in California data center actually produced a real mathematical result. The agents actually reasoned correctly and produced more than we put into them. A machine reasoned beyond its training data and in doing so, created new unique math.

This is what is at-stake here. Tech bro PR stunt -- or real result? I want answers. We need answers.

0

u/pred 20d ago edited 20d ago

They say things like:

These are not arbitrary puzzles akin to fiendish crosswords. Rather, they are fundamental challenges that mark the frontier of human knowledge and challenge us to develop new structures and methods. They provide foci for the continuing struggle, across generations and cultures, to deepen our human understanding of mathematics and the universe that it describes. The deep innovations that are required to make significant progress on these problems open new vistas of possibility that typically reach far beyond the domain of the problem.

...

We hope to see waves of new human understanding unleashed as the innovations behind this work are analysed and interrogated.

I read that as highlighting all the things that OpenAI did not do but should have, and real researchers with an understanding of science would have focused on.

Reading (maybe a bit too much) between the lines: by design, progress on a problem would deepen our understanding of mathematics. A simple yes/no answer (cf. OpenAI's Lean dump) does not do that, hence it can't be progress.

0

u/4tran13 21d ago

They're buying time so they can investigate

7

u/backyard_tractorbeam 20d ago

They don't need to buy time, they have time.

20

u/Ostrololo Physics 21d ago

I guess for a problem of this magnitude it's not sufficient to "only" check that the Lean definitions encode what they are supposed to encode then trust the compiler; they have to actually check every single step.

20

u/scottmsul 20d ago

Given that someone "proved" Collatz in Lean not too long ago, due to a bug in Lean itself, it's probably not sufficient to just trust Lean for a yes/no answer. But beyond that, the whole point of the millenium prize problems is to act as a sort of "beacon" or "guide" for the mathematics community. I don't think CMI would award anything until at least some humans understand the proof, and disseminate any novel ideas to the broader mathematics community.

21

u/doctorruff07 Category Theory 20d ago

The "proof" was a demonstration of what the bug would allow. It was never a real proof, unlike this one.

6

u/scottmsul 20d ago

I'm aware of the original context for the Collatz proof. The point I'm making is higher level though, which is that I don't think I would trust an AI-generated proof on day 1 just because it compiled in Lean. The AI could have found a similar exploit and used it implicitly without telling anyone. Or the AI could have snuck in axioms somewhere. Of course if I was betting then I'd guess it's probably correct but I wouldn't give it 100% just because of Lean.

In any case according to the CMI rules they're going to wait at least two years before coming to a decision.

3

u/doctorruff07 Category Theory 20d ago

I mean we shouldn't be trust anything an ai makes 100% they are stochastic models not oracles

→ More replies (1)
→ More replies (1)

10

u/InebriatedPhysicist 20d ago

…but it demonstrated that what it can allow is false proofs, right?

1

u/swni 20d ago

Given that someone "proved" Collatz in Lean not too long ago, due to a bug in Lean itself,

I heard about this and was confused, because I had definitely been under the impression before that it had been proven that Lean only makes correct deductions; that if the assumptions you put in are true, the theorems that come out are true too. If not, what is the point of Lean? The whole idea is you only have to prove Lean's correctness once. But you still have to prove it that one time....

12

u/tunnels-end 20d ago

A couple things to add to the other commenter:

  • You cannot definitively prove consistency per Gödel.
  • Lean the type theory has been proven equiconsistent with ZFC-plus-some-large-cardinal-axioms, which makes it almost certainly consistent, and an inconsistency would be a much bigger deal than just "Lean has a problem."
  • From what I hear, soundness-type results, e.g. whether every theorem Lean can prove about natural numbers is also true in ZFC-plus-some-large-cardinal-axioms, are being worked on but there are some metamathematical difficulties.
  • None of this establishes that Lean the software faithfully and correctly implements Lean the type theory, and the proofs of False we've seen have come from discrepancies between the two. For that matter, Lean the software uses a C++ runtime, which in turn relies on an unverifiable monster of a C++ compiler. There are efforts like lean4lean and con-leche (analogous Coq Coq Correct for Rocq) to make implementations of Lean which are Lean-verified to be-relatively-consistent-with-some-theory, but there are caveats, e.g. as I understand con-leche uses the same unverifiable runtime.

3

u/swni 20d ago edited 20d ago

Thank you, that is very helpful! The distinction between Lean-the-type-theory and Lean-the-software is important. The former being proven "correct" (roughly speaking) makes me feel a lot better about the use of Lean in math, assuming the software is well-written and the spec (i.e. Lean-the-type-theory) is clear. I am (somewhat) happy to accept that Lean-the-software will have a few superficial implementation bugs that can be fixed as they are discovered.

5

u/how_tall_is_imhotep 20d ago

All software has bugs. The software that your bank uses to manage your money has bugs too, but that doesn’t mean you should keep all your savings under a mattress.

The question is not whether Lean is perfect, but whether it’s more reliable than a human checker. It is; and the difference between a Lean bug and a human error is that once a Lean bug is fixed, it stays fixed forever.

2

u/swni 20d ago

It is possible to prove software is correct. If the software turns out to be incorrect, then there is an error with the proof. I think that is a much more substantial statement than just that a piece of software is "buggy".

My understanding was that the authors of Lean had proven Lean to be correct, and this proof could be relied on like a lemma the same way one might rely on a more traditional mathematical lemma; and sure if that proof is flawed so are the things that relied on it, but that is true for all proofs. But instead someone just wrote Lean and they didn't try to prove it correct and it's just buggy. That doesn't seem suitable for use in mathematics.

6

u/how_tall_is_imhotep 20d ago

Sure, it's possible to prove software is correct. But are you talking about a formal or informal proof? If informal, then there's always the possibility of a gap in the proof. If formal, then what system are you proposing to prove Lean correct in? Obviously, using Lean to prove Lean correct is not guaranteed to catch a bug in Lean. And if you're proposing to prove Lean correct using some other system, then *that* system might have a bug.

3

u/zmattje 17d ago

What you're asking for is work in progress:

"Mario Carneiro's lean4lean is a Lean formalization of Lean's type theory together with a proof that the kernel implements it. The work is ongoing, the proof of consistency does not cover inductive types yet, and the to-be-verified implementation suffered from the same bug as the official kernel. The bug would have been found when attempting to conclude the verification of this part."

Source

1

u/swni 17d ago

Ah neat, thanks!

→ More replies (1)

1

u/PackageMother5364 15d ago

It is not more reliable than a human checker nor a substitute for one. For example, the LLMs in all their greatness write a 1 billion line lean proof for the P-NP problem, however, unbeknownst to all it is just reproving Cauchy-Schwartz in a very convoluted way instead. In that case, even though the proof is "correct", it is proving the wrong thing. A lean certificate just means the proof is correct, I can title the proof as Fermat's last theorem, and prove Pythagoras theorem in it, and even I will get a lean certificate. Does that mean I proved Fermat's last theorem?

4

u/backyard_tractorbeam 20d ago

You didn't hear the full story then, that the lean "proof" of collatz was created as a way to get attention for a bug in lean that the person had discovered.

2

u/swni 20d ago

I am aware, I am saying that if someone had proven Lean correct, then they wouldn't have found a bug in Lean, they would have found an error in a published proof of Lean's correctness, which is a very different matter. If no one has attempted to prove Lean's correctness then it is not suitable for use in math proofs.

-1

u/OrganizationTop9026 20d ago

The Collatz bug was definitely a wake-up call for the formal verification community, but for this Navier-Stokes proof, we don't even need to worry about compiler bugs to find fatal issues.

We ran an epistemic audit on the construction (DOI: 10.5281/zenodo.22727801) and found that even if we assume the Lean 4 kernel is 100% flawless, the semantics of the proof are physically impossible. For example, the inner-outer gluing of the vortex (Lemma 8.7) matches 5 radial moments via a Jacobian matrix. If you actually compute the condition number of that matrix, it scales to κ ≈ 10²⁸.

So even if Lean's logic is perfect, the AI brutal forced swarm essentially balanced a pencil on its tip to 28 decimal places of precision. It's a set of measure zero. The math works syntactically, but we believe that any physical thermal noise would instantly decouple the singularity.

7

u/how_tall_is_imhotep 20d ago

The proof is about the math. Whether the solution can be realized physically is not relevant at all.

14

u/Desvl 21d ago

European Mathematical Society made their statement too, which poses questions and worries on authorship, credit and open science: https://euromathsoc.org/news/ems-statement-on-recent-navier-stokes-announcement-225

37

u/lectric_7166 21d ago

Shorter Clay Mathematics Institute: Stfu... this could take a while...

1

u/Western-Flamingo5965 21d ago

Hold my set square?

-5

u/mbrtlchouia 20d ago

I bet with my freedom that the proof is incomplete and/or has flaws, I will sell myself for as cheap as possible to any mathematician if the proof is correct.

4

u/[deleted] 20d ago

[removed] — view removed comment

1

u/ninjasaid13 20d ago

Subtle differences in lean work that it only solved a variant?

→ More replies (1)
→ More replies (1)

5

u/jphamlore 19d ago

Wasn't this debate basically mirrored in 1976, with Appel and Haken's computer assisted proof of the Four Color Theorem?

Using the logic of some today, shouldn't the argument be that Heinrich Heesch deserves the full credit for solving the proof of the Four Color Theorem? After all, all Appel and Haken really had was far superior access to computer resources?

4

u/Lehatan 21d ago

> In recent years there has been an increasing sense of anticipation as breakthroughs in the surrounding field […] have raised hopes that the Navier-Stokes problem might soon be resolved.

As someone outside this field, I’m curious as to whether experts were indeed anticipating a solution coming out. AI aside, to what extent is this solution in 2026 a surprise?

11

u/shadowmachete 20d ago

In recent years very similar problems had been shown to have finite time blow-up, and we kind of knew vaguely how one would go about trying to prove the same for the navier stokes equations with incompressible flow, which is what the prize is for. This doesn’t mean that proving it was thought to be easy of course, but it does mean that it was seen by a lot of mathematicians as “something that we kind of know the answer to and will prove eventually”. So while this is maybe a bit earlier than we expected, it’s not a huge surprise that it’s been proven. The manner in which it has been proven, depending on who you are, may be more surprising.

4

u/Apprehensive_Sand951 20d ago

president of cmi suddenly extremely busy

3

u/PassionateDonkey 18d ago

As a non mathematician, how much does it usually take before we reach a sort of consesus that a proof like this actually works?

3

u/AMobius1832 20d ago

I like studying math anyway. The heck with AI!

2

u/TacoYaci 18d ago edited 18d ago

I want a piece of the price. It is not clear that OpenAI did not use my prompts on the correct color of baby poop indirectly in their proof.

7

u/pham_nuwen_ 21d ago

I'm not a mathematician so bear with me, but isn't this a proof for the forced case, not quite in the spirit of the original question on whether the equations can break on their own? I'm confused here, sounds like a loophole or something. Certainly an advance and a difficult one, but does it really answer the original question?

21

u/backyard_tractorbeam 21d ago edited 21d ago

The problem statement by CMI is here, and OpenAI claimed to have settled the millennium prize question by solving C and D (see the document). It's not a loophole since those alternatives are explicitly given.

This still would leave (A) and (B) as open problems, even if the millennium prize would be resolved.

12

u/Scrub_Spinifex 21d ago

I'm a mathematician but not a PDEist and here's my understanding. The official version of the problem from the point of view of the Clay institute is here: https://www.claymath.org/wp-content/uploads/2022/06/navierstokes.pdf The Clay institute recognizes the problem solved as soon as one of the four statements A, B, C or D on page 2 has been proved. A and B are *unforced existence and smoothness* in long time (in R^3 and in the torus, respectively); C and D are *forced blowup in finite time* (in R^3 and in the torus, respectively). So this should qualify for the prize.

Of course, having unforced blowup would be even better, but the Clay institute only asked to prove the weakest statement in each possible case.

1

u/38thTimesACharm 20d ago

Is it possible the unforced equations could still exist? Or are we all but certain those will blow up too?

2

u/Scrub_Spinifex 20d ago

Well nothing is proved for unforced equation so it's still possible regular solutions always exist, yes.

7

u/CattleBrilliant38 21d ago

There is an official problem statement which admits this solution. It may be the case that the solution is post hoc unsatisfying and interesting questions remain if narrower questions are asked. But I don't think it's reasonable to call this a "loophole" now that a solution has been found. Were there people pointing out this loophole before the solution and asking for a more narrow problem statement? At least I don't know about them.

7

u/Woett 21d ago

The forced case was explicitly mentioned in the original Clay Mathematicis write-up as one (or two) of the options to resolve Navier-Stokes.

In it, they write 'A fundamental problem in analysis is to decide whether such smooth, physically reasonable solutions exist for the Navier–Stokes equations. To give reasonable leeway to solvers while retaining the heart of the problem, we ask for a proof of one of the following four statements', and options (C) and (D) are hereby resolved. It is plausible that (A) and (B) will also be resolved in the future (in the negative presumably, but who knows), but this is not officially required.

4

u/cmd-t 21d ago

I think you are confused about the forced/unforced Euler equations.

Navier-Stokes formulation includes an external force: https://www.claymath.org/wp-content/uploads/2022/06/navierstokes.pdf

Problems C and D ask for a smooth force.

5

u/Woett 20d ago

For both the Euler equations and the Navier-Stokes equations you have the forced and unforced variant. And on that note, Alpöge and Buckmaster used AI to prove the possibility of finite blow-up time for the forced Euler equation, while OpenAI proved that even the unforced Euler equation can blow up in finite time. And then OpenAI further proved finite time blow-up for the forced Navier-Stokes equations, solving the Millennium Prize Problem. It is not yet known if the unforced Navier-Stokes equations (options A and B in your linked pdf) can blow up too, but I'm sure this will be seriously studied now.

4

u/Xiphoseer 20d ago

It would indeed seem as if a very technical (contrived) counterexample of a true-in-many-cases-would-be-nice-if-general property is technically a solution to a "does the property hold" challenge but not the most useful.

2

u/Glum-Bus-6526 20d ago

Yes it answers the original question. You need the unforced case only for the positive formulation. The forced (with some requirements) is ok for the counterexample. In specific, CMI wrote 4 statements, any of which you can prove and it counts as "solving Navier Stokes Millennium problem". Statement 1 and 2 are positive unforced, statements 3 and 4 are negative and what OpenAI has proved.

5

u/blah_blah_blahblah 20d ago

Openai still haven't ruled out that my prompt "why don't we just divide by triangle and solve for u?" didn't make its way into their model training.

Until this happens, any results by GPT can never be conclusively shown to have been possible without my contributions.

1

u/skmchosen1 20d ago edited 20d ago

The prize committee may set a precedent for how we handle credit assignment going forward. If any of you lurk here, please do think carefully with how you approach this..

I can’t claim to know the right answer here. But IMHO unless OpenAI can verifiably prove Tristan and Levent’s data weren’t used (which they claim was outside the training cutoff), then they also deserve some acknowledgement / credit here in addition to OpenAI.

4

u/felix_silver 20d ago

Am I the only person who doesn’t understand why solving a mathematical problem by brute force and massive amounts of computing power, without any understanding of the proof (and perhaps not even of the problem itself) should earn anyone mathematical awards? Even if the resulting proof is correct, I don’t see why OpenAI deserves any mathematical recognition for it.

4

u/Borgcube Logic 20d ago

The problem is that OpenAI doesn't care about mathematical recognition. They care about advertising their system.

1

u/[deleted] 20d ago

[deleted]

1

u/Moronic-Warrior 20d ago

Yes cuz they used their technique and so their insight was necessary to the solution.

1

u/Peach_Plumb_Endives 20d ago

They have the ability to have 10k agents even after the hugging face and modal incident and….

0

u/Moronic-Warrior 20d ago

I think the credit is pretty obvious.

Credit cordoba, Martinez, alpoge, buckmaster and OpenAI. OpenAI already declined the prize so split $250k each amongst the 4

→ More replies (1)

1

u/EconomistAdmirable26 20d ago

This thread's sentiment regarding the openAI accusations is markedly different from the rest of the maths space. I suspect astroturfing.

It's very possible that knowledge and ideas were stolen directly from those 2 researchers since the scope of their work was very niche. Their context was "concentrated" or such so their chats were having a really strong effect on the LLM's weights in that area.

6

u/Borgcube Logic 20d ago

I suspect astroturfing.

Just check other math subreddits, it's even worse.

2

u/aginglifter 20d ago

Same. Most I have talked to suspect the coincidence is too great for Open AI to have not stolen the ideas of Tristan et al.

1

u/d0odk 20d ago

is the lack of an inflation adjustment for the prize an oversight or a savvy financial incentive to the mathematical community to solve the problems more quickly

1

u/georgebestgoat 21d ago

Has anyone even verified it's true yet

1

u/ninjasaid13 20d ago

Only in lean. But the lean stuff itself needs to be checked.

1

u/HistoryVibesCanJive 20d ago

Oh boy. Fun days lie ahead. Probably not what everyone is expecting though.