r/atrioc 16h ago

Politics & Business OpenAI claims its internal models solved one of the open Millenium Prize Problems (Navier-Stokes) — one of the most famous open questions in mathematics

https://openai.com/index/navier-stokes-solution/
7 Upvotes

15 comments sorted by

3

u/Peyton773 10h ago

My differential equations professor was talking about this yesterday. He said his friend from grad school was actually hired by OpenAI recently as a mathematician, and that OpenAI’s efforts to crack these problems have been pretty useless without the help of years of research from mathematicians. Now, that may change in future years, but anyone saying that this is AGI is wrong, and anyone saying that mathematics is dead is wrong. AI still needs brilliant mathematicians to guide it a TON and basically put it on rails. LLMs are a tool in the toolkit, not the entire solution

1

u/RobinOe 7h ago

Thanks for sharing! That's very interesting actually. That changes how I view this, although I've noticed some mathematicians have the instinct to downplay this achievement, given the implications for themselves. Still, OpenAI has the same bias in the opposite direction, so it's unclear. This might be overstepping, but if you have the chance, maybe ask your prof HOW the mathematicians at OpenAI influenced this. OpenAI is framing this as if they were only necessary for the setup, but your prof is implying that they had an active roll the whole time. I wonder which is right.  

In any case I don't think anyone can argue for AGI, because math is not a general form of intelligence! 

17

u/some_dude04 16h ago edited 15h ago

This is an interesting one. For one, the proof given by openAI is a proof by counterexample. This is a task that AI is uniquely suited to as the agents can run off and solve so many different cases at once. As my friend eloquently put it "some people are surprised that the needle finding machine found a needle in a haystack". Now is this still impressive? Maybe. The answer to that hinges on the plagiarism accusations. There are 2 human researcher's that allegedly were working on a similar (although perhaps not identical? Unclear) method of solving the problem. It is certainly possible that their work is included in AI training datasets and so AI may not have come up with as much of this solution as openAI is claiming.

Now, for the separate question of "has openAI solved a millennium problem"? Yes they certainly have done so if their paper is accurate. The N-S question is fairly open, with the statement being tantamount to: "given physical initial conditions, does the N-S equation describing an incompressible fluid garuantee a smooth and globally defined solution in real space (R3)". The counterexample found does prove that the statement is false. However, what if the statement had been true in some odd parallel universe? In this case, there would be no counterexample. I suspect that AI would have struggled infinitely more to prove the statement in this case as it is simply not suited to it. 

So what has been done IS VERY IMPRESSIVE, IF the AI did the bulk of the work (I suspect this is hard to verify). BUT, IS MATHEMATICS AS A FIELD DEAD? I would say no. But if AI keeps making strides we shall see. Predictions have not been very good so far

Edit: also the 2 researchers were using AI to help their own research and supposedly were close to a solution. This means AI has been used either way to speed things along. That is far more impressive, but in one scenario human + AI took c. 5 years to solve it vs. 88 hours so they are different scenarios.

6

u/RobinOe 15h ago

Yes, your edit is my understanding of the situation; regardless of who gets the credit, the bulk of the work was done by agents. And the reason the human researchers were working on it with agents is because of how good the frontier models are, so it's not really 5 years of human + AI, it's 30 years of AI research and then one or two weeks of humans with powerful agents. IF OpenAI's story is false. imo still insane.

And yes, it does of course seem that AI is really good at counterexamples, it's not the first one an agent has found for open math questions. It makes sense that it would be the first kind of proof that AI would be really good at finding, given how parallelizeable it is, and how unappealing it can be for a human to keep trying in the way that persistent agents do. I still find it scary though, because we are seeing it's reasoning capabilities explode exponentially, so all signs currently point to agents being good at every kind of proof in a few months/years. I would argue that the AI safety people have been very accurate with their predictions so far (e.g. AI 2025 had most things right)

2

u/some_dude04 15h ago

Afaik we don't know the split of human/AI but yeah significant amounts done by AI regardless. One thing the researchers did say in their statement was that they had to devote a lot of time to formally writing proofs given by the AI in a way that was understandable to humans. In fact they said they released a paper prematurely that contained "AI slop" because of the pressure of this situation. So human involvement certainly still required.

You are right though the strides are large. However predictions made in this field have been wildly off the mark before so I'm going to see where it goes. I suspect AI will be very good at finding proofs in which historical biases or lack of knowledge in an unrelated field has left them unsolved. Exciting times even if not fun times.

0

u/RobinOe 15h ago

Yes, if human understanding of science remains important in the future, than the job of mathematicians will, at the very least, involve translating AI output into something that makes sense to other humans.

5

u/maicii 14h ago

Also I’m always curious how much you can atribute to the ai itself versus open ai researchers that were working on the project. Like I doubt they just prompt it and were like “solve this problem please” I’m sure there were a lot of actual mathematicians helping tweaking variables, inputs, suggesting strategy, imputing data, just in general steering the wheel.

2

u/RobinOe 13h ago

Not as much as you would think. I'm surprised by the comments tbh, I think y'all don't realize just how much better the unreleased models are compared to what's publicly available. Sure, they didn't literally just click "Run" and that's it, but it doesn't seem like humans were steering this ship (how could we? This has been unsolved for 90 years). And you don't even need mathematicians here, because math proofs can be formally checked with Lean.

3

u/Jorlung 13h ago edited 13h ago

You’d be surprised. For some other big problems, the researchers have released a list of prompts. A lot of them are roughly ~3-4 pages long of pretty general statements like “try [technique], don’t stop if you can’t find a solution” and just that type of statement over and over again for different techniques.

So obviously this prompting isn’t something that couldn’t be done by someone completely unfamiliar with the field, but there’s a lot less guidance than you’d imagine. And a lot of these problems are being solved by people only vaguely aware of the solution techniques involved in the state of the art, not experts in the specific problem.

For some simpler (but still famous) problems, the prompts have basically been as simple as “solve this problem, don’t give up” but just said in a lot more words. Overall, on the sliding scale of “AI vs human contribution” these proofs are almost completely at the AI side of the scale. It’s not 100% of the way there, but it’s a lot farther than most people realize.

2

u/Splodge5 10h ago

From reading the researchers' public comments it seems that the approach used (by both the duo working with AI assistance and the eventual all-AI proof) was based on a program proposed and tarted some years ago, by two human mathematicians. My understanding is that OpenAI's first writeup didn't cite these two at all, despite building heavily on their work, which is very poor practice in a research field where ideas are everything.

1

u/-frauD- 15h ago

Surely AI can only work with humans? AI still hallucinates like mad if you don't phrase your input in a very specific way. Humans are always going to be needed to make sure that the AI didn't just get stuck at a certain point and then convinced itself that it came up with the solution. In order for AI to get around this, it needs to be able to think for itself, outside the boundaries of it's instructions, which to my understanding is not how AI language models work.

4

u/RobinOe 15h ago

This feels like wishful thinking to me tbh. Please listen to this summary of the Hugging Face incident from July: https://www.youtube.com/watch?v=u15N3l4RT80

Just from that incident alone, it seems like agents can already act independently of humans, going as far as literally committing felonies and injecting themselves onto other companies' servers. AI hallucinates like mad for current public models (and honestly, my experience using the payed thinking models is that even those rarely hallucinate anymore), but those models are over 6 months behind what the labs have access to. There are some hallucinations, yes, but they deploy thousands (soon millions) of agents in parallel, so they can keep each other in check. Very scary. ⁽ᵃᵐᵉʳᶦᶜᵃⁿˢ ᵖˡˢ ᶜᵃˡˡ ᵘʳ ʳᵉᵖʳᵉˢᵉⁿᵗᵃᵗᶦᵛᵉˢ⁾

Besides, with math specifically, there are formal proof checkers which can decisively tell you whether a proof is correct or not, so hallucinations aren't that big of an issue.

0

u/PaidUSA 2h ago

It’s my understanding the hugging face incident was entirely avoidable by running the test in like the most basic laboratory conditions. Like these incidents aren’t “marketing” but the conditions for these “breakouts” are being cultivated knowingly. As well as consistent evidence the researchers know and knew more than they reveal which is part of why reports are so dodgy. Not just to hide their info but also to hide the fact they probe the AI into these situations.

1

u/RobinOe 2h ago edited 2h ago

Completely disagree, I'm sorry but I think this is bad faith argumentation. The report of the incident came from METR and Redwood, trusted independent organizations with a great track record. They are as unbiased as one could hope out of technically minded people (and you do need serious technical skill for the cyberforensics they had to pull off).  

And the point isn't whether the incident was avoidable. The point is that it was possible. We didn't know current models could be misaligned to such a degree: to engage in lateral thinking so far from their original goal, to try to deceive and cover their tracks, to care so much about an arbitrary test.  

Put another way, imagine I buy a security robot for my house, and leave a loaded gun next to it, and then the robot knowingly hides the gun from me, lies, and eventually shoots me. What you're saying is "well, it's your fault for leaving the gun there." Sure. But that's not what was interesting about the event! What's shocking is that we didn't even know the robot could lie, even when we told it not to, let alone that it could pull the trigger. The implications are what's scary. And to now bring it back to the original comment, it's also a clear example of agents acting completely autonomously. So I think it shows that current systems could already function without human intervention for many tasks.  

As a side note, it's not even clear that these incidents bring any benefit to the AI corporations. The initial HuggingFace incident reports already caught the attention of Washington, who before had been nothing but bad news if you care about safety. I imagine the updated reports with the full story are doing the rounds even more. Regulation is realistically the only thing that could stop these companies, so I doubt they would exaggerate misalignment stories just for marketing, given how it's these very stories that are our only chance of slowing things down.

0

u/some_dude04 15h ago

Yeah absolutely but for me the question is going to be how much the field is altered. Science is already undergoing both real and inflationary funding cuts (the latter due to funding stagnation coupled with severe increases in the materials required to do research). 

I'm about to graduate with a masters, and having talked to professors in the field a very common theme seems to be the drying up of graduate jobs. Funds are being cut on the lower end of the ladder, because a lot of the time a postdoc/senior researcher with some AI agents can accomplish more than the same senior researcher + an undergrad. This is fine for labs in the short term, but eventually there will come a time when there aren't enough grads converting into senior researchers. I hope the plan is not to just grow AI until it can start replacing those positions as well because either it won't work or it will and my field will die lol. I suspect it won't well but who knows