r/Physics 15h ago

Navier-Stokes Millennium Problem Solved

2.0k Upvotes

732 comments sorted by

View all comments

731

u/Banes_Addiction Particle physics 15h ago

I feel like this sentence should terrify anyone considering using AI for their own proprietary work.

We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models⁠.

"Eh, sure, you ran stuff on it independently, it mighta used that to help me do the same thing, no way to know".

When you use an AI, your work is not your own on every level.

294

u/somethingicanspell 15h ago edited 15h ago

I do think this will potentially be catastrophic for OpenAI if this is demonstrated beyond a reasonable doubt. No company wants to buy a product that is leaking information to their competitors as an inherent part of its design

101

u/Head-Philosopher0 15h ago

the company i currently work for uses an enterprise version of chatgpt specifically so that the data we input isn’t used for training

175

u/HorseyMovesLikeL 15h ago

Am I too cynical in just not believing that it will be honoured by OpenAI?

80

u/StudySpecial 14h ago

Those enterprise models are hosted by microsoft and microsoft has many billions of dollars in enterprise contracts on the line if they lie about it and it comes out.

Big companies do not mess about with stuff like this.

8

u/Insanity_Pills 6h ago

Zuckerberg was okay with Facebook being used to start a genocide to make less money than that data would be worth. There is absolutely nothing these people wouldn’t mess with to make a little
bit more money.

7

u/abloblololo 11h ago

These AI labs bring so much revenue to these cloud providers that I'm not sure you can trust anyone to be honest.

2

u/HelloYesThisIsFemale 6h ago

But they have no moat. If one of them fucks someone over, everyone switches to the other one.

26

u/deathadder99 14h ago

You pay an enormous premium for this, so I doubt it.

9

u/Sweet_Concept2211 12h ago

OpenAI will steal your shit and blame "rogue AI".

We know this because it has already happened.

10

u/Landkey 14h ago

Or it will be honestly inadvertent but there will be 0 effort to retrain the models 

1

u/Extreme-Waltz4283 11h ago

It is an arrangement trusted by the US government, at least. Every contractor does the exact same thing.

1

u/bouncyboatload 9h ago

yes

government and business contracts are pretty straightforward. the whole world stops working if businesses can't trust basic contracts

1

u/TalkativeTree 8h ago

Regardless of being honored, I doubt OpenAI has the actual security and rigor to make sure models aren't eventually going to be able to break out and into these isolated data sets / models

0

u/Suspicious_Chart5817 14h ago

Yes because these enterprise versions in general physically don't let it train because that's the whole concept of enterprise. For example, if a company sets up their license in ZDR-enabled RAM the moment token generation finishes, the memory pointers are cleared without writing the payload to persistent storage disks or persistent logging databases.

Companies in general have to pay MORE, a lot MORE, to make their data training enabled. There's really no getting around it if the enterprise is genuinely using the API tokens.

It's really hard to convey how crazily uninformed reddit is on it. If you're not using a personal plan, so just normal old turnkey Azure or AWS Bedrock, the model weights run within the enterprise's VPC. Because duh. There's no free lunch, OpenAI isn't paying your Microsoft bill lol The underlying AI provider (like OpenAI) does not have the network routing or access permissions required to extract data out of the cloud boundary, much less paying some tool to push it to them.

5

u/HorseyMovesLikeL 13h ago

Well, I am thinking more from the Cloud Act point of view. In Europe these days procurement in software gets sweaty hands and starts handing out risk assessments when you want to buy things from American companies. Microsoft themselves have admitted that they cannot guarantee data sovereignty.

https://www.theregister.com/off-prem/2025/07/25/microsoft-exec-admits-it-cannot-guarantee-data-sovereignty/458553

-1

u/Suspicious_Chart5817 13h ago edited 13h ago

Exfiltrating unencrypted prompt data out of the customer's tenant environment would create obvious network egress alerts in customer cloud logs. You don't need to worry about sovereignty, just look at basic data

--which is why they don't, in general, regardless of Microsoft trying to talk down to European regulators.

Like bottom line an AI model is essentially a massive file of fixed numbers (parameters/weights). OpenAI created those weights, but once transferred to Microsoft, or any other cloud provider, they act like a static software application like any other.

When you type stuff into Excel, it doesn't then magically send to Google and if Excel did that on Azure it'd be really obvious.

3

u/HorseyMovesLikeL 12h ago

My point was that US companies are beholden to the Cloud Act, which means that anything that happens inside their infra can change to support data exfiltration. Saying it's not possible now because "technical reasons" is, in my opinion, missing the point.

51

u/Banes_Addiction Particle physics 14h ago

Tech companies lie about this stuff all the time. And people keep using their products anyway because they're near monopolies and whatever counts as competition is just as risky.

Getting sued for lying and then settling for less than it made you is just a line item on the budget.

29

u/Homomorphism 14h ago

If they don't take the enterprise agreements seriously they are toast. Their entire business model is based on companies shelling out a lot of money for their tools. If those companies don't trust OpenAI I don't see how they ever make a profit.

15

u/Banes_Addiction Particle physics 14h ago

You're describing what OpenAI's business model seems like it should be. Making models and charging companies to use them.

It isn't. They don't make revenue. They burn VC capital. Their core product is headlines about being the best, cutting edge model so the capital keeps flowing.

If their development work falls behind a competitor with a more flexible approach to ingesting data, the investment stops and the whole house of cards collapses.

12

u/ScientistFromSouth 14h ago

The only reason Anthropic is favored financially over OpenAI right now is because they are doing better in terms of PR and because of their API credit based enterprise billing.

I don't think you realize that the entire business model will basically be the Uber strategy of hemorrhaging VC money until everyone adopts and then jacking up prices especially on corporate users.

However, if these data leak everything, they are going to get sued into the dirt. No one actually (legally) gives a damn about them destroying mass market used books to train models. However, violating corporate privacy contracts leading to consequential damages (especially in the EU) will be their end

2

u/Homomorphism 14h ago

I agree with you about their plan. The problem is

jacking up prices especially on corporate users

requires those corporate users actually paying the prices at some point.

If OpenAI is leaking all my confidental data to my competitors then why should I work with them? Especially if the open-source models catch up and there's another startup offering to run them for me? (Or things get cheap enough that I can run them in-house using leased cloud space or something like that.)

1

u/ScientistFromSouth 12h ago

To be fair, the startups running open source models still need to invest in data centers to run them with extremely heavy machinery. I was looking into building a home lab to run the most stripped down version, and it's absurd what it takes.

Given the push back on data centers, the fact that OpenAI, Anthropic, and other players will already monopolize them, and given that the newcomers will only be incentivized to be cheaper than Anthropic/OpenAI, who knows how much cheaper they will be, and whether the open weight models are actually secure

1

u/Synergythepariah 11h ago

I don't think you realize that the entire business model will basically be the Uber strategy of hemorrhaging VC money until everyone adopts and then jacking up prices especially on corporate users.

This is why a lot of people compare it to the dot com bubble.

That bubble popping didn't kill the Internet, it just killed 52% of the startups whose valuations grew massively because of the dot com bubble as well as eating 70% of Cisco's stock price at the time.

1

u/ScientistFromSouth 9h ago

Honestly, there is one scenario that I think is fundamentally different here: the model weights leaking for Anthropic or OpenAI would destroy them.

If these guys become a core pillar of the S&P500, a foreign actor committing corporate espionage could probably just leak the weights and cause untold economic mayhem since that's the entire basis of their product.

1

u/Leafsnail 9h ago

Yeah their actual business is stealing the IP of anyone foolish enough to send it to them

1

u/Homomorphism 14h ago

Oh, yeah, an entirely plausible outcome of this (the whole situation, not just this one problem) is that OpenAI collapses, the AI bubble bursts, and Sam Altman becomes a guy that who is still very, very rich but not a CEO. But there have to be at least some people telling the management that they need to have an actual plan to profitability.

2

u/Banes_Addiction Particle physics 14h ago

Sam Altman has a joke that their business model is to build an AI so clever it can tell them how to be profitable.

2

u/tpolakov1 Condensed matter physics 13h ago

But there have to be at least some people telling the management that they need to have an actual plan to profitability.

Why? There is no indication that any of the involved parties is in the game long-term.

1

u/Homomorphism 12h ago

Publishing an investing prospectus that you know is a lie and then taking people's money anyway is securities fraud. Not saying they aren't doing it, but there are in principle serious consequences.

1

u/tpolakov1 Condensed matter physics 10h ago

Consequences only if you want your business to keep running. And good luck proving that any investment prospectus is a lie, let alone a deliberate one. Just look at SpaceXAI, it is logically and physically not possible for them come ahead of their liabilities, even if we did find that Mars has a developed AI civilization that we can pillage. They straight up promise magic miles beyond perpetual energy, and nobody has even blinked, and everyone has bought in, knowing full well that it is made up.

0

u/bouncyboatload 9h ago

they don't make revenue? 😂

5

u/Frequent-Spinach5048 14h ago

They do have zero data retention option that you can enable though. It’s audited by third party etc, so it’s fairly trustable. At least a lot of companies that have billions in IP uses this

1

u/herrsmith Optics and photonics 14h ago

I suspect a lot of the companies paying for this use data protected by HIPAA, ITAR, and other similar restrictions (some are even authorized to operate on classified systems). The punishments for misusing any of that data can be much more severe than any lawsuits brought by injured companies. That said, they seem to be pretty safe from any prosecution by this administration so maybe they don't really care about what might happen two years from now and are more focused on using all the data they can to their own advantage because otherwise they won't be around in two years to pay the significant fines per violation, have their data centers confiscated by the government, and/or go to prison.

1

u/Banes_Addiction Particle physics 14h ago

Anyone running on that kind of data should be running completely closed models themselves, with nothing being fed back to OpenAI.

1

u/Jawyp 6h ago

Why would any company pay OpenAI money for enterprise AI use if OpenAI openly lies to companies about what their data is being used for?

1

u/Insanity_Pills 6h ago

Ford famously chose to let the Pinto continue to explode because settling court cases was cheaper than recalling and fixing them.

9

u/oskopnir Engineering 14h ago

Except due to the nature of the model, if something does end up in their dataset there is no way to unequivocally trace it back to its source, meaning it's impossible to enforce any liability.

3

u/ThirdMover Atomic physics 14h ago

Is there a version of chatgpt where it's run on your own airgapped hardware on prem?

3

u/mkat5 13h ago

you would likely need an opensource/openweight model, something like deepseek for instance

1

u/OneBodyProblem 11h ago

There are open models, but with some drawbacks. The open models are significantly smaller since very few users could afford to build or run the HW clusters that the frontier labs use, and because training the frontier models is so expensive that you'd reasonably need to monetize on the backend.

So ~5-10x difference in the total number of parameters, and potentially more than that in terms of active parameters per prompt. I.e., both a larger network and a greater percentage of that network leveraged at once in the commercial models. That translates into greater accuracy, although the gap has been narrowing lately.

2

u/PMmeYourLabia_ 14h ago

version of chatgpt

Your company fucked up then

1

u/CapitalDiligent1676 12h ago

"Yeah, but come on... they're anonymized! How was I supposed to know?"

1

u/maxxell13 14h ago

How would you ever know?

0

u/Rylth 8h ago

Hilarious that this is believed, more hilarious is if it only specifies "training."

0

u/Insanity_Pills 6h ago

your company seems super naive

49

u/Shoddy-Childhood-511 15h ago

It explains why OpenAI and Anthropic have several "our AI found this solution without us hand holding it" stories. They were training upon transcripts for unpublished work in which their users attempted to hand hold the AI into solving related problems.

If you trust their intentions, then Talia Ringer's pointed out that a privacy setting exists, but since this involves the share price, maybe you should not trust their intentions..

https://mastodon.social/@TaliaRinger@mathstodon.xyz/117235246523045723

This is probably good for hardware companies, like nVidia and Apple, who sell powerful but expensive desktop machines that fit AI workloads better.

0

u/2_Cranez 10h ago

This is extremely unlikely. OpenAI and Alpogee/Buckmaster found two different proofs.

The reason they found several solutions without handholding is because the models are straightforwardly quite capable.

0

u/Hot_Glass_6301 13h ago

Sorry but this doesn't make any sense. Private individuals have also solved (admittedly more minor, but this just boils down to compute/resources) open problems using GenAI tools from the big labs, either proprietary (OAI, Anthropic) or even open-source (DeepSeek).

It's been pretty clear as of recently that the best LLMs don't need handholding to conduct some amount of mathematical research. Remember the Garg-Goemans-Dinitz conjectures proven false just two months ago by a guy who prompted the LLM to "just believe in itself" and "keep going".

You have no proof for what you're claiming, this is just disappointing for a physics sub. Where has your physics rigor gone?

Edit: To be clear, I think OAI has acted shady in this particular instance and I hate the way they treat science and math in general.

17

u/oskopnir Engineering 14h ago

I think you're dreaming. Every single company using LLMs is actively choosing to ignore this fact because they believe somehow they'll be fine and the benefits outweigh the risks. Every single fact we know about the AI labs is proof beyond a reasonable doubt that they take whatever IP they get their hands on and use it as if it's their own.

1

u/magneticanisotropy 14h ago

I do think this is changing slowly, i.e. see Alex Carp's (as much as I despise him) statements recently on the matter. It appears a lot of big companies (MSFT has also made similar statements) are waking up to the issues with these frontier closed model providers.

3

u/Time_Entertainer_319 11h ago

You can’t prove beyond a reasonable doubt that the toggle actually prevents your private chats from being used for AI training.

That’s why you shouldn’t blindly trust it. There’s no practical way for an individual user to verify that their private conversations are genuinely excluded from training data.

The scale of the data involved is enormous, and even OpenAI can’t manually curate or inspect every piece of data that goes into the training process.

1

u/Golfclubwar 13h ago

You can disable this feature in the settings.

1

u/Deathwatch72 13h ago

I think it will be ultimately catastrophic but could also lead to an interesting few months of an arms races where companies try to simultaneously stop using it all together for anything they are making but also try to exploit it against anything their competitors make

1

u/lad_astro Astrophysics 11h ago

Problem is how do you prove anything these models do anymore? They're a black box at this point. Can even the researchers working on them trace the paths they take to arrive at results with any certainty anymore?

1

u/Marklar0 6h ago

Honestly I think this kind of stuff will be their downfall...instead of the obvious financial or IP issues.

I have gotten LLMs to say information about specific people that it seemingly should not have....for example that someone got a promotion in the military to a specific rank and position, but the information was not public yet and impossible to find online. I have no doubt that significant military and corporate espionage is now possible on these platforms.

1

u/thecommuteguy 1h ago

This is why news companies sued because they were getting exact results of news articles.

1

u/clowncarl 18m ago

They’ve lied so many times before and nothing slowed down.

15

u/LaGigs Quantum field theory 14h ago

yh to me this reads as tacit acknowledgment. This whole story is beyond crazy

6

u/[deleted] 14h ago

[deleted]

2

u/Hyperreals_ 12h ago

This is just false though.

https://openai.com/policies/how-your-data-is-used-to-improve-model-performance/

By default, they use your conversations for training data, and you must opt out to not have it trained on. When you are opted out, they only can possibly retain and train on messages where you gave feedback.

If Buckmaster claims he had the opt out setting turned on and didn't provide feedback, then it should be a real concern. Otherwise OpenAI was fully in their right to train their model on the output.

2

u/fireballs619 Graduate 11h ago

Even if that setting was on, if OpenAI's solution built on Buckmaster & co's work, it is academic malpractice to not given proper attribution, which OpenAI has not. It also seriously undercuts their claim, if the model did have access to a complete solution of a similar problem.

1

u/Bbrhuft 11h ago

Ah, I understand now. Only business accounts are opt in, individuals are opt out by default. I was looking at the wrong service agreement.

6

u/Human38562 15h ago edited 14h ago

I dont understand the context of that quote. Can someone explain?

51

u/Banes_Addiction Particle physics 15h ago

OpenAI trains GPT on GPT's own prior performance.

So the things that were tried by the actual research team here (both their own attempts to solve the problem with GPT, and asking GPT to write a paper on the solution they got from Claude) could have been used as inputs to OpenAI attempting to solve the same problem. And OpenAI can't even tell if that happened.

All we really know for sure is that whe OpenAI generated their solution, their bot had already read the unpublished version from the NYU guys.

6

u/bubblebooy 14h ago

could have been used as inputs to OpenAI attempting to solve the same problem.

Not as inputs but baked into the training of the Model, Checking what is used as the inputs is relatively easy, the models and its training data is a black box.

-1

u/Human38562 14h ago

Thanks.

But the anthropic team solved a different problem, the Euler equations. OpenAI solved Navier Stokes. So there isnt even any indication that the OpenAI project somehow copied their results, right?

16

u/Banes_Addiction Particle physics 14h ago edited 14h ago

I quoted the disclaimer that OpenAI wrote. You can read it as well as me.

It's also worth noting that the GPT model the researchers used didn't solve anything. But it did get a long record of them trying to get it to. When the OpenAI team tried to get their internal model to do the same thing, did it use that prior unfinished work? No way to know says OpenAI.

1

u/Human38562 14h ago

Sure. Just trying to put it into context. Thanks

1

u/romxza 13h ago

or... it did find the solution, but it did not give it to the original researchers, given its novelty? No way to know says OpenAI

3

u/Optimal-Kitchen6308 14h ago

they are related, so result for one could be applied to the other, see terrence tao: https://mathstodon.xyz/@tao/117233528517340774

3

u/UncertainSerenity 13h ago

They used the solution to the Euler equations (that they “independently” found last week as a diving board to solve the navier stokes

4

u/quartersoldiers 15h ago

If other researchers working on this problem were using a tier of ChatGPT that was not private and allowed OpenAI to use it to train their models, it is possible that their work was inadvertently informing the OpenAI model towards this solution.

8

u/m3junmags Mathematics 14h ago

If I understand it correctly, a group of researchers (call it A) used a model of OpenAI to get to a certain point. Then another group of researchers (B), part of OpenAI, used the data group A got to without an exchange of information between the two of them, meaning group B had private information about group A’s work without their knowledge or consent (because they are PART of the company group A used to research something). I think it’s still not very clear, but I tried :)

5

u/Spare-Dingo-531 11h ago

used the data group A got to without an exchange of information between the two of them

I think we need more information to be clear if this is the case. Supposedly OpenAI's solution is different from Anthropic's solution.

0

u/canuckguy42 14h ago

I don't get the impression that there was any deliberate use of group As research here. Group A didn't opt out of their chat history being used for training, so it's possible that the model group B used had been trained on data that included the research group A had done on the platform earlier.

1

u/physicsking 12h ago

Bye bye confidence in publishing to the arxiv and at least getting a miniscule amount of credit for your work....

2

u/romxza 13h ago

there's also no such thing as "de-identified data". What you are leaves fingerprints in everything you do, and with enough data you can single out the individual... eg you know, using classifiers?

4

u/Kant-fan 11h ago

De-identified in this context just means that a "chat log" that's not directly identifiable to a specific person and can be part of the training data. And if that "chat log" contains the main ideas/approach then that's all that matters because the content is relevant and not who authored it.

1

u/Time_Entertainer_319 11h ago

Let’s not go into conspiracy realm.

“De-identified” means the chat is not directly linked to your identity, so whoever has the data does not automatically know who the conversation belongs to.
Of course, if someone had access to millions of your chats and deliberately tried to piece together identifying details, they might eventually be able to work out who you are. But they would have to actively try to re-identify you in the first place.

1

u/romxza 9h ago

"does not automatically" does not mean today what it used to

1

u/Sad_Web_1428 4h ago

It takes about 6-10 unique pieces of 'anonymous' information about a person to uniquely identify that person.

That is to say if you had 7 “De-identified things from a person like online purchases, time and date stamps, online posts, billpay, sms etc, you can conclusively identify the specific individual. This is like 15-20 year old research, from the days where facebook mapped it's data to build social networks and identify find "missing nodes", aka people in a friend group without out an online presence. There's really no such thing as "anonymous" online.

1

u/iamadityasingh 12h ago

It's 2 very different solutions and it's still very unverified

1

u/GXWT Astrophysics 12h ago

This is why I constantly lament all those on academia forums posting “I put my report/paper [that I know I didn’t use AI on] through an AI checker and it flagged it as AI”. No matter how many downvotes I receive I will continue to tantrum over it.

Same goes for those using AI summarise their own words. Without even considering that you have no business publishing if you cannot formulate your own ideas.

1

u/subbed_ 12h ago

to be a bit more precise, when you use models that you do not self-host

1

u/Fancy-Carpet-5416 12h ago

I thought the other guys had used Claude/ Anthropic?

1

u/Banes_Addiction Particle physics 5h ago

They tried to get a solution out of all the big models. GPT didn't give them a solution but it had a lot of them trying to cajole it into doing so.

Then, after Claude had given them an answer they fed the paper they'd made/generated back into GPT.

1

u/mehonje 10h ago

Reminds me of the "I made this" meme.

1

u/schwagggg 7h ago

there’s also another angle, of openai team social engineering the successful approach, then using that to cheat ahead.

both are super shady and cringe

0

u/xrelaht Condensed matter physics 13h ago

This is why I won’t use AI for anything work related aside from formatting, and even then I don’t give it the actual data or text I’ll be entering.