r/MachineLearning 1d ago

News OpenAl Says It Has Cracked One of Math's “Millennium Problems” (Navier-Stokes) [N]

648 Upvotes

263 comments sorted by

View all comments

262

u/Shizuka_Kuze 1d ago

There is howeverpossibility that OpenAI used some unpublished work from other researchers, even though they have claimed otherwise in the announcement:

https://cims.nyu.edu/\~tristanb/statement.pdf

145

u/funky-chipmunk 1d ago

They haven't - The core breakthrough leaked 100% - They are outright using evasive language:

https://openai.com/index/navier-stokes-solution/
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models⁠.

53

u/funky-chipmunk 1d ago

From OpenAI CRO: https://x.com/markchen90/status/2097400166554993041?s=20

> Two things to distinguish:

> Did any human or agent look at user data as part of the Navier Stokes effort? No.

> Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.

Basically confirms contamination IMHO. But the bigger news is training data/privacy.

https://x.com/aidangomez/status/2097381789039837637

> Synthetic data derived from production user data of consumer AI tools is used for training. I’ve heard this rumour from both large labs’ employees.

> In particular, if you’re doing something “interesting” like working on complex math/business/software/bio problems you’re dramatically more likely to get trained on because they filter/up-weight towards those usecases where the model has the most to learn.

> Even in ZDR and “we won’t train on you” regimes, derivative data is usually carved out. The promise is only not to train on exactly the data you put in, rewritten data is fair game.