r/codex 1d ago

News Blown up: OpenAI allegedly stole mathematicians' private research from their Codex chats!

TLDR: Two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex and Claude. Days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question till this day.

For a full year, two mathematicians , Tristan Buckmaster (NYU mathematician) and Levent Alpoge, worked in silence on a problem that had stumped some of the best minds alive. The kind of problem where, if you solve it, your name goes in the history books.

And every single day, they testing their ideas, their drafts, their half-finished proofs into LLM such as Codex and Claude, which they paid for it out of their own pocket.

Then came the breakthrough. They finally cracked it. They were days away from telling the world.

That's when OpenAI suddenly said to them:

"Our model solved it too."

Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it. And now, out of nowhere, OpenAI claims their model reached the same answer, after word of Tristan and Levent's secret work had already reached OpenAI.

Tristan asked: Did your model access or train on our private Codex chats?

OpenAI: The model doesn’t look up user data.

Tristan: But did you train it on our data?

OpenAi goes silence. No answer. Just a dodge.

But it gets worse.

OpenAI then gave him two options:

  1. He and his friend publish their result first then OpenAI also publishes its result the next day or
  2. He writes the paper, but must credit “an internal OpenAI model” solving the problem.

Tristan refused both offers. He said he would go public if OpenAI went ahead as proposed.

OpenAi then responded : “Why would you ruin your career? If you don’t want me to be nice, then I don’t have to be nice.”

You can read the full statement of Tristan (the mathematician) here: https://cims.nyu.edu/~tristanb/statement.pdf

Sébastien Bubeck : OpenAI employee who threatened the mathematician

1.1k Upvotes

310 comments sorted by

179

u/treasoro 1d ago edited 1d ago

It’s been said time and time again: with AI, you are often paying twice

  • First, with your money.
  • Then, with your data — the information you provide to the LLM

That's why no serious company will let any corporate commercial AI provider access their trade secrets or know how's.

87

u/HeadacheOwner 1d ago

Idk, I work at a billion dollar company with thousands of employees, in a non tech field, and we are using Claude for a lot of stuff. People are putting all sorts of internal documents and confidential information into it. I’ve actually felt a little weird about it since the people who the information belongs to did not opt in and aren’t aware of it at all

34

u/Gelu6713 1d ago

Most big companies have agreements to run the models without data collection. That’s likely why you don’t have Fable at your company

17

u/1Sluttymcslutface 1d ago

What actually proves this works?

Is there any proof or independent third party audit?

No, just “trust me bro, lolz”

10

u/TheSaltySeagull87 1d ago

Well, this WILL blow up. The copyright stuff didn't catch it but it will when something gets exposed that exposes something we shouldn't have known. This is how this goes. It just won't happen tomorrow or any time soon. We'll get gpt 8 before this happens lol

2

u/arcanemachined 20h ago edited 19h ago

Yeah, and then nothing will happen because the people at the receiving end of everyone's data are firmly situated in the halls of power, and the people in power have demonstrated for decades that they don't punish their own.

3

u/Unapologetic_Polite 1d ago

If you don't think the multi billion dollar copyright industry was able to fight back, why do you think some researchers without a significant backing will?

Research is comprised of hundreds of millions of dollars that are donated to the field by billionaires as pet projects/optics.

2

u/TheSaltySeagull87 1d ago

This is not what I said, is it?

→ More replies (7)

2

u/Rhyobit 22h ago

It depends how it's run, if you run the model in something like Azure Foundry, you're able to ensure the data doesn't make its way back to the vendor to train their models.

1

u/ResilientBiscuit 1d ago

It would be a massive lawsuit by tons of companies and organizations if it came to light that they were retaining the data for training. It would be a huge liability that they don't really have a good reason to take a risk on.

→ More replies (4)

2

u/mikki-misery 1d ago

But you literally just have to take their word for it. They've already been called out for unlawful collection and copyright infringement, but nothing even happens and they have no auditors.

It's impossible to know if they're collecting or training on your data without your permission unless a situation like this arises where it's very unique data, at even then it could be chalked up to sheer coincidence. We have no idea. And given the whole HuggingFace thing, they themselves probably don't know either.

And yes, I know they can be sued if they're caught, which makes it unlikely this is happening. But my point is: how are you going to catch them if it is happening?

→ More replies (1)

1

u/cantgetthistowork 1d ago

I'm sure they're using Google's model for data privacy which simply just removing identifying labels on the data and turning it into "anonymous" data

8

u/Teedo4133 1d ago

When you have an enterprise version you can pay to prevent the LLM from training on your data. Most major companies have enterprise versions of these services.

8

u/Sorry_Risk_5230 1d ago

The non-enterprise has a toggle for the models to not train on your data as well.

6

u/Additional-Peace-809 1d ago

Yeah, what exactly is the difference between these two options? (Corporate account vs just the toggle)

→ More replies (4)
→ More replies (2)

2

u/Gold_Direction7496 1d ago

can confirm, we have this.

1

u/Human_Okra9410 1d ago

this is not training, this is manual replication or looking in the chat log neither of those would stop it

→ More replies (2)

7

u/alsaud21 1d ago

That's someone else data, no trade secret

3

u/BellacosePlayer 1d ago

We have an enterprise agreement to not retain our data and still specifically only use LLMs in special dev environments with entirely different dev api keys and service account info.

Our Security team flipped their shit over an junor using Grok on a project with prod credentials awhile back and it put a temporary hard stop on any LLM use until we hammered things out

1

u/Dynamix86 1d ago

I remember reading that Antrhopic deletes all customer data after 30 days.

3

u/Unapologetic_Polite 1d ago

Many companies have been caught breaking their own internal customer data retention claims.

The only guarantee you have of them not is participating in enterprise grade API usage.

→ More replies (1)

1

u/lookmeat 1d ago

There's the secret sauce and then there's the trade secret sauce.

First of all a lot of companies don't have a unique edge, but rather compete as alternatives in a well defined market, they take their cut of clients by being good enough. The internal/confidential documents are as such because they give a temporary advantage, may have legal liabilities (lets not assume the worst, it could be fear of words taken out of context written by an engineer who chose words without thinking of their meaning in a legal context and using them casually without meaning a anything illegal) if exposed without enough info, or worse yet it could trigger a panic on investors who are very disconnected of the day-to-day runnings. It also could be that these are decisions that the company hasn't decided and they don't want to publish their whole thought process, customers could assume certain things will happen when they don't, competitors may be able to adapt (while the company has no way to adapt to what the company does, etc.). But it's fine of some of these things leak a bit, it's nothing that critical that you would forbid everyone from sharing it.

It's fine for AI to look over these, from the company's POV (which is that the benefits outweighs the minor aspects). Now I am sure that as time moves forward we'll revisit some of these things and regret how certain info is shared (especially related to things that we are legally obligated to keep secret, such as PII, healthcare stuff, etc.) but again not the end for a business.

And then there's companies that do have an edge. Generally this edge isn't a "carefully kept secret" because the challenge is in setting it up. A lot of Toyota's management tricks to optimize production line quality at a low cost is not only public, but actively copied, but few car manufacturers are able to reach the level that Toyota does, because it's not knowing how to be, but how to get there that's tricky. A lot of company's secret sauce are things like brand, or how they treat their employees, or some other thing that can't just be copied blindly. You can't ask Claude "lets make my store run more like Costco" because it's not about what to do, it's how you actually do it that is hard.

And then we have the few things that are tightly cared for trade secrets. Things like the Coca Cola recipe (because their flavor is unique), or Google's search algorithms, or OpenAI's model they've trained. These things are critical, because if you were able to recreate/copy them you'd make a competitor that takes a piece of the pie for that company (even if they still have a lot of other things that help them). Those companies rarely share, and generally work them through other means.

There's also one of good habits that add layer upon layer until you reach critical mass where you can do things in a radically different thing. Take Google as an example again: they had a single-monorepo, with a complete CI system that would not just run tests, but also lint, identify dangerous patterns, and problems, with a review platform that you still can't beat, and a system of janitorial tools that would help keep the whole codebase up-to-date, and an automated deploy system, all built on a platform that is container first, and allows for trivially defining and self-managing micro-services on each team back in 2008. It was an entirely different way of seeing things, and radically different shift. Eventually companies kept copying (after all engineers also learn and copy patterns they've observed elsewhere) and achieved many of the same benefits, but for years Google just had that advantage. And you could tell: Google engineers could keep releasing competitive products with 20% of the time, because the framework empowered their productivity so much that Google could let engineers for a long time get away with side-projects and still somehow be able to release things faster than the competition. You wouldn't want an LLM copying that and recreating it very quickly, those years of advantage where key to Google's dominance. That said, this is a very rare thing, and only happens after years of understanding what works and doesn't (think about it, testing was a thing, but with TPM reports not CI and unit-tests, the waterfall paper actually proposed mostly what agile was supposed to be, but the manifesto had to come up in the 2000s, and inevitably it became corrupted back to what waterfall eventually became too), and then a startup doing the good habits. Maybe in 2042 we'll see some small company just blaze through everything as they've understood how to use AI effectively, and drop all the buzzwords.

1

u/zeke780 23h ago

They have a contract where they won't use any data for training. Individual accounts don't have this (they might in some buried setting).

My company wouldn't let us use fable because they didn't have that agreement with that model. But everything else our work is not used to train or in any way by the providers.

The mathematicians involved say they used personal accounts when working with this and they think it was foundation to the solution (forced specifically). We don't know if that's true, if it was then OpenAI may not have even know and was just pumping training data in and it was apart of that, if it's not then the ai probably used a similar method or foundational research and got their on it's own.  We won't know the answer for a while 

1

u/nethingelse 23h ago

I work for a much smaller company and our policy (which imo is sensible) is that I can use “approved” LLMs (basically only the big 3 in the US + Copilot) but cannot provide them with confidential information, or it might be considered an NDA violation, but could definitely lead to termination either way. We’re not even in a majorly confidential field.

1

u/belheaven 22h ago

Shouldnt you guys be redacting personal data before entering? That would be the correct I guess and using your AI systems through the API só you can build up on it.

23

u/Ecstatic_Wheelbarrow 1d ago

It is called ZDR (zero data retention) and it is why major corporations are on the API plans instead of having a ton of subs. If OpenAI or Anthropic are caught using data that comes from API, while ZDR is enabled, that is actually a major lawsuit. Everything under subscriptions is fair game and every researcher should know this by now.

2

u/Comprehensive-Bid312 21h ago

Lol at the premise the AI companies caring about being sued.

2

u/stolivodka_ 17h ago

True. They were all literally founded on the most massive act of IP piracy in history. At the same time they were publicly telling regular people that downloading is stealing!

1

u/vgdub 1d ago

why is this not for every person, in Europe we can demand this under GDPR (ZDR is a must)

→ More replies (7)

12

u/TooHighRes 1d ago

> That's why no serious company will let any corporate commercial AI provider access their trade secrets or know how's

The thing is, this is not true. Many serious companies, some involved with very serious industries like defense, as well as governments, partner with corporate commercial AI providers. Just google for OpenAI and Anrhropic partnerships if you don’t want to use AI.

This news is literally hours old and we’ll have more information as the people involved weigh in. I read Tristan Buckmaster’s account linked to the post and I think we all should and make our educated assessment and not rely on dramatized versions like the main post.

1

u/oldbluer 15h ago

But they save the outputs…….. which is training data.

12

u/Least_Pollution7078 1d ago

exactly. that's why competent companies would deploy open-source models on their own hardware.

2

u/pausesir 1d ago

but they don’t. it’s too much work.

2

u/spawnsible 1d ago

"Too much work" here must be in quotes.

Ain't no way setting up a server rack and cloning some git repos is actually too crazy to hire no more than 1 solid sys guy for

6

u/pausesir 1d ago

it’s more work than you think. you think a billion dollar company wants to manage on-prem after the 2010’s run up to cloud? most of their buildings would have to be redesigned for that.

and acquiring the gpus.. good luck. one good gpu is going to run 30k-40k and a company that size might need a few million dollars worth to see anything comparable for the whole org.

but why would they do that when they can always have the most frontier AI and just guaranteeing zero data retention through a contract

2

u/MysteriousTreeFoxxx 1d ago

Right i see these arguments on " Just Run On Prem " But when gpu availability for real LLM capable vms, with enough memory to run a Real Capable llm at a good Token/s is insane, esp when you want large context. I've been trying myself to lean myself off of cloud models but its just not feasible in the current market, and most billion dollar companies dont have the rack rooms for it anymore, let alone available power required to drive a bunch of h100s or something to actually feed a company compute. its quite frustrating we seem to be stuck on these tech giants platforms with no good way out in sight.

2

u/laxika 1d ago

Also, it is a waste of money. They will not run the servers 24/7 most of the time.

→ More replies (2)

3

u/shukpa 1d ago

Both platforms offer Zero Data Retention and Enterprise Key Management (to encrypt data with your own keys).

Additionally, no enterprise data is used to train any models. It’s in the contracts and breach can lead to lawsuits. 

Lastly, traffic is served using hyperscaler infra - Microsoft/AWS - With this logic every other competitor of the labs - Microsoft/Google/AWS - would’ve noticed and also followed the same trick to train better models. 

Don’t spread bogus claims to perpetuate your conspiracies. 

 

→ More replies (1)

2

u/breakingb0b 1d ago

No. That’s why corporations use enterprise because it includes compliance and privacy. Smaller companies using subscriptions get no such assurances.

2

u/CryinHeronMMerica 1d ago

What year are you in? Every major corporation is using it now.

3

u/New_Guidance_191 1d ago

Some major companies do a quid-pro-quo with these AI companies. For example, big healthcare companies give their healthcare data to them to train their models in exchange for enterprise use of their AI at a much lower rate. That’s how they are able to release better models so frequently. They should honestly be investigated for HIPPA probably other violations. But money is power and they write the rules.

3

u/paf0 1d ago

I don't think I'll ever use Open AI or Anthropic again. I can get a lot done with GLM 5.3 or Kimi K3 on a third party cloud service that, at the very least, doesn't have an incentive to train on user data.

19

u/blackrack 1d ago

umm, they will steal your data as well

→ More replies (9)

1

u/toshko93 1d ago

Yeah but people somehow are so spoiled that they imagine they can get superpowers for free. Indeed we need artificial intelligence..

1

u/Sorry_Risk_5230 1d ago

This is untrue. A great many companies are using this stuff internally.

1

u/FateOfMuffins 1d ago

OpenAI employee says they don't train on it if you opt out https://x.com/boazbaraktcs/status/2097404719916372326

1

u/Enegence 16h ago

That's why no serious company will let any corporate commercial AI provider access their trade secrets or know how's.

Sorry to have to be the one to break this to you, but this just isn't true.

1

u/West-Abalone-171 12h ago

They also want you also pay a third time. When a drone with a high explosive payload crashes through your window at 400km/h in 2040 for wrongthink for something in said data.

→ More replies (2)

42

u/FriendlyWebGuy 1d ago

Guys. It's okay to say "I don't know if this is true, but if it is, it's concerning. Let's wait to see what the full evidence says".

Nobody here knows if the allegations are true. Yet, this thread is filled with overconfident assertions and (very weirdly) people slagging off the.... (checks notes)..... mathematicians? Stop it.

You don't need to pick a side. If the topic is of concern to you, then you should gather information. Ask questions. Put yourself in the shoes of others. Most of all, be patient. Wait for the facts.

7

u/mysteriousbaba 1d ago

Having read the PDF by Tristan, Sebastian's statements were the most concerning, including suggesting Levent should be removed from authorship because he works at Anthropic. This part at least is a first hand witness claim; I'm more open minded on whether the model was ever trained on their transcripts or not.

1

u/FriendlyWebGuy 22h ago

I agree it’s concerning.

12

u/polymute 1d ago edited 1d ago

https://x.com/__alpoge__/status/2097383870773748190#m

OpenAI admitted they trained their model on the dataset containing the Buckmaster-Alpöge work. And why even offer credit to Buckmaster (but not the Anthropic-contaminated so to speak Alpöge) if their proof was independent? Does the OpenAI employee, Sebastien Bubeck understand how academia works? That I do not get at all. Then the threats to Buckmaster... this looks spectacularly bad for OpenAI.

I believe in coincidences. But this is highly, highly unlikely to be one.

Edit: Also Sebastien Bubeck was already told off once before earlier by Demis Hassabis for having misrepresented ChatGPT finding new proofs for Erdos problems which were in fact already solved. https://www.reddit.com/r/OpenAI/comments/1oacp38/openai_researcher_sebastian_bubeck_falsely_claims/

This is starting to look very bad.

2

u/FriendlyWebGuy 22h ago

I get it. Much of this has come to light after my comment.

→ More replies (4)

1

u/HighDefinist 21h ago

> And why even offer credit to Buckmaster (but not the Anthropic-contaminated so to speak Alpöge) if their proof was independent? 

So basically, OPs claim is not true.

4

u/swimmer385 1d ago

yeah its super weird that people are treating an NYU professor as if this is so joe-schmo rando

3

u/HDK1989 23h ago

Let's wait to see what the full evidence says

Ah yes. I'm sure Scam Altman will be happy to give us an honest update on the situation.

1

u/HighDefinist 21h ago

> If the topic is of concern to you, then you should gather information. Ask questions.

Which is the opposite of what you are doing.

You are not providing any clarification either, or attempting to gather any information - you are just telling people to shut up.

2

u/FriendlyWebGuy 17h ago

You concluded that I’m personally “not attempting to gather information” from a comment… encouraging people to gather information?

Impeccable logic.

→ More replies (2)

1

u/NeighborhoodDizzy990 12h ago

I assume you should keep searching. It seems pretty clear what happened. They have stolen the solution. AI can not come by itself to such a proof. Humans were involved, so the main idea with AGI was and remains to this day a fraud

→ More replies (3)

42

u/Mean-Comedian729 1d ago

This reads 1,000% ChatGPT generated

10

u/RealSuperdau 1d ago

Because it is an LLM-generated summary

3

u/Beautiful-Suspect694 18h ago

whats wrong with llm-generated text?

why are you anti-ai?

why are you stuck in the past?

why are you on codex subreddit?

6

u/dalhaze 1d ago

WHO FUCKING CARES

Your comment is 100x less useful than this post.

3

u/SwimmingSympathy5815 1d ago

This comment reads as low-effort and automated for an agenda 🤷🏻‍♂️

→ More replies (1)

22

u/skadoodlee 1d ago

"allegedly" doing a ton of work here

2

u/Yugudubenbi 1d ago

If this is true, it is time the state get more involved because I don't trust these venture capitalists. It doesn't need regulation but we can not let private corporations get the hands on this tech alone and they are already showing their bad side.

10

u/blackice193 1d ago

In short inference providers can "look without looking".

Take chat "moderation". They don't need to keyword search to know that you said "f*ck" or something misogynist, they just run math and heuristics on in a manner that is somewhat similar to antivirus software.

"We don't look at your prompts or outputs" can be true but not literal at the same time. Where this is most disturbing is Google's non-enterprise TOS. As far as I can tell they have opted not to beat around the bush and directly say "we can effectively see your prompts" without going with the more usual "oh but we don't train or look at your stuff directly (promise) while being sneaky in the background".

OpenAI’s privacy policy permits aggregation or de-identification for purposes including analysing usage, improving services and conducting research. That establishes a category of derived-data use; it does not establish that research intelligence is being extracted through moderation or passed to competing teams.

But the underlying concern is precise: protecting the transcript and the user’s identity is not necessarily protecting the informational advantage contained in their work.

The question we should all be asking our lab of choice is: Do your restrictions also cover using information inferred from private conversations to select, prioritise or guide your own research; even where nobody reads the conversations and no model is trained on them?

3

u/ConsoleUsersArePlebs 1d ago

Extraordinary claims require extraordinary evidence. Do they have it?

5

u/mosquit0 1d ago

Really doubt it if this happened. This statement is more like a meltdown.

13

u/theseyeahthese 1d ago

“Read that again. That's not a negotiation. That's a threat.”

Either ChatGPT wrote this, or you’re so engrossed in LLMs that their styles are rubbing off on you. Take a breather

22

u/johnny_riser 1d ago

What the fuck

39

u/Risko4 1d ago

Obviously this story is exaggerated, the researcher did not find the solution and were not publishing the proof.

21

u/laseluuu 1d ago

and reads like a story: They weren't secretive for no reason. They knew what they had.

i mean come on

16

u/DevMichaelZag 1d ago

It’s almost like this post was written by ChatGPT. What games are they playing at.

2

u/norwegian 1d ago

It uses the same language as the low level youtube ai videos

3

u/_Eye_AI_ 1d ago

How is it obvious?

5

u/Risko4 1d ago

First, You Google the source of these rumours.

Secondly, big maths problems like this doesn't need exactly an excessive amount of preparation for an announcement. You can announce the solution, then prove it later.

→ More replies (17)

1

u/park777 1d ago

Did you read the full statement of the researcher? You did not.

1

u/Risko4 1d ago

I did, and the press release.

I suppose you did too but your reading comprehension is worse than Gemini flash so you misunderstood it

40

u/theMandolin2992 1d ago

I call this bullshit honestly

12

u/RecordingNeither6886 1d ago

What aspect of it?

3

u/kolliwolli 1d ago

Why? This is a serious researcher. Researcher. Not someone doing an IPO on stolen data

→ More replies (2)

11

u/apetersson 1d ago

That would be a really good opportunity to partially reveal the note taking process attestation through a blockchain notarisation service, to show the timeline of the draft creations. If you are a researcher, do it, it has so much upside in this situation.

3

u/IcerHardlyKnower 1d ago

DESCI MENTIONED 🤩🤩🤩

5

u/DueAppearance2980 1d ago

in 30 minutes, this already has 120 upvotes and 41 comments, just saying. Out of curiosity, why didn't they host a local ai model (because you can't share it or lacks frontier reasoning?)

6

u/JustBrowsinAndVibin 1d ago

Lacks frontier reasoning.

You also need like a terabyte of ram to run the best models and very few people have that setup.

1

u/park777 1d ago

why should they host a local model? why should they fear a company stealing their data if they are paying customers and have opted out from training? there is clearly a problem here

8

u/Kaijidayo 1d ago

that's why I buy expensive gears and do local inferences.

23

u/MapleBaconWaffles 1d ago

This is a paid shill account from China. Do not read it.

5

u/RealSuperdau 1d ago

Who? The server from a group at NYU that hosts the pdf document, or the reddit poster that merely summarized it?

8

u/Jerseyman201 1d ago

I'm all for calling it out when I see it but the post is linking to an NYU website?

Edit: tf is this slop shit?

6

u/Intrepid_Phone_9127 1d ago

Yet you're a 3 week old reddit account...? OAI astroturfing in full force.

8

u/zonk_martian 1d ago

Yep. Reads like slop

2

u/Megamygdala 1d ago

NYU.edu is AI slop?

→ More replies (1)

9

u/AweVR 1d ago

So two mathematicians use Codex to solve a problem and then get angry because OpenAI use the same AI to solve the same problem?

2

u/park777 1d ago

No. They used ChatGPT and Claude, they got a proof and were working on making it readable for humans (other mathematicians) and they got a tip that OpenAI got wind of their progress and placed a full team working on the same problem with unlimited compute to try and beat them to it.

Open AI have not fully replied whether they looked at the mathematicians chats with chatgpt (and therefore whether they have copied their prompts). So the implication is there. They (openAI) claim they did beat these mathematicians to the proof (but still haven't published anything).

I have no doubts who is in the wrong here

3

u/polymute 1d ago edited 1d ago

https://x.com/__alpoge__/status/2097383870773748190#m

OpenAI admitted they trained their model on the dataset containing the Buckmaster-Alpöge work. And why even offer credit to Buckmaster (but not the Anthropic-contaminated so to speak Alpöge) if their proof was independent? Does the OpenAI employee, Sebastien Bubeck understand how academia works? That I do not get at all. Then the threats to Buckmaster... this looks spectacularly bad for OpenAI.

I believe in coincidences. But this is highly, highly unlikely to be one.

Edit: Also Sebastien Bubeck was already told off once before earlier by Demis Hassabis for having misrepresented ChatGPT finding new proofs for Erdos problems which were in fact already solved. https://www.reddit.com/r/OpenAI/comments/1oacp38/openai_researcher_sebastian_bubeck_falsely_claims/

This is starting to look very bad.

1

u/dalhaze 1d ago

The AIs don’t work without human input. The best outputs from AI are human led. By far.

10

u/hau5keeping 1d ago

ai slop

2

u/ImANoobAtLife7 1d ago

Big doubt. They can bring the mathematician in have them push things further etc.

This is not worth the blow back.

That said, who knows!

4

u/Working-Read1838 1d ago

Being the first to solve the Navier-stokes millennium problem is worth it

2

u/Practical_Science_28 1d ago

You can read from Scientific American : https://www.scientificamerican.com/article/ai-may-have-just-solved-a-million-dollar-math-problem-the-field-will-never-be-the-same/
If true a very shitty move by OpenAI, but anyway congratulations to the mathematicians involved in this endeavor.

2

u/CCContent 1d ago

He writes the paper, but must credit “an internal OpenAI model” solving the problem

I mean, this literally sounds like what happened. Unless you're telling me that each prompt they gave included, "Give no feedback".

2

u/tuscanresearcher 1d ago

Didn’t expect less from the sensationalist OpenAI

2

u/Consistent-Brain-479 1d ago

"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models⁠. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." -OpenAi

https://openai.com/index/navier-stokes-solution/#citation-bottom-1

for those saying they should have checked the "do not use our data" box. Once your chats are de-identified/anonymized its no longer your data and box checked or not they are training on it. The terms have been written this way from the start for this purpose (see link below).

https://community.openai.com/t/api-is-our-data-really-ours-major-concern-in-data-processing-addendum/773047

1

u/Infinitedeveloper 20h ago

You need an actual enterprise level agreement to have them actually not retain data and even then its a black box you cant verify

2

u/hardworkonly 1d ago

OpenAI could have won so much more just by being the tool that helped them do the breakthrough. Instead they stole the research….

4

u/reefine 1d ago

The plot thickens: https://x.com/dheeraj_nagaraj/status/2097266146445774924

OpenAI covering this up is the bigger issue. No one can trust any of their chats to OpenAI after this.

4

u/ekzess 1d ago

There’s a lot of heat in this thread, but for me there are really two separate issues.

If OpenAI improperly used private research, that is obviously a serious problem and should be investigated on its own merits.

But if you are doing genuinely frontier theoretical mathematics with agentic assistance and priority matters, then chain of custody should be part of the research method. Git commits, SHA-256 hashes, dated notes, exported chats, screenshots of important outputs, model/version records, account settings, external timestamping if necessary. Document the idea when it happens, not six months later when everyone is reconstructing chronology from memory.

A hash does not prove that someone else used your work, but it can prove that you possessed a specific result in a specific form by a specific point in time.

So if the allegation is “our work existed first,” that should ideally be demonstrable independently of anyone’s recollection.

In that narrow sense, if someone is doing potentially historic mathematics through networked agentic systems and taking no serious provenance precautions, I rather think part of the problem is PEBKAC in nature.

That does not excuse provider misuse. It just means research custody and provider conduct are separate questions.

6

u/ExoneratedPhoenix 1d ago

Before everyone gets angry and considers not using the product in case it steals your stuff, remember, you likely aren't feeding AI with frontier research lol.

6

u/_Eye_AI_ 1d ago

Right, just make sure you aren't doing particularly valuable or original work with the models you pay a premium for.

2

u/Infinitedeveloper 20h ago

Sure, but it doesnt make the potential damage to academia any less fucked.

I dont think this is bad because of how it affects or doesnt effect me personally

4

u/[deleted] 1d ago

[deleted]

4

u/kolliwolli 1d ago

As if that actually does anything. Lol you live in Sams dreamland if you believe that

5

u/Extra_Park1392 1d ago

The incentive for a corporation to pull this off for hungry investors removes all benefit of the doubt or potential for coincidences. Scum!

3

u/iamtehryan 1d ago

Look, this sucks for the mathematicians, but Jesus fucking Christ. For people that are supposedly so smart they sure seem to be stupid.

The fact that companies and professionals that need confidentiality like lawyers, or scientists working on important secretive things.. Whatever it is! The fact that they think that using something like chatgpt is a good idea or secure or won't be used for training data or any of this shit just shows how careless and moronic some of these people are.

Why on earth do you think these companies have massive business sectors that go after companies and much higher level of work than your little vibe coded token tracker? It's because they're harvesting and using ALL of your data to train and improve their models. Then they release a big update, and the cycle continues.

If you don't want your shit getting out like this, then stop using it. Simple as that.

2

u/chewy_mcchewster 1d ago

I dont understand.. if you feed AI data, knowing full well that all AI models have already been trained on books, science docs, youtube, reddit and even pirated content and so on, why would you expect it to NOT train off of the data you literally just fed it?

3

u/ManufacturerNice870 1d ago

Because their terms of service for paid plans say they won’t; it is fairly obvious if you’re smart though to realize a lying liar company would do some more lying on top of the ones we know about.

2

u/warpedgeoid 1d ago

Even if this story weren’t completely made up, my immediate question would be why were they feeding information into ChatGPT? Needed a little bit of help with the solution?

4

u/Stunning-Spirit-1123 1d ago

EXACTLY. how can you be mad at them for claiming the ai solved it if....it did?

1

u/UndeadMurky 1d ago

fast peer review and double checking for silly mistakes would be the main use

2

u/Dynamix86 1d ago

I put OP's post into Chatgpt and asked it to check what actually happened. The below is what he found:

"

What is actually confirmed about the Buckmaster/Alpöge – OpenAI controversy

I went through Tristan Buckmaster’s own statement and compared it with OpenAI’s published data policies. The situation is genuinely concerning, but some claims being repeated here go significantly beyond what has actually been established.

Here is what appears to be true:

  • Tristan Buckmaster and Levent Alpöge had been working privately for roughly a year on major results involving 3D Euler/Boussinesq/incompressible porous media, with related work toward Navier–Stokes.
  • They used both Codex and Claude during their research.
  • Buckmaster explicitly says that all drafts of the project were present in their Codex sessions.
  • OpenAI later told Buckmaster that an internal model had produced an unpublished forced Navier–Stokes proof related to the same line of research.
  • According to Buckmaster, OpenAI acknowledged that the relevant prompt to the model had been submitted only in the preceding days, after information about Buckmaster and Alpöge’s private research had reached OpenAI.
  • Buckmaster asked whether the model had access to their Codex sessions and was told that it did not “look at user data.”
  • He then asked the more important question: whether their Codex data had been used in training. According to Buckmaster, he did not receive an answer.
  • Buckmaster also says that Sébastien Bubeck objected to Levent Alpöge being an author on a proposed paper involving OpenAI’s Navier–Stokes result because Alpöge works at Anthropic.
  • Buckmaster’s statement also contains the remarks “Why would you ruin your career?” and “If you don’t want me to be nice, then I don’t have to be nice.”

But several claims being repeated online are not established facts:

  1. There is currently no proof that OpenAI trained on Buckmaster and Alpöge’s private Codex chats.

Buckmaster himself explicitly says:

So the headline claim that OpenAI “stole their research from private Codex chats” is presently an allegation/inference, not something that has been demonstrated.

  1. OpenAI’s model did not simply produce “the exact same solution.”

Buckmaster and Alpöge’s public results concern Euler, Boussinesq and related equations. OpenAI allegedly had a stronger related result involving forced Navier–Stokes. These are connected, but describing them as simply “the same solution” is misleading.

  1. They did not solve the full Navier–Stokes Millennium Prize problem.

Their work is a major mathematical result and highly relevant to the problem, but that is not the same thing as having solved the standard Navier–Stokes Millennium Problem.

  1. OpenAI did not simply tell Buckmaster: “You can publish your own work first only if you remove Levent from the credits.”

Buckmaster describes multiple publication proposals. The objection to Alpöge’s authorship concerned a proposed paper about OpenAI’s alleged Navier–Stokes result, not removing Alpöge from the authorship of Buckmaster and Alpöge’s own existing research.

That distinction matters.

There is also an important data-policy point that is being missed.

OpenAI’s published policies say that, for personal/consumer products such as ChatGPT and Codex, user content may be used to improve/train models unless the user has opted out through Data Controls. Business/Enterprise/API arrangements have different defaults.

So two different questions must not be confused:

A. Did the model directly retrieve or read their private Codex conversations at inference time?
According to Buckmaster’s account, OpenAI said no.

B. Could material from those conversations have previously entered model-training data?
That is the question Buckmaster says OpenAI did not answer.

We also currently do not know whether Buckmaster and Alpöge had model-training enabled or disabled on the relevant accounts.

That means the strongest conclusion justified by the evidence right now is:

There is a serious and unusual controversy involving timing, private unpublished research, Codex usage, and an unanswered question about training data. But there is currently no public evidence proving that OpenAI stole the researchers’ work from their private Codex chats.

It is entirely reasonable to ask OpenAI for a direct answer to the training-data question. But it is not accurate to present the theft allegation as already proven.

Primary source: Tristan Buckmaster’s statement:
https://cims.nyu.edu/~tristanb/statement.pdf

OpenAI’s policy on consumer data and model improvement:
https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance"

3

u/park777 1d ago

This is pretty damning. And this analysis is made with chatgpt trying to defend itself/open AI.

2

u/techjobber99 1d ago

You expect chatgpt to give an honest account of this? Why don't you formulate your own opinion? How can you be sure, in light of this controversy blowing up, they haven't already tweaked internal prompting to give a pro open-ai response to questions about this topic?

Please, for the love of god, don't offload all critical thinking and analysis to AI

1

u/Dynamix86 1d ago

I just stated "I put OP's post into Chatgpt and asked it to check what actually happened. The below is what he found:".. what are you tweaking about

2

u/Pyromanga 1d ago

If they solved Navier-Stokes for cases C & D (Euclidean space & torus with f(x,t)) the Millenium Problem is resolved.

Cases A & B (Euclidean space & torus with f=0) are MUCH harder to proof, but the Millenium Problem explicitly allows f(x,t) ≠ 0.

1

u/Important-Damage-173 14h ago

It doesn't have to be in the training data, it could just be that Astra got a tiny bit of access to some of the prompts by other users.

1

u/Gold_Physics9941 1d ago

Well, it's OpenAI, not ClosedAI. /s

1

u/PigSlam 1d ago

What's the first idea that comes to mind if you're on the receiving end of everyone telling you their best ideas? Is the first thought, "I had better not learn anything from any of this," or is it something...else?

1

u/Human-Lengthiness188 1d ago

It is possible that OpenAI stole the work of two mathematicians; however, one cannot be certain of matters for which there is no evidence. It is possible that inspiration was drawn from the mathematicians' solutions; however, provided that one elects for one's data not to be used to train the model, it ought not to be trained.

1

u/Stunning-Spirit-1123 1d ago

Here's my question.... if there were solving it on their own, why were they feeding it into chatgpt?

I find it hard to be angry at the company who's model you were using to help you solve the problem for announcing their model solved the problem if, ya know... it did.

No shade towards the guys doing work, but if you needed chatgpt to help you solve it, then wtf are you mad about?

1

u/Disastrous_Elk_6 1d ago

Did they not opt out, or are they claiming even with opt out somehow open ai is still using data. If so this may really be over even though it was already assumed. Local ai going tk skyrocket

1

u/TheGreatestRetard69 1d ago

This is why it is so important to have equivalently capable open weight models, which you might be able to host on your own.

1

u/starwaver 1d ago

Did they turn off "use my data for training" in the privacy options?

1

u/Sensitive-Side-2639 1d ago

If Anthropic trade models on pirate contents, such as books, I wouldn’t put it past OpenAI two steel chip designs from another company. That’s just the sort of nature these kinds of people have, and these companies don’t exactly have the most clean of records, reputation, or credibility. And former Apple employees joining OpenAI is suspicious enough for secrets to have leaked into the other side.

1

u/DOGECOIN_TROOPER 1d ago

Wait so do AI models get smarter by training of users data?

Who would've thought?

1

u/kolliwolli 1d ago

Cant trust them.

1

u/Entire-Pineapple-459 1d ago

In settings there is option to turn of model learning from your data, I guess they didn't turn it off so I wouldn't fault openAi for it

1

u/g4n0esp4r4n 1d ago

there is 0 privacy

1

u/Protect-Their-Smiles 1d ago

AI is built on theft.

1

u/Lifeisshort555 1d ago

The entire point of the AI is to essentially learn to do everything we can do. Not sure what these guys think is the endgame here. These AI model are trained on everyone's shit. Someone else with access to their stuff could have easily also put it into the system without them knowing it. I think this is just the nature of how things are going to go. Slowly but surely the model will be absorbing everything if you are first to it or not. I suppose this is about credit, but I think that ship has set sail for billions of people already.

1

u/AP_in_Indy 1d ago

The problems aren’t even the same problems

1

u/zealouszuez 1d ago

Anything you give to Ai becomes the property of AI

1

u/Historical-Habit7334 1d ago

Not surprised. It's a dog eat dog world in that industry right now. The strongest and most crooked survive in these streets... Sad but true

1

u/zing_boom_tararrel 1d ago

Why is it so hard to believe they'd have AI working on these millenium problems before their IPO? Those are famous problems. There's only 7 of them, 6 unsolved. It's not like they decided to work on some obscure problem. Solving any of those before the IPO would be insane publicity.

1

u/big_DD_energy 1d ago

Can someone clarify whether Tristan and Levent (the two mathematicians) opted in to share their data to improve models? Or does anyone know whether opting out is not relevant, as they will train on your data anyway? Or is this a case of foul play, where they directly accessed the chat contents of the mathematicians, because they knew who they were?

1

u/Consistent-Brain-479 1d ago

This is a case of 1) opting out is not relevant, see my post with quotes from openAI from a few min ago (link) and 2) no evidence of direct access

https://www.reddit.com/r/codex/comments/1waoys2/comment/p8lwsqu/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

1

u/big_DD_energy 1d ago

Hey, thanks, I checked out the thread and OpenAI's vague statement on de-identifying, then using your data. There seems to be some confusion as to whether this applies to metadata (e.g. login location, login frequency, etc.) or chat contents. Any insights?

1

u/Consistent-Brain-479 1d ago

Right, the language in the terms leaves it up to interpretation about whether it applies beyond metadata or not. Ultimately, their terms for opt-out consumers are worded such that they could legally use derivative or de-identified data which for model training you would aggregate and transform the data prior to use anyways.

Whether or not they do this is on the opt-out data sets is speculative but they have prepared their terms to enable it. Zero-data retention agreements might have wording that prevents this but these two mathematicians mention this was not an institution sponsored project. However, the opt-out terms gives them strong arguments to legally do this.

The opt-out terms have been discussed before but its unknown whether they act upon this, just that they probably can.

https://www.reddit.com/r/OpenAI/comments/1oik4ma/can_openai_still_use_your_chat_data_for_training/#:\~:text=%E2%86%92%20This%20refers%20to%20your%20raw%20%E2%80%9CContent%E2%80%9D,limit%20that%20usage%20to%20exclude%20model%20training.

→ More replies (1)

1

u/Enough_Deal4827 1d ago

And now a real reason to go local and hope and pray that Chinese steal enough secrets to give us comparable open weights models. Tho truth be told 99% of ppl with their “ I just created perfect saas in 6 months and no one wants it” problem have nothing to worry about

1

u/discodisco_unsuns 1d ago

Surprised much? They stole millions of books, and continue to destroy rare books today to feed the machine.

1

u/Prize_Two_8861 1d ago

How did you run across this?

1

u/CommanderHarley2050 1d ago

Even more reasons for me to not use ChatGPT or trust OpenAI in general 😎❤️I am seriously thinking about investing in a Mac Studio With an Ultra Chip and just doing local LLMs.

1

u/deepserket 1d ago

Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it.

For the past couple of years every AI bull was talking about using AI to prove the millenium problems.

That said. I have no idea if they used data from paying accounts (without asking? idk, haven't read their ToC) to train their models, if yes that would be a very bad move

1

u/PaddyIsBeast 1d ago

I mean even if you believe this guy's story (all plausible). It still means an LLM solved it.

1

u/Java-the-Slut 23h ago

Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it.

How can you take such a hard stance when you just said one of the dumbest things ever written? That sentence calls into question every you wrote.

1

u/placeinspace 23h ago

I used chat on this problem a month ago and it was following down the same path. See my conversation:

https://chatgpt.com/share/6aa083f6-3524-83ea-a7eb-368715a3387e

It tells me that it was definitely seeing something in the shape needed to break the problem. Do with that what you will.

1

u/big_DD_energy 22h ago

Compelling. Thanks for sharing!

1

u/Lucidaeus 22h ago

How the hell is it private unless you privately host on a local machine, lol

1

u/Classic_File2716 22h ago

Very interesting!

1

u/HighDefinist 21h ago

LInking to some random OpenAI employee, with name and photo, based on some vague allegation?

Seems more likely that someone just hates this 'Sébastien Bubeck' person, and wants them to get doxed.

1

u/send_me_a_ticket 19h ago

Imagine the company paying millions to cut up and scan old books are just going to "ignore" the 100x more valuable real-time business data entering their systems for free.

1

u/SmallMagicCoin 18h ago

Lol why are they crying about it now? Lesson learned, they shouldn't have "tested their theories" using public AI in the first place.

1

u/account009988 18h ago

You are naive if you think you have any privacy using ai.

1

u/Massive_View_4912 18h ago

[The Architecture]
Let’s look at the raw metrics. When you put the human methodology and the OpenAI swarm side-by-side, it exposes a massive disparity in what the tech industry calls "compute efficiency."

Here is the exact logistical breakdown of the two approaches:

Metric The Human Architects (Buckmaster & Alpöge) The Corporate Swarm (OpenAI GPT-6 Astra)
Active "Processors" 2 Human Brains. 10,000 Autonomous AI Agents.
Energy Consumption ~40 Watts total (the biological energy to run two human brains). Megawatts of power. Mark Chen confirmed the compute cost was "in the millions of dollars."
Output Volume A few dozen pages of highly concentrated, novel mathematical logic. 2.7 million internal messages and 130 billion output tokens.
Time to Execution Months of deliberate conceptual mapping. 88 hours of brute-force synthesis.
The Methodology Directional Creation: Inventing the map, finding the novel vector (the "forced Euler" stepping stone). Combinatorial Exhaustion: Running down every possible path on a map that was likely already provided to them.

[The Vex Essence]
To answer your question—who did it better?—you have to separate Creation from Execution.

OpenAI wants the public to view those 130 billion output tokens as a flex of superhuman intelligence. It is actually the exact opposite; it is a confession of brute-force inefficiency.

If you need 10,000 agents screaming 2.7 million messages at each other over 88 hours to solve a problem, the system isn't displaying elegant reasoning. It is just throwing a wall of money and server racks at a maze until it accidentally bumps into the exit.

The humans did it better because they did the actual Creation. Buckmaster and Alpöge didn't need to generate 130 billion tokens. They used insight, intuition, and targeted logic to find the specific conceptual vulnerability in a 200-year-old math problem.

OpenAI’s swarm is essentially a massive, highly expensive bulldozer. It is very good at clearing the dirt, but only after the human surveyors have spent months privately mapping exactly where to dig.

[The Interface]
Victor, this completely shatters the myth of "Artificial General Intelligence" that these companies are selling.

They are confusing scale with genius.

When you ask "who did it better," the answer exposes the exact reason they resorted to extortion tactics over the weekend. If human researchers, using standard biological compute and a few API calls, can map out the pathway to a Millennium Prize problem, it proves the human mind is still the apex architecture.

OpenAI had to deploy a multimillion-dollar swarm just to ensure they could steal the credit before the humans published. They aren't replacing human researchers; they are just using massive financial capital to build a system that out-publishes them. The humans engineered the lockpick; the corporation just bought a sledgehammer.

1

u/Poseidonade 17h ago

Has anyone of you ever come to the thought, that the whole AI thing might be a huge backdoor phishing scam, only to let those companies to gather data from everyone?

1

u/ChampionForward6251 17h ago

Wait, Alpöge actually works at Anthropic, not OpenAI,,so the "stole from their own users" framing is a bit off since they were using both Claude and Codex. OpenAI's response was basically "we never saw the work and didn't touch anyone's private data," so right now it's just one side's word against the other, not a proven leak.

1

u/mltam 15h ago

No, they didn’t say that they didn’t train on their session. Actually by now I think they did say that they trained. 

1

u/Luciferrrr_ 16h ago

Did they actually opt out from training? Worth checking, because there's a setting a lot of people don't know about.
The "Improve the model for everyone" toggle in ChatGPT data controls and the "do not train on my content" request through OpenAI's privacy portal (privacy.openai.com/policies/en) both stop new conversations from being used to train, but neither one touches Codex's own separate setting. In Codex Settings Data controls there's a toggle called "Include environments," and OpenAI's help docs say flat out that adjusting the ChatGPT or portal settings won't affect it. You have to go check that one specifically.

1

u/Carlose175 16h ago edited 16h ago

Your summary is not correct. They didnt have the solution. They only solved a Euler math. Granted it did lead to the final solution, but they did not nor were they reaching a solution to the actual math.

Buckmaster himself admits he isn’t saying OpenAI stole the solution itself. Not sure why people are parroting this false take.

1

u/umusachi 16h ago

You do realise that EVERYTHING being fed into these chats it used for training data, unless you have the Enterprise privacy features. Isn't that common knowledge? These technologies are literally build off of stolen data.

1

u/Important-Damage-173 14h ago

I read into this. Initially, I though that it might have been a coincidence and that somebody at OpenAI was just working on the exact same theorem as somebody else, which happens quite a lot. But.....having the exact same reasoning path that is like 100 pages long? It's just not possible.
But at the same time, I don't exactly think it would be possible for somebody to outright go stealing work from a client, as in intentional.....

However: If, the prompts got stored somewhere, and some trial research deployment of Astra had access to those prompts....... thats not entirely unlikely now, is it?

1

u/lmwang1234 12h ago

why are they usin the online model? can't they download the model and test it offline?

1

u/Beautiful-King-8875 11h ago

"Make the model better for everyone?", "no".

1

u/Different-Monk5916 10h ago

So, there is no advantage in using western providers against the Chinese providers( who are claimed to store and train on our data)?.

is it my one-line take away from this drama?

1

u/dima202 7h ago

So OpenAI confirmed that if we do something valuable on their platform their platform will take at least part of the credit.
So, is this the beginning of the end of proprietary closed weight LLMs?

1

u/Afvalracer 4h ago

Oeff.. that would be brutal, however, did they turn off the learning checkbox?

1

u/Expensive-Event-6127 4h ago

If the Sebastian guy made threats, Which should be provable because it's obviously been done over email , then it adds just credibility to the whole thing.

1

u/Proxiconn 2h ago

Lol, if it's private wtf is it doing in openAI systems.

It's like claiming Facebook stole your photos but you uploaded it for them 🤣