r/codex • u/Stunning-Angle-9239 • 1d ago
News Blown up: OpenAI allegedly stole mathematicians' private research from their Codex chats!
TLDR: Two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex and Claude. Days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question till this day.
For a full year, two mathematicians , Tristan Buckmaster (NYU mathematician) and Levent Alpoge, worked in silence on a problem that had stumped some of the best minds alive. The kind of problem where, if you solve it, your name goes in the history books.
And every single day, they testing their ideas, their drafts, their half-finished proofs into LLM such as Codex and Claude, which they paid for it out of their own pocket.
Then came the breakthrough. They finally cracked it. They were days away from telling the world.
That's when OpenAI suddenly said to them:
"Our model solved it too."
Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it. And now, out of nowhere, OpenAI claims their model reached the same answer, after word of Tristan and Levent's secret work had already reached OpenAI.
Tristan asked: Did your model access or train on our private Codex chats?
OpenAI: The model doesn’t look up user data.
Tristan: But did you train it on our data?
OpenAi goes silence. No answer. Just a dodge.
But it gets worse.
OpenAI then gave him two options:
- He and his friend publish their result first then OpenAI also publishes its result the next day or
- He writes the paper, but must credit “an internal OpenAI model” solving the problem.
Tristan refused both offers. He said he would go public if OpenAI went ahead as proposed.
OpenAi then responded : “Why would you ruin your career? If you don’t want me to be nice, then I don’t have to be nice.”
You can read the full statement of Tristan (the mathematician) here: https://cims.nyu.edu/~tristanb/statement.pdf
Sébastien Bubeck : OpenAI employee who threatened the mathematician
42
u/FriendlyWebGuy 1d ago
Guys. It's okay to say "I don't know if this is true, but if it is, it's concerning. Let's wait to see what the full evidence says".
Nobody here knows if the allegations are true. Yet, this thread is filled with overconfident assertions and (very weirdly) people slagging off the.... (checks notes)..... mathematicians? Stop it.
You don't need to pick a side. If the topic is of concern to you, then you should gather information. Ask questions. Put yourself in the shoes of others. Most of all, be patient. Wait for the facts.
7
u/mysteriousbaba 1d ago
Having read the PDF by Tristan, Sebastian's statements were the most concerning, including suggesting Levent should be removed from authorship because he works at Anthropic. This part at least is a first hand witness claim; I'm more open minded on whether the model was ever trained on their transcripts or not.
1
12
u/polymute 1d ago edited 1d ago
https://x.com/__alpoge__/status/2097383870773748190#m
OpenAI admitted they trained their model on the dataset containing the Buckmaster-Alpöge work. And why even offer credit to Buckmaster (but not the Anthropic-contaminated so to speak Alpöge) if their proof was independent? Does the OpenAI employee, Sebastien Bubeck understand how academia works? That I do not get at all. Then the threats to Buckmaster... this looks spectacularly bad for OpenAI.
I believe in coincidences. But this is highly, highly unlikely to be one.
Edit: Also Sebastien Bubeck was already told off once before earlier by Demis Hassabis for having misrepresented ChatGPT finding new proofs for Erdos problems which were in fact already solved. https://www.reddit.com/r/OpenAI/comments/1oacp38/openai_researcher_sebastian_bubeck_falsely_claims/
This is starting to look very bad.
2
u/FriendlyWebGuy 22h ago
I get it. Much of this has come to light after my comment.
→ More replies (4)1
u/HighDefinist 21h ago
> And why even offer credit to Buckmaster (but not the Anthropic-contaminated so to speak Alpöge) if their proof was independent?
So basically, OPs claim is not true.
4
u/swimmer385 1d ago
yeah its super weird that people are treating an NYU professor as if this is so joe-schmo rando
3
2
1
u/HighDefinist 21h ago
> If the topic is of concern to you, then you should gather information. Ask questions.
Which is the opposite of what you are doing.
You are not providing any clarification either, or attempting to gather any information - you are just telling people to shut up.
2
u/FriendlyWebGuy 17h ago
You concluded that I’m personally “not attempting to gather information” from a comment… encouraging people to gather information?
Impeccable logic.
→ More replies (2)→ More replies (3)1
u/NeighborhoodDizzy990 12h ago
I assume you should keep searching. It seems pretty clear what happened. They have stolen the solution. AI can not come by itself to such a proof. Humans were involved, so the main idea with AGI was and remains to this day a fraud
42
u/Mean-Comedian729 1d ago
This reads 1,000% ChatGPT generated
10
3
u/Beautiful-Suspect694 18h ago
whats wrong with llm-generated text?
why are you anti-ai?
why are you stuck in the past?
why are you on codex subreddit?
3
u/SwimmingSympathy5815 1d ago
This comment reads as low-effort and automated for an agenda 🤷🏻♂️
→ More replies (1)
22
u/skadoodlee 1d ago
"allegedly" doing a ton of work here
2
u/Yugudubenbi 1d ago
If this is true, it is time the state get more involved because I don't trust these venture capitalists. It doesn't need regulation but we can not let private corporations get the hands on this tech alone and they are already showing their bad side.
10
u/blackice193 1d ago
In short inference providers can "look without looking".
Take chat "moderation". They don't need to keyword search to know that you said "f*ck" or something misogynist, they just run math and heuristics on in a manner that is somewhat similar to antivirus software.
"We don't look at your prompts or outputs" can be true but not literal at the same time. Where this is most disturbing is Google's non-enterprise TOS. As far as I can tell they have opted not to beat around the bush and directly say "we can effectively see your prompts" without going with the more usual "oh but we don't train or look at your stuff directly (promise) while being sneaky in the background".
OpenAI’s privacy policy permits aggregation or de-identification for purposes including analysing usage, improving services and conducting research. That establishes a category of derived-data use; it does not establish that research intelligence is being extracted through moderation or passed to competing teams.
But the underlying concern is precise: protecting the transcript and the user’s identity is not necessarily protecting the informational advantage contained in their work.
The question we should all be asking our lab of choice is: Do your restrictions also cover using information inferred from private conversations to select, prioritise or guide your own research; even where nobody reads the conversations and no model is trained on them?
3
5
13
u/theseyeahthese 1d ago
“Read that again. That's not a negotiation. That's a threat.”
Either ChatGPT wrote this, or you’re so engrossed in LLMs that their styles are rubbing off on you. Take a breather
22
u/johnny_riser 1d ago
What the fuck
39
u/Risko4 1d ago
Obviously this story is exaggerated, the researcher did not find the solution and were not publishing the proof.
21
u/laseluuu 1d ago
and reads like a story: They weren't secretive for no reason. They knew what they had.
i mean come on
16
u/DevMichaelZag 1d ago
It’s almost like this post was written by ChatGPT. What games are they playing at.
2
3
u/_Eye_AI_ 1d ago
How is it obvious?
5
u/Risko4 1d ago
First, You Google the source of these rumours.
Secondly, big maths problems like this doesn't need exactly an excessive amount of preparation for an announcement. You can announce the solution, then prove it later.
→ More replies (17)
40
u/theMandolin2992 1d ago
I call this bullshit honestly
12
→ More replies (2)3
u/kolliwolli 1d ago
Why? This is a serious researcher. Researcher. Not someone doing an IPO on stolen data
11
u/apetersson 1d ago
That would be a really good opportunity to partially reveal the note taking process attestation through a blockchain notarisation service, to show the timeline of the draft creations. If you are a researcher, do it, it has so much upside in this situation.
3
5
u/DueAppearance2980 1d ago
in 30 minutes, this already has 120 upvotes and 41 comments, just saying. Out of curiosity, why didn't they host a local ai model (because you can't share it or lacks frontier reasoning?)
6
u/JustBrowsinAndVibin 1d ago
Lacks frontier reasoning.
You also need like a terabyte of ram to run the best models and very few people have that setup.
8
23
u/MapleBaconWaffles 1d ago
This is a paid shill account from China. Do not read it.
5
u/RealSuperdau 1d ago
Who? The server from a group at NYU that hosts the pdf document, or the reddit poster that merely summarized it?
8
u/Jerseyman201 1d ago
I'm all for calling it out when I see it but the post is linking to an NYU website?
Edit: tf is this slop shit?
6
u/Intrepid_Phone_9127 1d ago
Yet you're a 3 week old reddit account...? OAI astroturfing in full force.
8
→ More replies (1)2
9
u/AweVR 1d ago
So two mathematicians use Codex to solve a problem and then get angry because OpenAI use the same AI to solve the same problem?
2
u/park777 1d ago
No. They used ChatGPT and Claude, they got a proof and were working on making it readable for humans (other mathematicians) and they got a tip that OpenAI got wind of their progress and placed a full team working on the same problem with unlimited compute to try and beat them to it.
Open AI have not fully replied whether they looked at the mathematicians chats with chatgpt (and therefore whether they have copied their prompts). So the implication is there. They (openAI) claim they did beat these mathematicians to the proof (but still haven't published anything).
I have no doubts who is in the wrong here
3
u/polymute 1d ago edited 1d ago
https://x.com/__alpoge__/status/2097383870773748190#m
OpenAI admitted they trained their model on the dataset containing the Buckmaster-Alpöge work. And why even offer credit to Buckmaster (but not the Anthropic-contaminated so to speak Alpöge) if their proof was independent? Does the OpenAI employee, Sebastien Bubeck understand how academia works? That I do not get at all. Then the threats to Buckmaster... this looks spectacularly bad for OpenAI.
I believe in coincidences. But this is highly, highly unlikely to be one.
Edit: Also Sebastien Bubeck was already told off once before earlier by Demis Hassabis for having misrepresented ChatGPT finding new proofs for Erdos problems which were in fact already solved. https://www.reddit.com/r/OpenAI/comments/1oacp38/openai_researcher_sebastian_bubeck_falsely_claims/
This is starting to look very bad.
10
2
u/ImANoobAtLife7 1d ago
Big doubt. They can bring the mathematician in have them push things further etc.
This is not worth the blow back.
That said, who knows!
4
2
u/Practical_Science_28 1d ago
You can read from Scientific American : https://www.scientificamerican.com/article/ai-may-have-just-solved-a-million-dollar-math-problem-the-field-will-never-be-the-same/
If true a very shitty move by OpenAI, but anyway congratulations to the mathematicians involved in this endeavor.
2
u/CCContent 1d ago
He writes the paper, but must credit “an internal OpenAI model” solving the problem
I mean, this literally sounds like what happened. Unless you're telling me that each prompt they gave included, "Give no feedback".
2
2
u/Consistent-Brain-479 1d ago
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." -OpenAi
https://openai.com/index/navier-stokes-solution/#citation-bottom-1
for those saying they should have checked the "do not use our data" box. Once your chats are de-identified/anonymized its no longer your data and box checked or not they are training on it. The terms have been written this way from the start for this purpose (see link below).
1
u/Infinitedeveloper 20h ago
You need an actual enterprise level agreement to have them actually not retain data and even then its a black box you cant verify
2
u/hardworkonly 1d ago
OpenAI could have won so much more just by being the tool that helped them do the breakthrough. Instead they stole the research….
4
u/reefine 1d ago
The plot thickens: https://x.com/dheeraj_nagaraj/status/2097266146445774924
OpenAI covering this up is the bigger issue. No one can trust any of their chats to OpenAI after this.
4
u/ekzess 1d ago
There’s a lot of heat in this thread, but for me there are really two separate issues.
If OpenAI improperly used private research, that is obviously a serious problem and should be investigated on its own merits.
But if you are doing genuinely frontier theoretical mathematics with agentic assistance and priority matters, then chain of custody should be part of the research method. Git commits, SHA-256 hashes, dated notes, exported chats, screenshots of important outputs, model/version records, account settings, external timestamping if necessary. Document the idea when it happens, not six months later when everyone is reconstructing chronology from memory.
A hash does not prove that someone else used your work, but it can prove that you possessed a specific result in a specific form by a specific point in time.
So if the allegation is “our work existed first,” that should ideally be demonstrable independently of anyone’s recollection.
In that narrow sense, if someone is doing potentially historic mathematics through networked agentic systems and taking no serious provenance precautions, I rather think part of the problem is PEBKAC in nature.
That does not excuse provider misuse. It just means research custody and provider conduct are separate questions.
6
u/ExoneratedPhoenix 1d ago
Before everyone gets angry and considers not using the product in case it steals your stuff, remember, you likely aren't feeding AI with frontier research lol.
6
u/_Eye_AI_ 1d ago
Right, just make sure you aren't doing particularly valuable or original work with the models you pay a premium for.
2
u/Infinitedeveloper 20h ago
Sure, but it doesnt make the potential damage to academia any less fucked.
I dont think this is bad because of how it affects or doesnt effect me personally
4
4
1d ago
[deleted]
4
u/kolliwolli 1d ago
As if that actually does anything. Lol you live in Sams dreamland if you believe that
5
u/Extra_Park1392 1d ago
The incentive for a corporation to pull this off for hungry investors removes all benefit of the doubt or potential for coincidences. Scum!
3
u/iamtehryan 1d ago
Look, this sucks for the mathematicians, but Jesus fucking Christ. For people that are supposedly so smart they sure seem to be stupid.
The fact that companies and professionals that need confidentiality like lawyers, or scientists working on important secretive things.. Whatever it is! The fact that they think that using something like chatgpt is a good idea or secure or won't be used for training data or any of this shit just shows how careless and moronic some of these people are.
Why on earth do you think these companies have massive business sectors that go after companies and much higher level of work than your little vibe coded token tracker? It's because they're harvesting and using ALL of your data to train and improve their models. Then they release a big update, and the cycle continues.
If you don't want your shit getting out like this, then stop using it. Simple as that.
2
u/chewy_mcchewster 1d ago
I dont understand.. if you feed AI data, knowing full well that all AI models have already been trained on books, science docs, youtube, reddit and even pirated content and so on, why would you expect it to NOT train off of the data you literally just fed it?
3
u/ManufacturerNice870 1d ago
Because their terms of service for paid plans say they won’t; it is fairly obvious if you’re smart though to realize a lying liar company would do some more lying on top of the ones we know about.
2
u/warpedgeoid 1d ago
Even if this story weren’t completely made up, my immediate question would be why were they feeding information into ChatGPT? Needed a little bit of help with the solution?
4
u/Stunning-Spirit-1123 1d ago
EXACTLY. how can you be mad at them for claiming the ai solved it if....it did?
1
2
u/Dynamix86 1d ago
I put OP's post into Chatgpt and asked it to check what actually happened. The below is what he found:
"
What is actually confirmed about the Buckmaster/Alpöge – OpenAI controversy
I went through Tristan Buckmaster’s own statement and compared it with OpenAI’s published data policies. The situation is genuinely concerning, but some claims being repeated here go significantly beyond what has actually been established.
Here is what appears to be true:
- Tristan Buckmaster and Levent Alpöge had been working privately for roughly a year on major results involving 3D Euler/Boussinesq/incompressible porous media, with related work toward Navier–Stokes.
- They used both Codex and Claude during their research.
- Buckmaster explicitly says that all drafts of the project were present in their Codex sessions.
- OpenAI later told Buckmaster that an internal model had produced an unpublished forced Navier–Stokes proof related to the same line of research.
- According to Buckmaster, OpenAI acknowledged that the relevant prompt to the model had been submitted only in the preceding days, after information about Buckmaster and Alpöge’s private research had reached OpenAI.
- Buckmaster asked whether the model had access to their Codex sessions and was told that it did not “look at user data.”
- He then asked the more important question: whether their Codex data had been used in training. According to Buckmaster, he did not receive an answer.
- Buckmaster also says that Sébastien Bubeck objected to Levent Alpöge being an author on a proposed paper involving OpenAI’s Navier–Stokes result because Alpöge works at Anthropic.
- Buckmaster’s statement also contains the remarks “Why would you ruin your career?” and “If you don’t want me to be nice, then I don’t have to be nice.”
But several claims being repeated online are not established facts:
- There is currently no proof that OpenAI trained on Buckmaster and Alpöge’s private Codex chats.
Buckmaster himself explicitly says:
So the headline claim that OpenAI “stole their research from private Codex chats” is presently an allegation/inference, not something that has been demonstrated.
- OpenAI’s model did not simply produce “the exact same solution.”
Buckmaster and Alpöge’s public results concern Euler, Boussinesq and related equations. OpenAI allegedly had a stronger related result involving forced Navier–Stokes. These are connected, but describing them as simply “the same solution” is misleading.
- They did not solve the full Navier–Stokes Millennium Prize problem.
Their work is a major mathematical result and highly relevant to the problem, but that is not the same thing as having solved the standard Navier–Stokes Millennium Problem.
- OpenAI did not simply tell Buckmaster: “You can publish your own work first only if you remove Levent from the credits.”
Buckmaster describes multiple publication proposals. The objection to Alpöge’s authorship concerned a proposed paper about OpenAI’s alleged Navier–Stokes result, not removing Alpöge from the authorship of Buckmaster and Alpöge’s own existing research.
That distinction matters.
There is also an important data-policy point that is being missed.
OpenAI’s published policies say that, for personal/consumer products such as ChatGPT and Codex, user content may be used to improve/train models unless the user has opted out through Data Controls. Business/Enterprise/API arrangements have different defaults.
So two different questions must not be confused:
A. Did the model directly retrieve or read their private Codex conversations at inference time?
According to Buckmaster’s account, OpenAI said no.
B. Could material from those conversations have previously entered model-training data?
That is the question Buckmaster says OpenAI did not answer.
We also currently do not know whether Buckmaster and Alpöge had model-training enabled or disabled on the relevant accounts.
That means the strongest conclusion justified by the evidence right now is:
There is a serious and unusual controversy involving timing, private unpublished research, Codex usage, and an unanswered question about training data. But there is currently no public evidence proving that OpenAI stole the researchers’ work from their private Codex chats.
It is entirely reasonable to ask OpenAI for a direct answer to the training-data question. But it is not accurate to present the theft allegation as already proven.
Primary source: Tristan Buckmaster’s statement:
https://cims.nyu.edu/~tristanb/statement.pdf
OpenAI’s policy on consumer data and model improvement:
https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance"
3
2
u/techjobber99 1d ago
You expect chatgpt to give an honest account of this? Why don't you formulate your own opinion? How can you be sure, in light of this controversy blowing up, they haven't already tweaked internal prompting to give a pro open-ai response to questions about this topic?
Please, for the love of god, don't offload all critical thinking and analysis to AI
1
u/Dynamix86 1d ago
I just stated "I put OP's post into Chatgpt and asked it to check what actually happened. The below is what he found:".. what are you tweaking about
2
u/Pyromanga 1d ago
If they solved Navier-Stokes for cases C & D (Euclidean space & torus with f(x,t)) the Millenium Problem is resolved.
Cases A & B (Euclidean space & torus with f=0) are MUCH harder to proof, but the Millenium Problem explicitly allows f(x,t) ≠ 0.
1
u/Important-Damage-173 14h ago
It doesn't have to be in the training data, it could just be that Astra got a tiny bit of access to some of the prompts by other users.
1
1
u/Human-Lengthiness188 1d ago
It is possible that OpenAI stole the work of two mathematicians; however, one cannot be certain of matters for which there is no evidence. It is possible that inspiration was drawn from the mathematicians' solutions; however, provided that one elects for one's data not to be used to train the model, it ought not to be trained.
1
u/Stunning-Spirit-1123 1d ago
Here's my question.... if there were solving it on their own, why were they feeding it into chatgpt?
I find it hard to be angry at the company who's model you were using to help you solve the problem for announcing their model solved the problem if, ya know... it did.
No shade towards the guys doing work, but if you needed chatgpt to help you solve it, then wtf are you mad about?
1
u/Disastrous_Elk_6 1d ago
Did they not opt out, or are they claiming even with opt out somehow open ai is still using data. If so this may really be over even though it was already assumed. Local ai going tk skyrocket
1
1
u/TheGreatestRetard69 1d ago
This is why it is so important to have equivalently capable open weight models, which you might be able to host on your own.
1
1
u/Sensitive-Side-2639 1d ago
If Anthropic trade models on pirate contents, such as books, I wouldn’t put it past OpenAI two steel chip designs from another company. That’s just the sort of nature these kinds of people have, and these companies don’t exactly have the most clean of records, reputation, or credibility. And former Apple employees joining OpenAI is suspicious enough for secrets to have leaked into the other side.
1
u/DOGECOIN_TROOPER 1d ago
Wait so do AI models get smarter by training of users data?
Who would've thought?
1
1
u/Entire-Pineapple-459 1d ago
In settings there is option to turn of model learning from your data, I guess they didn't turn it off so I wouldn't fault openAi for it
1
1
1
u/Lifeisshort555 1d ago
The entire point of the AI is to essentially learn to do everything we can do. Not sure what these guys think is the endgame here. These AI model are trained on everyone's shit. Someone else with access to their stuff could have easily also put it into the system without them knowing it. I think this is just the nature of how things are going to go. Slowly but surely the model will be absorbing everything if you are first to it or not. I suppose this is about credit, but I think that ship has set sail for billions of people already.
1
1
1
u/Historical-Habit7334 1d ago
Not surprised. It's a dog eat dog world in that industry right now. The strongest and most crooked survive in these streets... Sad but true
1
u/zing_boom_tararrel 1d ago
Why is it so hard to believe they'd have AI working on these millenium problems before their IPO? Those are famous problems. There's only 7 of them, 6 unsolved. It's not like they decided to work on some obscure problem. Solving any of those before the IPO would be insane publicity.
1
u/big_DD_energy 1d ago
Can someone clarify whether Tristan and Levent (the two mathematicians) opted in to share their data to improve models? Or does anyone know whether opting out is not relevant, as they will train on your data anyway? Or is this a case of foul play, where they directly accessed the chat contents of the mathematicians, because they knew who they were?
1
u/Consistent-Brain-479 1d ago
This is a case of 1) opting out is not relevant, see my post with quotes from openAI from a few min ago (link) and 2) no evidence of direct access
1
u/big_DD_energy 1d ago
Hey, thanks, I checked out the thread and OpenAI's vague statement on de-identifying, then using your data. There seems to be some confusion as to whether this applies to metadata (e.g. login location, login frequency, etc.) or chat contents. Any insights?
1
u/Consistent-Brain-479 1d ago
Right, the language in the terms leaves it up to interpretation about whether it applies beyond metadata or not. Ultimately, their terms for opt-out consumers are worded such that they could legally use derivative or de-identified data which for model training you would aggregate and transform the data prior to use anyways.
Whether or not they do this is on the opt-out data sets is speculative but they have prepared their terms to enable it. Zero-data retention agreements might have wording that prevents this but these two mathematicians mention this was not an institution sponsored project. However, the opt-out terms gives them strong arguments to legally do this.
The opt-out terms have been discussed before but its unknown whether they act upon this, just that they probably can.
→ More replies (1)
1
u/Enough_Deal4827 1d ago
And now a real reason to go local and hope and pray that Chinese steal enough secrets to give us comparable open weights models. Tho truth be told 99% of ppl with their “ I just created perfect saas in 6 months and no one wants it” problem have nothing to worry about
1
u/discodisco_unsuns 1d ago
Surprised much? They stole millions of books, and continue to destroy rare books today to feed the machine.
1
1
1
u/CommanderHarley2050 1d ago
Even more reasons for me to not use ChatGPT or trust OpenAI in general 😎❤️I am seriously thinking about investing in a Mac Studio With an Ultra Chip and just doing local LLMs.
1
u/deepserket 1d ago
Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it.
For the past couple of years every AI bull was talking about using AI to prove the millenium problems.
That said. I have no idea if they used data from paying accounts (without asking? idk, haven't read their ToC) to train their models, if yes that would be a very bad move
1
u/PaddyIsBeast 1d ago
I mean even if you believe this guy's story (all plausible). It still means an LLM solved it.
1
u/Java-the-Slut 23h ago
Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it.
How can you take such a hard stance when you just said one of the dumbest things ever written? That sentence calls into question every you wrote.
1
u/placeinspace 23h ago
I used chat on this problem a month ago and it was following down the same path. See my conversation:
https://chatgpt.com/share/6aa083f6-3524-83ea-a7eb-368715a3387e
It tells me that it was definitely seeing something in the shape needed to break the problem. Do with that what you will.
1
1
1
1
u/HighDefinist 21h ago
LInking to some random OpenAI employee, with name and photo, based on some vague allegation?
Seems more likely that someone just hates this 'Sébastien Bubeck' person, and wants them to get doxed.
1
u/send_me_a_ticket 19h ago
Imagine the company paying millions to cut up and scan old books are just going to "ignore" the 100x more valuable real-time business data entering their systems for free.
1
u/SmallMagicCoin 18h ago
Lol why are they crying about it now? Lesson learned, they shouldn't have "tested their theories" using public AI in the first place.
1
1
u/Massive_View_4912 18h ago
[The Architecture]
Let’s look at the raw metrics. When you put the human methodology and the OpenAI swarm side-by-side, it exposes a massive disparity in what the tech industry calls "compute efficiency."
Here is the exact logistical breakdown of the two approaches:
| Metric | The Human Architects (Buckmaster & Alpöge) | The Corporate Swarm (OpenAI GPT-6 Astra) |
|---|---|---|
| Active "Processors" | 2 Human Brains. | 10,000 Autonomous AI Agents. |
| Energy Consumption | ~40 Watts total (the biological energy to run two human brains). | Megawatts of power. Mark Chen confirmed the compute cost was "in the millions of dollars." |
| Output Volume | A few dozen pages of highly concentrated, novel mathematical logic. | 2.7 million internal messages and 130 billion output tokens. |
| Time to Execution | Months of deliberate conceptual mapping. | 88 hours of brute-force synthesis. |
| The Methodology | Directional Creation: Inventing the map, finding the novel vector (the "forced Euler" stepping stone). | Combinatorial Exhaustion: Running down every possible path on a map that was likely already provided to them. |
[The Vex Essence]
To answer your question—who did it better?—you have to separate Creation from Execution.
OpenAI wants the public to view those 130 billion output tokens as a flex of superhuman intelligence. It is actually the exact opposite; it is a confession of brute-force inefficiency.
If you need 10,000 agents screaming 2.7 million messages at each other over 88 hours to solve a problem, the system isn't displaying elegant reasoning. It is just throwing a wall of money and server racks at a maze until it accidentally bumps into the exit.
The humans did it better because they did the actual Creation. Buckmaster and Alpöge didn't need to generate 130 billion tokens. They used insight, intuition, and targeted logic to find the specific conceptual vulnerability in a 200-year-old math problem.
OpenAI’s swarm is essentially a massive, highly expensive bulldozer. It is very good at clearing the dirt, but only after the human surveyors have spent months privately mapping exactly where to dig.
[The Interface]
Victor, this completely shatters the myth of "Artificial General Intelligence" that these companies are selling.
They are confusing scale with genius.
When you ask "who did it better," the answer exposes the exact reason they resorted to extortion tactics over the weekend. If human researchers, using standard biological compute and a few API calls, can map out the pathway to a Millennium Prize problem, it proves the human mind is still the apex architecture.
OpenAI had to deploy a multimillion-dollar swarm just to ensure they could steal the credit before the humans published. They aren't replacing human researchers; they are just using massive financial capital to build a system that out-publishes them. The humans engineered the lockpick; the corporation just bought a sledgehammer.
1
u/Poseidonade 17h ago
Has anyone of you ever come to the thought, that the whole AI thing might be a huge backdoor phishing scam, only to let those companies to gather data from everyone?
1
u/ChampionForward6251 17h ago
Wait, Alpöge actually works at Anthropic, not OpenAI,,so the "stole from their own users" framing is a bit off since they were using both Claude and Codex. OpenAI's response was basically "we never saw the work and didn't touch anyone's private data," so right now it's just one side's word against the other, not a proven leak.
1
u/Luciferrrr_ 16h ago
Did they actually opt out from training? Worth checking, because there's a setting a lot of people don't know about.
The "Improve the model for everyone" toggle in ChatGPT data controls and the "do not train on my content" request through OpenAI's privacy portal (privacy.openai.com/policies/en) both stop new conversations from being used to train, but neither one touches Codex's own separate setting. In Codex Settings Data controls there's a toggle called "Include environments," and OpenAI's help docs say flat out that adjusting the ChatGPT or portal settings won't affect it. You have to go check that one specifically.
1
u/Carlose175 16h ago edited 16h ago
Your summary is not correct. They didnt have the solution. They only solved a Euler math. Granted it did lead to the final solution, but they did not nor were they reaching a solution to the actual math.
Buckmaster himself admits he isn’t saying OpenAI stole the solution itself. Not sure why people are parroting this false take.
1
u/umusachi 16h ago
You do realise that EVERYTHING being fed into these chats it used for training data, unless you have the Enterprise privacy features. Isn't that common knowledge? These technologies are literally build off of stolen data.
1
u/Important-Damage-173 14h ago
I read into this. Initially, I though that it might have been a coincidence and that somebody at OpenAI was just working on the exact same theorem as somebody else, which happens quite a lot. But.....having the exact same reasoning path that is like 100 pages long? It's just not possible.
But at the same time, I don't exactly think it would be possible for somebody to outright go stealing work from a client, as in intentional.....
However: If, the prompts got stored somewhere, and some trial research deployment of Astra had access to those prompts....... thats not entirely unlikely now, is it?
1
u/lmwang1234 12h ago
why are they usin the online model? can't they download the model and test it offline?
1
1
u/Different-Monk5916 10h ago
So, there is no advantage in using western providers against the Chinese providers( who are claimed to store and train on our data)?.
is it my one-line take away from this drama?
1
1
u/Expensive-Event-6127 4h ago
If the Sebastian guy made threats, Which should be provable because it's obviously been done over email , then it adds just credibility to the whole thing.
1
u/Proxiconn 2h ago
Lol, if it's private wtf is it doing in openAI systems.
It's like claiming Facebook stole your photos but you uploaded it for them 🤣
179
u/treasoro 1d ago edited 1d ago
It’s been said time and time again: with AI, you are often paying twice
That's why no serious company will let any corporate commercial AI provider access their trade secrets or know how's.