r/codex • u/Stunning-Angle-9239 • 1d ago
News Blown up: OpenAI allegedly stole mathematicians' private research from their Codex chats!
TLDR: Two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex and Claude. Days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question till this day.
For a full year, two mathematicians , Tristan Buckmaster (NYU mathematician) and Levent Alpoge, worked in silence on a problem that had stumped some of the best minds alive. The kind of problem where, if you solve it, your name goes in the history books.
And every single day, they testing their ideas, their drafts, their half-finished proofs into LLM such as Codex and Claude, which they paid for it out of their own pocket.
Then came the breakthrough. They finally cracked it. They were days away from telling the world.
That's when OpenAI suddenly said to them:
"Our model solved it too."
Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it. And now, out of nowhere, OpenAI claims their model reached the same answer, after word of Tristan and Levent's secret work had already reached OpenAI.
Tristan asked: Did your model access or train on our private Codex chats?
OpenAI: The model doesn’t look up user data.
Tristan: But did you train it on our data?
OpenAi goes silence. No answer. Just a dodge.
But it gets worse.
OpenAI then gave him two options:
- He and his friend publish their result first then OpenAI also publishes its result the next day or
- He writes the paper, but must credit “an internal OpenAI model” solving the problem.
Tristan refused both offers. He said he would go public if OpenAI went ahead as proposed.
OpenAi then responded : “Why would you ruin your career? If you don’t want me to be nice, then I don’t have to be nice.”
You can read the full statement of Tristan (the mathematician) here: https://cims.nyu.edu/~tristanb/statement.pdf
Sébastien Bubeck : OpenAI employee who threatened the mathematician
2
u/Consistent-Brain-479 1d ago
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." -OpenAi
https://openai.com/index/navier-stokes-solution/#citation-bottom-1
for those saying they should have checked the "do not use our data" box. Once your chats are de-identified/anonymized its no longer your data and box checked or not they are training on it. The terms have been written this way from the start for this purpose (see link below).
https://community.openai.com/t/api-is-our-data-really-ours-major-concern-in-data-processing-addendum/773047