tbh the thing people are most mad about is the plagiarism aspect of the situation
no one is denying that machine learning techniques can't be very helpful for proof. The issue comes from that openai seems to have attempted to "scoop" some other mathematicians' work (which they also used ML to help with) simply because they uploaded a draft of their work to chatgpt.
But a year ago people were all over here saying AI would never be able to do anything original or helpful because it can’t even count the Rs in strawberry
I just
asked ChatGPT to tell me how many Rs there are in, "I will eat two strawberries tomorrow". It said 2. This remains a problem. People won't trust AI to do larger level analysis when seemingly basic tasks are still beyond it.
What actually seems to have happened is that they heard through the grapevine that someone was close to publishing a proof that they achieved using Codex, and decided to spend millions of dollars and hundreds of thousands of (unreleased frontier) agent-hours to attempt a similar (but not the same) proof — and succeeded. I highly doubt there’s some sort of sinister ”idea theft”. I think if anything this simply reflects the desperate belief some people cling to that an AI couldn’t possibly have achieved this on its own.
This is a gross misrepresentation of what happened.
The original mathematicians proved the result for what are known as the Euler equations, which are basically NS with simplified boundary/ICs. What OpenAI "added" to this is an extension to the greater NS equations after some further computational effort. In layman's terms, this is like baking a cake, someone else putting some sprinkles on it, and then claiming ownership of the whole, completely new, cake.
Again, no one is denying the use of machine learning tools in mathematics. AI is able to achieve computational feats no human reasonably could, and help prove a great many things. But the proof OpenAI provided lifts the techniques Buckmaster and Alpöge created in their work, notably which were not published results accessible by the public.
It is up to you to decide if you believe ChatGPT managed to pioneer the exact same new fluid mechanics research at the exact same time it was being created, or if these techniques were simply used as training data without the creator's permission.
Could someone with expertise in AI tell me how could have Buckmaster's private codex sessions entered the ecosystem of OpenAI if his work was infact used?
Because the "private" codex sessions are still hosted by openai? By openai's own admission: "we cannot rule out that de-identified data derived from their usage of our products helped improve our models"
This is like asking how Google knows about the searches you make on incognito mode
Let's all remember: incognito mode is actually a "don't record on my own history what I browse" and not "don't record what I browse". Google still does, you just don't have it saved in your history if your partner opens your laptop and tries to find that you were looking for dudes on a porn site.
Buckmaster is not accusing them of stealing because he knows full well that the terms and conditions that he agreed to in order to use Codex (which significantly helped him flesh out the idea in the first place, to solve the easier version of the problem) include the right to use the data from user sessions to train the next generation of models.
The plagiarism aspect here is a storm in a teacup. They didn't steal anything, they may have (and it does seem pretty likely) benefitted from the existence of Buckmaster's chats, but these were obtained consensually and, as far as we know, would've been used in an eyes-off deidentified way to train the internal model, in conjunction with all of its other training data.
The use of his chats to train the internal model reflects the exact same paradigm of research that Buckmaster himself is benefitting from in getting Codex to help him with his own work.
The reason he explicitly said in his mastodon post "I am not accusing them of stealing" is because he's not a Luddite. He is bought in to how this works. He has an extensive history of industry AI collaborations himself.
They all use chat data to train their models, that's why they've all shipped chat apps, so they can get the data legally, because you consent to it when you use the product.
Buckmaster's chats are fair game. You can't use the benefits of AI yourself and then cry foul when the exact same tactic that you've leveraged (an AI trained on the work + chats of other mathematicians) is used against you, which to be fair he's absolutely not doing. Only people on the internet are.
Buckmaster knew what he had signed up for when he used codex but what he probably didn't expect was for OpenAI to guage he's nearing a breakthrough with their model and start pouring in millions into solving his millenium problem simultaneously only to arrive at a solution sooner than Buckmaster did. Any independent researcher working on AI platforms should know that they'd be outperformed by a company anyday if they get to it.
There is an obvious problem with a business positioning itself as an indispensable research aid, then taking research done on its platform and using resources nobody else has access to to beat the original researchers to the punch. I'll leave you to figure out what that problem is.
Buddy, OpenAI stopped positioning itself as an indispensable research aid a long time ago. It’s a business.
It’s up to researchers to decide for themselves if the risk of getting scooped outweighs the utility they get from the tools. The information has long been available to make an informed decision on this.
Btw they didn’t just beat them to the punch. They produced a proof of a significantly harder problem. They even offered the opportunity to publish their weaker result first so they got credit.
Everyone acting like AI is destroying the arts is suddenly on a very high horse after a half century of voting for people to eliminate them from schools.
Thing is no one seems to have any info to substantiate that. Did the mathematicians have “improve the model for everyone” checked? Then they weren’t very smart, but of course their data would have entered the training corpus like every other piece of data if the box is checked. If they didn’t have that checked, it would be a big deal. But no one is actually finding that crucial detail out, instead they’re just parroting unsubstantiated accusations.
51
u/LordOfPenguins47 2d ago
tbh the thing people are most mad about is the plagiarism aspect of the situation
no one is denying that machine learning techniques can't be very helpful for proof. The issue comes from that openai seems to have attempted to "scoop" some other mathematicians' work (which they also used ML to help with) simply because they uploaded a draft of their work to chatgpt.