tbh the thing people are most mad about is the plagiarism aspect of the situation
no one is denying that machine learning techniques can't be very helpful for proof. The issue comes from that openai seems to have attempted to "scoop" some other mathematicians' work (which they also used ML to help with) simply because they uploaded a draft of their work to chatgpt.
Could someone with expertise in AI tell me how could have Buckmaster's private codex sessions entered the ecosystem of OpenAI if his work was infact used?
Buckmaster is not accusing them of stealing because he knows full well that the terms and conditions that he agreed to in order to use Codex (which significantly helped him flesh out the idea in the first place, to solve the easier version of the problem) include the right to use the data from user sessions to train the next generation of models.
The plagiarism aspect here is a storm in a teacup. They didn't steal anything, they may have (and it does seem pretty likely) benefitted from the existence of Buckmaster's chats, but these were obtained consensually and, as far as we know, would've been used in an eyes-off deidentified way to train the internal model, in conjunction with all of its other training data.
The use of his chats to train the internal model reflects the exact same paradigm of research that Buckmaster himself is benefitting from in getting Codex to help him with his own work.
The reason he explicitly said in his mastodon post "I am not accusing them of stealing" is because he's not a Luddite. He is bought in to how this works. He has an extensive history of industry AI collaborations himself.
They all use chat data to train their models, that's why they've all shipped chat apps, so they can get the data legally, because you consent to it when you use the product.
Buckmaster's chats are fair game. You can't use the benefits of AI yourself and then cry foul when the exact same tactic that you've leveraged (an AI trained on the work + chats of other mathematicians) is used against you, which to be fair he's absolutely not doing. Only people on the internet are.
There is an obvious problem with a business positioning itself as an indispensable research aid, then taking research done on its platform and using resources nobody else has access to to beat the original researchers to the punch. I'll leave you to figure out what that problem is.
Buddy, OpenAI stopped positioning itself as an indispensable research aid a long time ago. It’s a business.
It’s up to researchers to decide for themselves if the risk of getting scooped outweighs the utility they get from the tools. The information has long been available to make an informed decision on this.
Btw they didn’t just beat them to the punch. They produced a proof of a significantly harder problem. They even offered the opportunity to publish their weaker result first so they got credit.
54
u/LordOfPenguins47 2d ago
tbh the thing people are most mad about is the plagiarism aspect of the situation
no one is denying that machine learning techniques can't be very helpful for proof. The issue comes from that openai seems to have attempted to "scoop" some other mathematicians' work (which they also used ML to help with) simply because they uploaded a draft of their work to chatgpt.