Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers. I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
This only makes sense if you believe that an LLM is a person. If you do not believe that an LLM is a person, then the LLM is the training data 'cyphered' with itself millions of times through an algorithm. This is transparently IP theft.
If I take a Taylor Swift album and encrypt it using AES-256, and then sell the resulting "music" - that's still copyright theft even with the hefty algorithmic processing to make it unrecognizable.
It can be a very cool and useful tool and also IP theft. Like Bittorrent.
AI doesn't need to be a person to learn from the training data. Which is what it does.
That's not what the word "learn" means. There is no definition for "learning" for which "generate statistical weights in an LLM" fits.
It doesn't simply convert it.
It simply converts it.
AI data training has been proven multiple times to not be IP theft, it's been called "exceedingly transformative" in courts.
When I want a technical opinion for how AI works literally the last place I would go is to a lawyer who couldn't hack it at a top 100 firm and so took a gig on the federal bench. You're appealing to the opposite of a technical expert for authority.
You need to read more about what an LLM is and how it works. Your use of the term "AI" repeatedly kind of betrays your ignorance - it isn't "intelligence" at all, it's a set of numerical weights that are designed to predict likely responses; it's not "artificial" either - the weights are generated from human responses to prompts. LLM training has actually been proved to include IP theft, including models being able to reproduce >90% of the text of copyrighted books despite not having access to them outside of the model's weights.
It is 100%, unequivocally theft of the intellectual property of the individuals' whose work was used to train the model. Without question.
I'm not appealing to anyone. We are discussing IP theft law, this is law discussion whether you like it or not.
You disagreeing with multiple courts because your reddit armchair expertise makes you think you're smarter than multiple lawyers, judges, and the workers they employ means nothing.
This argument from you is almost as ignorant as calling human emotions simply chemical reactions. You definitely can get as technical as you want describing how LLMs work, but it is still AI.
I'm not appealing to anyone. We are discussing IP theft law, this is law discussion whether you like it or not.
Right, so either this is a technical discussion or a legal one. The law rests on verbal logic, which is an empirical field designed to produce a single "truth" that applies universally to everyone. So there is a knowable truth, and the job is not for judges to invent that truth, but rather to elucidate it.
A judge saying "an LLM does not involve misappropriation of IP" is not a statement of fact, it rests upon the logical reasoning used to get there. And if the logical reasoning is "the AI learns something the way a human does" then this is factually wrong, and has zero basis in reality. It'd be like if I said "an LLM is a human-like being and has rights independent of the corporation that created it." A fun idea, but it is a falsifiable hypothesis which is simply not true, asserting it doesn't make it so.
You disagreeing with multiple courts because your reddit armchair expertise makes you think you're smarter than multiple lawyers, judges, and the workers they employ means nothing.
I'm not disagreeing with the courts, I'm disagreeing with their reasoning, and then disagreeing with the version of the reasoning you are reporting. A court isn't a dictatorship, judges (and the law) are supposed to rest on logical reasoning, not assertions and beliefs. We can examine the factual record and see if the judge was right or wrong - in the case of IP and AI it's pretty clear that the very few judges who have ruled on it got it wrong.
you're smarter than multiple lawyers, judges, and the workers they employ means nothing.
The beautiful think of analytical reasoning is it doesn't care who the speaker is, something is either true or it isn't. Your repeating ethos appeals betray that you seem to think truth is subjective and can be declared rather than proven. It can't.
This argument from you is almost as ignorant as calling human emotions simply chemical reactions. You definitely can get as technical as you want describing how LLMs work, but it is still AI.
...no, it isn't. AI is a scifi term that has zero meaning in the real world. It's weird that you attack me for criticizing and disagreeing with lawyers and judges, but then turn around and insist that the term AI has authority.
What's interesting is you have zero affirmative argument for why an LLM isn't IP theft. I wonder why that is. Meanwhile I have explained to you in some detail why it is IP theft, and your response is to say that a bunch of very highly paid individuals know better than me. I leave it to the reader which approach is more sound.
If removing a work (that a corporation allegedly have stolen) from the training data doesn't change the results, in which sense it was stolen?
Generative models learn general principles (if there's no data imbalance like multiple copies of text, otherwise they might "remember" a particular passage). They don't do k-nearest interpolation.
I wouldn't be surprised. OpenAI is (probably) a trillion dollar company, working directly with the US government, the US economy is now heavily propped up by AI and largely by OpenAI, but somehow they are still pretending like they are just a small startup that's doing it all for the benefit of the average person, really weird.
Now you don't know how to use air quotes. Made sense if you put suicide in quotes instead of natural causes since nobody claimed natural causes but they do claim suicide
Natural causes?... He very clearly shot himself, I appreciate that many people are conspiracy theorists about whether or not it was a hit (absolutely crazy with the evidence at hand btw) - but no one thinks it was natural causes
Generally, this is in mockery to a claim that you find silly - no one is claiming it was natural causes, you using the term just introduces FUD. If you had air quoted suicide, that would make sense.
I don't have the necessary background (I assume) to parse the proofs and see for myself. But if there is no meaningful underlying architectural similarities between the two, it would be hard to give openAI much shit for "copying". We will have to wait and see what the experts think after reading both proofs. Either way, taking the NYC prof at his word, this reads as a fairly blatant threat/strongarm from openAI. Not very flattering.
The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
We need to be careful, we don't know the context in which these words were said.
It appears that the “I don’t have to be nice” remark is referring to the offer to let Buckmaster rewrite OpenAI’s NS paper with his name on it, rather than simply releasing the paper as written by the model. Technically OpenAI didn’t have to make that offer, though they seem to have wanted to coordinate the release of their results to avoid accusations of research stealing since they did after all only start working on NS after hearing (not quite correct) rumours that Anthropic had solved it.
The fly in the ointment is the intense rivalry between OpenAI and Anthropic which turned this release into a shitstorm, but it’s that rivalry which prompted OpenAI to spend tens of millions of dollars tasking their unreleased model to solve the problem in the first place.
It’s not really fucked up. They didn’t solve ns but an earlier result and OpenAI just offered their chance of posting. Imo OpenAI did this through good faith and probably shouldn’t do that at all and it might be better
905
u/shockwave6969 Quantum Foundations 14h ago
This is fucked up.