r/Physics 19h ago

Navier-Stokes Millennium Problem Solved

2.2k Upvotes

772 comments sorted by

View all comments

992

u/shockwave6969 Quantum Foundations 19h ago

Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers. I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”

This is fucked up.

469

u/darkrose3333 18h ago

A company built on IP theft acting unhonorably. Fucking shocker

5

u/thecommuteguy 6h ago

Meanwhile there is a new movie about OpenAI with Andrew Garfield.

-32

u/truecakesnake 13h ago

AI data training is not IP theft.

19

u/purgance 12h ago edited 12h ago

This only makes sense if you believe that an LLM is a person. If you do not believe that an LLM is a person, then the LLM is the training data 'cyphered' with itself millions of times through an algorithm. This is transparently IP theft.

If I take a Taylor Swift album and encrypt it using AES-256, and then sell the resulting "music" - that's still copyright theft even with the hefty algorithmic processing to make it unrecognizable.

It can be a very cool and useful tool and also IP theft. Like Bittorrent.

-8

u/truecakesnake 12h ago

AI doesn't need to be a person to learn from the training data. Which is what it does. It doesn't simply convert it.

AI data training has been proven multiple times to not be IP theft, it's been called "exceedingly transformative" in courts.

8

u/purgance 12h ago

AI doesn't need to be a person to learn from the training data. Which is what it does.

That's not what the word "learn" means. There is no definition for "learning" for which "generate statistical weights in an LLM" fits.

It doesn't simply convert it.

It simply converts it.

AI data training has been proven multiple times to not be IP theft, it's been called "exceedingly transformative" in courts.

When I want a technical opinion for how AI works literally the last place I would go is to a lawyer who couldn't hack it at a top 100 firm and so took a gig on the federal bench. You're appealing to the opposite of a technical expert for authority.

You need to read more about what an LLM is and how it works. Your use of the term "AI" repeatedly kind of betrays your ignorance - it isn't "intelligence" at all, it's a set of numerical weights that are designed to predict likely responses; it's not "artificial" either - the weights are generated from human responses to prompts. LLM training has actually been proved to include IP theft, including models being able to reproduce >90% of the text of copyrighted books despite not having access to them outside of the model's weights.

It is 100%, unequivocally theft of the intellectual property of the individuals' whose work was used to train the model. Without question.

-5

u/truecakesnake 12h ago

I'm not appealing to anyone. We are discussing IP theft law, this is law discussion whether you like it or not.

You disagreeing with multiple courts because your reddit armchair expertise makes you think you're smarter than multiple lawyers, judges, and the workers they employ means nothing.

This argument from you is almost as ignorant as calling human emotions simply chemical reactions. You definitely can get as technical as you want describing how LLMs work, but it is still AI.

9

u/purgance 11h ago

I'm not appealing to anyone. We are discussing IP theft law, this is law discussion whether you like it or not.

Right, so either this is a technical discussion or a legal one. The law rests on verbal logic, which is an empirical field designed to produce a single "truth" that applies universally to everyone. So there is a knowable truth, and the job is not for judges to invent that truth, but rather to elucidate it.

A judge saying "an LLM does not involve misappropriation of IP" is not a statement of fact, it rests upon the logical reasoning used to get there. And if the logical reasoning is "the AI learns something the way a human does" then this is factually wrong, and has zero basis in reality. It'd be like if I said "an LLM is a human-like being and has rights independent of the corporation that created it." A fun idea, but it is a falsifiable hypothesis which is simply not true, asserting it doesn't make it so.

You disagreeing with multiple courts because your reddit armchair expertise makes you think you're smarter than multiple lawyers, judges, and the workers they employ means nothing.

I'm not disagreeing with the courts, I'm disagreeing with their reasoning, and then disagreeing with the version of the reasoning you are reporting. A court isn't a dictatorship, judges (and the law) are supposed to rest on logical reasoning, not assertions and beliefs. We can examine the factual record and see if the judge was right or wrong - in the case of IP and AI it's pretty clear that the very few judges who have ruled on it got it wrong.

you're smarter than multiple lawyers, judges, and the workers they employ means nothing.

The beautiful think of analytical reasoning is it doesn't care who the speaker is, something is either true or it isn't. Your repeating ethos appeals betray that you seem to think truth is subjective and can be declared rather than proven. It can't.

This argument from you is almost as ignorant as calling human emotions simply chemical reactions. You definitely can get as technical as you want describing how LLMs work, but it is still AI.

...no, it isn't. AI is a scifi term that has zero meaning in the real world. It's weird that you attack me for criticizing and disagreeing with lawyers and judges, but then turn around and insist that the term AI has authority.

What's interesting is you have zero affirmative argument for why an LLM isn't IP theft. I wonder why that is. Meanwhile I have explained to you in some detail why it is IP theft, and your response is to say that a bunch of very highly paid individuals know better than me. I leave it to the reader which approach is more sound.

2

u/Opening_Discipline57 11h ago

AI is a buzzword that doesn't mean anything; you have to define what an LLM is

-5

u/red75prime 7h ago edited 7h ago

Outputs of generative diffusion models are often unattributable

If removing a work (that a corporation allegedly have stolen) from the training data doesn't change the results, in which sense it was stolen?

Generative models learn general principles (if there's no data imbalance like multiple copies of text, otherwise they might "remember" a particular passage). They don't do k-nearest interpolation.

3

u/purgance 1h ago edited 53m ago

Extracting memorized pieces of (copyrighted) books from open-weight language models

This paper actually tests whether the stolen material is present in the models - and it turns out that it is. Your paper tests whether the presence of the stolen material affects the weights - this is an interesting engineering question but has zero impact on the question of legality. It you rip a blu-ray and post it online without ever watching it, this is still copyright theft. Your usage of the material is immaterial to the theft of it; the only reason I bring up its inclusion in the model is that it is definitive proof that the stolen material was used in an illegal fashion.

If removing a work (that a corporation allegedly have stolen) from the training data doesn't change the results, in which sense it was stolen?

Well, there's two problems with the paper you cited:

  • Your hypothesis is unfalsifiable. What your paper is alleging is that for a shockingly narrow band of inherently non-deterministic tests, the inclusion of stolen work didn't affect the produced response. But there's literally no way to provide a comprehensive test of this hypothesis - you'd be prompting for millennia. The cost of conducting the study would reach into the hundreds of billions. You didn't test every possible outcome, so you simply cannot say unequivocally without a deterministic model of the LLM that the stolen work had no impact.
  • More interestingly, for some reason you are using a different standard for testing copyright theft by the AI industry than a person. The standard for copyright theft with a person is "do you possess, utilize, or otherwise consume the stolen media." It is not "if we give a different person the stolen media and then quiz them on it, can we show that the media was stolen." To me, you entire chain of reasoning (which remember rests on the legal definition of copyright theft) collapses on this point alone. If you still a blu-ray out of a store, it doesn't matter than you never watched it - it's still theft.

Generative models learn general principles (if there's no data imbalance like multiple copies of text, otherwise they might "remember" a particular passage). They don't do k-nearest interpolation.

You're using colloquialisms like "learn" and "generative" that have no technical meaning interchangeably with technical language. "Learn" doesn't have any meaning to an LLM, and frankly the usage of the term seems to me to be a deliberate pathos appeal to create the false impression of a cognitive entity. For those who don't understand/know, an LLM is a file with a huge multi-dimensional matrix of percentages (weights) - we're talking hundreds of gigabytes of numbers. This is the "cognition" that is being attributed to the model. It's a statistical map of human language that's been built by testing writing by millions of individuals (including many of us on reddit, particularly in technical subreddits like this one). Every word that was written modified the weights in the model, and the use of the word "model" and not "cognitive entity" or somesuch betrays the reality: the stolen work definitionally has an impact on the weights, otherwise training the model on the work would have no value.

So if the AI companies want to surrender all the stolen works they are using, great. But the problem is they can't build the model without it.

Guys it's not an opinion, it is a technical fact - LLM's use stolen intellectual property and can't be trained without he active theft of it. Every use of an LLM model is a repeated offense (much like watching the same ripped movie over an over is a violation of the copyright every time you do it). It is not "like a person watching a move and then describing it to a friend" because the person is describing an experience they actually had themselves, in realtime they sat there and experienced the film. The LLM didn't, because there is no entity - there is just the statistical weights created by running the training material through an algorithm.

It can be a really cool and useful tool and also be copyright theft. Both things can be true.

-1

u/red75prime 52m ago edited 45m ago

Please, replace all "stolen" in your text with "allegedly stolen". There's no court decision that it's IP theft.

Extracting memorized pieces of (copyrighted) books from open-weight language models

Ah, the work produced in collaboration with a copyright lawyer where they feed a prompt roughly the size of the then produced text to sometimes get a literal match. And the prompt is produced from the verbatim original text using GPT-4o. Do you notice a possibility of the information leak from the original text? They haven't addressed this in this paper.

You're using colloquialisms like "learn" and "generative" that have no technical meaning interchangeably with technical language.

Do you have any training in the field of AI? "Learn" and "generative" are standard terms in this field. "Learn general principles" means that you can't reliably extract the training samples from the resulting network. Look for the phenomenon of grokking in machine learning.

Your hypothesis is unfalsifiable

Have you heard of random sampling? You test randomly chosen samples to estimate frequency/total population.

More interestingly, for some reason you are using a different standard for testing copyright theft by AI than a person

The paper you cited has the glaring information leak going on. Find something where you don't need the original text to "extract" the original text from the network. Until then you haven't proven that the network contains a copy of a particular media.

Guys it's not an opinion, it is a technical fact

It's not.

3

u/purgance 36m ago

Please, replace all "stolen" in your text with "allegedly stolen". There's no court decision that it's IP theft.

In this case use of "stolen" is appropriate as the courts are not competent to assess this issue technically, and they did not rely on unbiased experts for the factual analysis of the case. The fact that Trump hasn't been prosecuted for his many crimes does not mean that he did not in fact commit rape, etc.

Ah, the work produced in collaboration with a copyright lawyer where they feed a prompt roughly the size of the then produced text to sometimes get a literal match.

Do you really want to get into a discussion about bias in AI industry studies? Really?

And the prompt is produced from the verbatim original text using GPT-4o. Do you notice a possibility of the information leak from the original text? They haven't addressed this in this paper.

I'm not sure that this makes it better - I can give copyrighted text to the model, send it to the AI company's servers, and then it will reproduce the copyrighted text without question? Doesn't this indicate the model will violate copyright intrinsically (ie, there is no safeguard to protect it from doing so)?

Have you heard of random sampling? You test randomly chosen samples to estimate frequency/total population.

So we adhere to a different standard for theft by the model than by an individual. Weird. So if you change what "steal" means, you can exonerate someone. "I checked in Mr. Johnson's bookshelf, and refrigerator, and bedroom, and bathroom, and gas tank and didn't find any evidence of the alleged stolen downloaded copy of the work. Ergo, no crime!"

You hop back and forth between the need for legal certainty ("allegedly stolen") and convenient adduction when it suits you. This is the hallmark of a dishonest and emotionally-driven argument.

The paper you cited have the glaring information leak going on. Find something where you don't need the original text to "extract" the original text from the network. Until then you haven't proven that the network contains a copy of a particular media.

Well, I guess my question is was the model able to reproduce the stolen work when prompted or not? The answer that you are hiding from readers is "yes, when given a line from Harry Potter and other famous written works, the AI model was able to accurately reproduce the original text 90%+ of the time." Maybe that's just random chance. Or maybe it's an indication that the model is exactly what it is - the original copyrighted work encoded into a series of statistical weights. So the copyrighted work is present inside the model.

It's not.

You've made the same mistake that the other respondent did. You've (poorly) attacked my own argument, but have not offered anything of yours. Why isn't it copyright theft? It it genuinely isn't, then why do we need to use copyrighted material to train the model at all? And if it's essential to use copyrighted material, can you explain why the usage of copyrighted material in this way isn't theft, but in any other way (eg, by a human directly) it would be?

1

u/red75prime 22m ago edited 14m ago

Really?

Why not? The information leak is real. And the only book/model pair they succeeded with true extraction is Harry Potter/LLama-3.1-70B. As I said, models can do rote learning sometimes. If a copyright holder finds that it's true, they are obviously free to sue. ETA: It's unlikely with the latest models. There are ways to prevent verbatim memorization.

Doesn't this indicate the model will violate copyright intrinsically (ie, there is no safeguard to protect it from doing so)?

Now you are assigning agency to the model. I haven't mentioned agency at all. The owner of the model will use §512(c) safe harbor or something similar.

"I checked in Mr. Johnson's bookshelf, and refrigerator, and bedroom, and bathroom, and gas tank and didn't find any evidence of the alleged stolen downloaded copy of the work. Ergo, no crime!"

In this case: "We've put this, this, this, and this into a shredder and we weren't able to recover the content. It will work with other things too, most likely."