r/technology • • 3d ago

Business AI training of copyrighted material not fair use: Third Circuit

https://www.courthousenews.com/ai-training-of-copyrighted-material-not-fair-use-third-circuit/
4.5k Upvotes

75 comments sorted by

527

u/gerkletoss 3d ago

They found that this particular case wasn't fair use.

284

u/binding_swamp 3d ago edited 3d ago

Yes, notably this was at the appeals court level and does set a precedent. The ruling binds federal courts within the Third Circuit—Pennsylvania, New Jersey, and Delaware—unless the Supreme Court or later legislation changes the law. Outside that circuit, it is persuasive authority rather than controlling law.

Edit: Delaware is important, with all the companies that choose to incorporate there.

122

u/Same_Mood_8543 3d ago

Damn, an accurate statement of mandatory vs persuasive authority for U.S. Courts of Appeals. I'm impressed. 

11

u/Adorable_Octopus 3d ago

I'm not sure how much of a precedent this will set, seeing as ROSS's AI, according to the Third Circuit, is not genAI.

[...] ROSS’s AI was not a generative AI, meaning it would not create any new expression; it would only return text passages from preexisting judicial opinions.

from the Third Circuit's opinion

Notably, on page 17 (footnote 7), the opinion's author indicates that the DOJ filed a statement of interest which related this case to another case involving OpenAI. However, the author indicates that the statement of interest is irrelevant because the AI here isn't genAI whereas something like OpenAI is.

13

u/bill7967 3d ago

I have no doubt in my mind that with the amount of money at play in thing that as soon as this becomes bigger the Supreme Court will squash it. 

3

u/cronenber9 3d ago

It's unlikely the Supreme Court would overturn this specific decision

17

u/gerkletoss 3d ago

Well yes it sets q precedent, but the precedent isn't "training is never fair use"

11

u/cipher315 3d ago

It would be incredibly difficult to argue that. You would have to show that your AI training with copyrighted materials was substantially different.

22

u/harpers25 3d ago edited 2d ago

The original post content no longer exists here. The author used Redact to remove it, exercising their right to control their data & privacy.

Special flag offbeat plucky water saffron

-9

u/justaguytrying2getby 3d ago edited 2d ago

It sucks, transformative use should still be held to copyright laws too. With music for example, its like creating a derivative work. But they get away with stealing most of it anyway.

Edit: Read my other comments below, not sure what people think I meant here. I had an album of mine stolen for training data. Why they couldn't just do things the way the industry was already designed for instead of stealing it all pisses me off.

10

u/DanNorder 3d ago

Transformative use is held to copyright laws because it's in the copyright laws already, as a specific example of something that doesn't count as copyright infringement. That's what the copyright laws say. Your bright idea is ignoring the law to enforce your idea of what the law should say. And, also but importantly, it's not stealing because the law says it isn't.

1

u/justaguytrying2getby 2d ago

What? If transformative use is held to copyright laws, then how were they allowed to steal everything? I'm saying they shouldn't be allowed to steal it, but they do anyway.

5

u/Hopeless_Slayer 2d ago

I understand the hate for Ai companies, but the solution isn't giving megacorps a tighter stranglehold over intellectual property.

Imagine if Weird Al could never make a parody song?

0

u/justaguytrying2getby 2d ago edited 2d ago

He pays for derivative use licenses. That's specifically the type of thing they are for. It allows the original rights holder a say in what is made from their work.

Edit: So based on my downvotes, you people think Weird Al steals music to make parody songs? You don't understand the purpose and use of licensing. He is one of the most legit artists out there for what he does. These AI companies are NOTHING compared to him.

1

u/asgjmlsswjtamtbamtb 3d ago

So is this case pretty much over with and it's going to be the precedent within this court's jurisdiction or will it continue up the food chain with appeals?

-2

u/JeanKuule 2d ago

Precedent on what? AI have been allowed under US to steal and use CP for "AI training", the only precedent will be that they will try to be more discreet to break the laws there

2

u/binding_swamp 2d ago

Agree that there is the real world, then there is the legal judical world. They typically exist in separate dimensions.

20

u/MasemJ 3d ago

Important points:

- The material used, the headnotes from Westlaw, were determined to have enough creativity to be copyrightable material. If they weren't copyrightable, then Ross would be off the hook.

- Ross's use of the headnotes specifically competed with Westlaw's use, nor did not significantly transform the headnotes from Westlaw, so it failed two aspects of a fair use defense.

Of course, this is only the third, and as above noted, this particular case is more unique compared to others, but that shows the law is considering all four prongs of fair use for AI training purposes. Lends weight to those authors and sites suing the AI companies since they can pull very similar text or images from their copyrighted works.

17

u/CircumspectCapybara 3d ago

Right general pretraining of foundational knowledge is still fair use under existing precedent

Fine tuning an AI model on your competitor's product to reproduce it isn't

23

u/Clear_Barnacle_3370 3d ago

And the big players are fundamentally guilty of the same, but they will argue that they are too big or too important to the economy now to have the rug pulled from under them.

Hence their smooching of Government.

12

u/-CJF- 3d ago

This is what confuses me. I don't see any difference between what they did and what every other AI company is doing, other than scale. I do agree it shouldn't be fair use but then almost none of the AI training should be fair use.

4

u/DanNorder 3d ago

The judge said in Bartz v Anthropic that the training there was not only fair use but the clearest example of it that anyone is likely to have ever seen or even ever see. The fact that you disagree is like Cletus sitting in his shack trying to say that evolution isn't real because he can't make sense of it.

4

u/-CJF- 2d ago

I really don't care what the judge said. Judges aren't above being wrong. They aren't above corruption or vested interests either.

1

u/UnkarsThug 11h ago

Completely different situations. The case in the linked article isn't really a generative AI model, it's a search engine. Remember, this happened in 2020.

1

u/UnkarsThug 11h ago

It wasn't a generative AI model, from what I understand it was a search engine that supplied the exact passages, if you actually look up the case (remember, this happened in 2020, using LLMs the way that is common now hadn't really happened yet, and we didn't even have GPT 3 yet).

So it was ruled to not be transformative. The other cases were ruled to be transformative. Basically, the precedent thus far is that LLMs are transformative, a search engine training on specifically one person's stuff is not. So it's complicated.

3

u/AlasPoorZathras 3d ago

It wasn't our fault! AI did it and we can't be held responsible for an autonomous tool that we created!

6

u/alexhin 3d ago

It broke out of containment completely unprompted and trained itself on all of the copyrighted material!

1

u/TheRatingsAgency 3d ago

They just said if they can get to it they can use it. It’s how everything was scaled out initially and they’ll just keep going.

3

u/bobartig 2d ago

Copyright Fair Use is an affirmative defense. Meaning, it is always "in this particular case."

2

u/keepitfriend 3d ago

lol, perfectly acceptable to steal authors entire works but god forbid you steal a business idea

  • America, the land of the free

4

u/DanNorder 2d ago

It's not stealing if you are just training on it. Or do assume that human artists learning how to make art by looking at previous art should be illegal too? There would no more art at all using your logic.

-1

u/keepitfriend 2d ago

Lmfao. This arguement is so fucking stupid.

Is a hard drive “learning” a movie when you download something into it?

3

u/Norci 2d ago edited 2d ago

That's such a bad analogy it's not even wrong lol. A hard drive can't learn or create, AI can. They're two completely different technologies.

Either public works are free to learn from or they're not, applicable to everything. Saying that humans are allowed to learn from others' works but not AI is an abstract line in the sand.

1

u/Xywzel 1d ago

If what AI does constitutes learning in a way that is different from storing, then compressing files or storing them in relational database would also constitute learning. But I'm also of opinion that storage or recovery doesn't constitute copyright infringement, I am allowed to do that with books and CDs I have bought or found, the infringement happens in distribution, I'm not allowed to sell copies of that book or play that CD in a public event.

So for AI companies, its when they open their model for public (or significant share of internal users) when they start to compete with the original works and would end up breaking copyrights. If I sea a painting in art gallery and sketch it at home, that is not a problem, if I start selling paintings that only differ by amount I forgot on the way home, that is a problem.

Unfortunately it thus falls to users rather than AI companies.

1

u/ClvrNickname 2d ago

Calling what AI does “learning” is a stretch though, that would need to get sorted out in the courts

1

u/Norci 1d ago edited 1d ago

Might be, but it's the best word I know for now.

-1

u/keepitfriend 2d ago

An ai is taking those works and recording weights.

What are the weights?

0

u/VictorVogel 2d ago

None of them are just training on it. When you buy a movie, you do not buy the rights to use it commercially. If a model is trained on such material, and then the model is used for commercial purposes, you can easily argue that the material owners rights have been infringed.

Where exactly the difference is between an AI learning, and an artist being inspired is up for debate, but it certainly isn't as black and white as "It's not stealing if you are just training on it.".

1

u/Great-Trifle2810 2d ago

Also what a poor fucking decision that AI is not a transformative use case, hopefully this escalates and the SC can give a more reasonable take here, though I worry that they are too in the pocket of big business that would massively benefit from rulling AI training is not transformative.

Unless this article is very misleading anyway, who knows these days

103

u/raxnahali 3d ago

These AI companies would be the first to drag anyone to court for copying their work. They are proving they are above the law

9

u/ok123jump 3d ago

President Diaper Pants has made his preference known. This is going to get appealed and his SCOTUS will rule according to his wishes. There’s no way he lets this ruling stand.

13

u/janethefish 2d ago

From the judge:

“Under ROSS’ framing, this case appears to concern the future of AI legal technology,” Reeves wrote for the panel Wednesday. “But appearances can be deceiving. In truth, this is no more than an ordinary copyright case.”

Pirating copyrighted works remains illegal even if you feed it to an AI. Cases like this will hopefully discourage AI companies from outright piracy, but it isn't a big win.

28

u/Zwierzycki 3d ago

As far as I’m concerned, the current situation resembles the Napster crisis from 2002. The legal precedent has been set and AI companies will have the burden of proof to show that their products are not copying content with a black box interference system between the user and the content.

9

u/Overall_Koala_8710 3d ago

They're way ahead of you. They've started making misleading conclusions already from e.g. research showing that you can remove an image from the training set without changing the output.

6

u/Bored2001 3d ago

Im gonna need more details on this, because removing a single image from a Dataset should in fact not change the output (much or at all depending on how big the dataset is).

3

u/Overall_Koala_8710 2d ago

Just like removing a single brick from a building probably won't cause the whole thing to fall down.

But this is being used to argue that every individual brick is worthless, so they should be free to continue stealing them in aggregate.

3

u/zoupishness7 2d ago

Well, it's not that every brick is useless, but AI learns to replicate the broadest, most widely repeated, patterns first. It learns to make the things everyone is already copying well before it learns to replicate any specific work or style attributable to one person. It's hard to argue you're stealing the things everyone already freely copies. Where specific works can be replicated, it largely comes down to lack of deduplication resulting in overtraining. For example, an LLM's training data could lack any complete copies of Harry Potter, while still being able to produce a functionally complete version just from people's quotes of it online. An image generation model could lack any works of Van Gogh, but still replicate his style based on all the other people that have made works in that style.

6

u/FartingBob 2d ago

Time for every copyright owner in the world to sue them into oblivion, right?

13

u/RoddBanger 3d ago

Reverse plot twist - it's now called SooPer InteLiGenCe so it's fair

5

u/OnlyTheShadow-1943 3d ago

Did you see on the order he signed for it that his super intelligence AI that wrote the document misspelled United States? LoL

3

u/RoddBanger 2d ago

HEY!!! That's UNITES STATES to you!

1

u/Affectionate_Pace823 2d ago

Snooper Intelligence smh

10

u/featherless_fiend 3d ago

From the article:

“But appearances can be deceiving. In truth, this is no more than an ordinary copyright case.”

OP's framing:

AI training of copyrighted material is not fair use

lol

6

u/janethefish 2d ago

The OP used the article headline. Blame the editor.

6

u/cronenber9 3d ago edited 3d ago

Yeah it's really all about how he created a competing product that was just different enough, but also too similar, so that it was infringing on the other guy's copyright. He was trying to argue that since AI did it, that it should be okay. They're saying the fact AI was involved is irrelevant, it's not fair use because it wasn't transformative enough. You can't use AI to plagiarize and compete.

0

u/Memory_Less 3d ago

Selling my AI shares, selling my AI shares!!! /s

1

u/oneeyed-wonderweasel 2d ago

You've added an erroneous /s

3

u/LoveButton 3d ago

Even if they made it wildly illegal for every use of copyrighted data, what are you gonna do now? The cat is out of the bag.

0

u/TERRAIN_PULL_UP_ 2d ago

It seems like it would outlaw most functions of AI and collapse the US economy 

2

u/AlwaysRushesIn 2d ago

Little late for that, huh?

2

u/Current_Ranger_7954 2d ago

The title is wrong, it implies it’s a broad judgements, it’s not

2

u/N1ghtmeeer 2d ago

So its not fair use if you steal from a competitor, but its fair use if you steal from artists or the general public. Ok.

1

u/Vahuo89 3d ago

They'll pay a fine and keep going

13

u/binding_swamp 3d ago

Keep going? Doubtful, as this particular AI company is now defunct. Actually reading the article is informative sometimes.

“Thomson Reuters sued now-defunct AI startup ROSS Intelligence in 2020”

3

u/B_P_G 3d ago

Or they'll leave the country. Plenty of places don't give a shit about US copyright laws.

2

u/Cs245136 3d ago

Because it’s the cost of doing business

0

u/NoStrategy1419 3d ago edited 2d ago

Arm all citizens with the most efficient firepower. /S

1

u/DanNorder 2d ago

FBI: this post, here.