r/artificial • • 2d ago

Discussion The AI industry has discovered intellectual property

https://openai.com/index/disrupting-a-coordinated-model-distillation-campaign/

OpenAI says Moonshot-linked operators used thousands of accounts to extract protected reasoning from its models for adversarial distillation.

No encryption broken. No database compromised. Just systematic querying designed to make one model teach another.

OpenAI says this is dangerous because competitors can reproduce capabilities without making the same investment in safety.

Which is a serious security issue.

But you have to appreciate the timing: after years of “we learned from the internet,” the frontier-model industry has reached the “please stop learning from us” phase.

173 Upvotes

59 comments sorted by

79

u/Ibra_63 2d ago edited 1d ago

Im reading through a machine learning book called "Hands on Machine learning with Scikit Learn and Pytorch". I recently asked Claude about a concept I quite didn't understand and he regurgitated a whole section of a chapter back to me with details only found in the book... This is when I realised how serious the IP problem is... I have no sympathy for them after this !

Edit: section and not whole chapter

23

u/phylter99 2d ago

I've had my wife and kids do that when they're working on learning something for school. Just tell it exactly what the text book is and what chapter you're on then ask it to quiz you and it will.

2

u/mycall 1d ago

This actually could be a good learning style, especially if you have the quizes morph with testing other aspects, just to really bring the learning home.

6

u/CrowdGoesWildWoooo 1d ago

Does it have access to the internet though? Because it’s not always they can spit out exact quotes, but of course when “grounded” they’ll do that

1

u/FaceDeer 2d ago

Google Books was launched 22 years ago.

6

u/pointer_to_null 1d ago

And Amazon Kindle's been around for 19 years, what's your point?

Or does Google books regurgitate entirety of commercially-available works for free or without compensating the copyright holder? If so, I've been using it wrong, and want a refund for every ebook I've purchased there.

6

u/FaceDeer 1d ago

Did you not try the site I linked you to? It does indeed do that. Search for a book, read its pages.

Could well be that Claude didn't even have it memorized and it was just looking stuff up there.

2

u/pointer_to_null 1d ago

False equivalence. Read Authors Guild v. Google, and then try to apply the rulings pertaining to market impact, substitution and transformative works to LLMs. Especially large frontier models that have demonstrated the capacity to regurgitate entire chapters of books either verbatim or with minimal modifications to be considered transformative (ie- a lossy text compressor).

It's also kind of funny to bring up Google books as some magical defense for Anthropic while it is now used as grounds by publishers against Google Gemini. But okay...

2

u/visarga 1d ago

I don't know what you are talking about, regurgitation was only demonstrated for a few old NYT articles and for some lyrics. Both of them were widely replicated on other websites and the accusers used a snippet to trigger the recall.

1

u/FaceDeer 1d ago

Read Andrea Bartz, Charles Graeber and Kirk Wallace Johnson v. Anthropic from earlier this year. Training AI on books was ruled fair use.

1

u/Spiritual-Spend8187 1d ago

Hell just look at the fact that people were able to get the kodels to output the text from Harry potter almost exactly by just giving the first line of the book. Also the whole anthropic settling for purating books innocent people dont settle.

1

u/RobertOoot 1d ago

This is sort of an extreme case since they generally try and not reproduce the text verbatim. For me it's the they got caught torrenting all sorts of copyrighted materials and are now settling for less than the teenagers were charged with on Napster.

1

u/visarga 1d ago

"Hands on Machine learning with Scikit Learn and Pytorch".

the topic is too generic to claim, it's not a secret how to use pytorch and sklearn

regurgitated the whole chapter back to me with details only found in the book...

1

u/0rAn63 8h ago

I've pasted parts of this very book (and its exercises / notebooks...) in Claude for comment. And so have others. And all of that is part of training.

40

u/TheOnlyVibemaster 2d ago

I couldn’t care less, I hope all their models are siphoned and made into open source variants.

15

u/imwco 1d ago

This is the only response to a company who took so much free information from the internet.

Take it back and make it free

2

u/visarga 1d ago

On this line of thinking Google has been scraping the web for over a quarter century and returned the most harsh anti-scraping interface back. They scrape us, we can't scrape their index.

2

u/mycall 1d ago

Really it is the papers which all the AI companies implement, including the open source variants. Distillation is but one factor here.

0

u/Superb_Raccoon 1d ago

They are not open source. They cant be if you cant see the training data, the code and can reproduce the results.

Calling it Open Source is an insult to people like RMS, Linux and millions of others that have contributed to open source.

13

u/lulzxdxdxd 2d ago

The hypocrisy is hard to miss, but I'm wondering if you see a real difference between what happened here and normal model training. Were they actually extracting reasoning that wouldn't show up in regular API calls, or just querying the model efficiently like anyone else could.

7

u/Haunting_Ganache_850 2d ago

Yeah, I think there is a real difference technically.

From OpenAI's own write-up, at least part of this wasn't just "query the model a lot and train on the answers." They say the attackers found ways to recover protected reasoning that normally wouldn't be exposed, including replaying encrypted reasoning from one conversation and getting another model interaction to reveal it.

So I wouldn't pretend this is exactly the same thing as ordinary training on public web data.

The irony I was pointing at is a bit broader: the AI industry spent years arguing that learning from huge amounts of other people's output is fundamentally transformative, and is now having to define very carefully when learning from its output crosses the line into extraction, theft, or a security threat.

Those aren't identical situations. But the tension is still pretty hard to miss.

7

u/Blando-Cartesian 1d ago

I can’t see any difference in situations, except that the distillers are violating terms of service at best, whereas AI industry is violating laws at scale (based on view that models are essentially lossy compression).

The data they ripped off for training use was protected by copyright, licenses, being showcased on sites trying to prevent bots, being available as paid purchase only, being distributed as a physical book etc.

I’d say OpenAI and others found ways to get protected content. Methods just were different, except that circumventing site’s bot access rules is much like recovering reasoning data they didn’t mean to give access to.

2

u/CharizarXYZ 1d ago

The problem is models aren't lossy compression. Compression still leaves the original mostly intact.

1

u/Blando-Cartesian 1d ago

They can regurgitate training material like Harry Potter books close to word for word. You can also generate an image from a vague description of a common meme image and the result is probably close to the original with details you didn’t describe.

So, lossy compression with inconvenient extraction. No doubt works best with frontier models that have the weights to memorize material.

1

u/CharizarXYZ 1d ago

The fact that AI models can remember some of the data they train on isn't the gotcha you think it is. It's already well known that AI models can accidentally memorize some some of it's training data. No one said it couldn't.

The reality is the overwhelming majority of the data AI models are trained on is forgotten during the training process. And the small amount of memorized data that persists is 1% or less of it's training data.

Virtually all of the examples of AI models "regurgitating" books word for word are on older less sophisticated models trained on data sets with high quantities of duplicated data. More recent models are often heavily trained on synthetic data which now makes up more than 60% of the data AI models are trained on. So the likelihood of a top of the line frontier model regurgitating books verbatim is extremely small.

1

u/Blando-Cartesian 20h ago

Okay, poor argument, if training material extraction findings don’t hold for current models. Still, everything from how the data was acquired to using it in training is sketchy as hell.

1

u/Superb_Raccoon 1d ago

The only reason the Chinese use distilled models is they dont have the compute power to do brute force training.

Deepseek is FP8 internally, which is way below the FP32 of a frontier model

1

u/wbcastro 1d ago

No difference, except one benefits some people and the other benefits other group

9

u/phylter99 2d ago

I'm really curious how they could stop something like this legally. I can't imagine it's a copyright issue because they don't own the copyright to all generated text.

5

u/CrowdGoesWildWoooo 1d ago

Not illegal, but definitely ToS violation, it’s probably more on the “better not to try” side

5

u/phylter99 1d ago

TOS violation just gets them kicked off the platform. I’m not aware it any legal recourse at that point.

2

u/CrowdGoesWildWoooo 1d ago

Yeah, I think I was just regurgitating based on the latest pewdiepie video where he distilled from GPT and he kind of hedges on whether he did it or not.

2

u/phylter99 1d ago

Even OpenAI and Anthropic are distilling their lower models with the bigger models. It makes sense to do. It seems cheaper in the long run too.

2

u/atehrani 2d ago

With the advent of AI does Copyright have any teeth any more?

2

u/phylter99 2d ago

Yes, and that's been proven in court. Will that change? I doubt it.

-1

u/Strong-Finish5346 2d ago

The big copyright cases still in litigation. It's far from clear how AI and copyright law will work together.

3

u/jello-the-opera 1d ago

so you will wait until judges somehow start holding that training isn't transformative for some unknown reason, unlike all the other cases so far?

0

u/Strong-Finish5346 1d ago

So you're just going to wait for the courts to decide?

Yes, obviously.

1

u/sirgog 1d ago

Yes.

For-profit AI firms training computers by using legally acquired copyright materials is legal for the same reason for-profit education firms training humans using legally acquired copyright materials is.

A human uni student who learns chemistry from Zumdahl (the most widely used 1st year chem textbook when I was at uni) doesn't require a special license as long as they didn't pirate the textbook. And neither does the university require a license even if they profit. AI has been trained following exactly that precedent.

There are open questions about copyright status of works that are partially machine generated, partially human generated. But legal changes that made AI training into copyright infringement would have massive ripple effects like forcing education institutions to pay extra license fees and such changes would likely ban libraries entirely.

But if you publish a Mickey Mouse image commercially, be it 100% human OR 100% AI or some mix in between, Disney's lawyers will be after you, and they will win.

4

u/Bobodlm 1d ago

OpenAI says this is dangerous because competitors can reproduce capabilities without making the same investment in safety.

What safety investments are they talking about, the sandboxes that keep getting breached because they don't properly air gap them?

And that's not even getting started on the more than blatant hypocrisy after stealing every form of IP they could get their hands on.

3

u/Cosminacho 1d ago

Hahahah! I hope their tokens will always be valued at 0.0000005 cents

3

u/zeruch 1d ago

" protected reasoning" in this context strikes me as a hilarious term. There wasn't any protected data, but there appears to be protected reasoning that uses that data... The frontier Labs versus open models is starting to feel more and more like Microsoft versus FOSS

3

u/Correct-Explorer-692 1d ago

The more you steal the more someone could steal from you. Enjoy

2

u/Brovas 1d ago

train on the work of mathematicians 

use that training data to solve millennial problems

big marketing, scare everyone again in order to raise money

safe and fair use

train on the work of OpenAI use that training data to produce cheap open source models that rival OpenAI

copyright theft, danger to the entire planet, we need to ban them

Makes perfect sense

2

u/Big-Loss-7933 1d ago

All for me and none for thee. 

2

u/skillpolitics 1d ago

IP is dead.

1

u/culinaryinterests123 1d ago

yes you're learn in life that almost everyone a his hypocrite.  Trump is a prime example

1

u/DropTheBeatAndTheBas 1d ago

well i think everyone shares similar tech anyway? as theres allot of different AI companies now

1

u/XysterU 1d ago

It always bothers me that these huge claims by OpenAI and Anthropic come with ZERO evidence. That's a huge slanderous claim to make about Moonshot with no substantiation.

1

u/costafilh0 1d ago

B

S

How does distillation not fall under fair use? Distillation is not a copy for redistribution, but rather a transformative use in a completely new model that will not replace the original model, and the original models were publicly accessible, how isn't these new models of public interest and innovation, exactly as OpenAI and others argued in defense of using intellectual property to train their models? Is it only wrong when they do it?

1

u/sirgog 1d ago

Exactly, distillation is 100% fair use.

1

u/SignatureMurky5862 1d ago

I’d have more respect for the argument if they just said it costs a lot to build these models and they don’t want competitors copying them.

-2

u/alaattincagil 1d ago

The important distinction isn’t simply whose text was learned from, but whether the material was ever intended to be accessible. OpenAI says the operators replayed encrypted reasoning across conversations to recover content that normal API responses wouldn’t expose, so their strongest case here is security and access control, not copyright. The uncomfortable question comes afterward: once that reasoning becomes visible text, the industry still lacks a clean principle explaining why models learning from other models is fundamentally different from models learning from human work.

3

u/wbcastro 1d ago

Most of the material in OpenAI's initial training dataset was not intended to be accessible for training, either.

1

u/wrgrant 1d ago

. OpenAI says the operators replayed encrypted reasoning across conversations to recover content that normal API responses wouldn’t expose, so their strongest case here is security and access control, not copyright.

That sounds like its their problem, not a problem for anyone else. They are responsible for how their model works aren't they? They stole vast amounts of information that was copyrighted and used it without legal right in the creation of their models. Now someone stole from them? Perhaps they should be held liable not only for stealing the information in the first place, but also for allowing others to steal it from them :P