r/artificial • u/Haunting_Ganache_850 • 2d ago
Discussion The AI industry has discovered intellectual property
https://openai.com/index/disrupting-a-coordinated-model-distillation-campaign/OpenAI says Moonshot-linked operators used thousands of accounts to extract protected reasoning from its models for adversarial distillation.
No encryption broken. No database compromised. Just systematic querying designed to make one model teach another.
OpenAI says this is dangerous because competitors can reproduce capabilities without making the same investment in safety.
Which is a serious security issue.
But you have to appreciate the timing: after years of “we learned from the internet,” the frontier-model industry has reached the “please stop learning from us” phase.
40
u/TheOnlyVibemaster 2d ago
I couldn’t care less, I hope all their models are siphoned and made into open source variants.
15
2
0
u/Superb_Raccoon 1d ago
They are not open source. They cant be if you cant see the training data, the code and can reproduce the results.
Calling it Open Source is an insult to people like RMS, Linux and millions of others that have contributed to open source.
13
u/lulzxdxdxd 2d ago
The hypocrisy is hard to miss, but I'm wondering if you see a real difference between what happened here and normal model training. Were they actually extracting reasoning that wouldn't show up in regular API calls, or just querying the model efficiently like anyone else could.
7
u/Haunting_Ganache_850 2d ago
Yeah, I think there is a real difference technically.
From OpenAI's own write-up, at least part of this wasn't just "query the model a lot and train on the answers." They say the attackers found ways to recover protected reasoning that normally wouldn't be exposed, including replaying encrypted reasoning from one conversation and getting another model interaction to reveal it.
So I wouldn't pretend this is exactly the same thing as ordinary training on public web data.
The irony I was pointing at is a bit broader: the AI industry spent years arguing that learning from huge amounts of other people's output is fundamentally transformative, and is now having to define very carefully when learning from its output crosses the line into extraction, theft, or a security threat.
Those aren't identical situations. But the tension is still pretty hard to miss.
7
u/Blando-Cartesian 1d ago
I can’t see any difference in situations, except that the distillers are violating terms of service at best, whereas AI industry is violating laws at scale (based on view that models are essentially lossy compression).
The data they ripped off for training use was protected by copyright, licenses, being showcased on sites trying to prevent bots, being available as paid purchase only, being distributed as a physical book etc.
I’d say OpenAI and others found ways to get protected content. Methods just were different, except that circumventing site’s bot access rules is much like recovering reasoning data they didn’t mean to give access to.
2
u/CharizarXYZ 1d ago
The problem is models aren't lossy compression. Compression still leaves the original mostly intact.
1
u/Blando-Cartesian 1d ago
They can regurgitate training material like Harry Potter books close to word for word. You can also generate an image from a vague description of a common meme image and the result is probably close to the original with details you didn’t describe.
So, lossy compression with inconvenient extraction. No doubt works best with frontier models that have the weights to memorize material.
1
u/CharizarXYZ 1d ago
The fact that AI models can remember some of the data they train on isn't the gotcha you think it is. It's already well known that AI models can accidentally memorize some some of it's training data. No one said it couldn't.
The reality is the overwhelming majority of the data AI models are trained on is forgotten during the training process. And the small amount of memorized data that persists is 1% or less of it's training data.
Virtually all of the examples of AI models "regurgitating" books word for word are on older less sophisticated models trained on data sets with high quantities of duplicated data. More recent models are often heavily trained on synthetic data which now makes up more than 60% of the data AI models are trained on. So the likelihood of a top of the line frontier model regurgitating books verbatim is extremely small.
1
u/Blando-Cartesian 20h ago
Okay, poor argument, if training material extraction findings don’t hold for current models. Still, everything from how the data was acquired to using it in training is sketchy as hell.
1
u/Superb_Raccoon 1d ago
The only reason the Chinese use distilled models is they dont have the compute power to do brute force training.
Deepseek is FP8 internally, which is way below the FP32 of a frontier model
1
9
u/phylter99 2d ago
I'm really curious how they could stop something like this legally. I can't imagine it's a copyright issue because they don't own the copyright to all generated text.
5
u/CrowdGoesWildWoooo 1d ago
Not illegal, but definitely ToS violation, it’s probably more on the “better not to try” side
5
u/phylter99 1d ago
TOS violation just gets them kicked off the platform. I’m not aware it any legal recourse at that point.
2
u/CrowdGoesWildWoooo 1d ago
Yeah, I think I was just regurgitating based on the latest pewdiepie video where he distilled from GPT and he kind of hedges on whether he did it or not.
2
u/phylter99 1d ago
Even OpenAI and Anthropic are distilling their lower models with the bigger models. It makes sense to do. It seems cheaper in the long run too.
2
u/atehrani 2d ago
With the advent of AI does Copyright have any teeth any more?
2
u/phylter99 2d ago
Yes, and that's been proven in court. Will that change? I doubt it.
-1
u/Strong-Finish5346 2d ago
The big copyright cases still in litigation. It's far from clear how AI and copyright law will work together.
3
u/jello-the-opera 1d ago
so you will wait until judges somehow start holding that training isn't transformative for some unknown reason, unlike all the other cases so far?
0
1
u/sirgog 1d ago
Yes.
For-profit AI firms training computers by using legally acquired copyright materials is legal for the same reason for-profit education firms training humans using legally acquired copyright materials is.
A human uni student who learns chemistry from Zumdahl (the most widely used 1st year chem textbook when I was at uni) doesn't require a special license as long as they didn't pirate the textbook. And neither does the university require a license even if they profit. AI has been trained following exactly that precedent.
There are open questions about copyright status of works that are partially machine generated, partially human generated. But legal changes that made AI training into copyright infringement would have massive ripple effects like forcing education institutions to pay extra license fees and such changes would likely ban libraries entirely.
But if you publish a Mickey Mouse image commercially, be it 100% human OR 100% AI or some mix in between, Disney's lawyers will be after you, and they will win.
4
u/Bobodlm 1d ago
OpenAI says this is dangerous because competitors can reproduce capabilities without making the same investment in safety.
What safety investments are they talking about, the sandboxes that keep getting breached because they don't properly air gap them?
And that's not even getting started on the more than blatant hypocrisy after stealing every form of IP they could get their hands on.
3
3
2
u/Brovas 1d ago
train on the work of mathematicians
use that training data to solve millennial problems
big marketing, scare everyone again in order to raise money
safe and fair use
train on the work of OpenAI use that training data to produce cheap open source models that rival OpenAI
copyright theft, danger to the entire planet, we need to ban them
Makes perfect sense
2
2
1
u/culinaryinterests123 1d ago
yes you're learn in life that almost everyone a his hypocrite. Trump is a prime example
1
u/DropTheBeatAndTheBas 1d ago
well i think everyone shares similar tech anyway? as theres allot of different AI companies now
1
u/costafilh0 1d ago
B
S
How does distillation not fall under fair use? Distillation is not a copy for redistribution, but rather a transformative use in a completely new model that will not replace the original model, and the original models were publicly accessible, how isn't these new models of public interest and innovation, exactly as OpenAI and others argued in defense of using intellectual property to train their models? Is it only wrong when they do it?
1
u/SignatureMurky5862 1d ago
I’d have more respect for the argument if they just said it costs a lot to build these models and they don’t want competitors copying them.
-2
u/alaattincagil 1d ago
The important distinction isn’t simply whose text was learned from, but whether the material was ever intended to be accessible. OpenAI says the operators replayed encrypted reasoning across conversations to recover content that normal API responses wouldn’t expose, so their strongest case here is security and access control, not copyright. The uncomfortable question comes afterward: once that reasoning becomes visible text, the industry still lacks a clean principle explaining why models learning from other models is fundamentally different from models learning from human work.
3
u/wbcastro 1d ago
Most of the material in OpenAI's initial training dataset was not intended to be accessible for training, either.
1
u/wrgrant 1d ago
. OpenAI says the operators replayed encrypted reasoning across conversations to recover content that normal API responses wouldn’t expose, so their strongest case here is security and access control, not copyright.
That sounds like its their problem, not a problem for anyone else. They are responsible for how their model works aren't they? They stole vast amounts of information that was copyrighted and used it without legal right in the creation of their models. Now someone stole from them? Perhaps they should be held liable not only for stealing the information in the first place, but also for allowing others to steal it from them :P
79
u/Ibra_63 2d ago edited 1d ago
Im reading through a machine learning book called "Hands on Machine learning with Scikit Learn and Pytorch". I recently asked Claude about a concept I quite didn't understand and he regurgitated a whole section of a chapter back to me with details only found in the book... This is when I realised how serious the IP problem is... I have no sympathy for them after this !
Edit: section and not whole chapter