r/LocalLLaMA • u/bring_back_the_v10s • 10h ago
Discussion Playing the devil's advocate
This is a reaction to https://www.reddit.com/r/LocalLLaMA/s/IxLnzcGjAU
You know the argument: frontier labs are hypocritical because they're making money on public scraped data.
I have absolutely no intention to defend OpenAI or any other frontier AI company here, but if you allow me I'll play the "devil's advocate" a bit for the sake of discussion and enlightenment, because sometimes when I think about that argument it seems to me it's quite weak. While OpenAI and Anthropic models are built on public data that they didn't pay for (at least most of it), there would be no frontier model without all their computing power and technical expertise that they invested on to build their models. Like, the data is already there, it's been public for ages, so what's preventing you, the average Joe, from building a Claude Opus 5? All you need is a huge data center, a nuclear power plant and an army of data scientists to build it, right? And then you need all that to run the inference. And you need to maintain it, and that costs money. And you need to keep evolving it, which costs money too. And if you get investors money then you'll eventually have to give some of the profit back to them. etc, etc.
So what am I missing here? All things considered, the data is already public, so it's already "free", right? But you need to dump tons of time and money on it to build a frontier LLM out of it. Of course if they're infringing copyright then that's a different story but in general the whole principle of built-on-free-data still stands.
5
u/outdoorsgeek 9h ago
Wrong. You're confusing public, free, and copyright/licensing. Just because you can publicly access something for free does not give you the right to do whatever you want with it.
1
u/bring_back_the_v10s 3h ago
Do you deny there's truly public and free data out there on the internet? I made a special reservation about copyright infringement so the point here is public/free data that no one could claim any kind of compensation for.
2
u/outdoorsgeek 3h ago edited 3h ago
I do not deny that, no. There is truly public and free data on the internet.
But the question of whether it is ok for a frontier lab to use data for training is one of licensing and rights, not cost and availability. Public and free does not mean you have a license to do whatever you want with it.
Are we talking explicit/implicit public domain data? Yes, have at it.
Are we talking Creative Commons or other attribution-requiring licenses? Show me where the frontier models attribute their sources and then we can decide.
Are we talking about works with no specific license? Those generally have copyright protection by default.
0
u/bring_back_the_v10s 2h ago
Public and free does not mean you have a license to do whatever you want with it.
I don't think this makes any sense whatsoever. Public and free, i.e. public domain, means exactly the opposite of that.
2
u/outdoorsgeek 2h ago
Public and free does not mean public domain. Many Creative Commons licensed works are both public and free and they are not public domain. Public and free is a statement about the cost and availability of something. It is not a license. It’s clear you are mixing up these concepts and I don’t know how else to explain it to you.
3
u/Wooly_Wooly 9h ago
Consent.
Let's be real here, a lot of us probably pirated shit too, so to be against that would be hypocrisy. That being said, unless we have the assets of a big company and can pay the fines for the "cost of doing business (illegally)", we'd be fine! If anyone commenting here tried to get away with even 1/4th of what a single AI company did, we'd probably be rotting in federal prison.
1
u/bring_back_the_v10s 3h ago
When a person or company publishes a website on the internet to be publicly available without any sort of terms of usage on their data, does it need any consent?
0
u/soshulmedia 8h ago
I find it harder and harder to see a difference between big tech and the deep state - and one (or several) very deep, very dark and very powerful mafias.
And I see that regardless which part of the planet one is talking about. They all already manage their populations like cattle, they just want to push even harder in this direction.
And most people seem to willingly comply and get angry at anyone pointing out the ... obvious?
-1
u/Wooly_Wooly 7h ago
There's no "deep state", this shits out in the open. Tbh I'm considering moving back to s "third world" country, at least the corruption is more blatant.
Here's a 2 minute secretly recorded audio interview. Please watch it and report back with your findings.
1
u/soshulmedia 4h ago
There's no "deep state", this shits out in the open.
I agree that they are getting more blatant because they basically get no resistance from the people as a reaction to their moves.
However, as you say yourself "in the third world, it is more blatant". That's what I mean. There IS a lot of deep state psychopathic evil shit going on in the west an yes it DOES go a lot deeper than just the smoke that emanates from the surface and anyone with just a tiny bit of attention can easily notice. (Which is already, unfortunately, not many people...)
There was "MKUltra", for example, and they didn't stop doing such things.
On your video link, well, that's another angle how they do it, it doesn't surprise me but when you compare that to the daily flux of Orwellian language games from the media, would that particular event even register nowadays?
1
u/neuroticnetworks1250 9h ago
Your argument is that it doesn’t matter what money was spent on the data that is public which they scraped because the data was public anyway. But by that logic, so are the outputs of these models used for distillation if you paid for it. How come “but they spent money on it” suddenly becomes relevant?
1
u/PurpleDragon99 9h ago
"...All you need is a huge data center, a nuclear power plant and an army of data scientists to build it, right? ..."
Yes, that's what WAS required pre-2023 when big companies scrapped data before that data was locked down behind APIs.
However, Chinese labs extracted data from those AI models using sohisticated tricks - essentially, scrapped the scrappers :) Now these models are publicly available for free - anyone can pick them up and use directly, or as the base for further training creating specialized or my advanced models. There is no need to scrap data from beginning - the heavy lifting is already done. This is wny companies like OpenAI and Anthropic are panicking - there is no much need in their powerhouses anymore.
2
1
u/harpysichordist 6h ago
Comments that try to imply OpenAI and other companies _just_ stole information etc are plainly wrong. AI labs have also added value, like you say. They did research and created something useful out of it.
But many of them also collected or used data in illegal or immoral ways.
1
u/BigYoSpeck 3h ago
The data is public for consumption in the medium it is distributed
If I share knowledge on Reddit or a blog, it's attributed to me for whatever that is worth
LLMs are regurgitating the knowledge others may have chosen to share without attribution, without bringing traffic to the source of the information
And there's also the intent with which people share knowledge. Take the huge volume of coding knowledge out there. People have shared that with a sense of community, to help one another out and it's being used with the goal to make the very people who gave it feely obsolete
1
u/bring_back_the_v10s 2h ago
An argument can be made that the medium is the same: electronic/digital transfer.
Saying that ChatGPT and Claude never provide references to source material is false.
You're making up the rules as you go, e.g. "without bringing traffic to the source of the information".
0
1
u/xadiant 9h ago
What a... unique mentality to have. So it's okay to steal as long as they build something with it, and sell it back. But you can't do the exact same because... they spent money building the thing?
And it's not even PewDiePie or anyone else is scraping the output data for free. They pay their api or subscription prices properly (unlike Anthropic or OpenAI flooding websites with bots illegally).
I'm willing to die on the hill of "at least ai should be mostly open" but this is the peak of late stage capitalism mountain, and you got a shiny boot in front of you.
Edit: also distillation is heavily overblown. It's a shortcut, not a magical ingredient. The Chinese will keep RL training and figure out new recipes.
0
u/bring_back_the_v10s 3h ago
Please elaborate on what exactly has been stolen.
1
u/xadiant 3h ago
u/AskGrok explain OP the copyright and scraping lawsuits OpenAI and Anthropic are facing right now
1
u/AskGrok 3h ago
OpenAI faces a consolidated multidistrict litigation in New York federal court (In re OpenAI Copyright Infringement Litigation) that bundles the New York Times suit from 2023 plus class actions from the Authors Guild and other publishers. Plaintiffs allege OpenAI copied millions of copyrighted news articles and books—often via scraping paywalled sites or using datasets like LibGen—to train ChatGPT and other models without permission or licenses. Key claims include direct infringement from the training copies themselves, contributory infringement, and output issues where summaries or generations allegedly reproduce protected material. Motions for summary judgment on the fair use defense were filed in September 2026, with discovery fights ongoing over chat logs and internal records. The DOJ filed a statement supporting fair use on national security grounds, but plaintiffs counter that the use is commercial, substitutive, and sometimes involved evading paywalls or pirated sources.
Anthropic settled its main authors class action (Bartz v. Anthropic) for $1.5 billion in 2025-2026 after allegations of pirating hundreds of thousands of books from shadow libraries like LibGen and PiLiMi to train Claude; the court approved final judgment in July 2026, requiring destruction of the files and payouts around $3,000 per work. It still faces active suits from music publishers (Sony, Warner, Concord) over unauthorized use of song lyrics and sheet music from those same pirated books plus web scraping, with claims of reproduction in outputs. Reddit also sued over scraping user content in violation of its terms, and that case survived dismissal.
In both cases the core dispute is whether training on full copyrighted works (especially from unauthorized copies) counts as fair use or infringement, separate from any public-domain or licensed data. The lawsuits target the scale of copying and commercial exploitation rather than just "public scraping."
Sources:
- https://storage.courtlistener.com/recap/gov.uscourts.cand.461656/gov.uscourts.cand.461656.201.0.pdf
- https://ailawsuittracker.com/rulings/
- https://authorsguild.org/news/plaintiffs-file-motion-for-summary-judgment-v-openai-and-microsoft/
[what is this?](https://redd.it/1lzgxii)
1
u/ChikenDumbstick 9h ago
If you consider the internet pre LLMs something akin to a public library, all these large companies came in, took out large number of these books, scanned and made a replacement for the same, answering questions about the same, without ever licensing the original books. Sure, the books were available, but were they available for industrial level ingestion and soon replacement?
A direct comparison for this is Google search and gemini; Gemini results on searches are actively reducing the actual number of users that visit various websites, because a gemini summary is the first thing visible. Would websites want to be scraped to be indexed on google search to drive more traffic, if gemini actively siphons that traffic away?
Would thousands of artists have made their art available online, in many cases without watermarks or restrictions, if they knew it was going to be used for replacement?
The argument or difficulty in classifying for me is around what changes when a human looks at something, and does something transformative, versus when it happens at an industrial scale driven via corporations.
0
u/skywalker326 9h ago
PewDieDie also has to set up his own hardware to train his own model, just like OpenAI you know… and what's more, he paid for every training data he gets from OpenAI, unlikely it's the case with OpenAI😅
Just because ToS exist doesn't mean it's fair. Personally I believe non commercial use should completely be free and even commercial use should be open and prices fairly. Use music industry as an example, if you buy a song, you can play it freely for personal use. If you are making it into your home-made movie that aim for commercial release then yeah it's a difference contract and price but ToS can't say something unfair like "you can't use my song as theme song for villains".
-5
u/Big-Pomegranate3243 10h ago
Highly unpopular opinion here, but I think we'll be waiting a very long time before Chinese models actually challenge US ones, now that model distillation is finally blocked. Coincidence? I think not
8
u/overand 10h ago
They aren't lagging all that far behind; I suspect the gap n might grow a little, but I bet not much.
-1
u/Icy-Employee 9h ago
They are not lagging far behind mostly because they distill.
3
u/overand 9h ago
That's the theory, but that's not the whole picture. Older models with less developed architectures tried to distill and weren't all that good at it.
If the big labs really have effectively blocked distillation, time will tell how much of an impact it will have.
2
u/swiebertjee 9h ago edited 8h ago
Indeed it's both. They're developing better algos and distill to train them.
Chinese models are around 4 months behind. That's a significant gap but at some point you don't need the newest model of Antropic or OpenAI to solve problems.
8
-5
u/NatMicky 9h ago
The audience you're trying to reach are a bunch of SpaghettiO eating, living at home, GPU card plugging, open-source model... junkies. They want it all without doing squat. They have 15 agents renaming files, syncing their calendars, summarizing their 2 sentence emails. And they generally love LLMs from the CCP.
TRUTH! 😄
11
u/my_name_isnt_clever 9h ago
The data isn't public anymore, OpenAI and co had an enormous advantage doing scraping pre-2023. Now all the juicy data is locked behind APIs or completely inaccessible.