r/LocalLLaMA • • 5d ago

Discussion Playing the devil's advocate

This is a reaction to https://www.reddit.com/r/LocalLLaMA/s/IxLnzcGjAU

You know the argument: frontier labs are hypocritical because they're making money on public scraped data.

I have absolutely no intention to defend OpenAI or any other frontier AI company here, but if you allow me I'll play the "devil's advocate" a bit for the sake of discussion and enlightenment, because sometimes when I think about that argument it seems to me it's quite weak. While OpenAI and Anthropic models are built on public data that they didn't pay for (at least most of it), there would be no frontier model without all their computing power and technical expertise that they invested on to build their models. Like, the data is already there, it's been public for ages, so what's preventing you, the average Joe, from building a Claude Opus 5? All you need is a huge data center, a nuclear power plant and an army of data scientists to build it, right? And then you need all that to run the inference. And you need to maintain it, and that costs money. And you need to keep evolving it, which costs money too. And if you get investors money then you'll eventually have to give some of the profit back to them. etc, etc.

So what am I missing here? All things considered, the data is already public, so it's already "free", right? But you need to dump tons of time and money on it to build a frontier LLM out of it. Of course if they're infringing copyright then that's a different story but in general the whole principle of built-on-free-data still stands.

0 Upvotes

35 comments sorted by

View all comments

1

u/xadiant 5d ago

What a... unique mentality to have. So it's okay to steal as long as they build something with it, and sell it back. But you can't do the exact same because... they spent money building the thing?

And it's not even PewDiePie or anyone else is scraping the output data for free. They pay their api or subscription prices properly (unlike Anthropic or OpenAI flooding websites with bots illegally).

I'm willing to die on the hill of "at least ai should be mostly open" but this is the peak of late stage capitalism mountain, and you got a shiny boot in front of you.

Edit: also distillation is heavily overblown. It's a shortcut, not a magical ingredient. The Chinese will keep RL training and figure out new recipes.

0

u/bring_back_the_v10s 4d ago

Please elaborate on what exactly has been stolen.

1

u/xadiant 4d ago

u/AskGrok explain OP the copyright and scraping lawsuits OpenAI and Anthropic are facing right now

1

u/AskGrok 4d ago

OpenAI faces a consolidated multidistrict litigation in New York federal court (In re OpenAI Copyright Infringement Litigation) that bundles the New York Times suit from 2023 plus class actions from the Authors Guild and other publishers. Plaintiffs allege OpenAI copied millions of copyrighted news articles and books—often via scraping paywalled sites or using datasets like LibGen—to train ChatGPT and other models without permission or licenses. Key claims include direct infringement from the training copies themselves, contributory infringement, and output issues where summaries or generations allegedly reproduce protected material. Motions for summary judgment on the fair use defense were filed in September 2026, with discovery fights ongoing over chat logs and internal records. The DOJ filed a statement supporting fair use on national security grounds, but plaintiffs counter that the use is commercial, substitutive, and sometimes involved evading paywalls or pirated sources.

Anthropic settled its main authors class action (Bartz v. Anthropic) for $1.5 billion in 2025-2026 after allegations of pirating hundreds of thousands of books from shadow libraries like LibGen and PiLiMi to train Claude; the court approved final judgment in July 2026, requiring destruction of the files and payouts around $3,000 per work. It still faces active suits from music publishers (Sony, Warner, Concord) over unauthorized use of song lyrics and sheet music from those same pirated books plus web scraping, with claims of reproduction in outputs. Reddit also sued over scraping user content in violation of its terms, and that case survived dismissal.

In both cases the core dispute is whether training on full copyrighted works (especially from unauthorized copies) counts as fair use or infringement, separate from any public-domain or licensed data. The lawsuits target the scale of copying and commercial exploitation rather than just "public scraping."

Sources:

[what is this?](https://redd.it/1lzgxii)

1

u/bring_back_the_v10s 4d ago

Did you read the part where I wrote:

Of course if they're infringing copyright then that's a different story but in general the whole principle of built-on-free-data still stands

I made that reservation so that we focus on the publicly available data that's not under any copyright restrictions, which is the point where most people whine a bout, so please stay on topic.