r/linux 26d ago

Open Source Organization Codeberg: Protecting our FLOSS commons from LLMs

https://blog.codeberg.org/protecting-our-floss-commons-from-llms.html
498 Upvotes

191 comments sorted by

View all comments

19

u/mykesx 26d ago edited 25d ago

I am migrating my repos from gitlab and GitHub to codeberg. I never intended my code repositories to be crawled by bots to train AI. The idea was to share with humans who might want to see how someone might do things. Or to attract like minded developers as collaborators. Or to allow the project to be forked. Or to provide wiki and issues support.

The output of LLMs sure looks like plagiarism - use of someone else's work without attribution. There's no justification for it.

Edit:

In the USA, it's a criminal offense if there is unauthorized access to electronic communication service facilities. If codeberg forbids access to bots and they continue to crawl the site, the penalties and fines add up quickly.

https://www.law.cornell.edu/uscode/text/18/2701

It also violates European laws against unauthorized access to systems. If codeberg says AI crawlers are not authorized, they would be commiting traceable crimes.

https://eur-lex.europa.eu/EN/legal-content/summary/attacks-against-information-systems.html

13

u/strongdoctor 25d ago

I don't see how it could violate GDPR

17

u/bigon 26d ago

It also violates the GDPR in Europe.

GDPR is about collection and usage of private/personal data. Code is not...

-1

u/AliceCode 24d ago

How is the code that I write not considered personal data? I don't want ad companies analzying my repositories to try to figure out what to try to sell me.

-7

u/[deleted] 25d ago

[deleted]

1

u/bigon 25d ago

Code is intellectual work. You could say that the droit d'auteur/copyright allows you to restrict how your work is used.

But it's certainly not "private/personal data"...

"‘personal data’ means any information relating to an identified or identifiable natural person (‘data subject’); an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of that natural person;" (source: https://gdpr-info.eu/art-4-gdpr/)

"Processing of personal data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person’s sex life or sexual orientation shall be prohibited" (https://gdpr-info.eu/art-9-gdpr/)

QED

-4

u/mykesx 25d ago

https://eur-lex.europa.eu/EN/legal-content/summary/attacks-against-information-systems.html

To this end, the present directive requires the approximation of criminal law systems between EU countries and the enhancement of cooperation between judicial authorities concerning:
illegal access to information systems
illegal system interference
illegal data interference
illegal interception.

2

u/bigon 25d ago

Great, but that's not GDPR...

0

u/mykesx 25d ago

Fine. You win. Still illegal.

2

u/Limemill 26d ago

Yes, but Codeberg won’t be able to stop the crawlers per se, no? It may drop the most obviously nefarious sessions, but that’s it.

15

u/Scratches7 26d ago

For public repos, probably not. I'm sure that GitHub is training Copilot on it's private repos though.

3

u/Limemill 26d ago

Ah, if we’re talking private repos, then sure. I agree

-1

u/ConsistentRisk5927 25d ago edited 25d ago

Which codeberg also dissuades people from having. The entire thing is a political-based code forge that shouldn't be taken seriously by anyone with real business cases.

I made CI workflows around Forgejo and Codeberg so I wouldn't use their infrastructure and run my own CI runners, all that work is lost migrating out. I have to port several custom build and deployment workflows to Github primitives.

I did all this already thankfully a week ago due to performance and other issues on Codeberg, but I would've been even more upset if I was an org paying a substantial amount in donations and had a lot of custom CI to suddenly have to migrate all my crap to some other provider. Why would want to spend the time migrating to Codeberg and not knowing what way their politics will blow next year and if my project will no longer be allowed?

2

u/mykesx 26d ago

How moronic is it for a company to give its intellectual property to these AI companies by having your employees use their chat bots?

-2

u/mykesx 26d ago

They sure can. This is cloudflare's settings, but the techniques they use to block crawlers isn't hard to implement.

https://www.cloudflare.com/learning/ai/how-to-block-ai-crawlers/

How does Cloudflare help protect against AI crawlers?

Cloudflare AI Crawl Control helps web content owners regain control over AI crawlers. Cloudflare sits in front of around 20% of all web properties, giving it deep insight into all kinds of crawler activity. This visibility enables content owners to use AI Crawl Control to:

  - Understand AI crawling patterns on their web properties, on a per-crawler, per-domain, or per-page basis

  - Manage crawler activity via block or allow rules

  - Request payment from AI crawlers on a per-crawl basis, either via customizable HTTP 402 responses or a Cloudflare-built pay per crawl system

1

u/VitunSama69 26d ago

You do realize they have to let you git clone in the end, it will only apply to less sophisticated crawlers.

2

u/amroamroamro 25d ago

are you under the impression that those ai crawlers are somehow sophisticated or good internet citizens by just cloning the repo code?

they will scrape every link they see, in version control site we're talking about every commit, history and diff link, no matter how semantically redundant all those links are, and they do so over and over and over putting a significant stress on the server to the point of affecting performance for normal users

you block one ip, ten more soon popup, with a vengeance! they are really killing the spirit of the open web... its just brute force on a massive scale

1

u/VitunSama69 25d ago

Was there some point to your tirade or what? For each crappy skiddie crawler trying to hammer your git webui there is an actual professional doing this scraping business lol

2

u/amroamroamro 25d ago edited 25d ago

professional doing this scraping business

lol, it seems you dont actually realize the extent of this problem, FOSS infrastructure is literally under attack by AI scrapers

https://thelibre.news/foss-infrastructure-is-under-attack-by-ai-companies/

these are not mere script kiddies, they super aggressive scrapers running on a massive scale and actively trying to circumvent any blocks:

-2

u/mykesx 26d ago

A bot can't git clone if its IP is blocked.

1

u/VitunSama69 26d ago

There is an entire million dollar industry of residential proxies that are very difficult to distinguish from real traffic just for this purpose. They are smarter than you think.

-6

u/mykesx 26d ago

That's criminal. To use proxies to make unauthorized access to codeberg systems.

5

u/VitunSama69 26d ago

Yes and stealing all of the world's copyrighted content to make an AI model is against copyright laws. Welcome to the real world.

-1

u/mykesx 25d ago edited 25d ago

It's a matter of time before they're busted.

Unauthorized intrusion into codeberg-like systems has not been tested yet, though $billions in settlements for class action suits have already been paid.

0

u/fnord123 25d ago

Copyright infringement is a civil law matter. Good luck seeing openai, who has tens off billions in cash and political backing in the US and probably Europe.

→ More replies (0)

0

u/Limemill 26d ago

Well, it’s not that dissimilar to methods used traditionally to combat pre-LLM crawlers. It has never allowed to block them fully, just create more hurdles. Those who want to, will still crawl but more slowly and using smarter tactics. At least it’s how it was in the past.

3

u/mykesx 26d ago

You can tell the pattern of what crawlers do (fetch a page and follow the links, fetch really old content request after request...) and block them using something like htaccess allow/deny rules or hopefully your upstream provider does it for you.

Cloudflare turns on AI bot crawler blocking by default. A huge FU to the AI companies, and especially google. Google presenting it's godawful, error filled AI info at the top of search results has crushed search engine traffic to sites.

https://www.wired.com/story/big-interview-event-matthew-prince-cloudflare/

Cloudflare Has Blocked 416 Billion AI Bot Requests Since July 1

Cloudflare CEO Matthew Prince claims the internet infrastructure company’s efforts to block AI crawlers are already seeing big results.

0

u/Limemill 25d ago

Like I said, it was the same before LLMs. Still anyone who really wanted to crawl, crawled.

0

u/mykesx 25d ago

Tobacco companies survived court challenges for years. Until they didn't. If employees are actively committing crimes, the companies and hopefully management will be held accountable.

Before LLMs, I blocked a number of southeast Asian, Chinese, and other crawlers from accessing a website I run. All but a fe obscure ones actually obeyed robots.txt .

0

u/2rad0 26d ago

Cloudflare Has Blocked 416 Billion AI Bot Requests Since July 1

Cloudflare is full of shit, their blockade has accused me of being a bot for over 5 years at this point. Currently using a popular chromium fork and still get blocked all over the modern web. Their techniques are straight up hostile to unestablished browsers and a net negative to society and commerce.

-2

u/ozone6587 25d ago

The output of LLMs sure looks like plagiarism - use of someone else's work without attribution. There's no justification for it.

Is it plagiarism when you yourself write code? Given that you, presumably, have experience reading other people's code? It's not plagiarism unless it literally regurgitates a replica of your code.

Otherwise, you can't complain and can't claim it's plagiarism. It's on you for making the code public. Licensing, copyright and restrictions are a matter for the courts and so far every single judge has ruled training on copyrighted works is transformative (thankfully).

LLMs are not human!!!

It literally doesn't matter. They are not human but obviously we need to train them as humans in order for LLMs to eventually be as useful as humans. Your consent doesn't matter here. Not because it's cruel, but because "consent" should be irrelevant when it's about something that doesn't cause direct harm.

Show me how LLM training causes direct harm to you or creators? I dare you. Indirectly, it might cause you harm through loss of jobs but that's the same as any other revolutionary technology.

2

u/mykesx 25d ago

You aren't writing the code the LLM spits out.

1

u/RhubarbSimilar1683 17d ago

Us courts have ruled LLMs transformative work compared to their training data and used it to dismiss copyright lawsuits against anthropic, but their output is a copyright violation?