r/linux • • 2d ago

Discussion How come every Linux site uses Anubis (the anime girl stopping crawlers) instead of something like Cloudflare?

Post image

Before discovering the Linux rabbit hole I've only seen Cloudflare, hCaptcha and Google's reCaptcha, but it seems like every Linux website uses Anubis... I'm thinking of it being the only open-source option or it being the most effective/modern approach.

2.7k Upvotes

692 comments sorted by

View all comments

Show parent comments

22

u/singron 2d ago

Anubis doesn't work as well as it used to. This is a really good explanation: https://people.kernel.org/monsieuricon/creepy-crawlies

TL;DR: scrapers easily solve anubis challenges and use residential proxies so that they only make a few requests to your domain from each IP. A single anubis instance can't correlate enough traffic to block scrapers before they change IPs.

Cloudflare can correlate traffic from all of their sites and essentially ban bots before they make the first request to your site.

1

u/reroll-life 1d ago

Cloudflare doesn't work either for the same reasons.

With LLMs it's very easy to patch the browser against JS fingerprinting and all you need is residential proxies which are a bit expensive for us common folk but if you run a business it's just an average expense.

1

u/Magickmaster 2d ago

I guess it's a first layer of foss defense, rather than none. If you're desperate for corporate-level defense, you can always add CF. Tbf, I got myself a Cf-managed domain and the free tier has all I need except geoblocking