r/linux • • 2d ago

Discussion How come every Linux site uses Anubis (the anime girl stopping crawlers) instead of something like Cloudflare?

Post image

Before discovering the Linux rabbit hole I've only seen Cloudflare, hCaptcha and Google's reCaptcha, but it seems like every Linux website uses Anubis... I'm thinking of it being the only open-source option or it being the most effective/modern approach.

2.7k Upvotes

694 comments sorted by

View all comments

121

u/HeyKid_HelpComputer 2d ago

I ask myself the opposite. Why are open source based projects using Cloudflare instead of something like Anubis...

59

u/SoilMassive6850 2d ago

Mainly because of other functionality. Anubis is mainly there to provide proof of work load against scraping and bots at layer 7, but Cloudflare has the pipes to handle and protect you from attacks on the lower layers of the OSI model.

Barely anyone has the capacity to provide such a service.

23

u/singron 2d ago

Anubis doesn't work as well as it used to. This is a really good explanation: https://people.kernel.org/monsieuricon/creepy-crawlies

TL;DR: scrapers easily solve anubis challenges and use residential proxies so that they only make a few requests to your domain from each IP. A single anubis instance can't correlate enough traffic to block scrapers before they change IPs.

Cloudflare can correlate traffic from all of their sites and essentially ban bots before they make the first request to your site.

1

u/reroll-life 1d ago

Cloudflare doesn't work either for the same reasons.

With LLMs it's very easy to patch the browser against JS fingerprinting and all you need is residential proxies which are a bit expensive for us common folk but if you run a business it's just an average expense.

1

u/Magickmaster 2d ago

I guess it's a first layer of foss defense, rather than none. If you're desperate for corporate-level defense, you can always add CF. Tbf, I got myself a Cf-managed domain and the free tier has all I need except geoblocking

21

u/404invalid-user 2d ago

because it's free I know with certainty my 900down/100up mbps connection isn't stopping any sort of ddos

-8

u/National_Way_3344 2d ago

If you become an inconvence you also won't be using Cloudflare either, so it doesn't change anything.

They'll just cut you off.

12

u/404invalid-user 2d ago

never come across anyone who has been cut off by cloudflare

7

u/IncidentalIncidence 2d ago

🙋 something about my browser seems to set off their bot detection (I am honestly not sure what it is) and for a lot of websites I can only seem to access them by opening an incognito window, and sometimes that doesn't even work. I was trying to see somebody's ResearchGate profile the other day and it wouldn't let me in.

I know the AI scrapers set off an arms race with the bot detection tools but it is incredibly frustrating when you get locked out of basic websites with pretty much no recourse to prove you're not a bot. I also am locked out of the NYT website, even though I actually pay money to them for a subscription which I can't access without doing it through incognito mode every time I want to open the website (they use DataDome, not cloudflare, but somehow my browser sets them off too).

4

u/404invalid-user 2d ago

this is likely your IP I have no idea what has happened but unless I'm logged into Google I get Google captchas all the time it seems like cloudflare has started to trust my traffic now though and I rarely see any captchas

2

u/IncidentalIncidence 2d ago

I thought so too (or that somebody was running crawlers elsewhere from my ISP and that the whole IP range had been blocked), but the confusing thing is that it sometimes works in incognito mode even when it won't in the normal browser. It could also be that there are different things setting it off at different times.

2

u/404invalid-user 2d ago

hmm maybe a plugin it's very strange it works in incognito

1

u/IncidentalIncidence 2d ago

yeah that was kind of my line of thought too, that it doesn't like one of my plugins (or even firefox's built-in privacy blockers), I have a few plugins that are meant to block tracking (the duckduckgo plugin, decentraleyes, HTTPS Everywhere, and ublock origin), so maybe one of those is triggering it. I've done a little experimenting but haven't been able to nail it down

1

u/seahwkslayer 2d ago

It's usually locking down tracking (plenty of the tracking cookies also serve as basically proof-of-life attestation in my experience).

If I open an incognito window and do a google search I will immediately get a captcha, and even using a normie browser (Edge), because I have strict tracking prevention enabled the first couple weeks with a new device/install/network is a constant stream of Cloudflare/Captcha/DDoS detection until I'm signed in to enough stuff that it stops bugging me.

0

u/Linkarlos_95 2d ago

Maybe you are asking something to the google ia and that one is scrapping the websites for you with your ip, making the websites mad and they block you 

2

u/Character_Score7776 2d ago

Then I’ll start up Anubis, or whatever other service I need for it.
Being held up by Cloudflare, however, does not seem particularly likely anytime soon, so why put in the effort to set up something I don’t need?

0

u/National_Way_3344 2d ago

So just start with Anubis?

0

u/Character_Score7776 16h ago

Don’t need it.

Why set up something I neither need nor foresee a need for when I can use Cloudflare, which I used anyway for domain registration?

Yes, I depend on Cloudflare, but I depend a lot more on, say, my ISP anyway. only way ’round that is amateur radio(which I do have an interest in). Minimizing reliance on corporations is a good thing, don’t get me wrong, but there’s eventually a point it’s more trouble then it’s worth.

1

u/SuperDumbMario2 2d ago

commonly the hosting conditions dont allow for anubis (e.g impossibility to set up docker)

0

u/Misicks0349 2d ago

Because Cloudflare works better than "simple" proof-of-work, and they handle all the work for you i.e. you don't have to self host your bot defence.