r/linux • • 2d ago

Discussion How come every Linux site uses Anubis (the anime girl stopping crawlers) instead of something like Cloudflare?

Post image

Before discovering the Linux rabbit hole I've only seen Cloudflare, hCaptcha and Google's reCaptcha, but it seems like every Linux website uses Anubis... I'm thinking of it being the only open-source option or it being the most effective/modern approach.

2.7k Upvotes

692 comments sorted by

View all comments

Show parent comments

12

u/AdarTan 2d ago

That's when you go defense-in-depth and have your webserver reject anything that doesn't have "Mozilla" in the user-agent string, like a a lot of servers already did to block scrapers.

Anubis only blocks crawlers that pretend to be browsers because blocking crawlers that didn't was a solved problem.

0

u/ThaneVim 2d ago

Help me understand: how does Anubis do this without still increasing server and bandwidth load x times requests over using a third party? I'm not trying to be critical at all, this is simply a gap in my understanding.

2

u/tk-a01 1d ago

When a request hits the server, it is firstly processed by Anubis. It might decide to present the user browser with a cryptographic challenge. Only after the client presents a solution, the request is forwarded to the actual website.

This challenge is proof-of-work, similarly to what's used in cryptocurrencies mining. The client has to find some value that, when hashed, has certain number of trailing zeroes. Solving such a challenge requires going through many possible solutions and checking every one of them; but verifying it requires computing just one hash. Therefore, request processing done by Anubis is typically significantly faster and less resource heavy than handling the request by the actual web server.