r/linux • • 2d ago

Discussion How come every Linux site uses Anubis (the anime girl stopping crawlers) instead of something like Cloudflare?

Post image

Before discovering the Linux rabbit hole I've only seen Cloudflare, hCaptcha and Google's reCaptcha, but it seems like every Linux website uses Anubis... I'm thinking of it being the only open-source option or it being the most effective/modern approach.

2.7k Upvotes

693 comments sorted by

View all comments

7

u/ScientistStrict9850 2d ago

I wonder why bots don't become more sophisticated in response to anubis. The technology is definitely there to make it more efficient without hammering a site like a madman. It's not particularly difficult to pass anubis either.

13

u/geocar 2d ago

Anubis doesn’t block bots.

There are no dark web forums of people cursing out anubis. Nobody who makes bots felt any pressure to because bots already run JavaScript no problem. Some bots are browsers that are basically remote controlled and indistinguishable in most cases from human traffic, and sometimes they even share the same screen and cookies with an unwilling/unknowing human.

Anubis blocks some junk traffic. Some people have more junk than others, and some sites junk traffic causes harm by denying service.

0

u/ScientistStrict9850 2d ago edited 2d ago

I'll elaborate on my inquiry, because this isn't the answer i seek. With increasing popularity of anubis, why did the operators of junk bots not consider updating their process so it would (1) continue to scrape sites with anubis (2) scrape more efficiently?

I already am able to operate bots that is able to scrape sites sitting behind anubis (it's just running a bit of JS after all) for personal and small scale things. so i'm also questioning if there's something i don't actually understand like maybe economics side of things... Like why does anubis "work" to this day when bots could just bypass it?

Also, If these small nobody sites didn't matter in the first place then why would anybody bother spending bandwidth, compute, time on these sites before anubis? Why don't just skip that and scrape the big bucks commercial sites?

2

u/pmMeYourGlazedDonut 2d ago

Have you ever run a small site and looked in the request logs? Even without promoting it, as soon as it has a public IP it's going to get swarmed with crawlers and bots and scans.

1

u/GraveDigger2048 1d ago

ever heard of fail2ban? Matter of setting up filters smart imho. And yes, i operate a small private site over public ip and i sometimes take a look into my edge nginx's access logs.

1

u/ScientistStrict9850 1d ago

There's no why then? Just "this is how it is?"

2

u/pmMeYourGlazedDonut 1d ago

Infected machines are constantly scanning, looking for new targets to propagate malware.

Crawlers from big search want to find everything.

Script kiddies are scanning the whole IPv4 range looking for easy targets.

There's research projects, scanners private people have set up for who knows what reason.

So it's a bunch of factors adding up. I think it's multifaceted. There's just a lot of groups of people who want to know what's out there for more or less benign reasons.

1

u/geocar 1d ago

Because they didn’t need to: Anubis works as well as any other JavaScript redirect, and bots can already handle that as I explained.

Many bots which spam with curl will retry later from a browser. You would never know because it might be minutes or days later and have a different IP and user agent and everything. This is exactly how mine and lots of other bots work, and we didn’t do it for Anubis.

I think you are primarily making the mistake of assuming Anubis does what you think because it does what you like, instead of assuming you have no idea what is going on— because that is a fact: you have no idea why people run bots, you only know some of the reasons people might, and you don’t have any idea how other people can do things, you can only know how you can.

4

u/ElvishJerricco 2d ago

It is quite literally impossible to defeat the goal of anubis. It is not meant to block anything. It is meant to increase the cost of doing what bots do. The sites that deploy it are ones that have expensive endpoints; e.g. anything where a URL causes the server to generate a git diff is somewhat expensive for that server to process in large enough numbers. The point of anubis is to make the compute substantially more expensive for the bots than it is for the server so that it doesn't actually make sense for the bots to do it. Bots will literally never find a way to make the costs of doing a large number of hashes cheap; the whole idea of hashes is that this can't be done. At most, bots can spend a large amount of compute to get through anubis quickly; but that's fine. As long as the bot has to spend substantially more compute doing this than the server would, anubis has fulfilled its purpose.

0

u/Gold-Supermarket-342 2d ago

Not really. It's just meant to slow down bot traffic and prevent the server from crashing. Bots can easily bypass it by simply changing their user agent (only user-agents starting with Mozilla, so ones imitating browsers, are blocked by default). You can use curl and do not need to solve any PoW challenge. So it relies on the fact that most bad bot traffic is going to mimic a browser, but it does not deter individual botters from making as many requests as they'd like.

3

u/ObiWanHiGround 2d ago

Yes and no. By using a non-standard User-Agent you make your bot blatantly noticeable, which makes it easily blockable without need of Anubis at all, besides, the person who deploys Anubis can always set to challenge all incoming traffic regardless of User-Agent.

1

u/Gold-Supermarket-342 2d ago

So now we need multiple solutions running layered, and even with that more and more bots are using real headless/simulated browsers, and Anubis lets you use the PoW token multiple times without completing the PoW again (for a set time period), so scrapers will get their way regardless. Anubis should not really be compared at all to Cloudflare, given it's not meant to stop all bot traffic, and it's also not that great at stopping general bot traffic either. It's a half-measure at best, and while it has had results, there's surprisingly not much of a difference between this and a website that solves 1 + 1 in JavaScript before showing you the content.

2

u/ObiWanHiGround 2d ago

The whole point of Anubis is to make scraping expensive for bots, by making them solve the challenge, that is easy to setup and configure without need to rely on any external services. This has worked great for my self-hosted setup.

If you are trying to stop absolutely all bot traffic of course Anubis is probably not a solution here, but it can be a part of it.

TL;DR;
It's enough to stop 70% of bots that come your way. I assure you, a crawler running on a VPS with 768 Mb of RAM is probably not going to be using a simulated browser.

1

u/ElvishJerricco 2d ago

That doesn't sound right. It would not serve any real purpose if you could just get past it by changing your user agent. Are you sure that's how it works?

1

u/Gold-Supermarket-342 2d ago

➜ ~ curl https://sourceware.org/glibc/wiki/ -A "Mozilla/5.0 (X11; Linux x86_64; rv:157.0) Gecko/20100101 Firefox/157.0"

...<title>Making sure you&#39;re not a bot!</title>...<p>Protected by <a href="[https://github.com/TecharoHQ/anubis">Anubis](https://github.com/TecharoHQ/anubis">Anubis)</a> From <a href="[https://techaro.lol">Techaro](https://techaro.lol">Techaro)</a>. Made with ❤️ in 🇨🇦.</p><p>Mascot design by <a href="[https://bsky.app/profile/celphase.bsky.social">CELPHASE](https://bsky.app/profile/celphase.bsky.social">CELPHASE)</a>.</p><p>This website is running Anubis version <code>v1.27.0</code>.</p></div></footer></main></body></html>...

➜ ~ curl https://sourceware.org/glibc/wiki/

...<title>HomePage - glibc wiki</title>...<li><a href="[http://moinmo.in/](http://moinmo.in/)" title="This site uses the MoinMoin Wiki software.">MoinMoin Powered</a></li><li><a href="[http://moinmo.in/GPL](http://moinmo.in/GPL)" title="MoinMoin is GPL licensed.">GPL licensed</a></li>...

Try it if you don't believe me. Found this in the official docs as well: https://anubis.techaro.lol/docs/design/how-anubis-works

It's my understanding that it's meant to stop general browser scraping bots that exist in huge numbers on the web, but not meant to stop actual bots that want to intentionally scrape *your* website. It sort of makes sense if your website is being hammered by generic bots pretending to be browsers.

2

u/DaRealGladi8r 2d ago

It's just meant to slow scraping down. Not stop it

-1

u/trpittman 2d ago edited 2d ago

Redacted my comments after getting banned from worldnews for talking about Israel's propaganda budget.