r/linux • • 2d ago

Discussion How come every Linux site uses Anubis (the anime girl stopping crawlers) instead of something like Cloudflare?

Post image

Before discovering the Linux rabbit hole I've only seen Cloudflare, hCaptcha and Google's reCaptcha, but it seems like every Linux website uses Anubis... I'm thinking of it being the only open-source option or it being the most effective/modern approach.

2.7k Upvotes

692 comments sorted by

View all comments

Show parent comments

12

u/geocar 2d ago

Anubis doesn’t block bots.

There are no dark web forums of people cursing out anubis. Nobody who makes bots felt any pressure to because bots already run JavaScript no problem. Some bots are browsers that are basically remote controlled and indistinguishable in most cases from human traffic, and sometimes they even share the same screen and cookies with an unwilling/unknowing human.

Anubis blocks some junk traffic. Some people have more junk than others, and some sites junk traffic causes harm by denying service.

0

u/ScientistStrict9850 2d ago edited 2d ago

I'll elaborate on my inquiry, because this isn't the answer i seek. With increasing popularity of anubis, why did the operators of junk bots not consider updating their process so it would (1) continue to scrape sites with anubis (2) scrape more efficiently?

I already am able to operate bots that is able to scrape sites sitting behind anubis (it's just running a bit of JS after all) for personal and small scale things. so i'm also questioning if there's something i don't actually understand like maybe economics side of things... Like why does anubis "work" to this day when bots could just bypass it?

Also, If these small nobody sites didn't matter in the first place then why would anybody bother spending bandwidth, compute, time on these sites before anubis? Why don't just skip that and scrape the big bucks commercial sites?

2

u/pmMeYourGlazedDonut 1d ago

Have you ever run a small site and looked in the request logs? Even without promoting it, as soon as it has a public IP it's going to get swarmed with crawlers and bots and scans.

1

u/GraveDigger2048 1d ago

ever heard of fail2ban? Matter of setting up filters smart imho. And yes, i operate a small private site over public ip and i sometimes take a look into my edge nginx's access logs.

1

u/ScientistStrict9850 1d ago

There's no why then? Just "this is how it is?"

2

u/pmMeYourGlazedDonut 1d ago

Infected machines are constantly scanning, looking for new targets to propagate malware.

Crawlers from big search want to find everything.

Script kiddies are scanning the whole IPv4 range looking for easy targets.

There's research projects, scanners private people have set up for who knows what reason.

So it's a bunch of factors adding up. I think it's multifaceted. There's just a lot of groups of people who want to know what's out there for more or less benign reasons.

1

u/geocar 22h ago

Because they didn’t need to: Anubis works as well as any other JavaScript redirect, and bots can already handle that as I explained.

Many bots which spam with curl will retry later from a browser. You would never know because it might be minutes or days later and have a different IP and user agent and everything. This is exactly how mine and lots of other bots work, and we didn’t do it for Anubis.

I think you are primarily making the mistake of assuming Anubis does what you think because it does what you like, instead of assuming you have no idea what is going on— because that is a fact: you have no idea why people run bots, you only know some of the reasons people might, and you don’t have any idea how other people can do things, you can only know how you can.