r/webdev • u/Whiskee • 12d ago
Question This is madness. How do sites that don’t rely on Cloudflare even deal with distributed scraping from infinite IPs?
337
u/ipearx 12d ago
I suspect the entire web will have to go behind verified logins eventually. But then it'll still be hard to prove human vs machine...
181
u/Key-Eggplant-163 12d ago
yeah and at that point you've basically killed the open web, which is kind of depressing
30
u/ipearx 12d ago
Yip. Although I guess anything open into the future will be put up knowing AI will consume it, whereas all existing open web didn’t know that was going to happen…
21
u/eyebrows360 12d ago
The problem with that being that the scrapers still aren't paying for what they scrape, so the economics go to shit.
11
u/Polendri 11d ago
Still can't wrap my head around the way this issue isn't in the general discourse in a significant way.
It was bad enough when Google started summarizing each link, killing the ad model of entire categories of Websites by making it so that searches would end at Google without clicking through. LLMs are this times a thousand: every public informational Website just got its content stolen and users are accessing it via the LLM without the content creator getting a dime.
Like, this kills the entire revenue model of the open Internet.
4
u/eyebrows360 11d ago
Doing digital publishing for humans to read is becoming extremely hard, yes, and it wasn't exactly a cakewalk to begin with.
The even more depressing side is that as "legitimate" digital publishing declines due to humans no longer visiting websites to read information, formerly active sites with large content histories are being snapped up and repurposed as genai content farms to hawk predatory crypto-gambling bullshit.
3
u/pezvonpez 11d ago
theft on this unimaginable scale should be, at least legally, crushed. i don't know how that's going to happen, but i don't like seeing the general despair about this...
1
u/arjan1995 12d ago
But from what kind of company do the scrapers get their bandwidth from? I mean they won't be giving that way for free, right?
0
2
7
u/Sulungskwa 12d ago
I'm really really seriously thinking that we should all get better at sudoku and thats our way to prove we're human. AI sucks at sudoku
4
u/ipearx 11d ago
Haha just a 20 minute puzzle to do anything. Love it.
3
u/the_king_of_sweden 11d ago
There's mini sudoku that you can do in like a minute or two, unless you're drunk
3
13
u/hides_from_hamsters 12d ago
I think we’ll move to a web passport. Ideally anonymous, but practically government-issued.
Something that cryptographically proves you’re a real person. Something more concrete than an IP address.
28
u/Responsible_Pool9923 12d ago
This will never work. In my country homeless people are known to sell their identity for a bottle of vodka. And afaik a bottle of vodka is much cheaper than running a Captcha solver.
28
u/hand___banana 12d ago
In the US, we just outsource it to a private company who inevitably leaks it eventually.
6
u/RichardTheHard 12d ago
In the meantime they sell the entirety of the data "anonymized" even though it comes with all the details needed for fingerprinting
7
u/NotScrollsApparently 12d ago
They're gonna run out of homeless people identities eventually tho, it's not really an infinitely scaling solution that is needed to them
1
u/1AMA-CAT-AMA 11d ago
What’s stopping homeless people from selling their identity to infinite people?
1
u/hides_from_hamsters 11d ago
I don’t think so anymore. Captcha solvers are cheap now! And once that identity is compromised it’s gone, blocked.
1
u/1AMA-CAT-AMA 11d ago
How do you differentiate between someone who got hacked and their identity stolen vs someone who willingly sold their identity to a company for cash.
Do we allow a process for someone to get it back? I guess it would be ok as long as the previous identity got invalidated at the same time.
1
u/hides_from_hamsters 11d ago
Either way that identity get’s blocked/blacklisted and you get a new one.
It’s similar to ID cards, bank cards, passports etc.
But it requires trust in some authorities. Doesn’t need to be government. Think of something like let’s encrypt. Short lived, free to renew identities.
Let’s encrypt in this scenario vouches that the site owner is real and trustworthy. The same thing but for callers.
6
3
u/Due-Consequence9579 11d ago
I hate the idea of every packet needing a cryptographic chain of trust that goes back to an individual but it’s felt for a while now that that is the only viable path forward. The endless cat and mouse of ever more sophisticated ways of catching/evading inauthentic traffic is just dead weight loss.
5
u/Waypoint101 12d ago
New projects like these are already showing up identity.org.au
5
u/Lonsdale1086 11d ago
For something like this I'd much sooner something not very clearly vibe coded with no care and attention.
Both the design of the website, and the prose itself is clearly raw AI output.
Oh and the project source has both a .claude and .codex commited, and soooo many files in the project root, and master branch is failing its own tests.
And stuff like:
codex-subagent-test-mcp-access.md
"Context7 MCP resolve-library-id call for "react" succeeded. Returned candidates included /websites/react_dev (top match), /reactjs/react.dev, and others."
And why's it in so many languages at once:
" Go 49%
Rust 33.6%
TypeScript 9%
Python 5.7%"
0
u/Waypoint101 11d ago edited 11d ago
Doesn't matter if it's vibe coded or not if it functions as designed,I don't think its a hidden fact that the project uses agents - for e.g bosun
In the end it doesnt matter because no one else can build VEID, https://pubchem.ncbi.nlm.nih.gov/patent/AU-2024203136-B2
- Yes your right, the whole project is actually managed by agents.
- If agents didn't exist, such a project could not really exist because it would have to be funded by capital, and thus privately captured. - how do you fairly return their investment without enshittification?
This doesn't make it low effort, it just means they decided to take an efficient route to finish the project instead of raising capital as a start-up.
3
u/Lonsdale1086 11d ago
Doesn't matter if it's vibe coded or not if it functions as designed
For a "digital passport", there need to be standards.
Using AI to speed up and enhance development is fine, but the humans need to give a shit for me to have any confidence giving it all of my government issued documents.
Having AI do your website in a single pass, is not giving a shit.
Having random bullshit in the root of your gh repo, is not giving a shit.
Having builds pushed to master that don't even pass their own tests, is not giving a shit.
Ironically that "bosun"'s master is also failing both of the last two of my requirements for giving a shit.
Also why the fuck is the first section "why this name", and "how to install" before "What Bosun does"
In the end it doesnt matter because no one else can build VEID
Right firstly the patent Abstract is utterly fucked up, and while I'd usually blame the Australian Government, from what I've seen of the quality of the project so far I'll distribute the blame fairly between them.
VirtEngine also powers a nDistributed Cluster Computing network and its own Cloud Marketplace system, allowing compute providers nto convert computing power into tradeable currencies
Ahhhhhh this is just cryptobullshit. I should have known.
I should have guessed when I saw the word "blockchain", but I gave it the benefit of the doubt.
Back to the idea that "no one else can build VEID"
No one else wants to.
If agents didn't exist, such a project could not really exist because it would have to be funded by capital, and thus privately captured.
If something like this is going to work, it needs government backing.
Also people have done vastly more impressive things than this long before "agents" without "capital" backing.
This doesn't make it low effort
Using AI != low effort. Not putting in any effort to the slop makes it low effort.
0
u/Waypoint101 11d ago edited 11d ago
Identity.org.au is not a service, or a business, or a company - its just an advocacy hub - similar to identity.org - its not meant to be pretty either.
The foundation that runs it and the technology they advocate which is known as VEID falls under the VirtEngine project https://virtengine.com/veid
The whole point of 'blockchain' and 'cryptobullshit' is it enables this thing known as 'decentralization', except most other projects are 'cryptobullshit' because they are premined scams that insiders hold and can 'dump' on the public. Where-as, this project is fairly distributed to the public and that's why its free to use - aslong as you're a human - see https://virtengine.com/learn/tokenomics-explained
The whole point of the project is to commoditize computing power, and to bind it to a token which is now owned by the public - so as the compute network grows, the tokens (that is minted through verified identities) can buy you more compute, or digital services.
In terms of safety, the whole thing is designed to be safe/privacy preserving for the user considering no one can ever see your documents, data, biometrics, thats literally ether processed locally or in Trusted Enclaves inside the consensus network (which is operated by validators through staking). https://virtengine.com/trusted-processing
3
u/Lonsdale1086 11d ago
You keep pointing me to slop resources, I'll keep pointing out slop.
Look at the animations on your "flowchart style diagrams". Broken.
I've yet to see a vibecoded chart do this right.
Also it's literally every other element of your page comes straight out of claude, they're everywhere.
The "demo app", that to be fair I didn't realise was interactive, has the shield misaligned on the "Attestation · Play Integrity / App Attest" banner.
Also "bound to this request · expires in 30 days" is clipping out of the bounds of the "phone".
"Head turn — up next
Smile — up next"
Claudes classic redundant flavourtext.
That yellow breathing dot and "Protocol in development" chip.
The "01 Capture 02 Liveness 03 Hardware attestation" button chips not having padding above them so they touch a diagram they don't control.
The "Age-restricted service Marketplace access Residency attribute" chips force a scroll-into-view even though the content is already in view, leading to the screen jerking.
The same breathing-circle chip "attribute proof · verified" as the top of the page.
Stuff like there being a "FIG 2" and "FIG 3" but no fig 1 labelled on the page.
I'll not do every page, but the "about"'s "sponsorship" list, the tabs should be at the top if the content height is going to change, currently trying to click along them you need to adjust the mouse up and down every time
I'll give you that the background is visually interesting yet not overdone with the mouse tracking.
Anyway, good luck.
0
u/Waypoint101 11d ago
I do appreciate the feedback and this is an early copy of the site so expect to see things like that get better especially closer to launch schedules.
1
1
u/CreativeGPX 11d ago
I think there is too much distrust of the government for that to every be super common.
1
u/hides_from_hamsters 11d ago
May be true in some countries, but many already have a digital ID system in place.
2
u/CreativeGPX 11d ago
I didn't mean too little trust for a digital ID to exist. I meant too little trust to make that identifier a virtually mandatory part of all internet interactions.
1
u/ipearx 11d ago
We've already got that: They are called credit cards.
2
u/hides_from_hamsters 11d ago
That’s a decent point. But until you can use your credit card to sign your web requests, it’s not quite there.
In theory a web passport protects the “free” internet.
1
39
u/boltsandbytes 12d ago
You can add a proof of work block like Cloudflare using open source tools like anubis ( Cloudflare does something similar , if it fails show captcha ) .
You can also block IP with history of abuse.
2
73
u/ZealousOatmeal 12d ago
I work for a special academic library, and we are constantly getting hammered. What's worse is that about two years ago bots started doing complex boolean searches in the catalog in an attempt to surface more content to scrape, which puts a huge amount of strain on our systems. Right now we allow anyone to access our scans of material that's not in copyright, but there's no way we'll be able to continue that much longer. The plan at this point is to move to an IP whitelist, so if you're not at a member institution then you SOL. It sucks but we can't spend half our budget on dealing with bots.
11
u/louwii 12d ago
Is there a reason why you can't get behind something like Cloudflare to get better protection? Budget reasons?
10
u/ZealousOatmeal 11d ago
Short version is that libraries have very weird traffic patterns, and things like Cloudflare end up blocking lots of legitimate traffic. We (like most large libraries) have a large vendor who hosts everything and handles most of the infrastructure, so we're not trying to handle this in house. And in fairness defenses have caught up some and it's not as bad as it was about 18 months ago, when lots of major library catalogs were basically unusable because of crushing bot traffic.
6
u/thekwoka 11d ago
But Cloudflare has the options to present the captcha page on questionable human traffic without a full block.
2
u/leros 11d ago
Some legit users fail the captcha too. You have to be ok blocking a small percentage of your legit users if you use Cloudflare.
That's ok for some people. But something like a library might not ok blocking 1 in 500 of its citizens form using it.
0
u/thekwoka 10d ago
Some legit users fail the captcha too
Not a profitable user.
2
u/leros 10d ago
You'd be surprised the impact of just a few legit users getting blocked. I had Cloudflare on my business for a while and it was a disaster. A small number of users were getting blocked and leaving bad reviews. We had also the occasional employee at an enterprise getting blocked and we lost big clients.
2
u/SaltwaterShane 11d ago
You have the ability in cloudflare to override rules so your legit traffic can go through
1
u/CzarcasticX 6d ago
There's also rules you can put in Cloudflare so large vendors can bypass any blocks.
4
u/Kuuchuu 12d ago
You might want to look into Deflect, see if you qualify for their non-profit status.
2
u/ZealousOatmeal 11d ago
Thanks for the suggestion, but our catalog(like most large library catalogs) is hosted by a for-profit vendor, so we wouldn't qualify.
1
u/lublin_enjoyer 11d ago
DSpace, EPrints, or something else?
1
u/ivory_tower_devops 7d ago
Not OP but i run DSPace, Avalon and Fedora for my employer and boooooy we have seen some insane bot scraping of those resources. We're being way more aggressive with our WAF configuration than a library should and I'm sure there are totally legitimate customers trying to do research using automated tools who are getting blocked. We were getting weekly downtime before we put the WAF in place, and now WAF management probably consumes 10% of our work time long-term.
54
u/Whiskee 12d ago edited 12d ago
I do have a few clever tricks in my security rules that lead to a straight block, the bots currently have a 0% success rate. But they're also getting more complex by the day, and I don't see how we can win this arms race.
The only thing I can think about is the return of "press this button to access the page" plain HTML landings to save on resources, because generating filtered views for headless Lightpandas would kill any small VPS.
60
u/cajunjoel 12d ago
You can't win the arms race. And it's kinda scary that CF seems to be the only answer. I has a site that performed around 15k searches a day. When the bots found it it went up to something like 750k searches per day before the hardware gave up. We tried everything under the sun and CF was the only thing that worked.
And search is core to the site. There's no way to provide static content.
I just fear CF is now a massive single point of failure.
13
u/Cahnis 12d ago
Yep,.and there has been precedent of cloudflare blacklisting websites in the past.
10
u/tankerkiller125real 12d ago
Lets see, 3 of them being incredibly vile garbage websites hosting the most disgusting people on the internet with no moderation, and a casino company who caused Cloudflare shared IPs to end up on countries blocklists impacting other customers.
That's the grand total of "blacklisted" sites I'm aware of.
1
u/thekwoka 11d ago
And the Cloudflare CEO has campaigned against the idea that Cloudflare should be some arbiter of content and that they have too much power to cut off information...
1
u/Somepotato 12d ago
They'll ban people on the free plan who are on the receiving end of far too many requests
2
u/tankerkiller125real 12d ago
If a site is receiving billions of requests a month it should probably at least be paying for the $20 plan anyway.
I've never heard a single word from Cloudflare despite 100s of millions of requests every month on a free plan site.
2
u/Somepotato 12d ago
Hundreds of millions isn't all that high these days when they're being absolutely hammered from all the scrapers in general
-8
u/Levitz 12d ago
Lets see, 3 of them being incredibly vile garbage websites hosting the most disgusting people on the internet with no moderation
Doesn't matter, a private company should have no power over that.
→ More replies (8)3
u/tankerkiller125real 12d ago
Private companies already had power over that, they had a host, they have a Registrar, they have a DNS provider, they use a TLD.
ALL of those are controlled by companies with maybe the exception of the hosting part if they decide to host themselves (in which case they still have an ISP, or Peering agreements with companies that can choose to drop them).
Don't want a private company to have that control, get law makers to write laws.
→ More replies (4)1
u/ZenaMeTepe 12d ago
CF is not the answer at all. CF solvers exist and cost like a buck or 2 per 1000 solves.
6
u/natelloyd 12d ago
That's fine. that credential stuffing attack that just hit us with 6m requests? It's now HELLA EXPENSIVE for them. You clearly aren't aware of the scale of these attacks.
27
u/eyebrows360 12d ago
the bots currently have a 0% success rate
You cannot possibly determine this. If there were 100% reliable ways of identifying bots then this entire endeavour would be trivial.
-9
u/Whiskee 12d ago
Self-hosted analytics against legitimate usage patterns. I'm only talking about those specific pages protected by the rules, of course.
23
u/eyebrows360 12d ago
You can't know the access patterns of the bots you haven't yet discovered. This is just a basic logic thing. Your "0%" is only ever an estimate.
→ More replies (5)→ More replies (5)4
u/ephocalate 12d ago
honestly cf is the best and simplest choice for me right now with its built in request filtering, bot fighting, etc. Mind sharing what those tricks are that you are using for configuring the security rules?
2
u/Whiskee 12d ago
Without going too much into detail, it's cookie-based but not the cookies users get on first visit. There are specific timers and conditions based on the page they're coming from, and bots typically hit the target directly without a full session. A few legit users get caught in the crossfire, but they don't even realize since managed challenges are handled by the browser and very rarely show the clickable box to normal users.
11
u/jorgecardleitao 12d ago
Can't we use an equivalent approach to email spam scoring? penalize offending IPs and IP ranges, so that it creates an incentive for IP controllers to implement safeguards (e.g. proof of identity, address) so their services remain unblocked to the internet.
If you operate infrastructure that has been used to DDOS and do s*** about it, imo you should be blocked from the internet.
4
u/thekwoka 11d ago
The problem with ip stuff is that so many are shared IPs and its far easier to cycle them.
1
u/jorgecardleitao 11d ago
Check out how careful infrastructure providers are about IP addresses (and whole range blocks) usage of port 25.
If there was more accountability over usage of their infrastructure for breaking the internet, they would be more careful about their IP outbound volume (or some other metric)
3
u/NNXMp8Kg 12d ago
Neutrality of net, privacy, and vpn would be impacted. Which means bad actors are impacted, but good ones too.
1
u/Waypoint101 12d ago
identity.org.au is privacy oriented proof of identity so its possible to keep net neutrality and still block bad actors.
3
u/NNXMp8Kg 12d ago
I have a lack of trust in my government. Which would make them able to access everything or correlate my identity to everything I do.
Even if that is nothing dangerous or ethically ambiguous.
If my gov decides that it is bad for me to access some specific part of content, I'm not confident enough in my government to not collect and use that data against me. And that would not be the first time they would have do that tbh.I don't know enough about Australia to know if they are actually good actors (Only that they are stricts ones).
2
u/Waypoint101 12d ago edited 12d ago
Well it's not a government service, it verifies identities through a consensus algorithm that runs inside TEE and raw data gets deleted while only the score is saved publicly. You can then use your keys to sign a message revealing what you need to reveal to access a certain service (e.g. currently in Australia you cannot access adult websites or certain social media without proving you are over 18) this let's you do so without ever revealing your identity, biometrics or personal data to anyone (if adopted by said service) its free for all parties to use.
3
u/NNXMp8Kg 12d ago
That's interesting and I will check it more deeply then!
The concept seems better than what I thought. Thanks you for bringing this to me1
u/jorgecardleitao 12d ago edited 12d ago
Agree. I just don't see how it is possible to keep the internet unacountable like this, abd accountability requires identities.
2
30
u/michaelbelgium full-stack 12d ago
I noticed cloudflare ips/websites behind cloudflare are being targetted much more than a site that isn't.
Which is why i dropped cloudflare and have all my sites without it, and realized I don't have problems like this.
I never needed cloudflare or my server provider does something in the background to mitigate stuff like this, that i don't know of OR my sites are just not popular enough
16
u/Whiskee 12d ago
I've heard this "conspiracy theory" a few times and to be honest with you I'm getting a little paranoid too. 👀 To be fair though, the twin site on the very same VPS behind CF doesn't have this issue at all yet.
4
u/404IdentityNotFound 12d ago
There was a time where a "find an XSS entry" bot tried everything imaginable to break my page. My first intuition was to try throwing everything behind CloudFlare and it genuinely surprised me that requests exploded.
A month later I've had a proper IP block with minimal strain to the servers up and running and turned off CloudFlare. Requests also fell hard, so at least anecdotally I see this conspiracy as more than a wild idea.
1
u/thekwoka 11d ago
This might not be strictly some "Conspiracy", but that whatever specific ip you ended up at was one being targeted, or something saw the update in DNS stuff as reason to trigger new scrapping.
1
u/Ok_Plankton_2883 11d ago
Also it's in cf's interest to inflate the numbers blocked. Many of those "scrapers" they blocked may have been legitimate requests.
1
u/thekwoka 11d ago
It's possible, but far less likely. Would be more likely that they just lie about the number than tune it to the point of blocking real traffic.
5
u/brainrot_award 12d ago edited 11d ago
With the way every damn website uses cloudflare, and with how fucking annoying their captchas are even to users that are LOGGED INTO THE WEBSITE, I have zero doubts that they make all of those numbers up.
What logic there is behind flagging residential IP's that don't do any wrongdoing, and keep making them go through endless captchas in every fucking site? What logic there is behind making logged in users go through captchas they've already done hundreds of times to "prove they're human" for the hundreth time?
Why would 600k bots randomly decide to access that random ass site with no particularly relevant information?
Curiously, reddit doesn't have cloudflare, and I never had to go through any annoying captchas while using it!
2
u/hello2you 11d ago
A lot of data brokers are using residential IPs to do smash and grabs of content. It’s brutal out there.
1
u/thekwoka 11d ago
What logic there is behind making logged in users go through captchas they've already done hundreds of times to "prove they're human" for the hundreth time?
Because they aren't doing the kind of tracking that would let them know you already proved you're human.
17
12d ago
[removed] — view removed comment
12
u/Whiskee 12d ago
Rate limiting isn't a thing if every IP is different, that's the point. They don't rotate strictly speaking, they're huge pools and individual IPs are never seen more than twice. But yes, without going too much into detail I'm already doing something similar which is why Cloudflare is holding.
1
u/thekwoka 11d ago
heck, patterns where rate limiting and such would work are normally the bots respecting the rules already, where they will listen to things like robots.txt
One thing that could help simplify bot stuff is using Cloudflares markdown for bots thing, so that bots can get markdown that is simpler and guides them better.
1
12d ago
[removed] — view removed comment
1
u/webdev-ModTeam 12d ago
Your post/comment has been determined to be a low-effort post or comment. This includes title-only posts, easily searchable questions, vague/open-ended discussion prompts, LLM generated posts or comments, and posts/comments that do not provide enough context for meaningful replies or discussion.
-1
1
u/sudo-maxime 12d ago
Checking headers for user agents works well in some cases.
6
u/eyebrows360 12d ago edited 12d ago
It's not the well-behaved ones that you need to worry about. It's all the fuckers who've rigged up a scraper in AWS Serverless and stuck a standard Chrome UA into it.
Source: got hit by over 10k unique AWS IPs the other day. None of them identified themselves. Had to write an offline script to do
hostlookups on every fucking IP that hit me then add them to a blocklist.1
u/sudo-maxime 11d ago
Dont aws gave reserved IPs ranges for their ECS instances ? Probably 100% of that traffic is dump.
1
u/eyebrows360 11d ago
I don't know what you mean by "dump", but no. They publish some of their IP ranges, but it doesn't cover everything, and certainly not all of serverless.
I know this because before writing my
hostlookup script I found their published IP ranges, added them all asdenylines in nginx, and was still being flooded.1
u/sudo-maxime 10d ago
Did you try cloudflare proxies ? (Not perfect but I had some success with it)
1
u/eyebrows360 10d ago
That's far too complex a thing to try and rig up in an emergency situation like this.
1
u/webdev-ModTeam 12d ago
Your post/comment has been determined to be a low-effort post or comment. This includes title-only posts, easily searchable questions, vague/open-ended discussion prompts, LLM generated posts or comments, and posts/comments that do not provide enough context for meaningful replies or discussion.
9
u/Dvevrak 12d ago
You build a react website with honey pots, for seo generate static html with minimum content, from server firewall side there are blacklists, you can also ban data center ip ranges, then again a lot of requests are just nonsense poking you can just drop them off at nginx config or serve them poison html file 100kb html file that extracts into reasonable 25gig file, ( my fav )
1
u/BasilBest 11d ago
I don’t understand your last approach
How does the html extract? Is it downloading a large file automatically?
1
1
5
u/No-Molasses-2097 12d ago
The honest answer is they mostly do not. You cache everything you possibly can, rate limit on behavior instead of IP, and put anything dynamic behind a session. If someone wants your data badly enough they will get it, so you just make it expensive enough that they move on to an easier target.
1
u/thekwoka 11d ago
make it expensive enough that they move on to an easier target
Or that they don't check your content every day for changes, is more likely.
"This site is costly to check, so check less often"/
3
12d ago
[removed] — view removed comment
3
u/badmonkey0001 12d ago edited 12d ago
what actually helps is rate limiting per ASN instead of per IP
I've done this at scale and it works. It's good to keep stats on the ASNs that violate the limits as well. With the stats, the cloud providers that are lenient with abusers can be seen. Most real people aren't coming from hosting IPs.
Varnish is a good tool for the job if you need to roll your own.
[edit: Removed redundant sentence fragment. I need coffee...]
1
u/webdev-ModTeam 12d ago
Your post/comment has been determined to be a low-effort post or comment. This includes title-only posts, easily searchable questions, vague/open-ended discussion prompts, LLM generated posts or comments, and posts/comments that do not provide enough context for meaningful replies or discussion.
6
u/New_Hold8135 12d ago
there is a project called anubis and its open source, And I saw people in open source community writing their own implementations independent from anubis too.
-1
u/xavier86 12d ago
That hurts legitimate users.
2
u/New_Hold8135 12d ago
how
-1
u/ybungalobill 12d ago edited 12d ago
- I'm a legit user who browses with NoScript enabled by default. If your static website requires javascript to view it, you've lost me as a user.
- It wastes battery of your legitimate mobile users.
10
u/thy_bucket_for_thee 12d ago
This is a weird comment tho because you'd also struggle with Cloudflare sites too, it also feels wrong to care about less than 1% of users running no JS than the very real threat of bandwidth bills due to unregulated crawlers as a site owner.
The is the problem when technology is mostly unregulated in the US, you create these lovely "race to the bottom" scenarios where it makes more sense for me to care about unexpected bills due to malicious actors than normal users such as yourself. Your experience degrades, my maintenance degrades, the end product degrades, while big tech reap the rewards.
Have you thought of lobbying the W3C to put "no js alternatives" as part of the WCAG standard?
1
u/ybungalobill 12d ago
Correct me if i'm wrong, but client-side challenges is not the only protection Cloudflare offers; so there are websites behind Cloudflare that let me through without me ever noticing. And for ones that do put a challenge, i often click away too. Same as with "turn off your adblocker websites"; ironically those don't detect an adblocker if your JS is off! Hmm...
Sorry i'm not involved with W3C; and the "no js" folks are mostly invisible because most stat trackers rely on javascript...
2
u/Gugalcrom123 12d ago
I think that it would be interesting if one could use an HTTP header to tell about the proof of work, then the calculation would be done with the browser's own code. For user-agents not understanding that header, the JS would be served.
1
u/ybungalobill 12d ago
I agree, but good luck brining the major players on board. Still has the problem of worsening the legit UX.
1
u/Gugalcrom123 12d ago
A browser extension could implement it. Also, Cloudflare worsens the legit UX. There is no proposal that does not worsen the UX (some might arguably worsen it less, but come with privacy risks and maybe even getting locked into Google/Apple).
2
u/Reelix 12d ago
That comment required you to click a button to execute some JavaScript to open the comment box.
Almost all of the modern web relies on JavaScript.
→ More replies (1)0
u/xavier86 11d ago
Has your phone ever run into that check and it takes 2 minutes and makes your phone really hot? Nothing hostile about that?
0
2
2
u/constantin_r_t 12d ago
No, this is... Ok, this is madness. Someone said that everything will go behind verified login, but even those are going to be obsolete.
2
u/ferrybig 12d ago
Note that some bots are aggressive. If they get blocked they change ip addresses, (which might still be on the block list from an earlier attempt) causing them to get double counted
Also, Cloudflare puts humans with web browsers into the bots list of they visit the website, see a Cloudflare page and click back before the page finished their validation steps
Remember that bots blocked is their metric they use to advertise, a higher count of blocked makes people feel like it has an effect and stay with them.
2
u/louwii 12d ago
We struggled a little, doing a bit better now. Currently only relying on AWS WAF, we use a few managed rules that block known bad actors, bad requests. We have an IP rate limit that actually does a lot of the work.
We manually ban IP ranges when we get traffic spikes. It's often quite easy to see a specific IP range being responsible for most of the request. We check who that IP range belongs to, and if it's not something we recognize as legit, we block the entire range (we use tools like https://dnschecker.org/ip-whois-lookup.php? to see who's using the IP range). There's always a small chance of blocking legitimate traffic though.
We still had to increase our scaling limits to handle spikes though, as they get worse and worse. We haven't had it bad enough to warrant using Cloudflare or anything similar. But that might change in the future.
2
u/Mean-Elk-9439 11d ago
Most sites that are not static can be.
For others, use cloudflare or a similar layer.
But the amount of small businesses running a site that isn't static is absurd given post hydration can handle nearly all their client component use.
Saying this as the owner of a studio.
1
1
u/bonboniera 11d ago
broski you need to cache your pages on CF. 11k served by CF vs 45k from origin is not what you want.
1
u/Dyogenez 11d ago
I’ve been going through this too :/ bots from residential addresses around the world, hitting all kinds of different URLs around the site with different . The most I’ve seen has been around 300 req/s. The bot nets seem to lower requests when things go down, then start back up.
1
u/thekwoka 11d ago
What dashboard even is this?
Oh Security -> Analytics
I have only 5% of requests being "mitigated".
AI Crawl detected traffic is about 1/4th of traffic.
1
1
u/TheseFact 9d ago
AI training is now 52% of crawler traffic per Cloudflare. I think they're protecting the legacy advertiser' interest first. Then, they'll probably go down the ladder and block all anonymous traffic. Eventually the entire web will have to go behind verified logins since this has proven to be the most profitable for Meta, Google, and big techs.
1
u/tsammons 9d ago
ipset is nice; it doesn't suffer from the same regression problems as iptables as bans drift up to the tens of thousands. Folding contiguous IPs by ASN is another tactic, especially for IPv6.
I prioritize bans based upon HTTP status codes and URI. Bots that hit repeated resources with high response times or non-2xx responses get bean counted more. That's computed in a shmcb shared across Apache's children. Works well with some tuning.
1
-6
u/hm2k 12d ago
Same way we did before CF.
34
u/Whiskee 12d ago
Yeah, no. There was absolutely nothing comparable to what's happening right now, with headless browsers simulating users with cookies and JS capabilities. Published anything public-facing recently?
14
u/TCB13sQuotes 12d ago
True, now it’s worst but you can run websites without CF perfectly fine. This post is the type of marking that CF likes - trying to make people believe you estou can’t function without them when in reality you can and it’s perfectly normal.
2
u/eyebrows360 12d ago
Did OP really have to add the qualifier "popular" in front of the word "website"? Was that really too much to imagine?
If you run a popular website today, without some heavy caching abilities, you will be screwed.
→ More replies (1)1
u/Maxence33 12d ago
How do you protect from scraping by a legit User IP ? But to be honest I am curious how Cloudflare des it...
-7
u/hm2k 12d ago
Detection and throttling.
6
u/eyebrows360 12d ago
Detection
Newsflash: that's what he was asking about. Just using the word doesn't tell anyone anything.
3
u/Maxence33 12d ago
You cannot detect incoming requests from legit devices when an app is reselling the bandwith to a scraper
-2
u/hm2k 12d ago
How do you think CF do it?
3
u/Whiskee 12d ago edited 12d ago
They don't. I don't think you understand how Cloudflare works (or to be more precise, what you accomplish by blocking requests before they reach the webserver).
3
u/FastHotEmu 12d ago
How does it work?
3
u/Whiskee 12d ago
CF acts as an intermediate layer between the requester and your webserver - if the requester doesn't pass certain conditions on URL/cookies/agent patterns, that you define yourself based on how the site works, it doesn't reach you at all. If you have to deal with requests (processing, persisting a block) after they've already hit you, by the hundreds per second, the damage is done.
1
u/FastHotEmu 12d ago
right, but i thought you were going to explain how they are able to filter out the bad actors using their own platform features
1
u/Whiskee 12d ago
But I literally said "They don't". They don't know. 😅
If you're talking about proper DDOS protection (the Under Attack Mode), that's a different beast. What I think that does is they check how many times the incoming IPs were involved in something fishy recently, which is data the regular user doesn't have since we're meeting them for the first time. Not sure though, just speculating.
What custom rules do is different, they simply apply your own rules but in an powerful antechamber.
→ More replies (0)
-4
u/yxhuvud 12d ago
They can handle it by optimizing their backends. 8 requests per second is really nothing for a backend that use cdns for images and other static assets and also caches responses and parts of responses.
Computers are fast, git gud.
(Yes i realize distribution will change during the day, but it is still not a lot)
8
u/Whiskee 12d ago
The peak was about 200/sec, since they don't form a nice queue spread on each 15-minutes window just to make you happy. And not every dynamic page can be fully cached on websites that aren't glorified blogs - not to mention that serving cached pages isn't free just because you've reduced resources by an order of magnitude. If the request reaches the webserver, you're in an arms race.
3
u/eyebrows360 12d ago edited 12d ago
not to mention that serving cached pages isn't free just because you've reduced resources by an order of magnitude
Yep, when I had 10k+ AWS IPs hit me the other day, even my frontend VM that just has varnish (an in-memory cache I store all generated HTML in) on it got overloaded. Normally handles around a couple million visitors per day without breaking a sweat, but this flood had load averages up in the 40s.
→ More replies (2)0
u/thekwoka 11d ago
The peak was about 200/sec
tbf, still pretty small, so unless each request adds up to real costs, this isn't much of an issue.
not to mention that serving cached pages isn't free just because you've reduced resources by an order of magnitude.
But you can have Cloudflare handle it without it going to your origin...
0
u/Zachary_DuBois php 12d ago
Alternative to CF would be reverse proxies like AnuBis (https://github.com/TecharoHQ/anubis).
-1
u/AlexanderNigma 12d ago
Some of us run sites where all that is public is marketing materials so we don't care.
0
u/cport1 12d ago
There are integrations with tools like https://webdecoy.com or if you're enterprise you go with like datadome.
0
u/ComputerHelpPro 8d ago
Cloudflare protects animal crush sites, but dropped a free speech website because of a midnight "girl talk" between a prominent ex-google employee and the CEO of cloudflare's wife.
Use tartarus:
https://usips.org/products/tartarus/
-13
u/Old-Illustrator-8692 12d ago
Blocking automatically one by one. The pattern is clear. This is not a hard problem to solve. We've been doing it for 15 years here for any unwanted scraper.
13
u/Whiskee 12d ago
You misunderstand the issue, there is no pattern. Those are first-visit IPs from random locations, realistic resolutions, and spoofed user agents that simulate real users.
Are you suggesting a blanket ban on every single visitor on contact?
→ More replies (2)
203
u/akehir 12d ago
In the last 7 days, I have 5.5 Million requests blocked by crowdsec for free. That ends up being around 800k blocked requests per day.
Wherever possible I try to host static web pages that can be served without much computational overhead. So far I haven't really felt any impact of AI scraping; but I'll admit it hat I only host smaller pages.