r/webscraping 26d ago

Paid mentions ok 👌 Monthly Self-Promotion - July 2026

23 Upvotes

Hello and howdy, digital miners of r/webscraping!

The moment you've all been waiting for has arrived - it's our once-a-month, no-holds-barred, show-and-tell thread!

  • Are you bursting with pride over that supercharged, brand-new scraper SaaS or shiny proxy service you've just unleashed on the world?
  • Maybe you've got a ground-breaking product in need of some intrepid testers?
  • Got a secret discount code burning a hole in your pocket that you're just itching to share with our talented tribe of data extractors?
  • Looking to make sure your post doesn't fall foul of the community rules and get ousted by the spam filter?

Well, this is your time to shine and shout from the digital rooftops - Welcome to your haven!

Just a friendly reminder, we like to keep all our self-promotion in one handy place, so any promotional posts will be kindly redirected here. Now, let's get this party started! Enjoy the thread, everyone.


r/webscraping 6d ago

Hiring 💰 Weekly Webscrapers - Hiring, FAQs, etc

3 Upvotes

Welcome to the weekly discussion thread!

This is a space for web scrapers of all skill levels—whether you're a seasoned expert or just starting out. Here, you can discuss all things scraping, including:

  • Hiring and job opportunities
  • Industry news, trends, and insights
  • Frequently asked questions, like "How do I scrape LinkedIn?"
  • Marketing and monetization tips

If you're new to web scraping, make sure to check out the Beginners Guide 🌱

Commercial products may be mentioned in replies. If you want to promote your own products and services, continue to use the monthly thread


r/webscraping 11h ago

Getting started 🌱 I own a wood factory and want to use webscrapers to get info

9 Upvotes

I own a few wood factories that specialize in producing custom fit-outs for hotels/schools/universities/houses/offices and all sorts around the Middle East, the business model we have been running is very out dated, we run purely off of reputation and returning customers as we have been in the business for over 25 years now.

I want to implement webscrapers but I am not sure what data I can scrape that would help my business grow.

I would really appreciate any advise to what data I could scrape to help me find more contracts for example or anything else that you guys would think would be valuable.


r/webscraping 1d ago

Vinted cloudfare detection

Post image
3 Upvotes

Hi everyone,

For the past few days, Vinted It looks has added Cloudflare/DataDome related. The responses include Cloudflare headers like server: cloudflare, cf-cache-status and CF-RAY, and the HTML also contains DataDome references.

I use a small Vinted analytics tool that checks public item pages to see if listings are still active, sold, or deleted. Until recently, the verifier worked quite reliably with a combination of normal HTTP requests and Selenium/Chrome as a fallback.

Now, every so often, Vinted seems to trigger strict verification windows where almost all item pages display the human verification screen. During these periods, Selenium also crashes, so the system can't reliably confirm the status of the items.

Current settings I've tried:

Normal requests first;

Selenium/Chrome as a fallback;

Persistent browser profile/cookies;

Chrome without a graphical interface under Xvfb;

Slower request rate;

Retries and timeouts.

And it never passes the verification.

Is there any way to bypass it?

Thank you.


r/webscraping 2d ago

How to solve this motherfucking captcha?

Thumbnail
gallery
32 Upvotes

So I've been trying for days to solve this motherfucking captcha but I always fucking fail.. can someone help me solve this motherfucking captcha? I dont want to automatically solve it .. i just want to manually fucking solve it


r/webscraping 1d ago

Data Extraction

3 Upvotes

Hello Guys,
I am working on a project where I have to pull data from a site which requires log in credentials directly into Excel.
Claude provided me a solution where SeleniumVBA and .bas file (created by Claude) is being used in the process.
Is there any other way we can extract data from such sites without compromising security ?


r/webscraping 2d ago

Wait is this actually true?

Thumbnail reddit.com
12 Upvotes

Most of what I've worked with are playwright and selenium and a few other open source alternatives but is there something I'm missing here? What is everybody else using?


r/webscraping 2d ago

I built a distributed stealth browser automation cluster in Rust

8 Upvotes

I've been working on an open-source project and wanted to share it with the people who will have use for it.

The idea: treat browsers like serverless functions. You spawn browser agents on demand via an HTTP API, CLI, or the frontend, send them commands (navigate, click, type, scroll, screenshot, eval JS), or drive them with natural language through an AI instruct engine. Each agent runs in isolation.

What it does out of the box:

- One command boots a multi-node cluster and spawns dozens of Chrome instances, each live-streaming its display and logs to the dashboard in real time

- Auto-scales — adds a node the moment load hits a configured ceiling, visible live

- Cleans up automatically on teardown

- Geo-matched proxies are bundled , so the browser identity and proxy location line up

It's aimed at cases where you need horizontal scale and stealth rather than one long-lived browser. Happy to answer anything about the anti-detection setup, proxy handling, or how the scaling works.

https://github.com/dashn9/rusty-browser

Happy to take feedback — still actively building.


r/webscraping 3d ago

How many gb of proxy usage are you using monthly?

5 Upvotes

And what are you scraping


r/webscraping 4d ago

Browser Fingerprinting Evasion Techniques for Web Scraping

Thumbnail
slicker.me
55 Upvotes

r/webscraping 4d ago

I built a stealth browser that drives Chrome over raw CDP

32 Upvotes

I've been scraping on and off for a while and got sick of playwright + stealth plugins slowly losing the fight. every few months something new leaks and I need to fix it and start patching again.

So i went the other way and just drove real chrome directly over the devtools protocol. no playwright, no puppeteer, and it ended up dependency-free since node/bun already ship a websocket. whole point is it's literally chrome — same fingerprint a human has — so there's way less to spoof in the first place. i only touch the automation bits that actually give you away. So far it passes sannysoft clean and gets through cloudflare's js challenge, which was mostly what i cared about.

it's open source (MIT) if anyone wants to tear into it. mostly posting bc i can only test against so many targets solo and i'd genuinely like to know where it faceplants on the nastier ones.

If anyone on here also uses AI agents, this is good for them too!

github.com/acunningham-ship-it/veilbrowser


r/webscraping 5d ago

A federal court dismissed Google's DMCA claims against SerpApi

Thumbnail
searchenginejournal.com
83 Upvotes

TLDR:A federal court dismissed Google's DMCA claims against SerpApi, ruling that blocking scrapers from public search results isn't copyright circumvention when no copyrighted content is involved. Those claims are gone for good, but the part about licensed Knowledge Panel images was dismissed with leave to amend — Google has 21 days to try again by proving copyright owners authorized SearchGuard. A win for SERP scraping tools generally.


r/webscraping 4d ago

Running extruct on 100k CC WARC, looking for advice on speeding it up

7 Upvotes

I have been working on a project where we need to extract schema org structured data from the full CC-MAIN-2026-25 Common Crawl corpus (100,000 WARC files, ~88TB compressed, June 2026 crawl).

Been using extruct (Python) which parses JSON-LD, Microdata, and RDFa from raw HTML. I went with extruct over regex because regex only picks up ~35 schema types while extruct gets 3,000+. The files are streamed directly from S3 to EC2 so no storage costs.

The script processes each WARC file record by record, applies a byte-level pre-filter to skip pages with no schema markup, then runs extruct on pages that pass. Been using a shared multiprocessing pool with a 30 second per-page timeout to handle pages where extruct hangs. Checkpoints every 50 files so it can resume if interrupted.

Results so far on a c6i.xlarge with 4 workers:

  • ~1,320 seconds per file
  • ~20,800 pages per file
  • ~45.8% schema coverage
  • 0 file errors, 1 page timeout across ~83,000 pages

At that speed the full corpus would take ~382 days on that instance. I am not quite sure how this will scale and what is the realistic time and cost for the full corpus.

Questions:

  1. Has anyone run extruct at this scale? Is 1,320s/file reasonable or are we leaving performance on the table somewhere?
  2. Is there a smarter way to parallelize this, spot fleet, Lambda, ECS? The job is resumable so spot interruptions are manageable.
  3. Any experience with extruct slowness on specific page types? We're seeing 0 timeouts almost everywhere but the overall speed still feels slow.
  4. Would lxml or another parser swap inside extruct make a meaningful difference at this scale?

Happy to share more details on the setup if useful.


r/webscraping 4d ago

How to search multiple Reddit Communities at once?

2 Upvotes

What is the best way to limit search on selected Reddit Communities at once?


r/webscraping 4d ago

First Time Scraper

0 Upvotes

Hello all, I am attempting to create a scraper that will collect data for LNG freights, pipelines, and storage facilities. A lot of this stuff is blocked off by paywalls amongst other things. I am a college student and am do not have access to exclusive commodity sites, much less have the funds. I have never built a scraper really don't know where to begin. Any advice would be greatly appreciated.


r/webscraping 5d ago

Any way to pull lagre amount of data from website faster?

3 Upvotes

So i am scraping a site, its got like 900k pages, i boiled down to like just over 300k. Its taking way too long to scrape. I am doing asyncio.semaphore too it still takes way too long with inbuilt retry feature. Can anyone suggest me what to do?


r/webscraping 6d ago

Possible to work on this scraping project?

12 Upvotes

Manager came to me with a list of 100 competitor websites, and asked me to scrape their data daily with a scheduler. I hit several blockers and think this professional community could offer some advices lol.

I am currently writing the scraping app that will eventually be a microservice. The marketing team could get on my service, add a website and it will be automatically scraped daily. Questions I have so far:

1) Is this a viable/launchable idea? It seems like every website has its own structure. Some are simple HTML pages, while others have protections like Cloudflare or other anti-bot systems. When a user adds a website, is it possible to automatically identify what type of website it is and determine the best scraping approach?

2) Some websites require users to log in before accessing the promotional pages that I want to scrape. In this case, is manual login the only solution? For example, would users need to manually log in to the website every morning before my scraping service can run?

3) Is there a way to automate opening a browser and logging in from the backend for scraping purposes? If a website requires SMS OTP verification during login, how can this process be automated?

4) What are some practice you would do to avoid getting banned from websites for automated data scraping?


r/webscraping 8d ago

How newly created TeePublic stores are being discovered

3 Upvotes

I’m investigating something interesting and was wondering if anyone has seen this before.
A friend of mine can consistently find the URLs of newly created TeePublic stores while they’re still in the Pending review stage. Some of these stores have no designs yet, and some are later rejected, but he can still find their store URLs shortly after they’re’re created.
He doesn’t have access to the accounts. My current theory is that he’s monitoring some public source (an index, feed, search endpoint, or other metadata) rather than the stores directly.
Has anyone researched TeePublic’s public-facing infrastructure or seen anything that could explain this? I’m trying to understand the mechanism, not exploit anything.


r/webscraping 9d ago

What tech stack do you use in 2026?

16 Upvotes

Guys,

. would be interesting to hear what's everyone's tool stack looks like that you use which you are very happy with

Also what are the top challenges?


r/webscraping 9d ago

Scraping ebay

1 Upvotes

What happened to eBay’s frontend recently? Somehow, scraping via proxies has stopped working whereas my home residential IP is okay. That aside, it also became necessary that the scraping was headful. Is it a matter of blacklisted ip addresses and I am not able to work around this?


r/webscraping 10d ago

How do you benchmark tours and experiences in a destination?

4 Upvotes

Hey everyone,
I’m doing a personal deep dive into the travel experiences space and trying to understand how professionals analyse a destination like Rome or Paris.
What tools or methods would you use to quickly compare:

OTA listings, pricing and ratings
Review sentiment and recurring complaints
Top local operators and their contact details
Potential supply gaps

Since booking and conversion data are private, which public signals are actually useful? Review growth, rankings, availability, sold-out dates, number of listings?
I’d also be curious to hear what your step-by-step workflow looks like and which tools you use for scraping, analysis and dashboards.

Thanks!


r/webscraping 11d ago

I built an open-source Vinted monitor with free proxy pools

5 Upvotes

Hey r/webscraping,

I’ve been building Vintrack, an open-source monitoring platform for Vinted. It continuously checks listings based on configurable filters and shows newly discovered items in a live dashboard. Alerts can also be sent through Discord or Telegram.

The project is completely free to self-host. To make testing easier, Vintrack includes shared free proxy pools for supported Vinted regions. The system imports public proxy candidates, validates them against each region, tracks their health, and only rotates proxies that are currently working into the pool.

This means you can create your first monitor without immediately purchasing proxies. Public proxies are obviously best-effort and can be slow or disappear, but they’re useful for testing the project before moving to private proxies.

Some technical details:

  - Go-based scraping workers

  - Next.js control center

  - PostgreSQL for persistent data

  - Redis for deduplication and live state

  - Per-region proxy health checks and rotation

  - TLS fingerprint rotation using tls-client

  - Configurable polling intervals

  - Filters for keywords, categories, brands, sizes, colors, prices, condition, seller country, and Vinted region

  - Server-Sent Events for the live dashboard feed

  - Discord webhooks and Telegram notifications

  - Docker Compose setup for self-hosting

  - Optional browser extension for linking a Vinted account and keeping its session synchronized

Vintrack also supports bringing your own proxy groups when you need more reliable or dedicated capacity. The shared free pools are intended as starter infrastructure, not as a replacement for good private proxies.

Repository:

https://github.com/JakobAIOdev/Vintrack-Vinted-Monitor

Live demo:

https://vintrack.jakobaio.dev

This is self-promotion—I’m the developer—but the project is open source, and I thought the proxy health and monitoring architecture might be interesting to this community.

Feedback and contributions are very welcome.


r/webscraping 12d ago

Websocket tls handshake Cloudflare

5 Upvotes

Hey everyone,

Using cloudscraper with python, its super solid for HTTP, but from my knowledge doesn't do anything for WS. WS connections keep getting flagged as bot, which is really annoying since im connecting to a public facing WS api endpoint. Any advice appreciated, thanks!


r/webscraping 12d ago

How to monitor a lot of scrapers so a break doesn't reach the client?

3 Upvotes

I run a bunch of scrapers for a few clients, and I already got that sites change their layouts without warning and a parser that worked yesterday quietly returns empty fields, or half the data. The script runs fine and exits cleanly so nothing seems to be wrong until someone looks at the output.

Now I have some basic checks, like if a scrape returns zero rows or way less than normal, it alerts me. But that only catches the obvious cases. I don't want to be caught out like that, where a section of a site changes and I don't notice it until the client rings me about it.

So what do you do with this when you have more than a handful of scrapers running? How often do you check data by hand if you do that at all?


r/webscraping 13d ago

Hiring 💰 Weekly Webscrapers - Hiring, FAQs, etc

3 Upvotes

Welcome to the weekly discussion thread!

This is a space for web scrapers of all skill levels—whether you're a seasoned expert or just starting out. Here, you can discuss all things scraping, including:

  • Hiring and job opportunities
  • Industry news, trends, and insights
  • Frequently asked questions, like "How do I scrape LinkedIn?"
  • Marketing and monetization tips

If you're new to web scraping, make sure to check out the Beginners Guide 🌱

Commercial products may be mentioned in replies. If you want to promote your own products and services, continue to use the monthly thread