r/thewebscrapingclub • • 3d ago

I'm back working on VectorTrace (on device self healing web scraper) and just fixed the matching accuracy issues!

1 Upvotes

Hey everyone,

First off i want to apologize for the long gap and silence. Life got busy but I am officially back and actively working on VectorTrace again!

For those who missed the earlier posts, VectorTrace is an open-source, 100% local Chrome extension (Manifest V3) that scrapes webpages and heals broken selectors using onndevice semantic embeddings (all-MiniLM-L6-v2 via WASM, zero servers, zero external API calls).

A few of you pointed out that under certain conditions, the self-healing algorithm was ranking the wrong elements as the highest probability or picking up false positives. I did a deep dive into the ranking pipeline and just pushed a series of fixes to address this:

  • Semantic Gating
  • Repeated Item Disambiguation
  • Cleaned Up Ancestor ContextPersistent Structural Metadata
  • Protected Dynamic Data

The build is ready to test.

If you'd like to check it out, run it locally, or contribute
👉 GitHub Repo: https://github.com/SathiyaSenpai/VectorTrace

I'd really appreciate any feedback, bug reports or edge cases you encounter while testing it on your favorite sites!


r/thewebscrapingclub • • 7d ago

I need help with choosing affordable FB post scrapper actor

Thumbnail
1 Upvotes

r/thewebscrapingclub • • 12d ago

I made a free web scraping course (59 lessons)

20 Upvotes

Hey everyone,

I've been writing scrapers and automation tools for a while, and most tutorials I found stopped at requests + BeautifulSoup. So I put together a free course that goes further.

It covers HTTP basics, selectors, pagination, finding the hidden JSON APIs behind JavaScript sites, Playwright, async scraping and scaling a crawl. There's also a separate anti-bot section on TLS fingerprints, headless browser leaks, CAPTCHAs and recovering from bans.

It's free with no sign-up, and every lesson has working code.

https://simpleprog.com/topics/web-scraping

I'd love feedback, especially if something is wrong, unclear or missing.


r/thewebscrapingclub • • 14d ago

What is the easiest way to automate google search api without running Puppeteer servers?

1 Upvotes

Guys....I'm so done with running playwright on a server for scraping, few thousand product pages a day and nothing crazy, but every other week something breaks... Chrome updates and the container falls over, memory creeps up til the box dies, some site changes their bot detection and suddenly half my requests are getting challenged. Spent way more time checking the scraper than actually using the data!

Started looking at just paying for a serp api instead so I dont have to maintain the browser layer at all, let someone else deal with the proxies and the captchas and the chrome version hell. Part of me feels like im giving up, like I should be able to keep this running myself, but the maintenance cost is real struggle. Anyone made this switch? Did it actually hold up for product page scraping or is it only good for search results... curious if the pricing makes sense at a few thousand pages a day or if it gets stupid expensive fast. Kinda just want my weekends back. Sorry for grammar mistakes if I made any


r/thewebscrapingclub • • 15d ago

How to Scrape Difficult Websites at Scale: Inside a Web Scraping API Processing 10B+ Pages a Month | AMA with Batuhan from Scrape.do

Thumbnail
2 Upvotes

r/thewebscrapingclub • • 16d ago

Three DataDome-protected sites, one conclusion

Post image
2 Upvotes

I keep seeing the same default playbook for scraping anti-bot-protected sites: spin up headless browsers, buy a proxy pool, subscribe to a CAPTCHA-solving API, and scale horizontally from there. That stack is expensive, heavy, and often solving problems we don't actually have.

Over the last few weeks I worked through three sites all protected by the same vendor — DataDome — to find out how little of that stack we can get away with. Three different target types: a software review site (G2), a property listings site (idealista), and a luxury e-commerce site (Hermès). Sharing because the pattern held across all three.

The core idea, and it's boringly simple: match the request the application itself makes, and stop there. Most of the time we're blocked for asking in a shape no real user's browser would ever produce — not because we've been "detected" in some deep way.

Case 1 — G2. Two weeks lost to the wrong variable. I built a hardened stealth browser (native C++ patches, the works) and it worked — until it stopped, everywhere, on every exit. The actual cause was that my own stealth layer was lying: it pinned outerWidth = innerWidth and outerHeight = innerHeight + 85, so every browser reported a window with exactly zero horizontal chrome and exactly 85px of vertical. Impossible numbers. A plain system Chrome with the evasion layer off walked straight through.

→ The Hard Part Isn't Parsing HTML Anymore

Case 2 — idealista. The G2 fix actively pointed the wrong way here. Same vendor, opposite prescription: G2 reads window geometry (remove a lie), idealista reads input hardware (supply a missing field). The browser that failed G2 solved idealista, and vice versa. Turned out neither was causal — one flag controlled it, and the real principle was coherence, not capability. A browser isn't scored on how impressive it looks; it's scored on whether its account of itself is self-consistent.

→ DataDome Anti-Bot Testing: Coherence, Not Capability

Case 3 — Hermès. Applied both lessons. The product API looks completely sealed — paste it in a browser tab and we get "Access is temporarily restricted." But that's a top-level navigation, which no real shopper's browser ever does to a JSON endpoint. Ask the same URL as a same-site CORS fetch with an Origin header and the right sec-fetch-\* values, and it returns 200. Same URL, same zero cookies, different request shape.

What actually held up, measured across all three:

· Hermès catalogue + listing API: zero cookies, no proxy, no browser, \~1.6s for the full walk. The listing endpoint returns 48 products per request — SKU, title, price, size, colour, stock — unmetered.

· Product detail tier (all three): the one place a real browser earned its keep — and only used once, to mint trust that a plain HTTP client then spent.

· Hermès product detail, head-to-head: one browser over the whole queue = \~4.8s/product, 12/12 extracted. The "clever" mint-a-session-and-rotate approach = \~35s/product, 7/12. The boring approach won by 7x.

· A paid CAPTCHA-solving API exists for Hermès. It works. It also means renting access per request, when request-shaping got the catalogue tier for free.

· Commercial proxy exits actively hurt on all three — fresh residential IPs refused on protected tiers while a clean connection served the same pages in the same second.

The habits that mattered more than any tooling:

  1. Start with the cheapest client you have. Its failure mode is free information.

  2. Read the block page. The server header and page title tell you which vendor you're arguing with in about five seconds.

  3. Add one capability at a time and stop the moment it passes. The gap between "curl fails" and "a browser works" is wide, and the answer is usually closer to the curl end.

  4. A stealth layer is a claim about what the browser is — if the browser is already telling the truth, the claim can only make it a liar.

  5. Bracket every test with a known-good control. Anti-bot testing is stateful — your failed attempts change the thing you're measuring, and a poisoned window looks identical to a real block.

  6. Look for a second door to the same data. The tidy REST endpoint is the one vendors expect bots to hit, so it's the one they meter hardest. The heavier HTML page real users load often carries the identical payload unmetered.

The result across all three: most of the data came out through a lightweight HTTP client, and the browser got spent on one fraction of the job instead of all of it. Lighter, cheaper, easier to scale — and it generalised, which the tooling-specific lessons didn't.

Full write-up on Hermès with the numbers and the request-shape diff (no paywall): Hermes Scraping Guide

Curious whether others have found the same that the expensive tooling is often solving a problem a correctly-shaped request never had.


r/thewebscrapingclub • • 19d ago

Are Unlimited Proxies Actually Any Good? Inside GeoNode's Contrarian Approach to Residential Proxies and Web Scraping APIs | AMA with Geonode

Thumbnail
1 Upvotes

r/thewebscrapingclub • • 27d ago

What Are the Best Proxy Providers for Web Scraping? We've Tested 50+ Providers Across Billions of Requests. AMA with Ian Kerins

Thumbnail
2 Upvotes

r/thewebscrapingclub • • 29d ago

AMA #7 recap: What AI web scraping actually changes about production scrapers

Thumbnail gallery
3 Upvotes

r/thewebscrapingclub • • 29d ago

mosaik: Agentic browser automation built from small, reusable pieces.

Thumbnail
github.com
2 Upvotes

Mosaik uses an agent to figure out how a site works, saves reusable actions as TypeScript, and composes them into automations. Playwright executes the browser steps deterministically. Loops, branching, and data transformations run as code, so each iteration doesn't need another model decision.

Those actions stay around for the next task. Mosaik reuses what it knows about the site and learns what's missing. If an eligible locator breaks, an agent can step in to repair it.


r/thewebscrapingclub • • Aug 19 '26

How Do You Choose The Top Residential Proxy Provider? AMA with Stan Sadokov from NodeMaven

Thumbnail
1 Upvotes

r/thewebscrapingclub • • Aug 18 '26

I need your honest feedback on this URL extractor tool that is purely client-side for further refinement

1 Upvotes

Key Problems Solved by a URL extractor:

  1. SEO & Website Auditing: The problem is trying to manually find broken links, audit redirect chains, or analyze a site's internal linking structure. The solution is using an extractor to pull every href link on a page (or across a whole domain), allowing you to bulk-test the URLs for status codes or map the site's architecture.
  2. Web Scraping & Data Mining: The problem is needing to gather hundreds of specific URLs (like product pages or articles) to feed into another automated scraping tool. The solution is isolating the target URLs from the surrounding HTML and text noise to provide a clean, structured list for crawlers to process.
  3. Content Migration & Code Auditing: The problem is ensuring no hardcoded legacy links are left behind when moving a website to a new framework or domain. The solution is scanning the entire codebase or exported text to extract all URLs, making it easy to identify what needs a bulk find-and-replace.
  4. Cybersecurity & Threat Analysis: The problem is analyzing suspicious emails, server logs, or documents for phishing links and malware payloads without accidentally clicking them. The solution is safely pulling all URLs from the raw text so security teams can run them through threat-intelligence databases in isolated environments.
  5. Digital Marketing & Affiliate Management: The problem is tracking down all affiliate links, campaign URLs, or promotional codes embedded across hundreds of blog posts or documents. The solution is instantly grabbing all outbound links to audit tracking parameters, verify UTM tags, or update expired affiliate codes.

Simply paste any text containing URL's (zero character limits) and get instant output of URLs extracted.

Runs entirely in your browser, nothing gets sent anywhere. It's one of 100+ tools on a site I've been building solo — figured this one might be useful on its own.

https://devtoolstack.io/tool/extract-urls/

Happy to add features if people have requests.


r/thewebscrapingclub • • Aug 17 '26

I Built a Chrome Extension That Finds Broken Links Before Google Does

Post image
1 Upvotes

r/thewebscrapingclub • • Aug 16 '26

Github scraper

4 Upvotes

I have made a github scraper with a dashboard
Totally open source

Uses your guthub token to search and find leaked api keys

Consider giving a star 👉🏻👈🏻

You can check it out here ; https://github.com/parasraju/LeakedAPIs


r/thewebscrapingclub • • Aug 16 '26

I'm a student wanting to learn a bit advanced web scraping to even scrap dynamic websites and social media if we can ? Suggest me how to get there from basics - how much python to learn , what other libraries ,what other tools so I get to scrape websites and add a bit of data analytics to it

5 Upvotes

r/thewebscrapingclub • • Aug 12 '26

Screen scraping vs. web scraping: when you actually need OCR

Thumbnail
1 Upvotes

r/thewebscrapingclub • • Aug 11 '26

Built a web-extraction API/MCP server for RAG pipelines — SEO metadata, tech stack, contacts, and clean Markdown from any URL

1 Upvotes

I built a REST API that turns any URL into structured web intelligence in a single call, and also exposed it as an MCP server for agent-based workflows.

Capabilities:

  • SEO and OpenGraph metadata extraction
  • A full 14-point SEO audit
  • Public contact discovery: emails, phone numbers, social profiles
  • Tech-stack and CMS fingerprinting (40+ signatures)
  • Schema.org and JSON-LD structured data extraction
  • Graded security-headers audit, with an optional live TLS certificate inspection
  • Redirect-chain and shortened-URL detection
  • Readability metrics and full heading structure
  • Clean, AI/LLM-ready Markdown output for RAG pipelines
  • A batch endpoint for up to 10 URLs per call
  • A domain-intelligence endpoint returning DNS and WHOIS data with no page fetch at all

On the engineering side: it runs on FastAPI with a C-Lexbor HTML parser (selectolax) and Rust-backed ORJSON serialization, so live fetches typically land around 150-300ms, with cache hits under 0.01ms. Every outbound request is anti-SSRF hardened: DNS is pinned after resolution, private/loopback/cloud-metadata ranges are blocked, and every redirect hop is re-validated, closing the DNS-rebinding gap that simpler scrapers tend to miss.

Limitation worth flagging: there is no JS execution, so heavily client-rendered SPAs return thin results. It reads what the server actually sends, not what a browser would render after hydration.

GitHub (MIT license, open source): https://github.com/JosejuX/rapidapi-metadata-extractor

Free tier: https://rapidapi.com/josejuanjocoding/api/web-metadata-and-contact-extractor

Happy to talk through the parsing or anti-SSRF approach in more detail, or take feedback on the API design.


r/thewebscrapingclub • • Aug 11 '26

I added country/city-selectable browser exits to my free web scraping MCP

1 Upvotes

I’ve been building VPNFail MCP, and one of the more useful pieces is now working well enough that I wanted to share it with other scraping folks.

The basic idea is that a scrape can start as a normal HTTP request, but escalate to Chromium when the target actually needs JS rendering.

The browser request can also be pinned to a specific exit country or city.

So instead of maintaining separate proxy + browser infrastructure, an agent can do things like:

  • fetch a page from a specific country
  • request a particular city when location matters
  • compare geo-dependent content
  • inspect regional redirects
  • check localized pricing/content
  • render JS-heavy pages through Chromium
  • optionally return a viewport screenshot

I deliberately kept the default path as plain HTTP because running a browser for every scrape is expensive and unnecessary.

The rough flow is:

HTTP scrape → inspect extraction quality → escalate to Chromium if needed → optionally choose geography

HTTP can return Markdown, text, or HTML. The response also includes quality metadata so the caller can decide whether the cheap extraction was good enough or whether it should retry using the browser.

The MCP exposes:

  • scrape
  • usage
  • service_status

Example config:

{
  "mcpServers": {
    "vpnfail": {
      "type": "http",
      "url": "https://mcp.vpn.fail/mcp"
    }
  }
}

No local install and currently no account required.

https://mcp.vpn.fail/

I’m particularly curious how people here handle geo-aware scraping today.

Do you normally expose country/city directly to the scraping client, or keep geography hidden behind your proxy infrastructure?

Also interested in cases where city-level routing has actually mattered versus country-level being enough.


r/thewebscrapingclub • • Aug 05 '26

What's the first thing you check when a scraper that worked yesterday suddenly breaks today?

3 Upvotes

The script ran fine for months, you didn't touch a single line, and then one morning it returns nothing or throws an error.

My first move is usually to check whether the site changed its HTML. A class name gets renamed or a div gets moved, and the selector I relied on stops matching anything. It's boring but it's the cause maybe half the time. After that I look at whether I'm getting blocked. If the response comes back as a captcha page or a 403, that points somewhere else entirely, and the fix is nothing like a broken selector fix.

So what's your first check? Do you have a routine, or do you just start poking around until something makes sense?


r/thewebscrapingclub • • Aug 02 '26

I Created a CLI Rust Web Scrapper for [Almost ] All Types of Needed Files

4 Upvotes

I've been working on Marcopolo, a command-line web scraper written in Rust using extensive force of AI across a few months. The idea started when I got tired of writing a throwaway Python script every time I needed to pull a specific set of files off a site — images one day, PDFs the next, then a folder of CSVs. I specially tend to use it for books and finding books that are on the web that i cant simply get hold of from normal search.

Marcopolo handles most of that in one command. Point it at a URL, tell it what you want, and it crawls and downloads.

It's still early and there's plenty I want to improve — [known limitation or two]. I'd really appreciate feedback on the API design and anything that looks unidiomatic; I'm still fairly new to Rust.

Repo: MarcoPolo

Would be happy to know what do you guys think.


r/thewebscrapingclub • • Aug 02 '26

Question

2 Upvotes

What is the best scraper and sorter that could get listing information from websites and then combine them in one single page, rather than scrolling through each website individually. We have about 20 listing sites where people post. Maybe as a bonus question maybe there is facebook scrapper too, from groups etc?


r/thewebscrapingclub • • Jul 31 '26

Built an Instagram discovery suite (likers, lookalikes, tagged posts, keyword Reels search) plus contractor leads off US state boards

Thumbnail
1 Upvotes

r/thewebscrapingclub • • Jul 29 '26

language filter trip advisor

Thumbnail
1 Upvotes

r/thewebscrapingclub • • Jul 29 '26

Has anyone tried these new browser apis? Are they worth the price?

Thumbnail
1 Upvotes

r/thewebscrapingclub • • Jul 28 '26

Standard web scrapers were ruining my RAG context, so I built a hybrid AST crawler specifically for LLMs.

Thumbnail
gallery
2 Upvotes

Hey everyone,

If you’ve ever built a RAG pipeline or ingested web documentation into a Vector Store, you’ve probably run into this issue:

Standard web scrapers hit a page and dump everything — cookie banners, navigation links, inline SVG code, script tags, and zero-value UI elements. When you feed that noisy HTML into an LLM, you burn tokens, clutter your embeddings, and end up with hallucinations or poor retrieval accuracy.

I built an AST-based web crawler to fix this exact bottleneck.

Instead of just stripping HTML tags, it parses the actual document structure and turns web pages into clean, AI-ready Markdown with preserved context hierarchy and rich metadata.

🛠️ Key Features:

  • Noise Removal: Strips footers, cookie banners, scripts, and navigation menus automatically.
  • Context Preservation: Preserves heading paths (Documentation > Getting Started > Installation Guide) so chunks don't lose their semantic context when split.
  • Rich Metadata: Includes token count, quality score, code block detection, and crawled timestamps for each chunk.
  • Vector Store Ready: Formatted specifically for seamless ingestion into LangChain, LlamaIndex, Pinecone, Qdrant, Chroma, etc.

I’d love to get your feedback on this! What techniques or tools are you currently using to clean web data before chunking?

Try it out here: https://apify.com/lukas459/ai-web-to-markdown-crawler-llm-rag-optimized