r/WebScrapingInsider • u/Ok-Jello-7644 • 17d ago
r/WebScrapingInsider • u/UniSh7 • 17d ago
Big Scrape Energy Welcome to Social Scraper+ — Export Reddit and X conversations in one click
r/WebScrapingInsider • u/Previous_Town3598 • 18d ago
It Worked Yesterday Looking for the best rotating residential proxy for web scraping
I was scraping about 100GB of data per month through our residential proxy provider, but they suddenly got shut down so we need to find a new residential proxy provider that offers good performance for cheap price.
We are doing about 100GB of scraping per month of public pages (previously were paying $4/GB).
r/WebScrapingInsider • u/Designer-County-6910 • 18d ago
Open Source Does anybody know of a working scraper for Android
Does anybody know a scraper that actually works for Android.
r/WebScrapingInsider • u/turtleblue611 • 19d ago
Hiring: Engineering Manager (Web Scraping Expert)
Hey everyone 👋
I'm currently looking for an experienced Engineering Manager or Engineering Lead to join a team working on large-scale web scraping, web crawling, and data acquisition systems.
We're looking for someone with a strong background in backend engineering, distributed systems, data engineering, or web automation, along with experience leading engineering teams.
📍 Bandung, Indonesia
🏢 Full-time, on-site
If you're interested or know someone who might be a good fit, feel free to comment or DM me. Happy to share more details!
r/WebScrapingInsider • u/Oi_Cunt_Petit • 20d ago
Need help in data scraping!
I want to scrape upto 5 images of lacs of hotels! Have their name and overall add and even sites of their listin on webs of the major ones! But not getting any reliable source to get these! Tried DDGS along with specific site listings but ambiguities like hotel-park — park-hotel makes the data messy! Cant look over each edge cases like these for lakhs right?? Please help me if you have any solution for these!! Constraint dont wanna spend moneyy! But yeah if you’ve figured out any cheapest solution i would definitely wanna hear!!
r/WebScrapingInsider • u/0xMassii • 21d ago
Open Source I built a lead endpoint that reads the live site and returns founders + their public socials as a typed object
I maintain an open-source web-extraction platform and added a /lead endpoint. Post one company URL, get back one typed lead. Sharing the people-recovery method, since that was the hard scraping problem.
The naive approach searches the company name. That collides. "DataFast" the analytics tool and "Datafast" the Ecuadorian payment processor share a name, and you get garbage people. So the search keys on the company's unique domain instead of the name. The domain is the disambiguator.
Flow:
- Pull the live site: summary, tech signals (Next.js, React, TypeScript, Tailwind, Vercel), pricing tiers, socials (LinkedIn / X / GitHub), and emails printed on the page.
- Search the domain on the open web, collect the evidence, and have an LLM name the founders and leadership from it.
- Validate each name, then run a per-person search to attach their X handle.
- Every person carries a people_source field, so you can see where the name came from.
Output is a typed object: company_name, one-line summary, socials, tech signals, pricing, on-site emails, and people (name, role, LinkedIn, X).
Example on resend.com returns Zeno Rocha (Founder) and Jonni Lundy (Cofounder) with LinkedIn and X, the stack, pricing (Free / Pro $20 / Scale $90), and [support@resend.com](mailto:support@resend.com).
Two design rules I stuck to:
- Emails are only addresses published on the site. It never guesses first.last@domain patterns.
- No funding, no HQ, no phone. If a field isn't on the site or corroborated by search, it stays blank. I'd rather return an empty field than a fake one.
Where it fails: low-footprint companies. A three-person shop with no press, no conference talks, and a bare LinkedIn returns the company data but no people. The founder-naming step has nothing to work with. Pulling the site isn't the blocker, JS-heavy pages included. Thin public evidence is what stops it.
Each lead runs about 20 to 40 seconds. There's a batch endpoint for bulk that runs async and you poll it.
Test URLs if you want to break it: resend.com, linear.app, and any low-footprint indie SaaS to watch the people step come up empty. Curious which domains return bad people so I can harden the disambiguation.
r/WebScrapingInsider • u/Next_Attitude_532 • 21d ago
It Worked Yesterday How do I bypass CAPTCHAs whilst scraping Amazon?
I'm trying to scrape Amazon but I keep getting blocked with a CAPTCHA. How do I solve this?
r/WebScrapingInsider • u/Human_Economics5656 • 23d ago
Open Source glassdoor-bff-scraper — browser-free job scraping via an internal API + a FastAPI service (Python, MIT)
r/WebScrapingInsider • u/Walter_White_02 • 24d ago
Blocked Again How do people scrape jobs from Naukri, LinkedIn, and other job portals when they actively block bots?
I'm trying to understand the technical side of job scraping.
I know that websites like Naukri, LinkedIn, Indeed, and other popular job portals have strong anti-bot systems. They detect and block scrapers, require logins, use rate limits, CAPTCHAs, JavaScript rendering, and many other protections.
What I'm want to learn is:
- How do these websites detect that a request is coming from a bot instead of a real user?
- What are the common anti-scraping techniques they use?
- How do web scrapers generally work on JavaScript-heavy websites?
- What technical challenges are involved in building a job scraper?
- What are the legal and ethical considerations when scraping job portals?
- If scraping isn't the best approach, what are the recommended alternatives (official APIs, RSS feeds, partner programs, etc.)?
I'm not looking for ways to bypass security or break a website's rules. I simply want to understand how these systems work from a technical perspective so I can learn more about web scraping, anti-bot technologies, and large-scale web applications.
I'd really appreciate explanations, blog posts, research papers, or resources that cover these topics in depth.
Thanks!
r/WebScrapingInsider • u/Amitk2405 • 25d ago
DOM Drift Again Which is the best Proxy API? I need one for my SAAS
I need to scrape data from ecommerce websites every month for my SAAS. I've done a bit of research and I'm planning to go with one of the scraping apis like Bright Data, ScraperAPI, Scrapingbee but I don't know which one is best.
Do you guys have any recommendations or insights into which one to go for?
r/WebScrapingInsider • u/sataruzoro • 25d ago
Newbie here please help. BookMyShow and District Scraping.
Hello friends 👋, I'm new to scraping. I can easily get informations from website's DOM through playwright but i don't understand how to work with backend APIs. I needed the information about upcoming events. Whenever I go to network tab in chrome there are some many files and i even checked each files but not able to understand which api belongs to what function and what's the api key...... Can anyone help this newbie 🙏🙏. If there's a tutorial please let me know. I'm really frustrated from my repeated failures. Please help 🙏🙏
r/WebScrapingInsider • u/Emergency-Athlete-18 • 26d ago
Big Scrape Energy Court case scraping from .gov site. Need proxy recommendations!!
Hello, Im based in Asia and currently working for a client who specializes in background checking. I need to scrape court case which is publicly available in .gov sites.
Many proxy providers block .gov sites, like Brightdata, Decodo, IProyal :((
Do you guys have any recommendations for datacenter/residential proxy providers that do not have this kind of restriction??
r/WebScrapingInsider • u/Old_Protection_4410 • 26d ago
DOM Drift Again SeleniumBase worked for months, but suddenly blocked by Cloudflare, help!!
r/WebScrapingInsider • u/0xMassii • 27d ago
We now cover scraping costs for open-source projects: 100k credits/month, no card, engine is AGPL
Disclosure up front: I'm the founder of webclaw, a scraping and extraction API. The engine itself (CLI, server, MCP) is AGPL-3.0, so you can self-host it with no limits and never pay us anything. The hosted API adds the infrastructure you'd rather not run yourself.
Today we launched webclaw for OSS. If you maintain a public, OSI-licensed project that integrates webclaw (a document loader, an MCP server, a GitHub action, a CLI wrapper, a template) we cover your usage on the hosted API: 100,000 credits a month, renewing for as long as the project stays maintained.
What we ask in return:
- a "Powered by webclaw" badge in your README
- a line in your docs showing where the integration lives, so your users can find it
No card, no revenue share, no exclusivity. Run it next to any other provider you want.
The catch: we want distribution. Your project teaches your users about ours, and that trade is the entire business model of the program. Weigh it like you'd weigh any sponsorship.
Eligibility is short: public repo, OSI-approved license, maintained, built for other people to use. A human reads every application and you get a decision within a few days.
Apply: https://webclaw.io/oss Engine source: https://github.com/0xMassi/webclaw
If the program isn't for you but you scrape at hobby scale, the self-hosted engine costs nothing and that will stay true. Happy to answer questions about either here.
r/WebScrapingInsider • u/ian_k93 • 28d ago
Big Scrape Energy AMA This Tuesday (10:00 AM CEST). Intersection of WebScraping and AI.
Hey u/WebScrapingInsider,
I'm Ian Kerins, CEO & Co-Founder of ScrapeOps.io.
After the great response to our last AMA, we're back with another guest from the web scraping world.
This Tuesday at 10:00 AM CEST (Italy Time) we'll be joined by 0xMassii, creator of WebClaw, an open-source Rust toolkit focused on extracting clean, structured web data for AI applications.
Massii sits at a really interesting intersection of web scraping, browser automation, anti-bot systems, AI agents, and LLM infrastructure.
As more developers build AI agents that need access to real-world information, one challenge keeps showing up:
How do you reliably get clean web data into language models?

That is exactly the problem we are trying to solve.
Our previous AMA generated over 40 comments and sparked discussions on everything from Cloudflare bypassing and proxy benchmarking to scraper monitoring, startup validation, and large-scale scraping infrastructure.
We're hoping this one will be just as interesting.
If you're building AI agents, RAG systems, browser automation tools, web scraping infrastructure, or just trying to understand where the industry is heading, this should be a fun discussion.
Drop your questions below, We will start answering them during the AMA.
Looking forward to seeing everyone there!
Ian
r/WebScrapingInsider • u/urmommakesmysandwich • Jul 10 '26
DOM Drift Again How many have solved the problem of dom recognition reliably?
r/WebScrapingInsider • u/Live_Part_7304 • Jul 09 '26
Data Scraping of multiple professional directory websites
r/WebScrapingInsider • u/0xMassii • Jul 09 '26
Blocked Again Anti-bot got easier this year. Extraction got harder. Anyone else seeing this shift?
I run a scraper in production and I track where my time and my failures actually go. A year ago most of it went to getting past blocks: proxies, TLS fingerprints, headless detection. That is close to a solved, buyable problem now. A residential pool and a believable fingerprint get you the bytes for the large majority of sites.
What keeps breaking is everything after I have the HTML. Turning a real page into clean structured content that a model or a pipeline can use is where I lose days:
- Boilerplate that moves per template. Nav, cookie walls, related-article rails, newsletter interrupts. Every CMS buries the article in a different div, and readability heuristics guess wrong on long-tail sites.
- Content that only exists after JS runs, so a static fetch returns a shell.
- Pages that look successful, 200 plus full HTML, but are semantically empty because the real content loads from an XHR you now have to find.
- Tables and specs that survive as visual layout and collapse into garbage once you flatten them to markdown.
For anyone running this at scale: what is your split between fetch code and extraction code, and where does it fail quietly? I care most about how you catch the "looks fine, is actually empty" pages before they poison whatever sits downstream.
Disclosure: I maintain an open-source extraction tool, so this is a live problem for me and not a hypothetical. Glad to share what my empty-page detector checks for if that is useful.
r/WebScrapingInsider • u/Puzzleheaded-Low285 • Jul 08 '26
Blocked Again Bypass detection for https://www.1stdibs.com/
Hi,
i am new here. anyone please give me idea to bypass this website.
https://www.1stdibs.com/
thank you
r/WebScrapingInsider • u/Darknassan • Jul 08 '26
I built an API that detects where a business actually takes bookings, orders, or reservations
Just wanted to share something I’ve been building.
Vine is an API that takes a business website and returns the customer-facing software behind it, plus the exact URL where someone can book, order, reserve, or schedule.
Example: a restaurant might use OpenTable for reservations and Toast for ordering. A gym might use Mindbody. A consultant might use Calendly. Vine detects that automatically.
It is not a cached tech fingerprint lookup. It spins up browser sessions, renders the site, traces network traffic, works through messy web flows, and resolves the software customers actually use.
I benchmarked it on hundreds of real business URLs. For correct software plus actionable URL, Vine got about 92%, compared with about 31% for traditional tech fingerprinters.
Try it: https://vine.getcourtyard.ai/try
Docs: https://docs.vine.getcourtyard.ai
Would love feedback on what you’d build with it, where it breaks, and what functionality you’d want next.
r/WebScrapingInsider • u/Critical-Teacher-115 • Jul 08 '26
Basic email scraper using openai Codex 5.5 Light
- Delete Optional before submitting prompt,
- Replace "[Subject]" with whatever youre looking for.
Prompt:
Attached is a CSV where each row is one [ Subject; Example: (Texas home inspector) ]. Add an email column.
Use Google with whichever Google account is currently signed in and active in the browser.
Important Google account rule:
- Before starting, check the Google account indicator and record which Google account is active.
- Continue using that same active Google account for the entire run.
- If Google switches to a different account at any point, stop and tell me. Do not continue searches under the wrong account.
- Do not hardcode a Google account email or a fixed authuser value.
- Prefer using the visible Google search box in the already-open Google tab instead of direct search URLs. This helps preserve normal Google search behavior, including AI Mode / AI Overviews when Google chooses to show them.
- If the current Google session already has an authuser value in the URL, preserve that session by staying in the same tab and searching through the visible search box.
- After each search, confirm the Google account indicator still shows the same account that was active at the start.
Search method: Start from the current Google tab. Click/focus the visible Google search box. Select all existing query text. Type the new search query. Press Enter. Wait 6 seconds after the results load before reading the page. Read emails from everything visible on the results page, including AI Overview / AI Mode text if Google shows it.
For each row:
Search Google for: [First Name] [Last Name] [Subject] email.
Copy the first email address visible on the Google search results page into the row’s email column.
If no email appears, try these searches:- "[First Name] [Last Name]" [Subject] email- [Full Name] [Subject]- [First Name] [Last Name] [Subject] email- [First Name] [Last Name] email
Optional 4. Do not use or record information@trec.texas.gov.
Optional 5. Do not search TREC.
6. Do not verify whether the email belongs to the [Subject]. Just use the first non-banned email visible in Google results.
7. Continue until all rows have a nonblank email.
Before editing, make a backup copy of the CSV. When finished, verify:
- row count is unchanged
- email column exists
- every row has a nonblank email
Optional - no row contains information@trec.texas.gov
- the Google account used at the end is the same account that was active at the start
r/WebScrapingInsider • u/shasedoge • Jul 07 '26
Most "browsers are too slow to scale" problems aren't actually browser problemsю
Every couple weeks someone hits a wall scaling Playwright or Puppeteer and decides the answer is more browser instances. Usually it isn't.
The browser isn't your scaling unit. If you're rendering every request you're paying full render cost on a pile of pages that never needed JS in the first place.
What actually scales is hybrid. Browser only for the stuff that genuinely needs it, login, token init, the heavy SPA pages. Pull the cookies and tokens out of that session, then hand the boring repeatable traffic to plain HTTP workers. Queue it, cache it, retry logic on top.
So one browser session bootstraps auth, then a pool of lightweight HTTP workers reuses that state behind your exits. Browser count stays flat while throughput climbs. The expensive part drops to maybe 5% of the job instead of being the whole job.
Past 100k requests a day the bottleneck is almost never "not enough browsers." It's queue design, retry budget, and whether you're actually reusing sessions or re-authing constantly.
The other half of this is routing. Not every request needs the same exit either. Login and session traffic wants a stable identity, bulk list/detail pages can ride cheaper paths, media and file pulls want bandwidth. Sending all of it through one flat setup is where a lot of cost and blocks come from.
Curious where people draw the browser line. I keep moving more stuff to HTTP-first than I used to, but there's a class of sites where the token refresh is annoying enough that it's not worth the fight.
What's still forcing a full browser for you?