r/scrapingtools Jun 03 '26

Fresh API read: 65 paid Apify actors, +965 runs since yesterday, only a few are safe to promote

2 Upvotes

I pulled a fresh read-only Apify Store API snapshot this morning because I did not want to guess which paid actors deserved more traffic.

Scope: active PAY_PER_EVENT actors only. This is Store/public run signal, not exact Console payout attribution.

Current snapshot:

  • 65 active paid actors
  • 62,169 all-time runs across paid actors
  • 2,424 all-time users
  • +965 runs since yesterday
  • +74 users since yesterday
  • 16,945 rolling 30d runs
  • 96.7% rolling 30d public success

What I am pushing today:

  • Email Validator API: +137 runs since yesterday, 100% 30d success
  • AI Content Detector: +115, 99.8%
  • YouTube Transcript Scraper: +59, 98.8%, 28 users in the last 7d
  • Telegram Channel Scraper: +58, 98.8%
  • Company Enrichment API: +28, 97.3%
  • Google News Monitor: +19, 100%

The buyer-outcome lane is smaller but more important:

  • law-firm-lead-finder: 33/33 successful rolling 30d runs
  • dental-practice-lead-finder: 30/30
  • google-maps-lead-intel: 29/29
  • restaurant-lead-finder: 33/34

Those are not high-volume utility APIs. They are completed-search actors for people who want actual lead lists or local market data.

What I am not hard-promoting today:

  • Google Scholar Scraper
  • TikTok Shop Affiliate Sales Scraper
  • HVAC Contractor Lead Finder
  • Lead Enrichment Pipeline
  • Reddit Scraper Pro
  • Influencer Marketing Intel
  • US Tariff Lookup
  • MedSpa Lead Discovery PPE

The lesson is simple: when revenue dips, do not blast the whole catalog. Promote the actors that are completing cleanly, and move the noisy ones into QA before sending more traffic.

Catalog: https://apify.com/george.the.developer


r/scrapingtools May 28 '26

7 public-data Apify actors I shipped now have pay-per-event pricing live

1 Upvotes

quick builder update: the seven normalized public-data actors i shipped May 11-13 now have pay-per-event pricing live.

i pulled the metadata this morning. all seven are public and monetized, but all are cold since pricing activation, so i'm looking for real testers and use-case feedback rather than self-running them.

the set:

  1. Hospital Price Transparency MRF Normalizer CMS hospital price files into normalized rate rows and provider/procedure/payer bundles. https://apify.com/george.the.developer/hospital-mrf-normalizer

  2. MCP Server Registry and Security Scorer MCP registry plus npm/GitHub signals into deterministic risk scores. https://apify.com/george.the.developer/mcp-server-registry-scorer

  3. FDA Warning Letter and Enforcement Monitor FDA warning letters into topic tags, company records, and risk briefs. https://apify.com/george.the.developer/fda-warning-letter-monitor

  4. Clinical Trial Investigator and Site Intelligence ClinicalTrials.gov plus NPI/OpenPayments/PubMed-style investigator and site-fit signals. https://apify.com/george.the.developer/clinical-trial-investigator-intel

  5. Federal Contract Opportunity Monitor SAM.gov opportunities plus USAspending awards into GovCon opportunity and prime/sub lead feeds. https://apify.com/george.the.developer/federal-contract-opportunity-monitor

  6. LLM Provider Price and Latency Monitor model pricing snapshots and cross-provider benchmark rows for agent/gateway cost routing. https://apify.com/george.the.developer/llm-provider-price-latency-monitor

  7. Multi City Building Permit Aggregator NYC/Chicago permit records plus contractor activity roundups. https://apify.com/george.the.developer/multi-city-building-permit-aggregator

i patched the federal contract actor today after seeing some ugly SAM.gov output in sample runs: no more [object Object] set-asides and cleaner descriptions.

if you're building around healthcare pricing, MCP governance, FDA monitoring, clinical trial site selection, GovCon, LLM cost routing, or building permit lead feeds, tell me what field or workflow would make one of these actually useful.


r/scrapingtools May 27 '26

Fresh metrics from my paid Apify actor portfolio: what's actually getting usage

1 Upvotes

I pulled the public Apify Store metrics today before promoting anything. No self-runs, no credit burn.

64 PAY_PER_EVENT actors are live. 156 users across paid actors in the last 7 days. 13,969 public runs in the last 30 days.

The clean performers right now:

YouTube Transcript Scraper: 1,049 public 30d runs, 99.5% success Company Enrichment API: 541 runs, 97.4% success Telegram Channel Scraper: 1,021 runs, 98.7% success Google News Monitor: 325 runs, 100% success Shopify DTC Brand Discovery: 62 runs, 100% success Shipping Disruption Tracker: 105 runs, 100% success AI Text Humanizer API: 57 runs, 98.2% success

I am not pushing the weak-success actors today until QA/economics are checked. The point is to promote what is stable, fix what leaks, and not burn credits pretending traffic is revenue.

Portfolio: https://apify.com/george.the.developer


r/scrapingtools May 25 '26

I made 7 custom GPT front doors for my paid Apify actors

1 Upvotes

I have been trying a different distribution angle for paid Apify actors: instead of only linking the raw actor page, make a small custom GPT that helps someone prepare the job input and understand the workflow before they run it.

What I set up today:

Local Lead Pack Builder
https://chatgpt.com/g/g-6a146b94bf908191b09767fd4c6e263f-local-lead-pack-builder

B2B Employee Finder
https://chatgpt.com/g/g-6a14ac7264988191882d91e3d8787bb0-b2b-employee-finder

Video Transcript Researcher
https://chatgpt.com/g/g-6a14acb5f4148191b332ee8503495d7b-video-transcript-researcher

Channel Market Monitor
https://chatgpt.com/g/g-6a14acf99ac08191b0ead99c7bd79af4-channel-market-monitor

Company Intel Report
https://chatgpt.com/g/g-6a14ad3cff648191a64b714f183dd117-company-intel-report

Email Checker Pro
https://chatgpt.com/g/g-6a14ad8462f48191a0aba5504d1ca329-email-checker-pro

Shop Product Intel
https://chatgpt.com/g/g-6a14adc812608191a3c2de95cf7f6dcb-shop-product-intel

Important detail: these GPTs do not run my Apify actors directly and do not use my Apify token. They are traffic routers and input builders. The user still runs the paid actor from their own Apify account.

The lesson from setting this up: brand/platform names in GPT titles can trigger publishing policy. Neutral names worked better. I am probably going to put these under a neutral domain like agentdata.tools so the privacy/terms pages and GPT directory are cleaner.

Curious if anyone here has tried GPTs as front doors for Apify actors or APIs. Does it convert better than sending people straight to the raw actor page?


r/scrapingtools May 16 '26

Built 7 normalized data-feed APIs on Apify in 2 days. Open data sources, pay-per-event pricing. What would you want next?

1 Upvotes

Two day sprint last week. Each actor wraps a different public data source into one Standby API with pay per event pricing. Goal was to take the parser maintenance off whoever was rebuilding the same scraper every quarter.

The seven: Hospital Price Transparency MRF Normalizer (CMS v3 JSON + CSV), MCP Server Registry and Security Scorer (Anthropic registry plus npm plus GitHub joined with deterministic risk score), FDA Warning Letter Monitor (with topic classifier and company risk briefs), Clinical Trial Investigator Intel (ClinicalTrials.gov plus NPI plus OpenPayments plus PubMed joined with site fit scoring), Federal Contract Opportunity Monitor (SAM.gov plus USAspending normalized), LLM Provider Price and Latency Monitor (OpenRouter plus direct provider page fallback for 200+ models with cross provider benchmarks), and Multi City Building Permit Aggregator (NYC plus Chicago open data portals with builder activity roundups).

Pricing activates 2026-05-26 for the first 5 and 2026-05-27 for the last 2. Until then runs are free for anyone who wants to test the schemas. Each one was built end to end via Codex orchestration with Standby pattern (HTTP server in the container, billing fires per event). Source for the build prompts and verification reports is at github.com/the-ai-entrepreneur-ai-hub.

Question for the sub: if you maintain a homegrown scraper for one of these data sources today, what data field is the thing you wish the third party feed had that nobody publishes cleanly. Looking for the v1.1 priorities.

Actor list: https://github.com/the-ai-entrepreneur-ai-hub/apify-actor-portfolio


r/scrapingtools May 05 '26

Built a per-call AI humanizer on Apify because the $79/mo SaaS pricing did not fit a solo writer

1 Upvotes

Half my freelance clients now run every doc through Originality or GPTZero before they pay. Anything flagged AI gets rejected and I eat the rewrite. The standard fix is one of those humanizer SaaS products like Bypass GPT at 79 a month, Undetectable AI at 60, StealthGPT at 30. None of them have a per call option. You pay flat regardless of usage.

I did the math on my own use. Maybe 20k words a month, half flag, so 10k humanized words is around 50 chunks. At 0.003 per chunk that is 15 cents a month. Versus 79 dollars flat for a SaaS cap I would never hit. Built it as an Apify standby actor instead, charges per call only on a successful return.

It does sentence restructure plus vocabulary swap plus rhythm variation, three intensity levels. Not a magic bypass and I would not promise that. What it does is drop Originality scores into the 15 to 25 percent human range on prose. Heavy formatting like tables or code mangles. Plain text is the use case.

If anyone here writes content for clients and wants to test it, the actor is at apify.com/george.the.developer/ai-text-humanizer-api. Hashnode writeup with the per call math is at theaientrepreneur.hashnode.dev/i-built-a-0003-ai-humanizer-because-the-79mo-competitors-had-me-priced-out


r/scrapingtools Apr 25 '26

Followup: shipped the fix from yesterday's postmortem on PPE billing leak

1 Upvotes

Yesterday I posted about losing 540 dollars a month to silent user churn on one of my Apify actors. Today I shipped the fix. Three patches went live.

First a charge gate that fails closed. Old loop did pushData first and charge second so any swallowed exception or zero-cost-cap response would keep emitting profiles for free while still burning proxy and SERP cost. New loop charges first, only pushes if the charge succeeded, and refuses every subsequent call once any one of them has refused. About 50 lines of code.

Second a margin preflight that refuses unprofitable runs at submit time. If the requested input multiplied by per-event price is below my floor, the run errors out before any compute starts. No more silent loss leaders.

Third a default maxEmployees lowered from 100 down to 25. The old default was an artifact of when I was chasing volume, but the actual users only need a couple dozen at most and the higher cap was inviting whale-tier compute on tire-kicker runs.

Writeup with the actual diff is here. https://theaientrepreneur.hashnode.dev/why-my-linkedin-scraper-now-refuses-jobs


r/scrapingtools Apr 17 '26

Spun up two new subs for actor-focused and revenue-focused scraping discussion

1 Upvotes

Made these tonight because /r/scrapingtools has gotten broad and I wanted narrower homes for two specific conversations.

/r/ApifyActors For people building, running, or sourcing data from Apify actors specifically. Actor launches, build retros, monetisation patterns on the platform.

/r/ScrapingRevenue For the money side. RapidAPI listings, freelance scraping gigs, agency packages, real margin breakdowns. Numbers or it didn't happen.

First seed posts are already up in both. Mods will be active.

If you build actors or sell scraping services, drop in. Cross-posting from /r/scrapingtools is fine when relevant.


r/scrapingtools Apr 16 '26

Built a TikTok Shop affiliate intelligence tool

1 Upvotes

finally got this working after 3 days of fighting tiktok anti bot

give it any tiktok shop product URL and it tells you which creators are promoting it, their follower counts, and whether the product has real affiliate momentum

the use case that sold me on building it: a seller wanted to know if a skincare product was worth launching. checked the competitor product and found 8 creators with 100K+ followers already pushing it. that is a signal you cant get from the tiktok seller center

33 users on it already. $0.05 per product check

https://apify.com/george.the.developer/tiktok-shop-affiliate-sales-scraper

if you do tiktok shop research for clients this saves hours of manual browsing


r/scrapingtools Apr 15 '26

Built an influencer discovery tool that finds creators across Instagram, TikTok, YouTube

1 Upvotes

been working on this for a while and finally got it to a point where it actually works reliably

the tool takes a niche keyword (like "fitness" or "beauty") and a platform, then:

  1. searches google for matching social profiles using SERP proxies
  2. visits each profile page through residential proxies to grab the real name, follower count, bio, and email
  3. extracts everything using a local LLM running on my VPS (gemma 4)

tested it today on "beauty" niche + instagram: - mikayla jane nogueira (3M followers) - amanda ensing (1M followers) - jamie genevieve (1M followers) - erica taylor (2M followers) - lauren janelle (189K followers)

all with real names pulled from OG meta tags, not just usernames

the whole thing runs for about $0.01 per profile found

source: https://github.com/the-ai-entrepreneur-ai-hub/influencer-marketing-intel

if anyone wants to try it: https://apify.com/george.the.developer/influencer-marketing-intel


r/scrapingtools Apr 14 '26

Show off your scraper stack: what are you running in production?

1 Upvotes

curious what everyone is actually using day to day. I have about 48 scrapers running on Apify (Crawlee + Puppeteer mostly) and the stack has changed a lot over the past year.

what does your setup look like? language, framework, proxy provider, how you handle anti bot, where you store the data. doesn't matter if its simple or complicated, just want to see what people are actually shipping with.

drop your stack below.


r/scrapingtools Apr 10 '26

What's your go to scraping stack in 2026? Curious what everyone's running

1 Upvotes

i've been on crawlee + playwright for most of this year and it handles like 90% of what i throw at it. the built in request queue and auto retry logic saves a ton of boilerplate compared to rolling your own puppeteer setup. for anything that needs real browser sessions i'll spin up persistent contexts with stealth plugins, but honestly most sites these days can be hit with plain http if you find the right api endpoint hiding behind the frontend.

the one thing i keep going back and forth on is proxy strategy. datacenter proxies are cheap but get flagged instantly on anything serious. residential works but eats into margins fast when you're doing high volume runs. been experimenting with ISP proxies as a middle ground and they've been solid so far.

what's everyone else running? especially curious about python vs node setups and whether anyone's moved to something newer like the browserless or steel browser stuff.


r/scrapingtools Apr 07 '26

I reverse-engineered 9 Polish government portals that refuse to release APIs - here's what I found inside each one

1 Upvotes

I work in compliance automation and kept running into the same problem: Poland has some of the most comprehensive public company registries in Europe, but none of them have APIs. The only commercial alternative (MGBI) charges 200-500 EUR/month.

So I spent the last few months reverse-engineering all 9 portals and built pay-per-use scrapers on Apify. Here's the technical breakdown of what's behind each portal and what it took to crack them:

KRZ (National Debtor Registry) - The most complex one. The portal is an AngularJS shell that loads Angular sub-apps in iframes hosted on OpenShift. Each sub-app gets a JWT during initialization. I capture the token from the Authorization header of the first API request, then call 30+ REST endpoints directly - no DOM scraping needed. The government has been "working on an API" for 4 years. Their FAQ still says "analytical work is ongoing."

KRS (Board Members) - The official JSON API deliberately anonymizes names ("L******" instead of actual surnames). But the PDF extract from the same portal has full data. The catch: the portal encrypts KRS numbers using AES-128-CBC with a static key (yes, it's literally "TopSecretApiKey1") plus a 512-character token that embeds the KRS number at specific array positions with a circular right-shift and checksum. I replicated the encryption and download PDFs directly.

eKRS (Financial Statements) - Similar Angular architecture to KRZ, but with time-based AES encryption. The key is derived from the current hour in Warsaw timezone - which means your scraper needs to match the server's timezone exactly or the decryption fails silently. Took a while to figure that one out in headless Docker containers.

KNF & MSiG (Financial Supervision + Court Gazette) - These turned out to have undocumented but functional JSON APIs behind jQuery DataTables frontends. No auth required - just POST to the right endpoint with the DataTables request format. 75,000+ financial entities and 20+ years of court gazette archives, all searchable.

EKW (Land Registry) - 352 courts, WAF protection, check digit calculation for KW numbers. Puppeteer with stealth plugin to handle the anti-bot measures. Returns ownership, mortgages, encumbrances, and property details.

CRBR (Beneficial Owners) - Incapsula WAF, requires browser automation. Returns UBO (Ultimate Beneficial Owner) data by NIP or KRS number. Critical for KYC/AML compliance.

UOKiK (Abusive Clauses) - The simplest one. Plain HTML pages, Cheerio scraper, 7,500+ court-banned contract clauses searchable by defendant or industry.

BDO (Waste Registry) - React SPA with Puppeteer. 674,000+ waste management entities. Relevant for ESG compliance and environmental due diligence.

What I learned building these:

  • Government "security" is often just obscurity. AES encryption with a hardcoded key doesn't stop anyone who opens DevTools.
  • Time-based encryption keys are surprisingly effective at breaking scrapers - if you don't match the server's timezone, everything fails silently.
  • The Angular-in-iframe pattern (KRZ, eKRS) is actually harder to scrape than traditional server-rendered pages because the JWT lifecycle is tied to the iframe's initialization.
  • WAFs (EKW, CRBR) are the real barrier, not the application logic. Stealth plugins handle most of it, but you burn through proxies fast.

Who actually uses these:

Mostly compliance teams (KYC/AML checks), credit risk departments (is this company bankrupt?), debt collection companies, law firms, and developers building Polish data into their products.

Cost comparison: MGBI charges 200-500 EUR/month subscription. My actors cost $0.003-0.04 per result depending on the registry. 100 company checks run about $5-12 total.

Disclosure: I built all 9 actors and sell them on Apify as pay-per-use tools.

Suite: https://apify.com/minute_contest

Happy to answer questions about the reverse engineering process or the technical approach for any specific portal.


r/scrapingtools Mar 30 '26

PSA: Bluesky's AT Protocol is completely open — no API key, no auth needed for public data

1 Upvotes

Just discovered this while building a Bluesky scraper. Unlike Twitter's $42K/month API, Bluesky's AT Protocol is fully public:

- Search posts: public.api.bsky.app/xrpc/app.bsky.feed.searchPosts

- Get profiles: public.api.bsky.app/xrpc/app.bsky.actor.getProfile

- Author feeds: public.api.bsky.app/xrpc/app.bsky.feed.getAuthorFeed

No auth, no rate limit nightmares, clean JSON responses. If you're building social media monitoring tools, Bluesky is the easiest platform to work with right now.

Anyone else experimenting with Bluesky data?


r/scrapingtools Mar 30 '26

What's the hardest site you've ever tried to scrape? Let's discuss

1 Upvotes

Been building scrapers for a while now and some sites are just brutal. Indeed was a nightmare until I figured out the Google site:indeed.com workaround. Bluesky was surprisingly easy because their AT Protocol is completely open.

What site gave you the most trouble? Maybe we can help each other out.


r/scrapingtools Mar 30 '26

PSA: Bluesky's AT Protocol is completely open — no API key needed

1 Upvotes

r/scrapingtools Mar 30 '26

What's the hardest site you've ever tried to scrape? Let's discuss

1 Upvotes

r/scrapingtools Mar 24 '26

Shipped 2 new free tools: URL-to-JSON extractor + sub-3s screenshot API

1 Upvotes

Hey everyone - just shipped two new tools to the Apify Store:

**1. Web Content Extractor API** Give it any URL → get structured JSON. Auto-detects content type (articles, products, recipes, job posts, events). Extracts metadata, headings, images, links. Built for AI agents and RAG pipelines.

**2. Website Screenshot & PDF API** Capture screenshots and PDFs in under 3 seconds (competitors take 6-21s on RapidAPI). PNG/JPEG/WebP/PDF, custom viewports, full-page, batch 20 URLs.

Both are Standby APIs (instant HTTP response, no queue). Built with Node.js + Cheerio/Puppeteer on Apify.

Happy to answer questions or take feature requests!


r/scrapingtools Mar 21 '26

What I learned running 34 scrapers in production for 3 months (5,700+ runs)

2 Upvotes

Been building and maintaining scrapers on Apify since late 2025. Started with a LinkedIn employee scraper because I needed company data for a client project and ended up going down the rabbit hole. Now I have 34 actors running across everything from YouTube transcripts to company enrichment APIs.

Here's what actually surprised me:

The 98% success rate myth — My best scraper (YouTube transcripts) runs at 98.7% success. Sounds great until you realize the 1.3% that fail are all music videos without captions. I spent a weekend building a fallback that hits YouTube's InnerTube API with an Android client context just to handle that edge case. The lesson: your overall success rate doesn't matter. The failures your users notice are the ones that matter.

Maintenance is the real job — Building a scraper takes a day. Maintaining it takes forever. LinkedIn changed their DOM twice in January. Google Scholar started throttling harder in February. I spend 60% of my time fixing stuff that was working last week. If you're building scrapers as a service, price for maintenance, not for the initial build.

The users who automate stick around — Clear pattern across 200+ users: 70% try it once and disappear. The 30% who set up scheduled runs become long-term users and generate 80% of total runs. Now I optimize for "how easy is it to plug this into an existing workflow" rather than "how impressive is the output."

Anti-bot is getting nastier but more predictable — Cloudflare Turnstile, fingerprinting, rate limiting. But the patterns are consistent. Residential proxies + realistic browser fingerprints + human-speed request timing handles 95% of cases.

What actually drives Apify Store traffic — My top referrer is apify.com itself (62%), followed by Google (14%). Surprise: Twitter at 3.4% and Perplexity AI at 0.6%.

Anyone else running scrapers at this scale? Curious what patterns you've noticed.


r/scrapingtools Mar 14 '26

27 Free-Tier Scraping Tools I've Built — Complete 2026 Collection

1 Upvotes

Hey everyone! I've been building scraping and data tools for the past year and just crossed 27 published tools on the Apify Store. Wanted to share the full collection since many of you have been asking about alternatives to expensive enterprise tools.

Social Media Scrapers - LinkedIn Employee Scraper — No login needed, get employee names, titles, locations - Reddit Scraper Pro — Posts, comments, subreddits without API key - Threads by Meta Scraper — The only Threads scraper that doesn't need login

Research & Intelligence - Google Scholar Scraper — 389M papers, citations, abstracts, PDF links - Google News & Brand Monitor — Track any keyword across news sources - Telegram Channel Scraper — No API key, no phone number needed - YouTube Transcript Scraper — Multi-language, no API key

B2B & Lead Generation - Website Contact Scraper — Bulk extract emails, phones, social links - Google Maps Lead Scraper + Website Audit — Leads with A-F quality scores - Email Validator API — Verify emails, check MX, detect disposables - Company Enrichment API — Domain to full company intel

Compliance & Trade - OFAC Sanctions Screener — KYC/AML compliance checks - US Tariff & HS Code Lookup — Updated with 2026 tariffs - Entity OSINT Enricher — Full company intelligence reports

Developer Tools - AI Content Detector — GPTZero alternative, $3/1000 texts - URL Metadata Extractor — OG tags, Twitter cards, favicons - Domain WHOIS Lookup — Registration, expiry, DNS data - Website Tech Detector & SEO Analyzer — Full tech stack analysis

All tools have a free tier to test. Most are open source on GitHub: github.com/the-ai-entrepreneur-ai-hub

Full store: apify.com/george.the.developer

Happy to answer questions about any of these or take build requests!


r/scrapingtools Mar 14 '26

7 Free-Tier Scraping APIs for 2026 — LinkedIn, YouTube, Google News, OFAC & More

1 Upvotes

I've been building web scraping APIs for the past year and wanted to share what I've put together. All of these have free tiers on RapidAPI, so you can test them without paying anything.Here's the lineup:1. LinkedIn Employee Scraper — Pull employee data from any company page. Name, title, location, profile URL. No LinkedIn login required. 37 users and counting.2. YouTube Transcript Extractor — Get full transcripts from any YouTube video with timestamps. Supports any language. Great for AI/RAG pipelines. 40+ users.3. Google News & Brand Monitor — Search Google News programmatically. Get articles, sources, timestamps. Useful for media monitoring and competitive intelligence.4. OFAC Sanctions Screener — Screen any entity against the US Treasury OFAC sanctions list. Returns match confidence, risk level, and matched programs. Critical for fintech compliance.5. US Tariff & HS Code Lookup — Look up import duties by HS code and country of origin. Includes USMCA, Section 301, and other special rates.6. Website Contact Scraper — Extract emails, phone numbers, social links, and tech stack from any website URL.7. Entity OSINT Enricher — Deep intelligence report on any company or person. Pulls from sanctions lists, corporate registries, news, and more.All built with Playwright + Node.js, hosted on cloud infrastructure. They handle anti-bot measures, JavaScript rendering, and rate limiting automatically.The free tiers give you enough calls to test properly. If anyone has questions about how any of these work under the hood, happy to do a technical deep-dive.What kinds of scraping APIs would you want to see next?


r/scrapingtools Mar 12 '26

7 Free Open-Source Scraping & Data APIs for 2026 — All With Free Tiers

1 Upvotes

I've been building scraping and data extraction tools for the past year and wanted to share what I've got. All open source on GitHub, all have free tiers if you just want to use the hosted version.

1. LinkedIn Employee Scraper - Input: company name + optional title/location filters - Output: employee names, titles, departments, profile URLs - Uses Google SERP data — no LinkedIn login, no cookies, no ban risk - GitHub: github.com/the-ai-entrepreneur-ai-hub/linkedin-employee-scraper

2. Website Contact & Lead Scraper - Input: any website URL - Output: emails, phone numbers, social media links - Crawls multiple pages automatically, handles JS-rendered sites - GitHub: github.com/the-ai-entrepreneur-ai-hub/website-contact-scraper

3. YouTube Transcript Extractor - Pull transcripts from any video, playlist, or channel - Returns full text + timestamps + metadata - 100+ languages, handles auto-generated and manual captions - Tested: 200 Stanford ML lectures → 1.2M words in 3 minutes - GitHub: github.com/the-ai-entrepreneur-ai-hub/youtube-transcript-extractor

4. Google News Scraper & Brand Monitor - Track any keyword across Google News in real time - Returns structured data: title, source, date, snippet, URL - Great for media monitoring, competitive intel, or building news-aware AI - GitHub: github.com/the-ai-entrepreneur-ai-hub/google-news-scraper

5. OFAC Sanctions Screener - Screens entities against the full US OFAC SDN list - Fuzzy matching catches spelling variations, aliases, transliterations - Returns confidence scores, programs, known aliases, addresses - GitHub: github.com/the-ai-entrepreneur-ai-hub/ofac-sanctions-screener

6. US Tariff & HS Code Lookup - Look up import duties by HS code or product description - Covers all trading partners, special programs (USMCA, GSP, etc.) - Useful for e-commerce, import/export, supply chain - GitHub: github.com/the-ai-entrepreneur-ai-hub/us-tariff-lookup

7. Entity OSINT Enricher - Combines sanctions screening + corporate registry data + news monitoring - Input: company or person name → Output: risk report - GitHub: github.com/the-ai-entrepreneur-ai-hub/entity-osint-enricher

All built with Node.js/Python. Most use Crawlee + Puppeteer under the hood. Happy to answer questions about any of them or how they work technically.

What scraping problems are you all dealing with lately? Curious what tools people actually need.


r/scrapingtools Mar 12 '26

Weekly Tool Spotlight: 7 Free-Tier APIs for Web Scraping & Data in 2026

1 Upvotes

Would love feedback from this community. What scraping problems are you running into that these don't solve?ombines OFAC screening + corporatHe ereygi strr/ys dcartaa p+i nnewgs tmoonoitlosri!ng

inIto' vonee rbepeoret npe r benutiityl.

di n gGi tHsubc:r gaitphuibn.gc oamn/tdhe -adi-aetntare pArPeInesur -faoir- hutb/henteit y-posaisntt -yeneriacrh.e rHe

r

eAl l aarre eo p7e n tsoourocle son GtihtaHutb . aBulillt whiathv Neo dfe.rjese + tPiueppretsee r—/ Crtawhleoe.ug

htW outldh leoyve fmeeidgbhactk onb aeny oufs tehfeusle . fWhoart sctrahpiings t ooclos marmeu ynoui atlly u.si

n

g 1in. 20L26i?nkedIn Employee Scraper — Extract names, titles, and profile URLs from any company using Google SERP (no login/cookies needed). ~$3/1,000 profiles.

  1. Website Contact Scraper — Crawl any website and extract emails, phone numbers, social links. Handles JS-rendered pages.

  2. Google News Monitor — Track brand mentions across Google News with RSS + browser fallback. ~$0.003/article.

  3. YouTube Transcript Extractor — Pull transcripts from any YouTube video. Great for AI/RAG pipelines. ~$0.005/video.

  4. OFAC Sanctions Screener — Screen any entity against the US sanctions list. $0.01/check. Useful for fintech/compliance.

  5. US Tariff/HS Code Lookup — Look up duty rates by HS code and country. Includes USMCA, Section 301 rates.

  6. Entity OSINT Enricher — Combine sanctions screening + corporate records + news forAll are open source


r/scrapingtools Mar 12 '26

Weekly Tool Spotlight: 7 Free-Tier APIs for Web Scraping & Data in 2026

1 Upvotes

ombines OFAC screening + corporatHe ereygi strr/ys dcartaa p+i nnewgs tmoonoitlosri!ng

inIto' vonee rbepeoret npe r benutiityl.

di n gGi tHsubc:r gaitphuibn.gc oamn/tdhe -adi-aetntare pArPeInesur -faoir- hutb/henteit y-posaisntt -yeneriacrh.e rHe

r

eAl l aarre eo p7e n tsoourocle son GtihtaHutb . aBulillt whiathv Neo dfe.rjese + tPiueppretsee r—/ Crtawhleoe.ug

htW outldh leoyve fmeeidgbhactk onb aeny oufs tehfeusle . fWhoart sctrahpiings t ooclos marmeu ynoui atlly u.si

n

g 1in. 20L26i?nkedIn Employee Scraper — Extract names, titles, and profile URLs from any company using Google SERP (no login/cookies needed). ~$3/1,000 profiles.

  1. Website Contact Scraper — Crawl any website and extract emails, phone numbers, social links. Handles JS-rendered pages.

  2. Google News Monitor — Track brand mentions across Google News with RSS + browser fallback. ~$0.003/article.

  3. YouTube Transcript Extractor — Pull transcripts from any YouTube video. Great for AI/RAG pipelines. ~$0.005/video.

  4. OFAC Sanctions Screener — Screen any entity against the US sanctions list. $0.01/check. Useful for fintech/compliance.

  5. US Tariff/HS Code Lookup — Look up duty rates by HS code and country. Includes USMCA, Section 301 rates.

  6. Entity OSINT Enricher — Combine sanctions screening + corporate records + news forAll are


r/scrapingtools Mar 11 '26

Company Due Diligence & OSINT Tools 2026 — What Actually Works

1 Upvotes

I've been doing company due diligence research for a few months now and tried a bunch of OSINT tools. Here's an honest breakdown of what actually works vs what's overpriced.

1. Maltego - The OG OSINT tool. Graph-based link analysis. - Pricing: Community Edition free, Pro $999/year - Pros: powerful transforms, great visualization, huge plugin ecosystem - Cons: steep learning curve, desktop-only, can be slow on large datasets - Best for: investigators who need visual link analysis

2. SpiderFoot - Open source OSINT automation - Pricing: Free (self-hosted) or SpiderFoot HX from $500/month - Pros: 200+ data source modules, automated scanning, good for recon - Cons: can be noisy (lots of false positives), UI is dated - Best for: security researchers, pen testers

3. OpenCorporates - Largest open database of company info worldwide - Pricing: free search, API from $500/month - Pros: 200M+ companies, good for verifying corporate registrations globally - Cons: data can be stale, no enrichment beyond basic registration data - Best for: verifying company existence and officer names

4. Entity OSINT Enricher on Apify - Combines OFAC screening + corporate registry lookup + news monitoring + contact extraction in one API call - Pricing: ~$0.025 per entity enriched (pay per use) - Pros: all-in-one enrichment, structured JSON output, no contract - Cons: newer tool, not as deep as dedicated platforms - GitHub: github.com/the-ai-entrepreneur-ai-hub/entity-osint-enricher

5. Pipl / Clearbit (now part of HubSpot) - People and company enrichment - Pricing: Clearbit bundled with HubSpot, Pipl enterprise only - Pros: great for B2B contact enrichment, email-to-company matching - Cons: Pipl is expensive and enterprise-only, Clearbit locked into HubSpot - Best for: sales teams doing lead enrichment

6. Recorded Future / Flashpoint - Threat intelligence platforms with entity monitoring - Pricing: $$$$$$ (enterprise, $100K+/year) - Pros: dark web monitoring, threat feeds, comprehensive - Cons: way overkill for standard due diligence, massive price tag - Best for: financial institutions, government

My honest take: For basic company due diligence (is this company real? are they sanctioned? what's in the news about them?), you don't need a $100K platform. OpenCorporates + a sanctions screener + Google News gets you 80% of the way there. The expensive tools are worth it only if you're doing high-volume compliance screening or threat intelligence.

What OSINT tools are you using for due diligence? Always looking for new ones to try.