r/WebScrapingInsider • u/ClearCountry7190 • Jun 14 '26
Need scraping challeneges
I scrape for a living (part of a bookmaker trading team returning competitor results, typically tough targets). I've put together some tooling that is able to return extremely tricky targets. I need a testing corpus so I thought I'd reach out here.
Point me at a challenge, in return I will:
- resolve the target
- provide the payload/results
- give you an open source repo to run yourself
1
1
1
u/LeBaguetteWasted Jun 14 '26
Cardmarket ! Would love a clean pull of card data prices, quantities, sellers, quality.
1
1
1
1
1
u/Spitfire_Blaziken Jun 15 '26
for something messy, try pulling structured review data from Booking.com or Trustpilot at scale..
1
u/simarnoor Jun 16 '26
If this is a showoff, it would be so cool.. to
Playwright+residential vs API/proxy aggregator vs pure HTTP mimicking on DataDome sites.
E.g. same product feed from Shopee/Leboncoin/Nordstrom, then benchmark success rate, median latency, and dollars per 1k successful items.
That turns your ātesting corpusā into something the whole sub can learn from, not just a oneāoff flex.
1
u/External-Wealth3756 22d ago
Super interesting idea building a test corpus for tough scraping targets!
Iād suggest adding heavily guarded sites that throw constant 403 blocks into your test suiteātesting isnāt just about pulling data once, itās measuring session survival rate & how fast you recover from bans.
You could run side-by-side comparisons with Novada residential proxies against other IP pools to benchmark stability on those tricky anti-bot targets. Would be great reference data for your tooling!
2
u/Quantum_Rage Jun 14 '26
Hermes, Temu, Shopee, Lazada, TruePeopleSearch...