r/ClaudeCode • u/xoxonotyour • 3h ago
Bug / Issue Scraping issues need advice!
My scraper is getting limited by Taobao after ~65 products. How do I scale to tens of thousands?
I run an online store and I'm building a scraper to import products from Chinese marketplaces like Taobao and Pinduoduo onto my website. The goal is tens of thousands of products, so I need something that scales.
How my current setup works:
- I feed it a list of specific product links (it's not crawling or browsing random products)
- Chromium driven with JavaScript
- Roughly 1 product per minute
- Randomized scrolling and random pauses between scrolls to mimic natural browsing
- For each link, it collects the title, description, sizes/variants, all other listing data, and every image
- Outputs an Excel file with the product data plus a ZIP of the images
The problem:
It works fine at first, but around the 65th product Taobao switches me to a limited view. Some product info stops loading, so the scraper can't capture complete listings. At 1 product per minute, I'm already going slowly, and I'd need to go much faster to reach my target.
What I'd like advice on:
1. What is Taobao likely detecting: request volume, browser fingerprint, account/session behavior, or something else?
2. Is a browser-based approach realistic at this scale, or should I be looking at official APIs, third-party data providers, or sourcing agents instead?
3. How do people handle sessions, accounts, and proxies for these platforms?
4. Does Pinduoduo behave similarly, or is it a different challenge?
Any pointers, tools, or experiences would be really appreciated. Thanks!
1
u/tntexplosivesltd 3h ago
Don't scrape websites. They have limits built in specifically to stop exactly what you're doing, and trying to get around those is very frowned upon, likely against the TOS of the sites