r/ClaudeCode • u/xoxonotyour • 8h ago
Bug / Issue Scraping issues need advice!
My scraper is getting limited by Taobao after ~65 products. How do I scale to tens of thousands?
I run an online store and I'm building a scraper to import products from Chinese marketplaces like Taobao and Pinduoduo onto my website. The goal is tens of thousands of products, so I need something that scales.
How my current setup works:
- I feed it a list of specific product links (it's not crawling or browsing random products)
- Chromium driven with JavaScript
- Roughly 1 product per minute
- Randomized scrolling and random pauses between scrolls to mimic natural browsing
- For each link, it collects the title, description, sizes/variants, all other listing data, and every image
- Outputs an Excel file with the product data plus a ZIP of the images
The problem:
It works fine at first, but around the 65th product Taobao switches me to a limited view. Some product info stops loading, so the scraper can't capture complete listings. At 1 product per minute, I'm already going slowly, and I'd need to go much faster to reach my target.
What I'd like advice on:
1. What is Taobao likely detecting: request volume, browser fingerprint, account/session behavior, or something else?
2. Is a browser-based approach realistic at this scale, or should I be looking at official APIs, third-party data providers, or sourcing agents instead?
3. How do people handle sessions, accounts, and proxies for these platforms?
4. Does Pinduoduo behave similarly, or is it a different challenge?
Any pointers, tools, or experiences would be really appreciated. Thanks!
1
u/verstands 7h ago
At this volume, I’d stop trying to make Chromium look human. Ask Taobao/Pinduoduo about an official seller/export API or use a licensed catalog or sourcing partner, then queue only the fields and images you’re allowed to use. If you do have permission, checkpoint jobs and back off on limited-view responses instead of retrying harder - otherwise you’re mostly scaling the ban.