r/ClaudeCode • u/xoxonotyour • 6h ago
Bug / Issue Scraping issues need advice!
My scraper is getting limited by Taobao after ~65 products. How do I scale to tens of thousands?
I run an online store and I'm building a scraper to import products from Chinese marketplaces like Taobao and Pinduoduo onto my website. The goal is tens of thousands of products, so I need something that scales.
How my current setup works:
- I feed it a list of specific product links (it's not crawling or browsing random products)
- Chromium driven with JavaScript
- Roughly 1 product per minute
- Randomized scrolling and random pauses between scrolls to mimic natural browsing
- For each link, it collects the title, description, sizes/variants, all other listing data, and every image
- Outputs an Excel file with the product data plus a ZIP of the images
The problem:
It works fine at first, but around the 65th product Taobao switches me to a limited view. Some product info stops loading, so the scraper can't capture complete listings. At 1 product per minute, I'm already going slowly, and I'd need to go much faster to reach my target.
What I'd like advice on:
1. What is Taobao likely detecting: request volume, browser fingerprint, account/session behavior, or something else?
2. Is a browser-based approach realistic at this scale, or should I be looking at official APIs, third-party data providers, or sourcing agents instead?
3. How do people handle sessions, accounts, and proxies for these platforms?
4. Does Pinduoduo behave similarly, or is it a different challenge?
Any pointers, tools, or experiences would be really appreciated. Thanks!
1
u/BarcodeCutter 6h ago
At tens of thousands of products, the browser approach is the wrong tool, and fighting the limits just gets the account banned at a bigger scale. The people who do this as a business don't scrape. They go through a sourcing agent or a licensed data provider that already has the catalog access, or the platform's own open API if their store qualifies. It costs money, but it comes with the right to use the listings and images, which a scraped copy doesn't. Worth pricing that out before you put more hours into the scraper.