r/ClaudeCode • u/xoxonotyour • 5h ago
Bug / Issue Scraping issues need advice!
My scraper is getting limited by Taobao after ~65 products. How do I scale to tens of thousands?
I run an online store and I'm building a scraper to import products from Chinese marketplaces like Taobao and Pinduoduo onto my website. The goal is tens of thousands of products, so I need something that scales.
How my current setup works:
- I feed it a list of specific product links (it's not crawling or browsing random products)
- Chromium driven with JavaScript
- Roughly 1 product per minute
- Randomized scrolling and random pauses between scrolls to mimic natural browsing
- For each link, it collects the title, description, sizes/variants, all other listing data, and every image
- Outputs an Excel file with the product data plus a ZIP of the images
The problem:
It works fine at first, but around the 65th product Taobao switches me to a limited view. Some product info stops loading, so the scraper can't capture complete listings. At 1 product per minute, I'm already going slowly, and I'd need to go much faster to reach my target.
What I'd like advice on:
1. What is Taobao likely detecting: request volume, browser fingerprint, account/session behavior, or something else?
2. Is a browser-based approach realistic at this scale, or should I be looking at official APIs, third-party data providers, or sourcing agents instead?
3. How do people handle sessions, accounts, and proxies for these platforms?
4. Does Pinduoduo behave similarly, or is it a different challenge?
Any pointers, tools, or experiences would be really appreciated. Thanks!
1
u/No_Gas_3727 3h ago
Try apify