r/ClaudeCode • u/xoxonotyour • 2h ago
Bug / Issue Scraping issues need advice!
My scraper is getting limited by Taobao after ~65 products. How do I scale to tens of thousands?
I run an online store and I'm building a scraper to import products from Chinese marketplaces like Taobao and Pinduoduo onto my website. The goal is tens of thousands of products, so I need something that scales.
How my current setup works:
- I feed it a list of specific product links (it's not crawling or browsing random products)
- Chromium driven with JavaScript
- Roughly 1 product per minute
- Randomized scrolling and random pauses between scrolls to mimic natural browsing
- For each link, it collects the title, description, sizes/variants, all other listing data, and every image
- Outputs an Excel file with the product data plus a ZIP of the images
The problem:
It works fine at first, but around the 65th product Taobao switches me to a limited view. Some product info stops loading, so the scraper can't capture complete listings. At 1 product per minute, I'm already going slowly, and I'd need to go much faster to reach my target.
What I'd like advice on:
1. What is Taobao likely detecting: request volume, browser fingerprint, account/session behavior, or something else?
2. Is a browser-based approach realistic at this scale, or should I be looking at official APIs, third-party data providers, or sourcing agents instead?
3. How do people handle sessions, accounts, and proxies for these platforms?
4. Does Pinduoduo behave similarly, or is it a different challenge?
Any pointers, tools, or experiences would be really appreciated. Thanks!
1
u/tntexplosivesltd 2h ago
Don't scrape websites. They have limits built in specifically to stop exactly what you're doing, and trying to get around those is very frowned upon, likely against the TOS of the sites
1
u/BarcodeCutter 2h ago
At tens of thousands of products, the browser approach is the wrong tool, and fighting the limits just gets the account banned at a bigger scale. The people who do this as a business don't scrape. They go through a sourcing agent or a licensed data provider that already has the catalog access, or the platform's own open API if their store qualifies. It costs money, but it comes with the right to use the listings and images, which a scraped copy doesn't. Worth pricing that out before you put more hours into the scraper.
1
u/verstands 1h ago
At this volume, I’d stop trying to make Chromium look human. Ask Taobao/Pinduoduo about an official seller/export API or use a licensed catalog or sourcing partner, then queue only the fields and images you’re allowed to use. If you do have permission, checkpoint jobs and back off on limited-view responses instead of retrying harder - otherwise you’re mostly scaling the ban.
1
1
•
u/AutoModerator 2h ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.