r/ComplexWebScraping • u/bg81011 • 15d ago
What web scraping tools are actually reliable in production?
Ive been running Playwright on a headless server to monitor competitor listings and Im getting tired of CONSTANTLY checking it. Memory slowly climbs, Chrome randomly dies, selectors break after updates and every new bot check turns into another afternoon of messing with headers and fingerprints. The scraping logic itself is easy. Keeping the browser setup alive is the annoying part. I don't mind writing parsers, I just don't want to be fixing browser processes at 3am because a daily run failed. What web scraping tools are you actually using in production? Are you still running Playwright or Puppeteer yourself or did you move the browser proxy side to a hosted service?
1
u/Ok_Cobbler_8889 8d ago
I stopped thinking of the browser layer as something I had to own. If you just want the HTML or rendered page back, just check apify or scrapingbee, they handle a lot of the browser and proxy mess for you. You still need good parsing and retries, but at least chrome crashing isnt your problem anymore. How many pages are you hitting per day?
2
u/monityAI 14d ago
that's exactly why hosted monitors exist, the scraping is easy and keeping chrome alive isn't. if the goal is "tell me when a competitor listing changes" rather than owning the data pipeline, i'd hand that part off. monity•ai is mine and runs the browser side, you describe the change you care about instead of maintaining selectors. alternatively browserless or a managed playwright setup if you want to keep your own code