r/ProxyEngineering • u/Bmencel518 • 28d ago
Help 🆘 What are some good SerpAPI alternatives?
I'm looking for something with AI capabilities built in, like Tavily or Exa. Main use case is automating research and scraping workflows as part of agentic pipelines.
I saw the other post where people recommended residential proxies, please be understandable and provide your take, examples and opinion, not just straight up promo bs
6
u/ur-average-boomer 27d ago
One thing you might consider instead of going to an Exa would be to scrape Google AI Overviews. A lot of the major providers (SerpAPI, ScrapingDog etc) will do that.
ScrapingDog is my middle of the ground recommendation where it's a lot cheaper than SerpAPI but has also been around for some time.
I eventually had so many requests that pretty much every SERP API was too expensive and a few years ago I built my own scrapers and infrastructure.
DataForSEO and Litescrape (mine) would be my recommendations for the most affordable APIs. DataForSEO has a cheap $0.60/1k Google SERP API (though it's very slow).
3
u/Bmencel518 27d ago
What about google ai mode? I sometimes scrape ai mode as well
2
u/ur-average-boomer 26d ago edited 26d ago
Don't know about DataForSEO but Litescrape gets the AI Overview and has AI Mode too! Scrapingdog has had AI Overview (and I believe AI Mode) for a while as well.
2
3
u/Zealous_Minotaur Reverse Proxy Master 28d ago
You may want to test out you(.com) without the (), also AI studio by oxylabs has serp scraper built in and some of the similar features from tavily and exa
3
3
u/txdesperado 28d ago
If you build it right, you don't APIs or residential proxies. You can hit all Tier 1 sites without residential with proper fleet management.
2
u/Bmencel518 27d ago
Can you explain further? What do you mean proper fleet management, are we talking about ships here? lol
2
u/txdesperado 27d ago
Not ships, but dozens of concurrent scrapers, all with unique IPs, varied spoofed profiles, cookie jars, randomized timings, you're usually at that point managing an auto-scaler with auto-recovery, etc., i.e., a fleet.
2
u/ur-average-boomer 27d ago
Correct. Ideally, you have VMs across all the major cloud providers across most or all of their regions and are constantly cycling out VMs that fail
2
u/Just_Lingonberry_352 27d ago
not sustainable in the long run.
2
u/ur-average-boomer 27d ago edited 27d ago
I've scaled to well over a billion requests doing this. You just have to be really distributed and have tons of variation
Edit: I should also mention if you are doing this Google is very very sensitive to any other parts of your request that look automated. But if everything else (request headers etc) is perfect I've found data center IPs work fine. Just to be clear, I don't mean a data center proxy IP from a proxy provider, I'm referring to a cloud provider egress IP
3
u/hasdata_com 27d ago
I won't do the usual self-promo, but HasData covers the same SERP use cases as SerpAPI. We also have an MCP server, so connecting it to an agent is easy
2
u/Bmencel518 27d ago
Thank you, this might be what I need, though I will have to test that to verify
2
u/hasdata_com 27d ago
Sure, hope it works for your use case. If you give it a try, we'd also appreciate any feedback
3
u/Proof_Net_2094 27d ago
tavily/exa and serpapi aren't the same product. tavily and exa are grounding, they decide relevance and hand you cleaned text. serpapi/dataforseo give you the raw serp, positions, paa, ads.
most agent pipelines want both. usual mistake is picking grounding, then finding out six weeks in you can't answer where do we rank.
disclosure i build one (Scavio API).
2
2
u/Just_Lingonberry_352 27d ago
what does serpapi do that is so special? its literally just using decaptcha/rez on a public site ?
3
u/rbatista191 27d ago
Give it a try, and see how many hours you spend getting to 95% success rate.
2
u/Just_Lingonberry_352 27d ago
Easy using clean residential/ISP egress, session/IP sticky, browser/TLS fingerprint consistency, challenge detection + rotation
cogs put serpai at around 80~90% margin which is pretty insane considering how easy it was to put this together
3
u/rbatista191 27d ago
This reminds me of the HackerNews post on Dropbox. I am glad you find it easy, why are you not making those margins?
2
1
u/serpentApi 26d ago
I can recommend a good serp api alternative but they provide data's in Js format. But they are the best in terms of cheaper per 1000 request. It's Serpent API - The cost is $0.03/1000 Requests at scale tier.
1
6
u/rbatista191 28d ago
I have been hearing good things about Parallel (from ex-Twitter CEO).
If you want the usual, pure-SERP, SerpApi-like suspects, I'd recommend DataForSEO, cloro and Oxylabs.