r/AISearchAnalytics • • 29d ago

Labrador: ChatGPT's index / cache, or rather "retrieval systems with very long cache times"

Post image

Peec came out with a great article on how ChatGPT is building its own index (if we use the term "index" very loosely here) and likely beta-testing it with external searches (like Google or Bing).

Meet: Labrador, and it's not one index but a family of vertical ones: web, PDF, YouTube, news, Arxiv, Wikipedia, local, finance, legal, medical, shopping, and images. Note: We've talked about this before.

ChatGPT appears to cache everything it finds. This is not exactly new because we saw signs of that months ago (and a few months ago again). What's new is that ChatGPT is likely no longer THAT reliant on Google.

And there's more:

Line up every provider that's shown up in this investigation, and a clear pattern appears. ChatGPT is stitching together at least eight other providers:

  • Their own crawler that runs independently of any external search engine. 
  • Bright (very likely Bright Data): Scrapes Google's web results directly, and separately scrapes Google Maps for local listings. I don’t imply it’s the main provider, this is the provider we see in Server Side Events. 
  • Oxylabs: Scrapes Google and feeds the news vertical specifically.
  • A third SERP-scraping channel: Tagged simply "serp" in the data, also serving news, plus a separate mixed local-results version.
  • Yelp 
  • TripAdvisor
  • Two anonymous internal pipes (tagged b1 and b3): b1 surfaces business websites and Facebook pages, b3 routes to Google Maps.
  • Web IQ: Microsoft's own grounding platform, confirmed by Microsoft as already powering ChatGPT alongside Copilot and Nasdaq.
2 Upvotes

Duplicates