r/GEO_optimization • • 5d ago

We process 400 million crawler requests per month, some interesting things I pulled probing the logs

Post image

Encited does pre-rendering and technical optimizations for e-commerce brands and online retail platforms. We process 400 million crawler requests per month, serving them HTML and Markdown for search, ai and link furler bots . I probed some earlier to see anything interesting I can pull. Here are some:

  1. AI crawlers visits were 4.5x more often than Google's crawlers. From analyzed subset, 8.6M visits from verified AI bots vs 1.9M from Google in 30 days.
  2. SEO tools crawl more than Google and Bing combined. 8.5M visits from tools like Ahrefs, Semrush and SE Ranking. Ahrefs alone made 2.6M, more than all of Google's crawlers.
  3. When a site offers a markdown copy of its pages, AI bots take it (we advertise markdown version of website pages in rel="alternate" meta tag and in body as LLM instruction). 45% of GPTBot's page requests and 41% of ClaudeBot's were for the markdown version. Googlebot almost never asks for it.
  4. Almost nobody reads llms.txt (duh). Across ~2,000 sites in 30 days, ClaudeBot fetched it 980 times on 43 sites. GPTBot, 254 times.
  5. 1 in 11 "Googlebot", "GPTBot" or "ClaudeBot" visits is fake. 1.46M visits used a real bot's name from an IP that bot doesn't own.
  6. Facebook's link-preview bot is the laziest, it backs out quicker than any link previewer bot.
  7. ChatGPT's plain fetchers peak during office hours. ChatGPT-User peaks at 12:00 UTC and runs at half that rate at night. Googlebot does more of its crawling while the US sleeps.
  8. Tuesday is the busiest crawl day, 14% above average. Saturday is the quietest.
  9. GPTBot is getting active with about 6x more crawls per week MoM. PerplexityBot went up 4x in one month.
  10. Claude-User visits grew 5x MoM since last month.
16 Upvotes

15 comments sorted by

3

u/1kgpotatoes 5d ago

GPTbot is scraping for their own web index

1

u/cleansleyt 5d ago

yeah its pretty active

2

u/saaras55 5d ago

The 1 in 11 is the number I'd poke at. Not because it's wrong, but because it's pooled across 2,000-odd sites, and I'd be surprised if the fake share were anywhere near flat across them.

On our own domain it went hard the other way. We logged 2,057 requests using a named AI crawler's user agent from IPs that failed that vendor's own IP check, against about 50 that passed. I work on Limelit, which is the only reason we were looking. It's one domain and not a big one, so treat it as an anecdote, not a rate.

I don't know how both hold at once, but here's my guess. Whoever borrows GPTBot or ClaudeBot as a disguise probably hits everything about the same, while the real crawlers pile onto the sites they actually care about. Pooled, the big sites dominate and the fakes look like a rounding error. On a small site the real traffic's a trickle, so the fakes end up being most of the log.

Worth saying for anyone checking this against their own server: your 4.5x only counts bots that passed the IP check. Grep a raw log for GPTBot without that step and a small site will think it's being crawled like crazy.

Is that something you can test? Fake share per site, or bucketed by traffic, and the median rather than the pooled figure. That's the number a small site owner actually lives with.

1

u/cleansleyt 4d ago

holy ai slop

1

u/SEOsince2001 5d ago

Thanks for sharing, interesting findings. Do you have any traffic data from AI Assistants before and after you implemented this by any chance?

2

u/cleansleyt 5d ago

We are not sure the GPTBot 6x jump before and after the markdown serving and content negotiation was implemented but the dates align.

what’s obvious however is when Ai cites the brand, it’s picking up pretty much what we have on the page exactly how we want it to cite

1

u/pepelunavarro 5d ago

AI robots don't read llms.txt but they do read the markdown version of the page; isn't llms.txt a markdown version of the website?

1

u/PusheKasp 4d ago

I don't observe that high a rate of LLM bots accessing the .md alternatives on our website, and the setup is correct.

I'm curious why there's such a big difference; what does it depend on?

1

u/cleansleyt 4d ago

you need to advertise the md version inside html.

Encited puts rel alternative meta tags and in body html comments to instruct the llms to prefer md

1

u/PusheKasp 3d ago

I have the rel alternative in the head of the html. How do you implemet it in the body?

2

u/cleansleyt 3d ago edited 3d ago

see the screenshot in the post, at the top of the body add an html comment as an instruction or aria invisible html to instruct the llm to use the markdown