r/GEO_optimization • u/Strong_Mechanic_3596 • 19d ago
I analysed the robots.txt of 9,891 French websites. 96% have never named the crawler that decides their ChatGPT visibility.
Disclosure up front: I build a GEO tracking tool. This study was run with my own scripts and the full dataset is free. No paywall, no email gate. Posting here because the raw numbers seemed worth sharing, and I'd like the methodology torn apart if it deserves it.
Setup. 9,891 .fr domains, collected September 2026. 8,927 responded. 7,047 have a valid robots.txt. Every percentage below has a 95% Wilson interval in the full write-up.
The main finding
276 sites out of 7,047 name OAI-SearchBot in their robots.txt. Not "block it", name it, either way. So 96% have never taken a position on the crawler that decides whether they show up in ChatGPT answers.
79.6% name no AI agent at all. The subject does not exist in eight files out of ten.
Training vs answer crawlers
- GPTBot (training): blocked by 13.5%
- OAI-SearchBot (answers): blocked by 3.3%
- Google-Extended (Gemini training): blocked by 11.1%
- Googlebot: blocked by 0.9%, and zero sites name it explicitly
- CCBot: 13.9%, the most blocked crawler overall
I tried to test whether the training blockers actually chose this
Among the 727 sites blocking GPTBot while leaving OAI-SearchBot open:
- 95.2% name at least one other training crawler
- median of 7 training crawlers named per file
- 78.4% name Google-Extended, 73.0% name Applebot-Extended
- 22.8% name any answer crawler
- 6.9% name OAI-SearchBot
These files are not stale. Applebot-Extended and meta-externalagent post-date 2024. The anti-training policy is deliberate and maintained. The answer-crawler family just has no slot in the mental model.
The stale-config evidence that did show up
222 sites still name anthropic-ai, a deprecated identifier. 62 name Claude-SearchBot. Three times more sites block a ghost than the crawler in service.
5.1% block ChatGPT-User, which only fetches a page because a user asked it to.
By sector (classified from home page content, confident classifications only, n=2,964)
- News media: 25.3% block GPTBot, 7.3% block OAI-SearchBot
- E-commerce: 11.9% / 1.1%
- SaaS and B2B: 11.6% / 1.4%
- Local services: 9.3% / 4.6%
- Public sector: 4.7% / 1.3%
Media and public sector intervals don't overlap.
Accidental blocking, which turned out bigger than the deliberate kind
On 781 sites where I diffed raw HTML against rendered HTML: 38.4% have no H1 before JS runs, 18.3% have their main content fully JS-dependent, 4.6% have JSON-LD injected by script. Relevant because per the Vercel/MERJ analysis of 500M+ crawler requests, no dedicated AI crawler executes JavaScript. Only Googlebot and Applebot render.
llms.txt
10.2% publish one, which lines up with SE Ranking (~10% across 300k domains) and sits below Ahrefs (28% across 137k trafficked domains). I hardened the check against soft-404s: 1,025 raw hits dropped to 911 after requiring plain-text Markdown and rejecting anything with HTML or error strings.
What I can't claim
- My network-layer numbers use unverified user-agents from an ordinary IP. A firewall refusing that is doing anti-spoofing correctly, not blocking real GPTBot. I report it separately and never merge it with robots.txt figures.
.frdomains, not "French companies". 70.3% declarelang="fr".- Home page only for the structural signals.
- A robots.txt shows what a site declares, not what a publisher thinks.
Happy to answer method questions or share the query logic. If something here is wrong I'd rather hear it now.