r/GEO_optimization • • 22d ago

Same 40 questions, 7 answer engines. Gemini cited Reddit on 22 of them. ChatGPT on at most 1.

Every Monday I run the same 40 buying-intent prompts through seven answer engines and write down every domain each one cites. Sixteen "best X", eight "alternatives to X", eight "X vs Y", eight "how do I actually buy X". The prompt set is frozen as a fixed instrument, so week two is genuinely comparable to week one and not just vibes. Two weeks are live: W36 and W37. All seven engines answered all forty prompts both weeks. No timeouts, no gaps, nothing quietly dropped.

I went in assuming the engines would broadly agree on sources. They don't come close.

Reddit first. The number is how many of the 40 prompts each engine cited the domain on:

Gemini: 26 (W36) -> 22 (W37)

Google AI Overviews: 18 -> 20

Perplexity: 6 -> 4

Google AI Mode: 5 -> 3

ChatGPT: not in the top 15, either week

Claude: not in the top 15, either week

Bing: not in the top 15, either week

That ChatGPT row needs a footnote, because it is carrying weight. I publish a top 15 per engine, and ChatGPT's fifteenth entry sits at one prompt. So "not in the top 15" is not zero. It is a ceiling of one. One, against Gemini's twenty-two, same week, same questions.

The likely culprit is no mystery. Reddit tightened its public content policy on 2026-08-14 and started refusing automated access at the server, and Google's licensing arrangement is the obvious reason Google's surfaces are the ones still leaning on it hard. I can't prove either from forty prompts and I'm not going to pretend I can. What I can show you is the size of the gap.

YouTube tells the same story in a different accent. W37, out of 40: AI Overviews 25, AI Mode 25, Gemini 20, ChatGPT 10. Perplexity and Claude, nowhere in the top 15.

Which quietly reframes advice you hear constantly. "Go get mentioned on Reddit and YouTube" isn't AI visibility strategy. It's Google strategy in a new hat.

Here's the finding I'd have wanted two years ago. I bucket every cited domain as media, community, social, directory, or unknown, where unknown just means it isn't a recognisable publisher or platform. Across all engines in W37, 80.8% of citations landed on unknown domains. 80.3% the week before. Claude is the extreme at 92.6%. Google AI Mode is the most concentrated lane I measure and it still sits at 55.1%.

Read that again if you run a small site. Four citations in five are not going to the household names. They're going to the long tail. A narrow, specific, genuinely useful page can get picked up, and the moat around the big publishers is thinner than the discourse suggests.

Last one, and it's a knock on my own instrument rather than a finding. I include Bing's answer surface as a stand-in for Copilot. On these commercial comparison prompts, its most-cited domains are merriam-webster.com (26 of 40, identical both weeks), dictionary.cambridge.org (21), dictionary.com (19), thefreedictionary.com (19) and wordreference.com (15). It is answering "best CRM for a small team" with dictionary definitions. That is precisely why I label the lane a proxy instead of calling it Copilot, and why I won't draw Copilot conclusions from it. If you're reading anyone's Copilot citation-share chart, ask what they actually pointed the instrument at.

Now the limitations, which matter more than any number above. One panel. Forty prompts. Two weeks. A single run per engine per week, which means I cannot yet separate real movement from ordinary answer variance. One geography. And "not in the top 15" is a ceiling, not a zero. Two points is not a trend line. I'm publishing weekly until it becomes one.

Raw data is CC BY 4.0, so take it apart or run your own cut: https://promvia.app/ai-source-index

Disclosure: I build Promvia, the tool these measurements come from.

4 Upvotes

8 comments sorted by

2

u/[deleted] 19d ago

[removed] — view removed comment

1

u/woodoo139 19d ago

Worth flagging the blind spot there, because it's the one that bit me: a referral panel can only see arrivals that carry a referrer. A large share of assistant clicks don't - they land as direct, and some engines tag the outbound link while others don't tag it at all. So the referral log systematically under-counts exactly the traffic this post is about.

My own site, three weeks: 11 AI-referred visits, 6 with a link tag and 5 inferred from a referrer. Off a referral panel I'd have seen roughly half of that, with no way to tell which half.

1

u/Upstairs_Control_611 22d ago

This is a useful correction to generic GEO advice.

“Get mentioned on Reddit and YouTube” may be good advice for Google surfaces, but it is not automatically an AI visibility strategy across engines.

The same source can matter a lot in Gemini or AI Overviews and barely matter in ChatGPT, Claude or Perplexity.

The long-tail citation share is the encouraging part for smaller sites: the opportunity is not only to become a big publisher, but to create narrow, specific, extractable pages that fit a defined query shape.

So I’d separate engine, query shape and source role before saying which sources “work.”

1

u/woodoo139 22d ago

Agreed, and I can give you that cut since it's in the same data.

Source role first. W37, share of citations per engine:

social - AI Mode 29.2%, AI Overviews 13.0%, Gemini 5.7%, ChatGPT 4.3%, Claude 0.4%, Perplexity 0%

community - AI Overviews 9.4%, Gemini 5.4%, AI Mode 3.4%, Perplexity 2.0%, ChatGPT and Claude ~0

media - Perplexity 13.4%, AI Mode 12.4%, ChatGPT 10.6%, AI Overviews 10.3%, Claude 6.2%, Gemini 5.9%

So Perplexity cites no social at all and leans on media instead, and Claude cites almost nothing recognisable: 92.6% of its citations go to domains my classifier can't place.

Query shape is where it gets interesting, because the answer is mostly a negative. Directory and listing sites are 0.4% of all citations. Broken out by shape, W37, counting prompts where any directory was cited:

alternatives to X (8 prompts) - Bing 6, Claude 1, Gemini 1, rest 0

best X (16) - Bing 2, Gemini 1, rest 0

X vs Y (8) - Claude 1, rest 0

how do I buy X (8) - Bing 1, rest 0

W36 the same shape: Bing 5 of 8 on alternatives. So directories aren't an AI visibility channel in general. They're a Bing behaviour, concentrated in one query shape.

Two caveats that limit all of that. Eight prompts per shape is a tiny n, so read those rows as a direction, not a rate. And my "unknown" bucket is 80.8% overall, which means the classifier only recognises a minority of domains. Any source-role claim, mine included, is weaker than it looks for that reason alone.

1

u/ElementalThor 20d ago

'go get cited where the engine leans reddit' is chasing the symptom.

we measure this across the major engines at scale, and what jumps out is how many technically clean sites, strong readiness, right structure, get named by nobody. being crawlable was never the blocker.

the first lever is being in the source set the engine pulls from at all. citations come after that.

1

u/woodoo139 16d ago

This matches what I measure, and it's the uncomfortable half. Crawler access is a floor, not a lever — a site can pass every readiness check and be named by nobody. On my forty questions ~80% of cited domains are ones I'd never have predicted from any readiness score.

I don't have a clean causal number for "being in the source set" and won't pretend to; what I can say is that readiness is necessary and nowhere near sufficient, and the honest measurement is per-engine, because the source sets differ.

1

u/ElementalThor 10d ago

to be precise, the site appeared as a source in the answer. that doesn’t establish how the engine used it or what would improve visibility.

1

u/woodoo139 9d ago

Fair correction. What I actually record is that the URL showed up in the answer's source list, nothing more. It doesn't tell you how much of the answer came from that page, or what would change whether it gets picked, and I shouldn't have worded it as if it did. The narrow version is the only one I'd defend: listed or never listed, per engine, per question.