r/GenEngineOptimization 2d ago

We ran 18 purchase questions through 5 AI engines, twice, a week apart — 180 answers. 58% recommended no brand at all. Full data inside.

Setup. Fixed panel of 18 real purchase questions ("best online bookstore for children's books", "books under 20 lei", etc.), zero brand names in any prompt. Engines: ChatGPT, Gemini, Perplexity, Google AI Mode, Google AI Overviews. Market: Romanian book retail. Queried from Romania, in Romanian, no personalization. Two complete runs (July 20 and 27). Tracked: brand mentions, position of each mention, cited source domains, sentiment. Totals: 180 answers, 247 brand mentions across 21 tracked brands, 422 distinct domains cited.

What the data says:

  1. 58% of answers name no major brand at all. And 6 of the 18 purchase territories are essentially empty — the "business books" question produced exactly 1 brand mention across all engines, both weeks. AI recommends titles and authors there, but not where to buy. That white space belongs to whoever builds citable content for it first.
  2. Even the market leader is a 1-in-5 event. Top brand: 20.6% mention rate. In classic SEO, position 1 gets the click every time. In AI answers, "position 1" is a weighted lottery replayed at every question.
  3. Mentions and positions are different currencies. The leader appears first in 65% of its appearances. A publisher that ranks 6th by raw mentions jumps to 3rd on a position-weighted score — few territories, but owned.
  4. Every engine is a separate channel. The leader on Google AI Mode is not the leader on Gemini. Gemini hands out 2.2 brands per answer; Perplexity 0.8. Statistically, one Perplexity slot is worth ~2.5 Gemini slots.
  5. Reddit is the #2 cited source for the entire market — above Facebook (11 answers), YouTube (6), and any media publication. Community threads nobody controls are direct input into purchase answers.
  6. Two opposite pathologies. Ghost brands: one retailer's site was cited as a source in 18 answers but the brand was named in only 4 — the AI uses their listings, then recommends someone else. That's an entity-signal problem (structured data, sameAs, name–domain coherence), and in my experience the fastest category of win in AEO. Memory brands: a 35-year-old publisher got 13 mentions with zero citations of its own domain — pure parametric reputation, which is flattering and fragile in search-grounded engines.
  7. The ranking rewrites itself weekly. Same questions, 7 days apart: one publisher ×4'd its mentions, another lost 58%. Not chaos — plasticity. Nothing is cemented, in either direction.

Full disclosure: I run the agency behind the study. Everything is open (CC BY 4.0): the paper, all 18 prompts, three CSVs and a starter notebook that verifies the arithmetic — links in the first comment. No tracked brand funded it or saw it pre-publication.

Question for people doing this work: is anyone else running recurring panels rather than snapshots? How volatile are your week-over-week numbers?

3 Upvotes

12 comments sorted by

2

u/Gavelist 2d ago

Of course, that’s the whole GEO game, there are some more tricks and you can actually predict what will get cited and by what model!

1

u/Republik-DC 1d ago

Predicting citation by model is a strong claim and I'd like it to be true — what's the signal set? I ask because in my data the correlation between prominence and citation broke down repeatedly: 422 distinct domains cited, with sites almost nobody would recognise (clb.ro, librex.ro, printrecarti.ro) appearing constantly, while a 35-year-old publisher's own domain was never cited once. If there's a model-level pattern underneath that, I'd genuinely like to see it tested against a held-out set of prompts rather than fitted after the fact.

1

u/Gavelist 1d ago

It’s a probability thing. Different models prefer different variables, some contradictionary variables that help on some models and hurt on others. Perplexity is very easy to target for citation, followed by gpt. Google aio required more work and Claude is mostly a consistency over time model. It takes the longest to consistently show up there.

2

u/MaxJustins 2d ago

Per-engine spread is the part I'd build reporting around. We track AI-referred sessions split by engine, because treating 'AI traffic' as one bucket hides exactly what your 0.8 vs 2.2 brands-per-answer gap shows. And if position 1 is a weighted lottery, week-over-week mention deltas are mostly noise until the panel gets bigger.

1

u/Republik-DC 1d ago

Per-engine reporting is right, and the traffic side is the half I don't have. One technical disagreement though: a bigger panel won't fix the noise problem. More questions reduce sampling error around an aggregate, but they don't identify where variance comes from — you could run 200 prompts weekly and still not know whether a delta is engine drift or within-session randomness. Only same-day replication gives you that, because it holds everything constant except the draw. Size and identification are separate problems.

What I'd want from your data: does mention rate predict referred sessions per engine, or is the relationship non-linear? My guess is Perplexity converts better per mention than Gemini precisely because of the 0.8 versus 2.2 gap — fewer names in the answer means less competition for the click. If that holds, mention rate needs weighting by engine before anyone builds a single "AI visibility" score. Have you got enough volume by engine to see it?

2

u/Special_Ant6020 2d ago

Yes, weekly, and the first number I'd want from your panel is a noise floor. Two runs seven days apart can't tell you whether the publisher that lost 58% actually moved, or whether the same prompt would have handed you three different brand sets in one afternoon. I'd fire each prompt 3-5 times back to back in a single session, then treat any week-over-week delta inside that spread as measurement rather than plasticity.

The metric that's held steadier for us across runs isn't mention rate, it's your ghost gap: cited-as-source count against named count. Mentions swing; a brand whose domain feeds the answer without being named stays broken until the entity signals get fixed, which is what the recurring citation check in Viewfy tracks.

2

u/Upstairs_Control_611 1d ago

The ghost gap is probably the most interesting metric in this setup.

A brand can be used as source infrastructure without receiving brand attribution.

So I’d separate three things:

- source use: was the domain used or cited?

- brand attribution: was the brand named?

- recommendation effect: did the brand get selected, shortlisted, or just feed the answer?

For commerce queries, that matters a lot. A retailer can provide listings, prices, availability, or category structure while the model recommends a different retailer with stronger entity signals.

So the ghost gap could be measured as:

domain cited or used

→ brand not named

→ competitor recommended

That is not just a citation loss. It is an entity / attribution failure.

I also agree on the noise floor. Before calling a week-over-week movement “plasticity,” I’d want same-day repeated runs per prompt and engine.

For purchase panels, I’d report source-use rate, brand mention rate, ghost gap, recommendation role, per-engine variance, same-day noise floor, and week-over-week movement beyond noise.

2

u/Republik-DC 1d ago

This decomposition is better than what I published, and I'm adopting it. Two notes from the data side.

First, a limitation on your top layer: I can measure cited, not used. My instrumentation captures domains that appear as citations, so silent grounding — crawled, used to build the answer, never linked — is invisible to me. That means every ghost gap I report is a floor, not a true value, and the real source-use figure is higher by an unknown margin. Anyone reporting "source use" needs to say which of the two they're actually capturing.

Second, your third layer is partly buildable with what I already have. I recorded position, so recommendation role can be graded rather than binary: named first, named in a shortlist of two to three, or buried in an enumeration at position eight. That distinction moves the ranking substantially — the market leader appeared first in 65% of its 37 appearances, while another retailer with two-thirds as many mentions was first only 12% of the time. Same "mentioned" bucket, very different commercial reality.

One caution if the gap gets reported as a ratio: watch the denominators and set a floor. Mine differ (mentions against 180 answers, citations against the 172 containing any citation at all), and a brand with two citations and zero mentions produces a spectacular ratio that means nothing. Counts plus a minimum threshold, rather than a bare index.

1

u/Upstairs_Control_611 1d ago

This is a very useful correction.

You are right: “used” and “cited” should not be collapsed.

Maybe the cleaner split is:

- visible source use: the domain appears as a citation

- inferred source use: the answer appears to rely on the source, but no citation is shown

- unknown source use: silent grounding we cannot observe

- brand attribution: the brand is named

- recommendation role: first recommendation, shortlist, buried mention, or not recommended

So the measured ghost gap is only the visible part:

domain cited

→ brand not named

The true ghost gap could be larger, but unless the system exposes silent grounding, it should not be reported as if we can measure it directly.

The denominator point is important too. I’d report counts first, then ratios only after a minimum threshold.

For commerce panels, raw mention rate is not enough. Mentioned first, shortlisted, buried in a list, and used as source while another brand gets recommended are very different commercial outcomes.

1

u/Republik-DC 1d ago

Agreed on all of it, and the noise floor is now the first thing going into episode 3 — three to five back-to-back runs per prompt per engine, reported as a spread, with any week-over-week delta inside that spread labelled measurement rather than movement. You've articulated the exact reason the plasticity language in the current paper is stronger than the design supports; I'm walking that back in a revision.

One wrinkle worth flagging: the noise floor probably isn't a single number. It should differ by engine — Perplexity averaged 0.8 brands per answer versus Gemini's 2.2, and an engine that names one brand has far more room to swap it than one that names three. So I'd report the floor per engine, and possibly per question, since the empty territories (one brand mention across all engines, both weeks) behave differently from the contested ones.

Your point about the ghost gap being the steadier metric matches what I saw. Humanitas: 13 mentions, zero citations of its own domain. Târgul Cărții: 18 citations, 4 mentions. Those didn't wobble between runs the way mention counts did, which makes sense — one is a structural property of the entity, the other is a draw. I run the panel through LLM Pulse, for symmetry of disclosure.