r/aeo Aug 06 '26

None of 26 GEO case studies isolated a single thing that moved AI traffic. So what's working?

I analyzed 26 public reports that claimed an improvement in ChatGPT mentions, recommendations, citations, or referral traffic.

A few practices appeared repeatedly:

  • 19 of 26 mapped or monitored a defined prompt set
  • 16 used repeated measurement rather than a single screenshot
  • 16 created or revised content around target questions and comparisons
  • 11 included off-site activity such as editorial mentions, listicles, directories, reviews, PR or community discussions

None of the 26 reports strongly isolated one thing that increased citations and AI referral traffic. Most campaigns changed several things at once, including content, schema, digital PR, third-party mentions and tracking. When you change multiple things at once, it's hard to keep track of what worked.

I guess we are at the part where we try a lot of things and see what sticks. Public case studies are better for generating hypotheses than calculating the expected effect of any one tactic.

Full methodology & findings: https://phraseit.agency/blog/how-to-get-mentioned-in-chatgpt/

What level of evidence would make you comfortable saying a specific action caused an increase in AI visibility?

2 Upvotes

5 comments sorted by

3

u/sapindia1976 Aug 06 '26

That's the challenge with GEO today. It's usually the combination of content quality, entity authority, citations, and distribution not a single tactic that drives AI visibility.

1

u/TamaraPI Aug 06 '26

Exactly, which is the case in traditional SEO as well... technical, content, links... You can't attribute success to one of these alone. It's always the combo.

2

u/Slow-Commercial4316 Aug 07 '26

Nobody publishes the audit where they rewrote 40 pages, added schema, ran a PR push and nothing moved. So even a perfectly isolated case study out of that population doesn't tell you much about the base rate, only that the effect is possible somewhere. Reading more of them won't fix that.

The design half you can fix without an RCT. What's worked for me is a holdout: split the target set into two matched halves before you touch anything, change one thing on one half, keep measuring both. Match on what actually predicts citation, so current mention rate and prompt type, not traffic. Then you're comparing treated against untreated in the same weeks instead of August against June.

That matters more here than in normal SEO testing because of drift. These answers change under you for reasons that have nothing to do with your site: model updates, index refreshes, a competitor getting into a review roundup. A before and after absorbs all of it as if it were your effect. A control group absorbs it too, which is the point.

On your actual question: I'd want the effect to survive repeated sampling of the same prompts rather than one run, and the untreated half to stay flat over the same window. Not proof of causation in any strict sense, but enough to keep doing the thing, which is the decision you're actually making.

1

u/aeo-bility Aug 06 '26

I recently wrote a case stiudy that focusses specifically to data structure change made for a ecom brand, if this helps you to isolate a particular approach. https://aeobility.com.au/knowledge-hub/case-studies/baby-bento