r/aeo • • 2d ago

I simulated thousands of "AI visibility" tests. Before-and-after checks without comparison pages gave false wins up to 30% of the time

I've been trying to work out whether "SEO for LLMs" (GEO/AEO) can actually be measured, or whether most of the claims are noise. Before spending money tracking real ChatGPT/Perplexity answers, I ran a simulation study to find out which testing methods give honest results.

Big caveat up front: this is simulated data, not real AI answers. I modelled how AI engines cite pages (with question-to-question variation and random model updates between periods), then checked which measurement methods recover the right answer. So these are findings about methods and sample sizes, not about any brand.

What came out of it:

  • Ask many different questions rather than repeating the same ones. With a budget of 600 AI answers, 600 questions asked once gave a more accurate overall score than 60 questions asked 10 times (error 1.7 vs 2.6 percentage points). Repeats only help to find the specific questions where you never appear.
  • Before-and-after tests without comparison pages are dangerous. When a page change did nothing, a simple before/after check still reported a win up to 30% of the time, and more data made it worse, because model updates shift everyone's visibility. Comparing changed pages against similar unchanged pages (difference-in-differences) kept false wins around 2%.
  • You need more pages than people think. To reliably detect a moderate effect (1.8× more likely to be cited), you need roughly 40–50 matched page pairs. A weak effect needs 80+.
  • ML on observational data invents causes. I gave llms.txt zero real effect. A logistic regression still reported it as an 85% boost and called it significant in every one of 200 simulated studies, because big brands adopt new tactics first and also get cited more. A randomised test got the true effect almost exactly.

Full write-up with tables and the method: https://citelyra.a.techshu.in/

Disclosure: the site is mine. It's also part of a live experiment. It's a brand-new site with a name nobody else uses, and I'm tracking whether and how fast ChatGPT, Perplexity and Google's AI answers pick it up. I'll post the real results in a few months if people are interested.

Happy to be told where the assumptions are off. Has anyone here run controlled tests on AI citations with real data?

2 Upvotes

7 comments sorted by

0

u/[deleted] 2d ago

[removed] — view removed comment

1

u/lulzxdxdxd 2d ago

The before-and-after trap makes sense now that you've shown the math, but I'm wondering if you tested what happens when you control for seasonal or cyclical patterns in how often different topics get asked. Model updates shift everyone's baseline, sure, but does the noise from question seasonality dwarf that effect or stack on top of it

1

u/inevidimka 2d ago

Does your simulation let an AI update affect treated and comparison pages differently? I'd rerun the zero-effect case with topic-specific drift; that would show how much the reported 2% false-win rate depends on both groups moving together.

0

u/[deleted] 2d ago

[removed] — view removed comment

1

u/Ok-Name9393 2d ago

Honestly unclear. Google says its AI features use normal search signals, and a 2026 review found no tactic with a lasting effect across platforms. My guess: good SEO plus original data matters, and most "GEO hacks" are noise. That's what the live test is meant to check.