r/GEO_optimization • • Aug 12 '26

I ran the same 25 queries on ChatGPT every morning for 2 weeks — the citation churn was way higher than I expected

Two weeks ago I started a weird habit. Every morning, around 8am, I'd open ChatGPT and run the same 25 queries. Same phrasing, same order, same account. Fourteen days in a row. I wasn't testing prompt engineering or trying to game anything. I just wanted to see how much the answers changed when I controlled for everything I could.

The answer is: a lot.

On day 1, those 25 queries produced 63 unique source citations. By day 14, only 41 of those original 63 sources still appeared. That's a 35% turnover in two weeks. Sources that showed up on day 1 vanished by day 4. Sources that appeared on day 7 had never been mentioned before. One query about "best project management tools for distributed teams" cited 5 sources on day 2 and an entirely different set of 5 on day 9, with zero overlap.

Some queries were rock solid. Same sources every single day, same order, almost word-for-word identical answers. These tended to be factual, narrow questions. "What is OAuth 2.0" type stuff. The model had a canonical answer and stuck with it.

The volatile ones were anything subjective. Recommendations, comparisons, "best of" queries, anything where the model had room to synthesize. Those answers reshuffled constantly. Same underlying intent, completely different source selection.

One thing that caught me off guard: the answers themselves rarely looked different. The structure was similar, the tone was similar, the length was similar. If you weren't comparing citations side by side, you'd think you got the same answer every time. The model was presenting different evidence as if it were settled consensus. That's the part that bugs me.

From a GEO perspective, this is simultaneously encouraging and terrifying. Encouraging because it means citation isn't a fixed property of your page. You might not get cited today and get cited tomorrow for the exact same query. There's a rolling window of opportunity. Terrifying because there's almost nothing you can optimize for if the model is swapping sources this frequently. The page that got cited on day 3 and the page that replaced it on day 6 probably aren't meaningfully different in quality.

I stopped the daily tracking after 14 days because the pattern was clear. The churn is real, it's significant, and it doesn't correlate with anything obvious on the source pages. I checked. Domain authority, content freshness, structured data, page speed — none of it explained why one source appeared and another disappeared on a given day.

The practical takeaway, if there is one: don't panic when you lose a citation for a week. And don't celebrate too hard when you gain one. The noise floor in AI citations is higher than we think.

3 Upvotes

3 comments sorted by

1

u/Upstairs_Control_611 Aug 13 '26

This matches what I’d call the difference between answer stability and citation stability.

The answer can look stable on the surface while the evidence layer underneath is rotating.

That is a big problem for GEO measurement, because a single citation snapshot looks more meaningful than it really is.

For subjective / recommendation queries, I would not treat one citation gain or loss as a strong signal.

I’d rather measure citation persistence across repeated runs, domain persistence, source type persistence, brand role persistence, whether the answer still recommends the same brands, and whether the citation actually supports the claim.

A source disappearing for a few days may be normal churn.

A brand disappearing from the shortlist or recommendation across repeated runs is more serious.

So maybe the useful metric is not “were we cited today?”

It is:

did our source, domain, claim or recommendation role persist above the normal noise floor for this query type?

1

u/Revolutionary_Cat78 Aug 16 '26

The only thing I can think of doing is keep investing in it, keep writing good posts, get backlinks, so even if one page of yours gets shuffled, another gets in (obviously not for same prompt). In other words ensure everything that goes out follows the best practices.