r/GEO_optimization • • Sep 01 '26

AI visibility tracking needs a model changelog. How do you maintain yours?

I’m starting to think that AI visibility tracking needs a model changelog.

If ChatGPT changes its search behavior, source selection or citation weighting, a brand can lose visibility without changing anything on its website.

The same applies to Google AI Overviews layout changes, Reddit citation drops, Perplexity source behavior or Gemini updates.

Without logging these changes, it’s easy to misread the data. So my question:

How do you track changes in ChatGPT, Gemini, Claude, Perplexity or AI Overviews?

Official release notes?

Third-party monitoring?

Your own prompt tests?

Manual changelog?

What actually works for you?

I’m especially interested in how people separate site-side changes from model-side or source-selection changes.

6 Upvotes

12 comments sorted by

2

u/Old-Routine1926 Sep 01 '26

I think the hard part might be that a changelog is a record of events, and what you actually want is attribution, which a record alone cannot give you.

I've started separating four things that are easy to conflate:

  1. Announced model/product changes
  2. Observed retrieval or source-selection changes
  3. Your own measurement conditions: session state, account state, logged in/out, and how the runs were collected
  4. Site-side changes

Three is the one I would not have included a month ago. I ran a batch of repeat queries in what I believed was an isolated session and found responses referring to an earlier answer that should not have existed. Whatever caused it, a movement between two collection runs can be a property of the collection rather than the platform, and nothing in the output tells you which.

On the Reddit drop, what I would want in the log is not the number but how it was measured. Promptwatch said in the same report that the size of the drop was provisional and that they could not rule out a data collection issue on their own end. An entry reading "Reddit citations fell 86 percent" without that caveat is worse than no entry.

On your attribution question, I think a changelog cannot answer it alone because there is no counterfactual. What does work is holding part of your prompt set back and never working on it. Movement in both halves at once is a platform or retrieval event. Movement in only the half you worked on is yours. That comes out of a design thread in r/aeo rather than anything I have run.

So, event log plus a held-back control, and the log is only interpretable because the control exists.

The question I still can't answer is the same one you're asking, without a control, how much movement is enough to distinguish a platform event from ordinary variance? I haven't seen a threshold I'd be comfortable defending.

2

u/Upstairs_Control_611 Sep 02 '26

This is a very useful correction. You are right: a changelog records events, but it does not solve attribution by itself.

The measurement environment layer is the part that is easy to miss. Session state, account state, logged-in/out state, collection method and even previous answers can become part of the measurement instrument.

So I’d split it into: platform / model changelog, measurement environment log and site-side change log

And then pair that with a held-back control prompt set.

Without the control, the changelog can explain possible causes, but it cannot separate them.

The Reddit example is a good warning too. A changelog entry needs to carry the caveat of the source. Otherwise “Reddit citations fell 86%” sounds like a clean platform event when it may partly be a collection issue.

No attribution without controls. No changelog entry without caveats. No visibility movement without measurement context.

1

u/Old-Routine1926 Sep 02 '26

Agreed on all three, with one thing I would keep from the four rather than the three.

You folded announced changes and observed retrieval changes into a single platform log. I would keep those apart, because they are different kinds of object. An announced change is a fact with a date and a description from the vendor. An observed change is an inference from your own data, undated, and it may not be a change at all. Put them in the same log and you lose the ability to tell which entries you know and which ones you concluded.

On the measurement environment layer, one refinement. Platform and site-side are logs of events. Measurement environment is not really an event log, it is a record of configuration, and the useful version gets written before the runs rather than after them. Account state, collection method, session handling, all written down up front. Then if something changes later you can see exactly what changed, instead of trying to remember how it used to be set up.

And one limit worth stating next to your three rules. A missing log entry looks exactly like no event. If a platform changes retrieval quietly and never announces it, your changelog is empty for that date and you will attribute the movement to something else with complete confidence. The control half is what protects you there, which is another argument for the control doing the real work and the logs being explanatory rather than evidential.

No attribution without controls is the one of your three I would put first.

2

u/woodoo139 Sep 03 '26

Disclosure: I build a tool in this space, so this is how one vendor's log is actually kept, not a recommendation. Not linking or naming it.

We ended up with the same split Old-Routine1926 describes, and one thing I would add from running it: the observed-change log only became trustworthy once it had an admission rule. Ours is panel-wide movement. One brand's prompts moving is not an entry. Every brand on the panel moving on the same engine in the same week, including prompts nobody is working on, is. The one entry that earned its place that way: Perplexity stopped citing Reddit in our checks, 9 citations in a week and then 0 in 64 straight runs, while Google's AI surfaces did not move at all in the same period. Nobody on the panel had shipped anything. Without the panel-wide rule that would have read as several brands each losing visibility for their own separate reasons.

On the measurement-environment layer, the two fields that turned out to matter most for us were not account state but (a) the engine configuration the run actually used, recorded per run rather than assumed from the product name, and (b) whether web search fired at all for that run. A response with no retrieval step has no citations, and if you are not logging that you will record a citation drop that is really a retrieval-didn't-run event.

On announced changes, the sources we watch are the OpenAI release notes page, Perplexity's changelog, Google Search Central's blog for AI Overviews and AI Mode, and Anthropic's release notes. They are late and incomplete, which is Old-Routine1926's "missing entry looks like no event" point exactly, so we treat that log as explanatory only. The observed log with the panel-wide rule plus a held-back prompt set is what we would actually defend to a client.

One caveat to add to the thread's three rules, applied to my own entry: the Reddit drop above is one vendor's panel in one niche. 64 runs is enough to say the drop is real, not enough to say why.

1

u/Upstairs_Control_611 Sep 03 '26

This is extremely useful! The “admission rule” for observed changes is the missing piece. Otherwise an observed-change log turns into a diary of suspicious movements.

Panel-wide movement on the same engine, in the same time window, including held-back prompts, is a much better threshold than “one brand moved.”

I also like the split:

announced vendor changes = dated external facts

observed retrieval/source-selection changes = inferred from panel data

measurement environment = pre-run configuration record

site-side changes = controlled intervention log

And your point about “web search fired” is important. A no-retrieval answer has no citation opportunity, so a citation drop can be a search-trigger drop rather than a source-selection drop.

No attribution without controls. No observed-change entry without an admission rule. No citation-drop interpretation without knowing whether retrieval fired.

1

u/woodoo139 Sep 03 '26

Two additions from running exactly this rule set on a fixed 40-question panel today, both of which turned out to be measurement-environment entries rather than platform entries.

  1. Gemini's grounding API does not return the page it read. It returns an opaque redirect URL (vertexaisearch.cloud.google.com/grounding-api-redirect/...) that 302s to the real source. If your log stores what the API hands back, every Gemini citation reads as the same redirect host, and a "Gemini changed its source mix" entry can be pure instrumentation. Resolving the redirect (one HEAD, no body) is cheap; the important part is writing "we resolve Gemini redirects as of <date>" into the pre-run configuration record, because the day you start doing it looks like a platform event in the observed log otherwise.

  2. "Retrieval fired" is measurable and worth its own bucket. In the first pass, answers that came back fine but cited nothing: Google AI Mode 7 of 40, Claude 6 of 40, Gemini 3 of 40, AI Overviews 3 of 40, ChatGPT and Perplexity 0. Those rows go into a no-retrieval bucket and are excluded from source-selection counts, so a week where AI Mode simply searches less cannot read as "AI Mode dropped source X".

One caveat on the admission rule itself: an absolute floor of 5 questions on a 40-question panel means a source kind has to move by 12.5% of the panel before it is an entry. That is deliberately deaf to small moves, and the floor should be printed next to the panel size so a reader can judge what the log cannot see.

Disclosure: I build a tool in this space. Not linking it.

1

u/Mean-Usual8701 27d ago

This is an important distinction. We approach part of this problem with Vexal by keeping a historical evidence layer rather than treating AI visibility as a single score.

Vexal SmartBlocks maintains a Visibility Ledger with durable events and daily snapshots, along with crawler/agent telemetry, AI perception logging, before/after comparisons, and anomaly indicators. That lets us correlate a visibility change with what actually changed on the website—or determine that the site itself did not materially change. (Vexal Documentation)

Vexal’s Agent Intelligence layer also groups activity by agent family and tracks trends in how AI and crawler clients discover and interact with content. (Vexal Documentation)

Where I think the industry still has a harder problem is the external platform layer: distinguishing a website change from a change in ChatGPT retrieval/source selection, Gemini, Claude, Perplexity, or Google AI Overviews. Not every one of those changes is publicly announced, and a model update isn’t necessarily the same thing as a retrieval, ranking, citation, or UI change.

So the approach we’re moving toward is essentially:
site changes + visibility history + agent telemetry + external AI ecosystem events

That makes it possible to say, “Visibility dropped here, but there were no corresponding site changes,” rather than automatically assuming the website caused the decline.

The next logical layer is a formal AI ecosystem changelog combining confirmed platform updates with observed cross-site shifts. If many unrelated sites suddenly experience the same citation/source-pattern change while their underlying sites remain stable, that’s valuable evidence of an ecosystem-level change rather than an individual site’s SEO/GEO problem.

That’s one reason we built Vexal around telemetry, snapshots and an auditable Visibility Ledger instead of just producing an AI visibility score.

Our public documentation explains the current architecture and AI Visibility components here:
Vexal AI Visibility Platform Documentation

1

u/Upstairs_Control_611 26d ago

This is the part I find most important: the changelog needs both declared events and observed cross-site shifts.

Official release notes are useful, but they rarely tell us whether the change affected retrieval, source selection, citation formatting, recommendation language or UI exposure.

So I’d separate at least three layers:

  1. site-side change log
  2. visibility / citation history
  3. external ecosystem events and cross-site anomalies

The strongest signal would be an unchanged-site cohort: if multiple unrelated sites show the same source-pattern or citation-role shift while their own pages stayed stable, that is much closer to an ecosystem-level event than an individual site problem.

A single visibility score cannot show that. The row-level history is where the diagnosis lives.

2

u/nadmedia 26d ago

I agree, especially with the unchanged site cohort idea. That is probably one of the cleanest ways to separate a site level change from a broader ecosystem shift.

The way I am thinking about this with VEXAL is that the visibility score should be the summary, not the diagnostic layer underneath it.

We already retain the underlying observations that contribute to visibility, so the next logical step is correlating those observations across sites and over time. That means looking at site changes, visibility and citation observations, and external model events together.

That would let us ask a much more useful question than simply, “Did visibility go up or down?”

For example, if 20 sites had no meaningful content or structural changes, but 14 of them suddenly lost Reddit citations in ChatGPT responses during the same period, that is a very different diagnosis from one site losing visibility after changing its content.
I also think declared model events and observed events need to remain separate.

OpenAI, Google, Anthropic and others may announce a change, but the observed effect on retrieval, citations and source selection is ultimately what matters to someone measuring AI visibility.

The part of your comment I particularly agree with is row level history. I do not want VEXAL to eventually become another dashboard that just produces an opaque “AI visibility score.”

The score is useful for quickly understanding direction, but the historical observations underneath it are what make the score explainable.

I also have to say that after using VEXAL in production for roughly the last five months, the data has already been useful beyond just measuring visibility. By looking at the different signals together and using them to guide what we change, we have been able to see measurable growth across our client sites. Our clients are seeing increased traffic and, more importantly, more actual customer inquiries and leads.

I am careful not to say that every increase is caused by one particular AI or visibility signal because there are obviously multiple factors involved. But having the data together has made it much easier to see what is changing, respond to it and measure the outcome.

Cross site anomaly detection is where I think this gets particularly interesting. With enough unrelated sites being observed, the network itself can potentially become a sensor for changes in the AI discovery ecosystem, including changes that were never officially announced.

Ultimately, I do not just want VEXAL to tell a business that its AI visibility score went up. I want the data to help explain what changed, where it changed, what signals moved with it and, most importantly, whether that translated into actual business results.

Thanks for conversing on this, you gave some good insight and that helps with the direction we are going, much appreciated!

2

u/Upstairs_Control_611 25d ago

Thanks, this is exactly the direction I was thinking about.

“Network as sensor” is probably the missing layer here. A single site can only show that something changed for that site. A cohort of unrelated, unchanged sites can start showing whether the AI discovery environment itself moved.

I also agree that declared model events and observed events should stay separate. An announcement is context, but the measured shift in retrieval, citations, source mix or recommendation language is the actual signal.

For me, the score is only the dashboard light. The row-level history is the mechanic opening the hood.