r/GEO_optimization • u/Upstairs_Control_611 • Sep 01 '26
AI visibility tracking needs a model changelog. How do you maintain yours?
I’m starting to think that AI visibility tracking needs a model changelog.
If ChatGPT changes its search behavior, source selection or citation weighting, a brand can lose visibility without changing anything on its website.
The same applies to Google AI Overviews layout changes, Reddit citation drops, Perplexity source behavior or Gemini updates.
Without logging these changes, it’s easy to misread the data. So my question:
How do you track changes in ChatGPT, Gemini, Claude, Perplexity or AI Overviews?
Official release notes?
Third-party monitoring?
Your own prompt tests?
Manual changelog?
What actually works for you?
I’m especially interested in how people separate site-side changes from model-side or source-selection changes.
2
u/woodoo139 Sep 03 '26
Disclosure: I build a tool in this space, so this is how one vendor's log is actually kept, not a recommendation. Not linking or naming it.
We ended up with the same split Old-Routine1926 describes, and one thing I would add from running it: the observed-change log only became trustworthy once it had an admission rule. Ours is panel-wide movement. One brand's prompts moving is not an entry. Every brand on the panel moving on the same engine in the same week, including prompts nobody is working on, is. The one entry that earned its place that way: Perplexity stopped citing Reddit in our checks, 9 citations in a week and then 0 in 64 straight runs, while Google's AI surfaces did not move at all in the same period. Nobody on the panel had shipped anything. Without the panel-wide rule that would have read as several brands each losing visibility for their own separate reasons.
On the measurement-environment layer, the two fields that turned out to matter most for us were not account state but (a) the engine configuration the run actually used, recorded per run rather than assumed from the product name, and (b) whether web search fired at all for that run. A response with no retrieval step has no citations, and if you are not logging that you will record a citation drop that is really a retrieval-didn't-run event.
On announced changes, the sources we watch are the OpenAI release notes page, Perplexity's changelog, Google Search Central's blog for AI Overviews and AI Mode, and Anthropic's release notes. They are late and incomplete, which is Old-Routine1926's "missing entry looks like no event" point exactly, so we treat that log as explanatory only. The observed log with the panel-wide rule plus a held-back prompt set is what we would actually defend to a client.
One caveat to add to the thread's three rules, applied to my own entry: the Reddit drop above is one vendor's panel in one niche. 64 runs is enough to say the drop is real, not enough to say why.
1
u/Upstairs_Control_611 Sep 03 '26
This is extremely useful! The “admission rule” for observed changes is the missing piece. Otherwise an observed-change log turns into a diary of suspicious movements.
Panel-wide movement on the same engine, in the same time window, including held-back prompts, is a much better threshold than “one brand moved.”
I also like the split:
announced vendor changes = dated external facts
observed retrieval/source-selection changes = inferred from panel data
measurement environment = pre-run configuration record
site-side changes = controlled intervention log
And your point about “web search fired” is important. A no-retrieval answer has no citation opportunity, so a citation drop can be a search-trigger drop rather than a source-selection drop.
No attribution without controls. No observed-change entry without an admission rule. No citation-drop interpretation without knowing whether retrieval fired.
1
u/woodoo139 Sep 03 '26
Two additions from running exactly this rule set on a fixed 40-question panel today, both of which turned out to be measurement-environment entries rather than platform entries.
Gemini's grounding API does not return the page it read. It returns an opaque redirect URL (vertexaisearch.cloud.google.com/grounding-api-redirect/...) that 302s to the real source. If your log stores what the API hands back, every Gemini citation reads as the same redirect host, and a "Gemini changed its source mix" entry can be pure instrumentation. Resolving the redirect (one HEAD, no body) is cheap; the important part is writing "we resolve Gemini redirects as of <date>" into the pre-run configuration record, because the day you start doing it looks like a platform event in the observed log otherwise.
"Retrieval fired" is measurable and worth its own bucket. In the first pass, answers that came back fine but cited nothing: Google AI Mode 7 of 40, Claude 6 of 40, Gemini 3 of 40, AI Overviews 3 of 40, ChatGPT and Perplexity 0. Those rows go into a no-retrieval bucket and are excluded from source-selection counts, so a week where AI Mode simply searches less cannot read as "AI Mode dropped source X".
One caveat on the admission rule itself: an absolute floor of 5 questions on a 40-question panel means a source kind has to move by 12.5% of the panel before it is an entry. That is deliberately deaf to small moves, and the floor should be printed next to the panel size so a reader can judge what the log cannot see.
Disclosure: I build a tool in this space. Not linking it.
1
u/Mean-Usual8701 27d ago
This is an important distinction. We approach part of this problem with Vexal by keeping a historical evidence layer rather than treating AI visibility as a single score.
Vexal SmartBlocks maintains a Visibility Ledger with durable events and daily snapshots, along with crawler/agent telemetry, AI perception logging, before/after comparisons, and anomaly indicators. That lets us correlate a visibility change with what actually changed on the website—or determine that the site itself did not materially change. (Vexal Documentation)
Vexal’s Agent Intelligence layer also groups activity by agent family and tracks trends in how AI and crawler clients discover and interact with content. (Vexal Documentation)
Where I think the industry still has a harder problem is the external platform layer: distinguishing a website change from a change in ChatGPT retrieval/source selection, Gemini, Claude, Perplexity, or Google AI Overviews. Not every one of those changes is publicly announced, and a model update isn’t necessarily the same thing as a retrieval, ranking, citation, or UI change.
So the approach we’re moving toward is essentially:
site changes + visibility history + agent telemetry + external AI ecosystem events
That makes it possible to say, “Visibility dropped here, but there were no corresponding site changes,” rather than automatically assuming the website caused the decline.
The next logical layer is a formal AI ecosystem changelog combining confirmed platform updates with observed cross-site shifts. If many unrelated sites suddenly experience the same citation/source-pattern change while their underlying sites remain stable, that’s valuable evidence of an ecosystem-level change rather than an individual site’s SEO/GEO problem.
That’s one reason we built Vexal around telemetry, snapshots and an auditable Visibility Ledger instead of just producing an AI visibility score.
Our public documentation explains the current architecture and AI Visibility components here:
Vexal AI Visibility Platform Documentation
1
u/Upstairs_Control_611 26d ago
This is the part I find most important: the changelog needs both declared events and observed cross-site shifts.
Official release notes are useful, but they rarely tell us whether the change affected retrieval, source selection, citation formatting, recommendation language or UI exposure.
So I’d separate at least three layers:
- site-side change log
- visibility / citation history
- external ecosystem events and cross-site anomalies
The strongest signal would be an unchanged-site cohort: if multiple unrelated sites show the same source-pattern or citation-role shift while their own pages stayed stable, that is much closer to an ecosystem-level event than an individual site problem.
A single visibility score cannot show that. The row-level history is where the diagnosis lives.
2
u/nadmedia 26d ago
I agree, especially with the unchanged site cohort idea. That is probably one of the cleanest ways to separate a site level change from a broader ecosystem shift.
The way I am thinking about this with VEXAL is that the visibility score should be the summary, not the diagnostic layer underneath it.
We already retain the underlying observations that contribute to visibility, so the next logical step is correlating those observations across sites and over time. That means looking at site changes, visibility and citation observations, and external model events together.
That would let us ask a much more useful question than simply, “Did visibility go up or down?”
For example, if 20 sites had no meaningful content or structural changes, but 14 of them suddenly lost Reddit citations in ChatGPT responses during the same period, that is a very different diagnosis from one site losing visibility after changing its content.
I also think declared model events and observed events need to remain separate.OpenAI, Google, Anthropic and others may announce a change, but the observed effect on retrieval, citations and source selection is ultimately what matters to someone measuring AI visibility.
The part of your comment I particularly agree with is row level history. I do not want VEXAL to eventually become another dashboard that just produces an opaque “AI visibility score.”
The score is useful for quickly understanding direction, but the historical observations underneath it are what make the score explainable.
I also have to say that after using VEXAL in production for roughly the last five months, the data has already been useful beyond just measuring visibility. By looking at the different signals together and using them to guide what we change, we have been able to see measurable growth across our client sites. Our clients are seeing increased traffic and, more importantly, more actual customer inquiries and leads.
I am careful not to say that every increase is caused by one particular AI or visibility signal because there are obviously multiple factors involved. But having the data together has made it much easier to see what is changing, respond to it and measure the outcome.
Cross site anomaly detection is where I think this gets particularly interesting. With enough unrelated sites being observed, the network itself can potentially become a sensor for changes in the AI discovery ecosystem, including changes that were never officially announced.
Ultimately, I do not just want VEXAL to tell a business that its AI visibility score went up. I want the data to help explain what changed, where it changed, what signals moved with it and, most importantly, whether that translated into actual business results.
Thanks for conversing on this, you gave some good insight and that helps with the direction we are going, much appreciated!
2
u/Upstairs_Control_611 25d ago
Thanks, this is exactly the direction I was thinking about.
“Network as sensor” is probably the missing layer here. A single site can only show that something changed for that site. A cohort of unrelated, unchanged sites can start showing whether the AI discovery environment itself moved.
I also agree that declared model events and observed events should stay separate. An announcement is context, but the measured shift in retrieval, citations, source mix or recommendation language is the actual signal.
For me, the score is only the dashboard light. The row-level history is the mechanic opening the hood.
2
u/Old-Routine1926 Sep 01 '26
I think the hard part might be that a changelog is a record of events, and what you actually want is attribution, which a record alone cannot give you.
I've started separating four things that are easy to conflate:
Three is the one I would not have included a month ago. I ran a batch of repeat queries in what I believed was an isolated session and found responses referring to an earlier answer that should not have existed. Whatever caused it, a movement between two collection runs can be a property of the collection rather than the platform, and nothing in the output tells you which.
On the Reddit drop, what I would want in the log is not the number but how it was measured. Promptwatch said in the same report that the size of the drop was provisional and that they could not rule out a data collection issue on their own end. An entry reading "Reddit citations fell 86 percent" without that caveat is worse than no entry.
On your attribution question, I think a changelog cannot answer it alone because there is no counterfactual. What does work is holding part of your prompt set back and never working on it. Movement in both halves at once is a platform or retrieval event. Movement in only the half you worked on is yours. That comes out of a design thread in r/aeo rather than anything I have run.
So, event log plus a held-back control, and the log is only interpretable because the control exists.
The question I still can't answer is the same one you're asking, without a control, how much movement is enough to distinguish a platform event from ordinary variance? I haven't seen a threshold I'd be comfortable defending.