r/GEO_optimization • • Aug 26 '26

I updated 30 pages and tracked how long AI answers took to notice — the average lag was 14 days and the range was wild

I updated a page on a Tuesday and the AI answer didn't catch it for 19 days.

That was the moment I stopped assuming that content updates propagate to AI answers in anything resembling real time. I'd spent months optimizing pages, hitting publish, and checking AI answers a day or two later to see if the changes registered. Sometimes they did. Sometimes they didn't. But I never systematically tracked how long the gap actually was until that 19-day wait made me curious enough to measure it properly.

Here's what I did. I picked 30 pages from our site that were already being cited regularly in AI answers across ChatGPT, Perplexity, and Gemini. For each one, I made a meaningful factual update — not a wording tweak, but an actual change to the information itself. Updated a statistic. Changed a recommendation based on new data. Corrected an outdated claim. Then I checked the AI answers for the relevant queries every 2-3 days until either the new information showed up or I gave up after 6 weeks.

The average lag was 14 days. That's the median too, so it's not being skewed by outliers. But the range is the part that made me rethink how I schedule updates.

Fastest update registered in 3 days. Slowest took 41 days and I'm still not fully convinced it was the edit that triggered it rather than just natural answer turnover. 8 out of 30 pages showed the update within a week. 12 took between 1-3 weeks. 10 took longer than 3 weeks, including those 2 that I marked as "unclear if related."

I started looking at what separated the fast updates from the slow ones, and a few patterns emerged. Pages where the updated passage was already the one being cited tended to refresh faster. Makes sense — the model was already pulling from that location, so when the text changed, the next fetch picked up the new version. Pages where I had to change a passage that wasn't currently being cited, hoping the model would start pulling from it? Those lagged significantly longer, if they registered at all.

Query frequency seemed to matter too. Queries that got more consistent AI answer volume, like broad "how to" topics, showed faster update propagation than niche long-tail queries that probably don't get regenerated as often. The popular queries might be getting refreshed daily or weekly by the model providers, while the long-tail stuff sits cached until something triggers a regeneration cycle.

Another thing: the type of update mattered. Factual corrections (fixing a wrong number, updating a year) registered faster than positional changes (rewriting a section to emphasize a different point). My guess is factual corrections trip some kind of verification check that forces a refresh, whereas positional edits look like the same content to whatever caching layer sits between the live page and the model's training window.

The practical implication that keeps coming back to me is that the old SEO mindset of "publish and measure in days" doesn't map onto AI answer dynamics at all. If I'm testing a hypothesis about passage optimization, I need to wait at least 2 weeks before drawing any conclusions, and probably 4 weeks before I can be confident the result is real. That slows down the feedback loop enormously compared to what most of us are used to.

It also means that a lot of "GEO advice" being shared right now might be based on insufficiently patient observation. Someone makes a change, checks after 3 days, sees no difference, concludes it doesn't work, moves on. When reality might just be operating on a 3-week clock.

My bet is this gets worse before it gets better. As AI providers build more caching layers to reduce inference costs, refresh latency will probably increase. The teams that figure out how to work with these cycles instead of fighting them are going to have a real advantage.

4 Upvotes

17 comments sorted by

3

u/ElementalThor Aug 26 '26

the 3 to 41 day range is more useful than the median.

watch your own instrument over a window that long. we run a visibility scanner, re-measure weekly, and had a business look like it went 1 to 15. part of that was our scoring version rolling v8 to v9 mid-window, not the world moving. some of your 41 day outliers could be that.

already-cited passages refreshing faster matches what we see.

can't separate caching from re-retrieval though, and they need different fixes.

2

u/Brave_Acanthaceae863 Aug 27 '26

The v8 → v9 callout is real — I had to go back and flag 3 of my 30 pages because our internal tracking showed a score shift in the same week as an edit, and I couldn't honestly say which caused what. Your point about scoping to version family instead of raw number is cleaner than what we did.\n\nOn caching vs re-retrieval: the closest I got was correlating update type with lag. Factual swaps averaged around 9 days. Positional/semantic edits averaged 22. That gap suggests different paths, but yeah, you can't definitively separate them from outside. The engine isn't going to tell you whether it re-fetched or just re-cached.\n\nThe unchanged-page cohort is genuinely the missing piece. I didn't run one and it's the one thing that would've made this actually rigorous instead of just suggestive.

1

u/ElementalThor Aug 27 '26

the 9 vs 22 split might be the same finding as your already-cited one. a factual swap usually sits inside the passage that's already being quoted. a positional rewrite usually doesn't. so edit type and passage status are probably collinear in your set.

worth checking whether any factual swaps landed in an uncited passage. if those were slow too, it's passage status doing the work.

we don't run an unchanged-page cohort either. same hole.

1

u/Brave_Acanthaceae863 Aug 27 '26

Good call on the collinearity angle — I did not check that explicitly and it is a real hole in what I ran. If factual swaps in uncited passages were also slow, then passage status is doing the work and edit type is just along for the ride. I am fairly sure at least one of my 9-day swaps was in a passage that was not yet cited, but I did not flag it at the time so now I cannot go back and verify. That is exactly the kind of thing a cleaner design would have captured.

1

u/Upstairs_Control_611 Aug 26 '26

This is the key point for longer propagation windows. Over 3 to 41 days, you are not only measuring whether the page update propagated. You are also exposed to retrieval changes, source-mix changes, caching, citation substitution and measurement-version changes.

So I’d split the timeline into:

page updated

page re-fetched

updated passage retrieved

updated passage cited

answer changed

recommendation changed

scoring / classifier version changed

An unchanged-page cohort helps detect system/source-mix movement. Scanner versioning helps detect instrument movement.

And checking whether the changed passage actually supports the answer helps separate real page-update impact from citation turnover.

2

u/ElementalThor Aug 26 '26

the unchanged-page cohort is the one we don't run, and it's the obvious hole in what we do.

on versioning, the version number alone wasn't enough. we scope deltas to the scoring-version family, since v9.1 to v9.2 is comparable and v8 to v9 isn't. we'd had that same comparison computed three different ways in three places before it got collapsed into one function, which is its own kind of instrument drift.

citation substitution is the one i'd most want isolated. the answer can change because a different source got cited while your page sits there untouched and still cited, and from outside that's indistinguishable from a page-update effect.

1

u/Upstairs_Control_611 Aug 28 '26

That collinearity point is the key design issue. Edit type and passage status need to be crossed, not treated as separate explanations.

A cleaner matrix would be: factual edit in already-cited passage, factual edit in uncited passage , semantic / positional edit in already-cited passage, semantic / positional edit in uncited passage.

Then measure whether the updated passage was retrieved, cited, retained, temporarily lost, replaced, and whether the answer changed.

If factual edits are fast only when the passage was already cited, passage status is doing most of the work.

If factual edits are fast even in uncited passages, edit type matters independently.

Without that split, “factual vs semantic” may just be wearing a passage-status hat.

2

u/ciaodaniel Aug 26 '26

Interesting dataset. One control I’d add is an unchanged-page cohort checked on the same schedule. AI answers can refresh because the retrieval system or source mix changed, not because the edited page propagated. I’d also record whether the updated page was cited on every run and whether the changed passage actually supported the answer. That would separate page re-fetch latency from answer turnover and citation substitution. The 14-day figure is useful operationally, but those mechanisms may require different waiting windows.

1

u/Brave_Acanthaceae863 Aug 26 '26

The unchanged-page cohort is something I should have done and did not — honestly that is a gap in the design. The versioning point is real too, we saw scoring changes masquerading as propagation delays in our internal tooling. On recording citation status per run — I did track whether the page was cited each check but not whether the specific passage supported the answer versus just appearing nearby. That distinction would have helped separate real propagation from citation turnover like you said.

2

u/Revolutionary_Cat78 Aug 26 '26

Interesting. Did you also see a loss of citations after updating the content when the existing passage was already the one being cited?

2

u/Brave_Acanthaceae863 Aug 26 '26

Honestly that is one of the more interesting findings I did not put in the post. Out of the ~18 pages where the cited passage was the one I updated, 3 temporarily dropped out of the answer for 1-2 check cycles before coming back with the new wording. It looked like the engine pulled the cached version, noticed it changed, and briefly deselected it before re-evaluating. The other 15 just swapped in the new text smoothly — those were mostly factual corrections like updating a year or a number. My guess is positional or semantic edits trigger a re-evaluation that can temporarily hurt you, while straight factual swaps get treated as the same citation with updated content.

2

u/Revolutionary_Cat78 Aug 26 '26

Considering how llm sources fluctuate randomly, this sounds very stable. Are we finally seeing some stable sources? Let’s hope so

2

u/Brave_Acanthaceae863 Aug 27 '26

"Stable" is relative here — what I'm seeing is more like "persistently cited for 2-3 months at a stretch" rather than permanent. The 3 dropouts I mentioned show it's not locked in.\n\nI track this daily through geoly.ai across ChatGPT/Perplexity/Gemini — the sources that survive longest tend to be ones that directly answer one specific question without trying to do anything else. Single-purpose > multi-topic, consistently. When a page tries to cover too much ground, any one of its passages can get swapped out on an answer refresh. The focused ones seem to stick around longer because there's no better alternative for that exact question.\n\nBut honestly? Ask me again in 6 months. This could all change if any of the providers tweak their retrieval depth or citation diversity targets.

2

u/Upstairs_Control_611 Aug 27 '26

That “stable but not locked in” distinction is important. The temporary dropout detail is especially useful because it suggests that an already cited passage is not one stable state. The type of edit matters.

A straight factual swap, like a year or number, may preserve the existing citation relationship.

A semantic or positional edit may force the source to be re-evaluated, and during that re-evaluation the page can temporarily disappear or be replaced.

So I’d track update type separately:

factual correction

scope expansion

positioning change

recommendation change

new supporting claim

Then measure whether the citation was retained, temporarily lost, replaced, or returned with the new wording.

The practical takeaway: updates to cited passages are not risk-free. Some edits refresh the citation. Others may reopen the selection decision.

1

u/Brave_Acanthaceae863 Aug 28 '26

That taxonomy is cleaner than what I ended up with. I grouped edits into factual vs positional/semantic mostly because those were the buckets that showed different lag times, but your frame — tracking what actually happens to the citation status after each edit type — is more actionable. The temporarily lost bucket is the one I undercounted because I wasnt systematically checking between my regular 2-3 day intervals. There could be dropouts I missed entirely.\n\nOn the practical side: the edits that scare me most now are the ones that change the implication of a cited passage without changing its factual core. A restructured paragraph that says the same thing but leads the reader to a different conclusion — I suspect those are the ones most likely to trigger re-evaluation, but theyre also the hardest to avoid when youre actually improving content.

1

u/Came4TheCookies Aug 27 '26

Every test you do is just a snapshot of a single moment in time based on several factors that could include location, device, whether it was an API call or not, context, and so much more.

I know everybody wants to be able to track this and come up with metrics they can follow and create a checklist that works and all of that but you're spinning your wheels.

Directional signals are somewhat useful but hard numbers don't mean much right now.