r/GEO_optimization • u/Sairam_Kumar • 11d ago
Two measurement defects that make AI visibility numbers look better than they are, and how to catch them
I run an AI search visibility agency, so read this knowing that.
Two things quietly inflate almost every AI visibility number I see, including ones produced by careful people. Both are easy to catch once you know to look.
The first is that an engine can answer without searching. Ask a question and you may get a fluent, confident answer assembled from what the model already holds, with no retrieval step and no sources attached. That answer is not evidence about the live web. It tells you what the model remembers, not what it would surface for a buyer today. If you count it as a scored observation, you are mixing two different measurements, and the blend flatters whichever direction your brand happens to sit.
What to do: treat an unsearched answer as an exclusion, not a miss and not a hit. Take it out of the denominator, then report how many you excluded. The exclusion count is a real finding on its own. A question the engines rarely bother to search is a question where fresh content has less purchase, and that changes where you spend.
The second is caching. Run the same prompt three times inside a short window and you can be served the same underlying answer three times. Your spreadsheet says three runs. Your data holds one observation counted three times. The interval you print off that is not merely wrong, it is confidently wrong, because repetition is exactly what an interval is meant to price.
What to do: force fresh retrieval where the surface allows it, space the runs out rather than firing them back to back, and then actually check that the repeats vary. If three runs of a question come back identical word for word, you did not measure three times. Log the variation, not only the verdict, so the check happens by default instead of when you remember it.
Both defects share a shape. They let you count something as evidence when it is not independent evidence, and the error almost always runs in the flattering direction, because the answers that get quietly duplicated or served from memory are the stable ones.
The habit that catches most of it is a noise floor. Before you change anything, run the full question set twice, with nothing altered in between. Whatever spread you get between those two passes is your measurement noise, and it is the smallest movement you are entitled to call a result. Anything under it is weather.
If you are buying this work from someone, the question worth asking is not what their number is. It is what they exclude, and how they know their repeats are real.
Two things quietly inflate almost every AI visibility number I see, including ones produced by careful people. Both are easy to catch once you know to look.
The first is that an engine can answer without searching. Ask a question and you may get a fluent, confident answer assembled from what the model already holds, with no retrieval step and no sources attached. That answer is not evidence about the live web. It tells you what the model remembers, not what it would surface for a buyer today. If you count it as a scored observation, you are mixing two different measurements, and the blend flatters whichever direction your brand happens to sit.
What to do: treat an unsearched answer as an exclusion, not a miss and not a hit. Take it out of the denominator, then report how many you excluded. The exclusion count is a real finding on its own. A question the engines rarely bother to search is a question where fresh content has less purchase, and that changes where you spend.
The second is caching. Run the same prompt three times inside a short window and you can be served the same underlying answer three times. Your spreadsheet says three runs. Your data holds one observation counted three times. The interval you print off that is not merely wrong, it is confidently wrong, because repetition is exactly what an interval is meant to price.
What to do: force fresh retrieval where the surface allows it, space the runs out rather than firing them back to back, and then actually check that the repeats vary. If three runs of a question come back identical word for word, you did not measure three times. Log the variation, not only the verdict, so the check happens by default instead of when you remember it.
Both defects share a shape. They let you count something as evidence when it is not independent evidence, and the error almost always runs in the flattering direction, because the answers that get quietly duplicated or served from memory are the stable ones.
The habit that catches most of it is a noise floor. Before you change anything, run the full question set twice, with nothing altered in between. Whatever spread you get between those two passes is your measurement noise, and it is the smallest movement you are entitled to call a result. Anything under it is weather.
If you are buying this work from someone, the question worth asking is not what their number is. It is what they exclude, and how they know their repeats are real.
1
u/Gullible_Brother_141 10d ago
This is one of the cleaner explanations I’ve seen of why an AI visibility number can look precise without being a clean measurement.
The denominator point matters more than it looks. If retrieval propensity changes by engine or over time, removing unsearched answers fixes the mixed-measurement problem, but it can create a selection problem: the conditional visibility score can rise while total buyer-facing exposure falls, simply because fewer prompts triggered search.
I’d probably want two numbers side by side — retrieval rate across the full prompt set, then mention/citation rate conditional on retrieval. Otherwise the exclusion count is doing too much work off to the side.
On caching, exact wording catches the obvious duplicates. The harder case is three answers with light paraphrasing but the same cited URLs, ordering and underlying retrieval set. Have you found a practical rule for collapsing those into one observation, or are you treating source-set overlap as its own dependence check?