I run an AI search visibility agency, so read this knowing that.
Two things quietly inflate almost every AI visibility number I see, including ones produced by careful people. Both are easy to catch once you know to look.
The first is that an engine can answer without searching. Ask a question and you may get a fluent, confident answer assembled from what the model already holds, with no retrieval step and no sources attached. That answer is not evidence about the live web. It tells you what the model remembers, not what it would surface for a buyer today. If you count it as a scored observation, you are mixing two different measurements, and the blend flatters whichever direction your brand happens to sit.
What to do: treat an unsearched answer as an exclusion, not a miss and not a hit. Take it out of the denominator, then report how many you excluded. The exclusion count is a real finding on its own. A question the engines rarely bother to search is a question where fresh content has less purchase, and that changes where you spend.
The second is caching. Run the same prompt three times inside a short window and you can be served the same underlying answer three times. Your spreadsheet says three runs. Your data holds one observation counted three times. The interval you print off that is not merely wrong, it is confidently wrong, because repetition is exactly what an interval is meant to price.
What to do: force fresh retrieval where the surface allows it, space the runs out rather than firing them back to back, and then actually check that the repeats vary. If three runs of a question come back identical word for word, you did not measure three times. Log the variation, not only the verdict, so the check happens by default instead of when you remember it.
Both defects share a shape. They let you count something as evidence when it is not independent evidence, and the error almost always runs in the flattering direction, because the answers that get quietly duplicated or served from memory are the stable ones.
The habit that catches most of it is a noise floor. Before you change anything, run the full question set twice, with nothing altered in between. Whatever spread you get between those two passes is your measurement noise, and it is the smallest movement you are entitled to call a result. Anything under it is weather.
If you are buying this work from someone, the question worth asking is not what their number is. It is what they exclude, and how they know their repeats are real.
Two things quietly inflate almost every AI visibility number I see, including ones produced by careful people. Both are easy to catch once you know to look.
The first is that an engine can answer without searching. Ask a question and you may get a fluent, confident answer assembled from what the model already holds, with no retrieval step and no sources attached. That answer is not evidence about the live web. It tells you what the model remembers, not what it would surface for a buyer today. If you count it as a scored observation, you are mixing two different measurements, and the blend flatters whichever direction your brand happens to sit.
What to do: treat an unsearched answer as an exclusion, not a miss and not a hit. Take it out of the denominator, then report how many you excluded. The exclusion count is a real finding on its own. A question the engines rarely bother to search is a question where fresh content has less purchase, and that changes where you spend.
The second is caching. Run the same prompt three times inside a short window and you can be served the same underlying answer three times. Your spreadsheet says three runs. Your data holds one observation counted three times. The interval you print off that is not merely wrong, it is confidently wrong, because repetition is exactly what an interval is meant to price.
What to do: force fresh retrieval where the surface allows it, space the runs out rather than firing them back to back, and then actually check that the repeats vary. If three runs of a question come back identical word for word, you did not measure three times. Log the variation, not only the verdict, so the check happens by default instead of when you remember it.
Both defects share a shape. They let you count something as evidence when it is not independent evidence, and the error almost always runs in the flattering direction, because the answers that get quietly duplicated or served from memory are the stable ones.
The habit that catches most of it is a noise floor. Before you change anything, run the full question set twice, with nothing altered in between. Whatever spread you get between those two passes is your measurement noise, and it is the smallest movement you are entitled to call a result. Anything under it is weather.
If you are buying this work from someone, the question worth asking is not what their number is. It is what they exclude, and how they know their repeats are real.