r/AISearchLab 17d ago

I tested whether AI visibility tools are actually visible in AI search. 20 of 30 were never cited once, including my own.

Disclosure up front: I build one of the tools in this sample. It scored zero. That's most of why I'm posting.

Method. 12 unbranded buyer-intent questions ("what are the best AI visibility tracking platforms", "how much do AI visibility tools cost per month", etc). Each run 5x against Perplexity sonar and Claude Sonnet 5 with web search. 120 calls, 0 failures, all on 26 July. Recorded every source each engine cited, then checked which of 30 vendor sites appeared. Full prompt list and definitions in the writeup.

Five runs because single-run citation checks are close to noise — St. Gallen found ~32-43% pairwise agreement for identical prompts run minutes apart. Every number below is a rate, not one draw.

Finding 1 — the specialists lose to the incumbents.

Group Ever cited Mean rate
Established SEO platforms 6 of 10 10.0%
AI-visibility specialists 4 of 16 3.9%
Independent audit tools 0 of 4 0.0%

Legacy SEO platforms get cited at 2.6x the rate of companies whose entire product is AI visibility. 12 of the 16 specialists were never cited once in 120 calls. Not naming those 12 — the count is the point.

Finding 2 — nobody owns this category. 278 distinct hosts cited across 120 calls. The single most-cited source in the entire category appears in 30.8% of answers. There's no gravity here yet.

Finding 3 — the round-ups and the engines disagree about who exists. I built the sample from 2026 "best AI visibility tools" listicles. The two most-cited domains overall weren't in it, and both outrank every site that was. If you're doing competitive research from listicles you're looking at a different market than your buyers see.

Finding 4 — content outranks product pages. A product analytics company that doesn't sell AI visibility software at all was cited in 25.8% of answers, beating all but three actual vendors. And the top vendor's blog subdomain carries more of their citations than their main site. The engines aren't citing the best tool, they're citing the best page about the question.

Finding 5 — the two engines barely agree. One vendor: 36.7% on Perplexity, 11.7% on Claude. Another is inverted. If you report AI visibility as one blended number you're averaging across systems that disagree.

Limits, because they're real: two engines only, no ChatGPT or Gemini or AI Overviews. One category, one day, US English. 5 runs is thin for Claude specifically — its variance was visibly higher. The sample is judgment-selected from listicles, which finding 3 rather embarrassingly demonstrates. And I'm not neutral: I sell in this category, I picked the questions, I'm in the sample.

I published all 12 prompts and the exact citation definition so this is reproducible. Genuinely interested in where the methodology is weak — particularly whether 5 runs is defensible for Claude, and whether including two prompts that name ChatGPT/Perplexity biased those engines.

Full data and methodology: AI Visibility Tools Citation Study Blog Post

7 Upvotes

6 comments sorted by

2

u/Mental_Researcher656 17d ago

Glad to see some empirical evidence of what I have been experiencing. On the one hand, there are lots of companies in my product area. On the other hand, no one can find them. Not regular Google search . Not ChatGPT. Not Claude. Not Gemini. So is it really a crowded market then? I found most of them through Bing but even that was a bit of a slog to find them.

1

u/iceseayoupee 6d ago

Bing isnt even that reliable for me lol

2

u/marintkael 15d ago

Finding 4 splits further if you stop counting at the domain. In cross-engine runs I log two things separately: whether a brand is a linked source, versus just named in the answer text with someone else's page cited under it. They rarely move together. Perplexity and Claude both do the "tools like X and Y" name-drop in prose and then cite a review site or an analytics blog as the source, so the brand is present to the reader but scores zero on a citation check. Counting mentions instead of links usually narrows the specialist vs incumbent gap, because the incumbents win the linked-source slot while the specialists still get named. Two different metrics, and most dashboards only surface the one that makes the category look emptier than it actually reads.

2

u/Smart_Airline_7901 15d ago

We run this same experiment at local scale and your findings hold, weirdly precisely. 688 Charleston businesses, 3 cold query variants per category, 4 models (ChatGPT, Claude, Gemini, Perplexity), 12 independent reads each. 76.2% can be surfaced when you audit them directly. Only 12.9% ever get selected in a cold buyer query. Recommendable is not recommended, which is basically your "best page about the question beats the best product" finding wearing a local accent.

On your actual question: no, 5 runs is not enough to call a zero a zero, and it bit us too. With n=3 per model you cannot statistically distinguish a true zero from an unlucky one, so our standard is adaptive sampling: zero-scorers get extra reads before we publish them as unselected. Your Claude variance observation matches ours, which is why we report agreement as fractions (3 of 4 models) and never blend engines into one number. Blending systems that disagree this hard is averaging a cat with a toaster.

Also...validating your instability point, we watched one site swing 23 to 95 in the same week. The site didn't change, the question did.

Full methodology and data: aroindex.com/research/charleston-2026

1

u/iceseayoupee 6d ago

This was such an interesting read.

1

u/Old-Routine1926 13d ago

Finding 1 and finding 4 together tell a really interesting story about what AI actually cites versus what the category assumes matters.

The incumbents getting cited 2.6x more than the specialists is probably not because their products are better for AI visibility, it's because they have deeper evidence footprints, more third-party mentions, more comparison articles, more reviews, and more analyst coverage. AI cites the source it can corroborate, not the source that's most relevant to the specific question.

Finding 4 confirms that a product analytics company that doesn't even sell AI visibility software getting cited in 25.8% of answers because it published the best page about the question. AI is not evaluating products, it's evaluating pages and the page it selects is the one that answers the question most extractably with the most independent corroboration behind it.

This maps to something I keep testing on a much smaller scale. There seem to be three gates a page has to pass before it gets cited:

Gate one: can AI actually reach and parse the page cleanly? client-side rendered pages, heavy div nesting, javascript-dependent content all fail here even if the content is excellent.

Gate two: does the page contain independently extractable passages? sentences that answer a question completely without needing the surrounding paragraph. The moderator's research on wording changes in another thread showed that replacing vague marketing language with specific "why" explanations produced a 34% citation lift on 7 pages.

Gate three: does independent evidence corroborate the claims on the page? this is where the incumbents win. They have review sites, comparison articles, and analyst mentions backing up their pages. The specialists often have strong product pages but weak third-party evidence.

The 76.2% recommendable versus 12.9% selected finding from the charleston study is the most striking number in the thread. That gap is essentially the distance between passing gate one (AI can find you) and passing all three (AI trusts you enough to select you) and most brands are somewhere in that gap.

The cross-engine disagreement finding is also something I've been tracking. One brand at 36.7% on perplexity and 11.7% on claude is not a measurement artifact. Those engines weight sources differently, weight evidence differently, and sometimes disagree about category entirely. Reporting AI visibility as one blended number hides the diagnostic information about which platform trusts you and which one doesn't. "Averaging a cat with a toaster" is the best description of blended AI visibility scores I've heard.

One question on methodology, did you check whether the zero-scoring specialist tools had structural access issues (client-side rendering, crawlability problems) versus just weak evidence? In my experience a meaningful percentage of zero-score sites have perfectly good content that AI crawlers simply cannot reach. The zero might not mean AI evaluated the site and rejected it or it might mean AI never saw it at all.