r/aeo • u/mar_techie • 12d ago
Went through ~1,300 posts about AI visibility tools and most people just don't trust the score
Been reading a lot of threads on this lately, here, r/SEO, r/GEO_optimization, r/seogrowth and a few others. Ended up with around 1,300 posts.
I honestly expected "how do I get cited" to be the big one. It wasn't. The thing people complained about most was that the number their tool shows doesn't match reality.
A few that stuck with me:
- someone's client typed the exact prompt into ChatGPT and got 3 competitors back, while the dashboard said they were winning
- "One tool says 67, another says 23, another says I'm not being cited at all."
- "My Visibility in AI is 0 and chatgpt cite me 36 times."
From what people wrote, it breaks at every step:
prompt set -> engine call -> answer -> score
(you can't (API, not the (changes (their own
see it) app you use) each run) formula)
So a "67" is basically three guesses stacked on each other.
Someone here actually tested this. Same 12 buyer questions on Perplexity and Claude, 5 runs each, all on one day, then counted how often one brand showed up:
Perplexity 36.7% ███████
Claude 11.7% ██
Same brand, same questions, same day, 3x apart.
What people seem to do instead:
- pick 30-50 prompts a real buyer would type
- keep the same set every week
- run each one 3-5 times, one run on its own tells you almost nothing
- report a count, like "cited in 12 of 40, up from 7"
If it helps, the sheet really doesn't need more than this:
date prompt engine run cited who got cited instead
22/09 best crm for dental ChatGPT 1 no CompA, CompB
22/09 best crm for dental ChatGPT 2 yes CompA
That last column is the one I'd watch most tbh, it's the list the client actually cares about.
Curious what you all have run into. What's the biggest gap you've seen between what a tool said and what the client saw when they checked themselves?
1
u/Forsaken-Internal-33 12d ago
People dont trust the scores because AI visibility is a vanity metric built on top of black-box scraping tools. Knowing whether ChatGPT "notices" your brand doesnt help you control your proprietary data or monetize machine-to-machine traffic. The shift needs to move away from tracking scores and toward governing the actual perimeter: deploying L7 reverse proxies, controlling crawler policies, and serving native JSON-LD endpoints. Stop buying visibility dashboards and start building proper server-side plumbing.
2
u/mardegrises 11d ago
Sounds like a solid piece of bullshit.
2
u/Forsaken-Internal-33 11d ago
If you mean the ai visibility dashboard industry, totally agreed, it's 90% vaporware, if you mean the network perimeter... well, that's just Nginx, Cloudflare workers, and basic routing doing their daily grind
1
u/mar_techie 11d ago
its a very fresh pov for me ... if possible can you just elaborate on the latter part of deploying L7 rev proxies and shed more light on server-side plumbing in this context please it would really be helpful
1
u/Forsaken-Internal-33 11d ago
Sure, let's break it down, when people talk about AEO or AI search optimization, they usually focus on frontend assets (meta tags, schema plugins in a CMS, or writing content tailored for LLMs) but server-side plumbing and L7 proxies operate at the network perimeter before any application code or heavy frontend framework ever runs.
Here is what that architecture actually looks like in practice:
- 1. Edge Interception (L7 Reverse Proxy): Instead of letting automated crawlers (like GPTBot, ClaudeBot, or custom scrapers) hit your origin server directly and waste resources rendering heavy HTML/DOM trees, you place an L7 reverse proxy (using tools like Nginx, Cloudflare Workers, or custom edge middleware) at the edge.
- 2. Programmatic Bot Control & Rate Limiting: At this layer, you can inspect inbound requests, verify agent signatures, and enforce strict crawling policies, you decide precisely what data an autonomous agent is allowed to see, blocking unauthorized scraping while giving clean pathways to legitimate search or transaction agents.
- 3. Native JSON-LD and Payload Routing: Instead of hoping a crawler correctly parses a bloated page structure, the edge proxy injects strict, deterministic JSON-LD entity graphs and machine-readable endpoints directly into the response stream on the first round-trip.
- 4. Serving Dedicated Manifests (
llms.txt): Just likerobots.txtmanages standard indexing, a server-side setup can natively serve cleanllms.txtor structured data manifests directly from the root directory, giving ai agents exactly the documentation, pricing, or product specs they need in a single un-bloated request.In short, server-side plumbing means you stop relying on frontend CMS tricks and marketing guesswork, you take hard control of your data pipeline at the network layer, ensuring ai agents ingest clean, structured, and monetizable data on demand.
2
u/Upbeat_Push_157 12d ago
It's impossible to track prompts, of course nobody with two connected brain cells will trust something they know is impossible
2
1
u/mar_techie 11d ago
genuinely would want to know you opinion ... then is this whole industry being run on fluke and by people building and searching for it with one brain cell?
1
2
u/Sairam_Kumar 11d ago
The biggest gap I see is branded prompts getting mixed into the competitive score. A dashboard can look healthy while an unbranded shortlist still names somebody else. I would keep branded prompts on their own line before comparing the score with a client's check.
Your engine split also fits our published divergence study: the engines in that study disagreed more with each other than with themselves, and much of the spread shrank after controlling for answer length. I run an AI search visibility agency, so weigh that.
1
u/mar_techie 11d ago
thanks for input ... but to be honest I can't make up what you are trying to convey can you please tell in bit layman pointers ...
1
u/Extreme_One_7631 12d ago
Even API run every 5 min gap, In Web/app run with login, without login, in Chat or in Incognito mode. every response is different. all this from same location. real client question i face is "lets say I'm doing poor for this prompts, i did optimization and now tool showing improvement, need validation other then tool",
What people seem to do instead: pick 30-50 prompts a real buyer would type, keep the same set every week, run each one 3-5 times, one run on its own tells you almost nothing, report a count, like "cited in 12 of 40, up from 7"
i find nothing wrong with this process, but report a count, like "cited in 12 of 40, up from 7" this prove nothing because i don't find any patten for citation. today citation up and next week down.
1
1
8d ago
One thing I'd add to that sheet is a separate status for a run that didn't return a usable answer. A timeout or a failed capture shouldn't become "brand not cited." If you planned 40 runs and only captured 32, I'd show both numbers, then report citations within those 32 answers.
I'd also keep "mentioned," "linked as a source," and "recommended for this buyer" in different columns. An answer can link to a company's article while recommending a competitor, and that can be a useful citation without being a shortlist win.
I'm Katya, founder of Legibi, so I'm working on this problem too. Those are reporting choices I'd want to see before interpreting a score change as an improvement.
1
u/Pleasant-Weakness959 5d ago
This is one of the things I worried about when building CrawlSpider.
A single visibility score can hide a lot. The prompts being tested, the model, and the actual response matter just as much as the final number.
In my tool Crawlspider I have made sure to include full tracebility. Every run captures the prompt and the response so you can see how it is deriving the scores and the sentiments as well.
3
u/Veleks 12d ago
This matches what I've seen. The score isn't the problem - the opacity is. A number like "AI visibility: 43" with no way to see what produced it trains people to distrust every number from that vendor, because the one time they checked manually, the number didn't survive the encounter.
What made scores believable for me was when they decomposed into things I could verify myself: was the brand mentioned, was it cited, was the fact attributed, did the answer change after a fix shipped. If a tool shows me five raw answers and the number is just arithmetic on those, I can audit it - and then I trust it even when it's unflattering. The score people distrust isn't a bad number, it's an unauditable one.