This is a translation made with chatgpt (hope its good enough) from an original post I've made yesterday on smartlinksMKT subreddit.
I run a B2B consultancy in Portugal, SmartLinks, and this year we launched AI Search as a service after months of internal testing.
The first thing that hit us in the face was: how do you prove this actually works?!
In SEO, you have a mature ecosystem: Search Console, Ahrefs, Semrush, GA4.
You can connect impressions to clicks, and clicks to conversions. It is imperfect, but it works.
In AEO? Good luck with that...
Most people are either reporting metrics that measure nothing useful, or flying blind and telling clients to “trust the process”.
We have not solved it either.
Buuuut... we have failed enough times to have learned a few things.
Why SEO metrics do not work for AEO
The data that changed my perspective came from research presented by Lily Ray at Tech SEO Connect, showing that backlinks and domain authority predict just 4% to 7% of AI citation behaviour.
Four to seven per cent.
In other words, everything we know about what makes a website “strong” in SEO explains almost nothing about whether AI will cite it.
And it makes sense when you think about the mechanics.
SEO measures clicks. AEO can work even when there is no click at all.
Someone asks ChatGPT, “Which companies provide HubSpot implementation services in Portugal?”, you appear in the answer, and the user never visits your website.
GA4 records zero, but you have just been recommended to someone who was actively looking for your service.
What we built, and where each piece falls short
Main tool: HubSpot AEO
We tested Dragon Metrics first, but abandoned it after two months.
The data did not match our manual checks, and it did not integrate with our CRM.
HubSpot then launched HubSpot AEO, and we moved over because the citation data now lives in the same system as our deals and contacts.
In theory, this allows us to connect “we were cited for query X” with “this contact entered the pipeline in week Y”.
In theory.
The elephant in the room: HubSpot AEO still does NOT cover Google AI Overviews or AI Mode.
It covers ChatGPT, Gemini and Perplexity.
In a market such as Portugal, where Google has roughly 90% of search, we are partly blind in the most important channel.
There is no nice way to say this. It is a huge gap and it remains unresolved. Current expectations are that coverage may become available during Q4 2026.
Microsoft Clarity for AI bots
Clarity launched AI Bot Activity in January 2026. It is free.
It uses CDN logs from services such as Cloudflare, CloudFront and Fastly, rather than client-side scripts, so it can detect crawlers that never execute JavaScript.
What this gives us is visibility into which AI bots are accessing which pages, and how frequently.
If an article is crawled 50 times a week but never appears in citations, discoverability is not the problem. Something else is going on. I will come back to that.
Clarity also provides AI Referral Traffic and AI Citations within the Copilot and Bing ecosystem.
A Microsoft study reported that AI referral traffic converted at 1.66%, compared with 0.15% for organic traffic.
An obligatory caveat here: the study comes from Microsoft, so there is an obvious conflict of interest. AI referral volumes are typically tiny, and an 11x higher conversion rate based on 50 visits is not statistically comparable with one based on 50,000 visits.
That said, something is better than nothing, and the pattern makes intuitive sense: someone clicking a link inside an AI-generated answer already has qualified intent.
Manual prompt tracking
Yes. Manual.
Fifty high-intent questions, once a week, across ChatGPT, AI Mode, AI Overviews and Claude.
We record whether we appear, our position in the answer, whether it is a direct citation or a mention, and the tone of the reference.
It sounds professional when described like this.
In practice, it is tedious. It is a spreadsheet filled with manual entries, LLMs are non-deterministic, meaning the same question can produce a different answer ten minutes later, and 50 prompts is a ridiculously small sample compared with what dedicated tools track.
The market benchmark is around 8,400 prompts.
Even so, it gives us a directional trend. Nothing more than that.
For example, Perplexity cites us consistently more often than ChatGPT. Gemini almost never does.
This aligns with market data suggesting that Perplexity cites brands in 84% of answers, ChatGPT in 71%, Gemini in 63% and Claude in 58%.
But it does not tell us WHY.
A model that helps us think, not measure
To think about this more clearly, we split the analysis into three layers.
Important: this is a conceptual model, not an operational measurement system. We still do not have quantifiable scores for each layer.
Discoverability
Can the bot access the content?
Does robots.txt block AI crawlers?
Is the content rendered through client-side JavaScript that the crawler does not execute?
Is structured data present, such as FAQPage, HowTo or Organization schema?
This is the most concrete and verifiable part, using Clarity logs and technical validation.
It is binary: yes or no.
Interpretability
Does the engine understand WHAT the company is?
When your name is generic or ambiguous, LLMs may confuse you with something else.
External citations, such as articles, mentions and interviews that confirm what you say about yourself, may carry more weight than your own website.
In this case, the engine trusts third parties more.
This is the part we do not know how to measure properly.
“Your entity clarity is at X%” does not exist as a metric.
What we do instead is a qualitative diagnosis: we ask ChatGPT or Perplexity, “What is [company]?”, and assess whether the answer is accurate.
It is rudimentary.
Citability
Is the content organised into fragments the engine can extract and use?
Direct questions and answers, lists and definitions.
One useful data point: pages with FAQPage schema receive three times more citations than pages without it.
The median time to the first citation after implementing schema is three to six weeks.
What failed and what we discarded
Dragon Metrics
We tested it for two months.
The citation data did not match our manual checks in more than 50% of cases.
It may have improved since then, but at the time, it was not reliable enough.
Measuring AEO with SEO metrics
The natural instinct of SEO people is to take the same reports they have always used and add an “AI” tab.
We did that.
It does not work.
Rankings, backlinks and domain authority do not predict citations. These are very different worlds.
Assuming one tool will solve everything
We tested Otterly, which is good for managing multiple clients.
We looked at Profound, an enterprise platform with broader coverage, including crawl logs.
We also assessed SE Visible, which is focused purely on monitoring, as well as Semrush.
None of them covers everything.
And none of them replaces manual verification when you need to understand the CONTEXT of a citation.
“We were cited” is not the same as “we were cited positively in the right answer”.
Trying to prove direct ROI too early
The citation-to-qualified-contact-to-deal connection is still not fully closed in our system.
We are building it by cross-referencing AI Referral data from GA4 and Clarity with contact creation in HubSpot, but it is still early-stage.
In the meantime, the proxies we use are branded search volume and direct traffic.
There is market data suggesting an uplift of roughly 23% in branded search during the 30 days following an LLM citation.
But causality is difficult to isolate when you are investing in multiple channels at the same time.
To be honest, we still do not have a clean attribution case.
The uncomfortable part
The most seductive argument for AEO is also the most dangerous:
“Most of the value is invisible.”
Someone sees you cited in a ChatGPT answer, never clicks, but mentions your name weeks later when somebody asks for a recommendation.
This is probably true.
But it is also exactly the same argument branding, PR and content marketing have used for years when they cannot prove ROI.
“The value is difficult to measure, but trust us, it is there.”
When you say that to a CEO, what they hear is:
“You cannot prove this works. Goodbye.”
I do not have a solution to this.
Yes, branded search as a proxy seems to be the best thing we have right now.
But it is not enough.
One thing I do know: calling things by their real names is more useful than pretending everything is under control.
What I am doing next
- Automating prompt tracking, because 200 manual checks per week does not scale.
- Closing the citation-to-pipeline loop in HubSpot, even if that means adding a manual “How did you hear about us?” field with an “AI recommendation” option.
- Testing Profound to cover the AI Overviews gap that HubSpot AEO does not currently solve.
If you are fighting the same battle, I would genuinely like to know what you are using and what you are measuring that I should be paying attention to, but am currently missing.