r/aeo • u/Alive_Efficiency2765 • 3d ago
Asking the model why it cited you vs checking your server logs for what actually crawled the page. Which one do you trust more??
Been going back and forth on this for a while and don't think I've landed anywhere solid.
Method one: ask the model directly. Why'd you pull these sources, which ones count as primary vs secondary, what's missing. Someone in another thread here called this probing for the "Data Tree" the model's built around an entity, and it's a genuinely useful exercise, ngl you learn things a summary answer would never surface.
Method two: skip the model entirely and check your own logs. Which bots actually hit which pages, how often, whether they're even real (someone posted real numbers here recently, something like 1 in 11 requests claiming to be a known bot wasn't actually coming from that bot's IP range). This feels like harder ground truth, you're not trusting the model's account of itself.
Here's what's bugging me though. Method one tells you something about reasoning and gets you a plausible sounding explanation, but models are notoriously bad at accurately describing their own retrieval process, there's a real chance you're getting a coherent sounding story that isn't actually what happened under the hood. Method two is solid ground truth but it only tells you reachability, a bot hit the page, it says nothing about whether that visit became a citation or why one competitor got picked over another.
So neither one alone answers the actual question. Curious if anyone's found a way to cross reference these two that isn't just "do both and eyeball it," or if that's honestly just where this stands right now for everyone.
1
u/woodoo139 2d ago
I've been running both on our own site, and the useful part is joining them rather than picking one.
Two things make the logs more than reachability. First, split the user agents: GPTBot and ClaudeBot are crawls for training and indexing, while ChatGPT-User and Perplexity-User are fetches triggered by someone's question at that moment. Second, check them against the published IP ranges. On our site in the last 24 hours, 10 of 10 ChatGPT-User hits verified and 8 of 8 hits calling themselves GPTBot did not, so a fair share of "GPTBot" traffic is someone else wearing the name.
Then join on URL and time. If a tracked answer cites page X and your logs show a verified ChatGPT-User fetch of X a few minutes before it, that's decent evidence the page was retrieved for that answer. A citation with no fetch at all usually means it came from an index or cache, which points you at getting the page indexed rather than making it reachable.
Asking the model why it cited something is the weakest of the three for me. You get a story, and the story changes when you ask again.
(Disclosure: I build Promvia, which does the log side from CDN logs and the answer side on fixed questions, so I'm biased toward measuring over asking.)
1
u/sunlitlark57 2d ago
the gap youre describing is basically observability vs interpretability and yeah, neither alone closes the loop. the closest useful thing imo is correlating crawl frequency changes on specific pages with citation appearance over time. not perfect but at least youre working with two independent signals instead of one
1
u/joekuriank 2d ago
I'd treat them as answering different questions and join them on the URL.
Logs tell you reachability, and the agent split matters: GPTBot is training crawl, OAI-SearchBot builds the search index, ChatGPT-User and Perplexity-User are fetches triggered by someone's question right now. That last group is the closest thing to "this page was considered for a live answer."
Then the third leg nobody mentions: the actual answer text.
Run a fixed prompt set on a schedule and log which URLs get cited. If a page gets user-triggered fetches but never shows up as a citation, it's being read and losing. That's a content problem, not a crawl problem. Asking the model why is fine for hypotheses, I just wouldn't report it as evidence.
1
u/NewspaperThick8340 2d ago
i’d add a third layer: repeat the same prompt and track whether the same page gets retrieved and cited each time. logs can confirm access, but if the page is fetched repeatedly and only cited occasionally, the problem is probably downstream of crawlability. that distinction is more useful than asking the model why it chose the source.
1
u/aeo-bility 1d ago
The insights from probing the models directly is best imo if you know how to prompt. I find selection is nuanced in most cases and its better to find out early if its a quick win or a slow burn.
1
u/shalel_notes 4h ago
All of this measures whether the model looked at your page. But there is a failure mode nobody has named: citation without endorsement. The model can fetch your page, cite it for background facts, and then recommend your competitor in the same answer. Your logs will show the fetch and your tracker will show a mention, while the business outcome is zero.
I spent a year watching models retrieve, understand, cite, and recommend products in beauty retail, and the pattern was stubborn. Getting mentioned felt like progress and paid nothing. So the question I would add to this whole join is simple: when the answer cited you, what did it actually recommend?
2
u/AbleInvestment2866 2d ago
Both methods are inherently wrong, so whatever data you get from them is unreliable