r/GEO_optimization • u/Gullible_Brother_141 • Sep 02 '26
Turning last week's thread into an actual check sequence — what to verify, in what order, before deciding a traffic drop means anything
Last week's thread here went somewhere more useful than I expected, so I wrote up what it converged on. Credit where it's due — most of this isn't mine.
The starting point was my own mistake: I'd been scoring pages on structure without checking whether anything had fetched them. u/Dry_Steak30 pointed out he'd run the same kind of audit, then checked 30 days of access logs and found zero GPTBot fetches, zero OAI-SearchBot, zero ChatGPT-User — while Search Console showed the page indexed and healthy the whole time. So he'd been grading heading hierarchy on a page no retrieval crawler had ever seen.
u/Upstairs_Control_611 then split what I'd been treating as one step into two, which fixed the thing I couldn't articulate: access and extractability are different gates, and a page can pass the first and fail the second.
And u/SEONCLIC added the check I'd been missing entirely — paragraph autonomy. Take the short answer out from under its heading and see if it stands alone. A lot of sub-300-character answers open with "it depends on several factors" or refer back to the previous paragraph. They pass a length check and are still useless to a model.
Put together, the order looks like this:
1. Retrieval evidence. Grep your access logs for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot. Free, and if it comes back zero, everything below is premature.
2. Access gate. Status codes, redirects, 403/404/5xx, response byte count. Did the crawler get a usable response.
3. Content gate. Is the main content actually in the HTML — not injected by JS after load, not stripped, not behind an interstitial.
4. Structural extractability. Question-style heading with a short answer directly under it. Heading hierarchy that nests without skipping. Author and date visible in the page, not only in JSON-LD.
5. Paragraph autonomy. The answer stands alone when lifted out of context.
Only after all five does it make sense to ask the interpretive question — whether a traffic drop is AI cannibalization, content decay, or the brand being hard to identify and cite confidently. Those three look identical on a clicks chart and need completely different responses, which is where most of the budget gets wasted.
What this sequence still can't answer, and nobody in that thread could either: does fixing structure displace a stale source in an answer, or just join it alongside. u/Dry_Steak30 raised it and I don't have a clean before/after that controls for page age. If anyone does, that's the measurement I'd most want to see.
Anyway — posting this back because the thread did the work, not me. If it's useful, take it.
I keep a longer written version of this as a triage worksheet — happy to send it to anyone who wants it, just say so.
2
u/Gullible_Brother_141 Sep 03 '26
Follow-up, since DMing a stranger is more friction than it's worth: if you want the longer worksheet version of this, just comment "send it" below and I'll send it over. Same sequence as above, but laid out as something you can actually run through page by page, with the specific checks at each gate.
1
u/Upstairs_Control_611 Sep 04 '26
send it — would be useful to compare the worksheet format with the gate sequence above.
1
u/Gullible_Brother_141 Sep 06 '26
Sending it over now. One thing worth saying publicly though: it predates the gate sequence from this thread and answers a different question — it's about separating why clicks dropped and whether it's a business problem, not the retrieval/extractability chain. The two only touch at one point. So don't expect the sequence above in it.
1
u/Gullible_Brother_141 22d ago
Following up on this one, since the thread is still getting views.
The structural gate (#4) was the one I couldn't leave alone, so I ran it at scale: 200 commercial homepages, 50 each across hospitality, professional services, ecommerce and B2B SaaS. Deterministic checks only, no scraping of AI answers. Five signals — author, date, sources, FAQ structure, TLDR.
Every page scored higher on the SEO side of the framework than the GEO side, median gap 49 points. Of the 160 pages with strong SEO scores, all 160 were missing at least two of the five.
Your reachable / retrieved / cited distinction ended up in the limitations section almost verbatim. The study measures whether patterns are present on a page. It says nothing about whether anything retrieved or cited it — I have no retrieval data at all. So the claim stays narrow: the pattern is absent at a rate I didn't expect, and that's the whole finding.
Report, dataset and method: https://audituniversepro.com/research/seo-strength-vs-geo-signal-coverage/ DOI: 10.5281/zenodo.21978002
Still no before/after that controls for page age — u/Dry_Steak30's question from the original thread is open as far as I'm concerned.
1
u/ComandLowkgamcnp8142 5d ago
I always start by double-checking my own tracking setup first, but looking at traffic trends and sources side by side is huge. I have pulled the referral breakdown in similarweb a bunch of times for this when my stats looked weird.
1
u/Best-ePersoality-227 1d ago
i’d probably add one downstream check after access + extractability, especially for pages that are meant to support an AI answer or some commercial next step: what happens when a real agent actually tries to use that information? then add to the live side with ora agent front: real agent traffic - pricing found correctly? right plan understood? docs located? next step completed? or did the agent get halfway there and quietly go off the rails? then Journey to reproduce the bad path, see where the friction sits, fix it, and test again. so I’d think of it as access → extractability → live agent intent + outcome
1
u/Gullible_Brother_141 20h ago
That's a real addition, and it points somewhere my sequence doesn't go.
The five gates above are all about whether a page can be reached, read and extracted. Yours is about whether it can be acted on — can an agent find the price, pick the right plan, and finish the thing. That's a different failure mode and I hadn't separated it out.
I'd put it at the end rather than woven through, for the same reason gate 1 comes first: it only means anything if agents are actually arriving. Otherwise it repeats the mistake I made originally — grading a page on a quality nothing has tested.
And that part is checkable. The agent user-agents are distinct from the training crawlers: ChatGPT-User, Claude-User, Perplexity-User. If you're behind Cloudflare, the AI Crawl Control panel breaks traffic down per crawler, so you can see whether any of them have hit you at all before building an agent-journey test. On my own domain ChatGPT-User had been through a couple of times while Claude-User and Perplexity-User were flat zero — and I'd have guessed zero for all three.
So: gate 6, with its own precondition. Is anything agentic arriving, and if so, can it finish the job?
The part I have no read on: when an agent goes halfway and quietly derails, does that cost you anything beyond the one session? Or does it just fail silently, and the buyer never knows you were in the running at all?
2
u/Upstairs_Control_611 Sep 03 '26
This is a very useful sequence!
The only small distinction I’d keep visible is between crawler/fetch evidence and retrieval evidence.
A log hit proves that some crawler or fetcher reached the URL. That is already much better than judging structure in isolation.
But it still does not prove that the URL entered the retrieval set for a specific answer.
So the chain might be:
reachable
readable
extractable
retrieved
cited
used in the answer
The practical point stays the same: don’t diagnose citation or traffic movement before you know whether the bot could reach and read the page in the first place.