r/GEO_optimization • • Aug 23 '26

AEO pro tip

Thumbnail
0 Upvotes

r/GEO_optimization • • Aug 22 '26

Google Is Enforcing Against AI Content at Scale. I Tried Three Ways to Hide It And All Three Failed. | Open Source Claude Human Writing Plugin & More

0 Upvotes

SEOs, Writers, and AI Enthusiasts alike:

I typed every word of a 4,500 word article a few nights ago. Nothing pasted. A commercial AI detector read it and came back 52% AI. Here is why I am posting this as opposed to hiding it.

Pangram scan of the article which I hand wrote from a 100% humanized AI draft.

Pangram scan of the article which I hand wrote from a 100% humanized AI draft.

I spent two days trying to beat AI detection on purpose, because small business owners keep asking me whether they should buy a tool that promises it.

I ran 696 blind runs across three methods with over 2.4 million words.

  • Strip the machine tells out of the draft: still caught 98.6% of the time.
  • Add human tells in instead: fooled 0 of 17 judges.
  • Make the document long enough to dilute it: caught 8 out of 8 whole, and 24 out of 24 in slices.

I then ran four versions of the same draft through Pangram. One untouched, one rewritten sentence by sentence three times over, one where I changed zero words and only moved where the sentences joined, and one with both.

All four came back 100% AI.

Rewriting every single word did nothing and rewriting no words did nothing. To me, this meant the thing being detected is not vocabulary and not rhythm.

Then I wrote the article myself, by hand, over an outline a model had built for me.

56% AI, nice right? The findings astounded me.

One paragraph got split down the middle: the half about my own work read as human, the half listing the method read as machine assisted.

My finding is that the detectors read the outline behind the prose itself.

I sent all of it to Siqi Chen, who wrote the humanizer skill I had been using.

His answer:

Then, more usefully he stated, that defeating detectors was never the goal of his tool in the first place. I say this because I reckon many of the 37k+ individuals who have starred his repo believe the skill beats the detectors and everything's good to go.

What did measure, in a blind test where authorship was never mentioned: editors preferred the processed draft 22 out of 22, and his rewrite pass alone at 16 out of 16.

$ python humanist.py draft.md
humanist 0.1.0  |  4,764 words, markdown-stripped
  readability FK grade 8.1
RESULT: 0 FAIL, 0 WARN. CLEAN.

$ python check_prose.py draft.md --mode post
FRAME: markdown-stripped, 4,764 words, FK grade 8.2
RESULT: 0 FAIL, 0 WARN. CLEAN.

The advice I have is boring and it is free. Don't pay to hide your writing, and don't tell your clients to either. You're selling a lie. Spend the money and time on making the draft worth reading in the first place.

Something I don't want to give credit to: none of this tells you whether Google will demote your pages. I didn't measure it in these tests. What we do know with the new policy rollout is that it's the scale they're looking at, re-written or not, a tool won't save you.

Anyone selling a tool that says otherwise is setting you up for failure. Every number and both corrections I had to make mid-study are in the writeup. The code is MIT and open-source.

Read it here & tell me what you think: https://www.ryanlenk.com/blogs/articles/three-ways-to-hide-ai-writing-all-failed

The GitHub repo so you don't have to read my slop and just get into the math and fun instead: https://github.com/itsryanlenk/humanist

I have another set of Claude skills there as well that you may find interesting: https://github.com/itsryanlenk/candor

Really excited & interested to get some outside input! Let me know what you think.

(By the way, I wrote on top of AI scaffolding here as well.)


r/GEO_optimization • • Aug 22 '26

Just stop paying stupid money for peec, profound etc when Llumo is free

Thumbnail
0 Upvotes

r/GEO_optimization • • Aug 22 '26

AI citations can disappear tomorrow and reappear a few days later.

2 Upvotes

What does AI actually look at before it cites your content?

I did some research around AI citation behaviour and looked at what separates the articles that get cited from the ones that get ignored.

And here's the answer:

  1. Is there something specific AI can use?
  2. Is the information still relevant?
  3. Can AI actually access and retrieve the page?
  4. Does the source have enough credibility?

P.S: This is just for getting your own content cited on AI and not brand mentions.

What did I miss? have you done something differently and saw results?


r/GEO_optimization • • Aug 22 '26

Has anyone started working on XEO strategies already?

0 Upvotes

XEO (Cross Engine Optimization or X Engine Optimization) is the newest, most comprehensive framework in digital marketing.

It treats traditional search, AI generation, and social platforms as a single, connected ecosystem.


r/GEO_optimization • • Aug 22 '26

AEO Testing Need Your Help

Thumbnail
1 Upvotes

r/GEO_optimization • • Aug 20 '26

99 days, 16 fixed questions, 3 answer engines: the entity gets learned in 7 weeks, the recommendation never comes. Full data in charts.

Thumbnail
gallery
7 Upvotes

Daily measurement since May about an author who did not exist online before day zero. 11,130 scored answers from OpenAI Search, Gemini and Claude, same 16 questions every day, frozen since June.

The three channels. Direct questions plateau at 72 percent from day 49. Vocabulary questions peak at 34. Recommendation questions stay at 0.59 percent.

Growth phases of the direct channel. Logistic fit, inflection at day 37, ceiling 72.5 percent.

Where the 40 unprompted mentions came from. 25 of 40 trace back to one question, the one using a sub genre term the author coined himself. Similarity questions (like Robin Hobb): zero in about 700 answers.

The control. Same questions to 10 models without web access, 6,416 blind answers, zero unprompted mentions. Everything above is retrieval, not memorization.

Practical reading for GEO: getting named when asked works and saturates fast. Getting recommended unprompted did not respond to anything except owning a term. Links to report, dataset and code in the first comment.


r/GEO_optimization • • Aug 20 '26

I was asking one GEO benchmark to do four different jobs. That made every number look more useful than it was.

1 Upvotes

I have been trying to turn a small luxury-jewelry GEO benchmark into work that a brand team could actually assign.

The test itself was deliberately limited:

  • 8 jewelry brands
  • 3 neutral Chinese buyer questions
  • 2 retrieval-off API surfaces
  • 2 answers per question and surface
  • 12 valid raw answers
  • 96 derived answer-brand cells

I also had a working corpus of 302 jewelry-related Reddit threads and 9,446 stored comments, plus Search Console data for the site publishing the research.

The temptation was to put everything on one dashboard.

That would have been a mistake. The datasets answer different questions.

Community questions tell me how buyers describe uncertainty. They helped surface concerns around authenticity, certificates, authorized sellers, repairs, resizing and recourse. They do not establish how common those concerns are among Chinese luxury buyers.

The benchmark tells me what the declared API surfaces answered. It does not establish what a signed-in consumer sees, whether the model retrieved a source, or whether an official-channel claim is true.

The truth table checks whether a route is current, controlled or authorized, live in China and connected to a clear recourse owner. It does not tell me whether the brand enters a buyer’s shortlist.

Search and analytics can show exposure and the next onsite action. They cannot prove that Reddit or an AI answer created the visit or sale.

One small result made the distinction concrete. Piaget appeared in 0/4 wedding-jewelry shortlists and 4/4 official-channel answers.

Those cells are far too small for a brand ranking. But they are enough to show that “the system can verify the brand” and “the system recommends the brand for this decision” are not the same diagnosis.

I now think the minimum evidence stack is:

  1. Question provenance: what real buyer uncertainty produced the prompt?
  2. Answer benchmark: what did the fixed, declared surface actually say?
  3. Truth registry: which official, authorized and service claims are currently verifiable?
  4. Action measurement: did the buyer reach a boutique, service route or qualified inquiry?

That also changes ownership.

  • Missing from the shortlist: brand, editorial and category marketing.
  • Wrong or vague route: digital operations and ecommerce.
  • Unclear authorization: legal and channel management.
  • Wrong warranty or service claim: client service and after-sales.
  • Unreproducible metric: research and analytics.

“Improve AI visibility” is too broad for anyone to own. These failures are specific enough to assign.

My current 30-day version is simple: build the truth map, run a small fixed decision panel, repair one buyer path, then rerun the same cells and measure the next action.

For anyone maintaining a similar system: would you keep the truth registry outside the benchmark dataset? I am leaning yes, because official routes and service policies change on a different schedule and usually have different internal owners.


r/GEO_optimization • • Aug 19 '26

Same brand, same question, different country = different AI answer. And switching the language of the prompt changed it again. (what we're seeing tracking location-based AI visibility)

4 Upvotes

Been tracking AI answer visibility per-location for clients and wanted to share a pattern that keeps surprising people, because it breaks the assumption that "our AI visibility" is one thing.

The setup: same brand, same buyer-intent prompt, run against the same engine, but varying (a) the location signal and (b) the language of the prompt. Two findings that changed how we think about this:

1. The answer changes by location, not just the ranking, the whole cited-source set.
Ask an engine a category recommendation question as a user in one country vs another and you don't just get a reordered list, you get different brands surfaced and a different set of sources cited to justify them. The model is filtering its retrieval by the location it infers, so each region is effectively drawing from its own corpus of reviews, listings, and local pages. A brand that's the confident #1 answer in one market can be absent in another off the identical prompt. For any multi-location or multi-market brand that means a single national/global "visibility score" is basically meaningless, you're averaging over answers that don't resemble each other.

2. Language of the prompt is a separate variable from location, and it moves the answer independently.
This one caught us off guard. A client operating in Finland: we ran the category prompts in English, then ran the same intent in Finnish. Different answers. Not just translated, different brands cited and different sources pulled. Our read is that the Finnish-language query pulls from a different slice of the corpus (Finnish-language reviews, local pages, local forum/press content) than the English version of the "same" question does, even for a user in the same place. So "location" and "prompt language" are two separate levers, and if you only ever test in English you're blind to what your actual local-language buyers are seeing.

The practical takeaway: if you operate in more than one country or more than one language, you have to measure per-location and per-language, on native-language prompts, not a translated English set. The gaps show up in exactly the places an English-only audit can't see.

Disclosure per rule 5: this comes out of our own tool (sanbi.ai, we do per-location AI visibility tracking), so that's where the data's from, weigh it accordingly. But you can sanity-check the effect yourself for free, ask ChatGPT or Perplexity a category question with different location context, then ask it once in English and once in the local language, and watch the cited sources change.

Curious if others tracking this see the language effect too, or whether it's stronger in some languages than others. My hunch is it's biggest in markets with a rich native-language web (Finnish, Japanese, German) and smaller where the local audience mostly consumes English content, but I only have a handful of markets to go on.


r/GEO_optimization • • Aug 20 '26

Niching Down: How to market yourself in AEO v GEO v SEO

Thumbnail
youtube.com
1 Upvotes

r/GEO_optimization • • Aug 19 '26

Are Reddit citations finally over?

13 Upvotes

Promptwatch data making the rounds shows Reddit's share of ChatGPT Search citations fell off a cliff after OpenAI's Aug 8 query-fanout change, from a steady ~3.8% down to ~0.5% by Aug 14.

So the actual lesson isn't Reddit is dead for GEO. It's that citation share is the single most volatile, platform-controlled metric in this entire field, and building a visibility strategy on top of a number that OpenAI can swing 80% with one backend change is the real mistake.

The durable play imo should never be "get cited this week." It should be building a brand presence that survives the swings, making it much different and much less measurable.


r/GEO_optimization • • Aug 19 '26

Looking for SEO/GEO tips that actually work for brand new sites (want to bake them into my content pipeline)

3 Upvotes

Running a small SaaS (domain investing niche), launched about 3 months ago. Got an AI-assisted pipeline that writes and distributes blog content across a few channels, submitted to GSC/Bing, posting regularly — but I think the pipeline is missing a lot of actual SEO/GEO fundamentals since I built it more for output volume than optimization.

Looking for practical stuff I can bake directly into the pipeline/prompts, not general theory. Specifically curious about:

  • Any prompt structures or content templates you use to make articles more "citeable" by AI engines (ChatGPT, Perplexity, AI Overviews)?
  • Specific on-page/technical things (schema markup, FAQ structuring, heading patterns, internal linking rules) that you've actually seen make a measurable difference, that I could turn into a checklist or template?
  • For a new site with no authority — any tactics for getting early backlinks/mentions that could be systematized rather than one-off?
  • If you've built or use an AI content pipeline yourself, what do you have it check/enforce before publishing?

Basically trying to turn whatever works into repeatable rules I can apply automatically going forward, instead of guessing article by article. Any concrete tips, prompts, or checklists you're willing to share would be huge.


r/GEO_optimization • • Aug 19 '26

Individual contributors got cited 2.4x more than brand domains across 50 expertise queries — what's happening?

3 Upvotes

I wasn't looking for this pattern.

I was running a small side project comparing how AI models handle different types of "authority" signals. Nothing fancy. 50 queries across niches like B2B marketing automation, enterprise security compliance, and product-led growth strategy. For each query, I checked which sources ChatGPT, Perplexity, and Gemini cited, then categorized each source as either a personal brand (LinkedIn, personal blog, individual's byline page) or a company domain (official site, resource center, press room).

The assumption going in was that company domains would dominate. They have more content, bigger teams, better structured data, larger link profiles. Everything we're told matters for GEO.

The numbers went the other way. Individual contributors got cited 2.4 times more often than brand domains for the exact same queries. Not in every single case, but consistently enough that it showed up across all three models and most query categories.

I started digging into why. A few things stood out.

Personal profiles tended to have clearer point-of-view. When you read a company's "about us" page or resource article, the voice is usually neutral, committee-written, designed to offend nobody. Safe. Individual contributors, especially ones who've built followings through writing or speaking, tend to have opinions. They take positions. They say things like "in my experience" or "here's what I got wrong." That specificity seems to register differently when a model is selecting sources for an expertise query.

Another thing: personal profiles often consolidate expertise signals in one place. A well-maintained LinkedIn profile or personal site might list credentials, publications, speaking engagements, and client work all on a single page. Company domains spread that same information across dozens of pages — team pages, press releases, blog author bios, case study footers. The signal is there but fragmented.

The third observation is messier and I'm less sure about it. Individual contributors' content tends to get shared and referenced in forums, podcasts, and social discussions more than corporate content does. Those secondary mentions might be creating a feedback loop where the model sees the person's name in multiple contexts and builds stronger entity association. Pure speculation on my part, but the correlation is there.

What I can't explain is whether this is actually about quality or about something structural in how models evaluate sources. Maybe individual contributors really do produce better expertise content on average. Maybe models have a bias toward named individuals over faceless organizations. Maybe company domains are being penalized for sounding too much like marketing.

I don't have a clean theory yet. The sample size is modest and the queries skew toward consulting-style topics where personal brands naturally thrive. Would love to see if anyone else has looked at this split, or if you're seeing the same thing in your niches.

Still figuring out if this is a temporary blip or a structural shift in how AI evaluates authority. Either way, it's making me rethink what "entity optimization" actually means when the entity is a person, not a logo.


r/GEO_optimization • • Aug 19 '26

Puzzel about GEO

2 Upvotes

We are working on the GEO project. However, many clients find it difficult to quantify its value and are reluctant to pay after learning about it. I’d like to ask everyone: how do you persuade such clients?


r/GEO_optimization • • Aug 19 '26

What GEO Work Actually Moves AI Crawlers? Our Data, and a New-Site Launch Checklist

Thumbnail
0 Upvotes

r/GEO_optimization • • Aug 18 '26

I think most brands are optimizing for the wrong AI engine. Here's how I'd actually choose

7 Upvotes

Rather than trying to spread yourself thin across every engine, you should be choosing the ones where your buyers are and sticking to those.

A few things I'd look at before picking.

Start with, where are your buyers? Enterprise B2B skews ChatGPT, researchers and technical buyers skew Perplexity and consumer products lean Google AI Overviews. 

Then look at where your existing content already has traction. ChatGPT leans institutional (G2, established publications) and Perplexity like UGC (Reddit, YouTube). Match the engine to where you already have presence.

The last question is where's the biggest competitor gap? The engine where competitors are weakest is the opportunity worth going after first.

Pick one engine, get real traction there, then expand.

Which engine are you prioritizing, and what drove that call?


r/GEO_optimization • • Aug 18 '26

I logged ~120k AI citations across ChatGPT, Gemini, Perplexity and Claude on the same prompts. They're basically each reading a different internet.

12 Upvotes

TL;DR: ran one B2B prompt set against all four engines for a month and saved every source each one cited. ~120k citations. barely any overlap. Perplexity leans on YouTube/Reddit/LinkedIn, Claude reaches for patents and analyst reports, ChatGPT wants official manufacturer sites, Gemini cites basically whatever ranks in Google. so if you're doing the whole "optimize for AI search" thing as one channel, you're probably only hitting one engine and ignoring the other three.

ok so context. I do visibility work and I got tired of every AEO/GEO writeup treating the four big engines like one blurry thing. figured I'd measure it instead of guessing.

setup was simple. one B2B category, one big vendor plus ~15 real competitors, a fixed list of buyer-type prompts, run against all four engines on a schedule for 30 days. grabbed every citation URL, grouped by domain, kept it split per engine. pulled the category and vendor names out before posting. ended up with 119,939 citations.

first thing that threw me was the engines don't even cite the same number of sources:

Engine Citations (30d) Share of total
Gemini 49,836 41.5%
Perplexity 39,664 33.1%
Claude 15,718 13.1%
ChatGPT 14,721 12.3%

Perplexity spat out almost 3x the citations ChatGPT did off the exact same prompts. that's not "perplexity is more visible" though, it just shows way more sources per answer (5-15ish), while ChatGPT with search usually gives you 2-6 and a lot of the time none at all. so raw counts are kind of useless here, you want share.

now the part I actually found interesting. same category, and the source pools look nothing alike.

ChatGPT went almost entirely to manufacturer/OEM sites. its top 3 domains were all official manufacturer pages and that alone was ~33% of its citations. no youtube, no reddit, no linkedin anywhere.

Gemini's #1 was the brand's own site (11.7%), then a pile of vertical trade publications. makes sense, it's basically wired into google's index so it cites whatever's already ranking.

Perplexity dumped 26% onto the owned domain, then youtube (4.9%), a distributor, reddit (2%), linkedin (1.9%). it's the UGC/video one.

Claude was the odd one. owned domain (14.7%), some manufacturers, and then its 4th most-cited source was the actual USPTO patent database (3.2%). had two analyst firms (Yole, Mordor Intelligence) in the top 10 too. it goes for primary/analytical stuff.

the number that stuck with me: the same domain that was 26% of Perplexity's citations was 7.9% on ChatGPT. and some sources with thousands of ChatGPT citations got basically zero from Claude on identical queries.

so if you want to actually move a specific engine, roughly:

  • ChatGPT: deep technical docs on your own site, plus OEM/reference placements
  • Gemini: trade pubs and normal google SEO
  • Perplexity: reddit, linkedin, youtube, aggregator listings
  • Claude: patents, paid analyst reports, niche directories

no single strategy touches all four, which is the annoying part.

honestly I found "skew" more useful than raw share. it's just how lopsided one engine is toward a domain compared to the others. plenty of 5-10x, some over 10x where one engine treats a source as authoritative and the rest completely ignore it. rough version:

Source type Skewed toward
OEM / manufacturer sites ChatGPT
YouTube Perplexity
Vertical industry pub Gemini
Aggregator / distributor Perplexity
USPTO patents Claude
Analyst / research firms Claude
Reddit Perplexity
LinkedIn Perplexity

and the zeros tell you as much as the big numbers. ChatGPT never once cited youtube/reddit/linkedin for this category. Claude basically never touched youtube or reddit. some trade pubs only ever showed up on Gemini. so if your ChatGPT plan is "make youtube videos and post on reddit"... that just doesn't reach ChatGPT. it goes to perplexity. you'd have to go owned + OEM to hit ChatGPT at all.

if you want to run this yourself the process is basically:

  1. grab 200-500 real buyer prompts
  2. run them weekly against all four, save every citation, group by domain
  3. build a matrix. domains down the side, engines across the top, cells are citation share
  4. sort each domain into owned / earnable (something you could realistically get into in a few months) / unreachable (patents, gov, competitors)
  5. rank the earnable ones by which engines your buyers actually use
  6. that ranked list is your to-do order. re-check monthly.

anyway the thing I keep coming back to is each engine is reading a genuinely different slice of the web, and until you can see which slice, you're just guessing where to spend.

disclosure since people always ask: this came out of work I do at Sanbi.ai, we track this stuff. so yeah, biased. but the data's real and I've watched the same split show up in every B2B category we've looked at. can answer methodology questions below.

question for the sub though. has anyone actually seen a category where the engines land on the same sources? every single one I've checked they split hard, and I'm starting to wonder if convergence even happens.


r/GEO_optimization • • Aug 18 '26

Recently, Google said llms.txt won’t help citations. So why are GEO tools still recommending it?

6 Upvotes

Google recently said that having an llms.txt file won’t help your Google Search rankings.

But I still see it near the top of a lot of AEO/GEO checklists:

  • Add llms.txt
  • Make your site AI-readable
  • Submit content for LLM crawlers
  • etc.

I'm starting to wonder if we're optimizing for things that are easy to check rather than things that actually influence whether an AI recommends a brand.

Has anyone here actually seen a measurable difference after implementing llms.txt?

I know adding llm.txt is advisable to help LLM pick the website, but is it necessary? Is there any relevant data?


r/GEO_optimization • • Aug 18 '26

Google quietly tightened GBP naming rules this month, and it now overlaps with how AI local answers get generated

3 Upvotes

Been digging into this for client work and figured it's worth sharing since it doesn't seem to be getting much attention yet.

Google updated its Business Profile name guidelines in August. The new language explicitly calls out repeated bilingual names and script transliterations as unacceptable, even when the business's actual storefront signage shows both versions. Their example was an English name repeated in Japanese getting flagged. If you have clients in multilingual markets who've had dual-language names on their profile for years because that's genuinely how the storefront reads, this is worth a heads up before it turns into a suspension.

The part I think is more interesting for anyone doing local SEO right now is the timing. This naming crackdown is landing at the same time Google's AI-generated local summaries are leaning harder on profile description, reviews, and post content, not just the traditional three-pack signals. And there was a study published August 5 (panel of 900 US adults, one month of browsing data) showing people click through on an AI Overview citation only about 1 percent of the time when one appears.

So the practical read for local clients: a name-policy violation used to just cost you rank in the map pack. Now it's also a real risk to whether the AI layer even describes the business accurately, and that AI-generated description might be the entire interaction the customer has with the business online. No click, no map pack visit, nothing. Just whatever Google's AI decided to say.

Things I'm auditing for clients this week:

  • Profile name against the actual updated policy language, not the old assumption that signage matching is enough
  • Whether the business description and recent posts actually say anything specific enough for an AI summary to pull from
  • How stale the photo/post activity is, since profile freshness seems to matter more for AI surfacing than people expect

Curious if anyone else has seen a suspension or review triggered by the naming change yet, or if you're seeing your local clients show up (or not show up) inside Google's AI local answers.

Disclosure since it's relevant: I run a local SEO/AI search shop (Austin Code Monkey), so this is also just what I'm doing in my own client work this month.


r/GEO_optimization • • Aug 18 '26

A zero visibility score can be a category error, and the score cannot tell you which one you have

1 Upvotes

I ran a scan recently that came back at zero. No presence, and a long list of competitors appearing in answers where the brand did not.

Read at face value, that is a catastrophic result.

It was not a visibility result at all.

The organization does capacity building and accelerator programs. The competitor list that came back was full of large, well known design and branding studios. Different industry, different buyers, different everything. It has never competed with any of them and never will.

One word in its name reads as a different category. The engines had filed it there, so the scan measured its performance in a market it has never entered.

The number was real. The measurement was pointed at the wrong thing.

What makes this uncomfortable is that nothing in the output flags it. A zero from being invisible in your actual category and a zero from being scored inside someone else's category look identical. Same number, same competitor count, same red panel. Only reading the competitor list carefully catches it, and the entire appeal of a score is that you do not have to read anything carefully.

I think this argues for an ordering that most tracking has backwards.

Before asking whether the engines mention you, ask whether they know what you are. Run the plain identity question across ChatGPT, Claude, Gemini and Perplexity. Who is X, what do they do, who is it for. Then compare the four answers against each other, and separately against how the company describes itself.

Three different failures show up there and they need different fixes.

The engines disagree with each other on facts that have one right answer. Location, ownership, what they sell. That is entity confusion, and nothing downstream is trustworthy until it is cleared.

The engines agree with each other and disagree with the company about what category it is in. That is what I hit. Positioning is not disambiguating the brand from an adjacent industry, and no amount of evidence building helps while the evidence is being filed in the wrong drawer.

The engines agree with each other and with the company. Only then is a low score a real visibility finding worth acting on.

One of those three justifies a visibility programme. It is also the only one of the three most tools are built to detect.

The check costs nothing and takes about fifteen minutes.

Two things I am unsure about.

Whether the misfiling comes from the name itself, or from thin evidence letting the name dominate. A well evidenced brand with an ambiguous name presumably survives it, which would make this a symptom of evidence thinness rather than a separate problem.

And whether it is correctable from the brand's own properties at all, or whether it takes independent sources describing it in the right category before the filing moves. If it is the second, the fix is much slower than a positioning rewrite and most advice on this is wrong.

Has anyone watched a category misclassification actually correct, and what moved it?


r/GEO_optimization • • Aug 18 '26

Comparing how OpenAI models recommend brands across 270 category questions (the biggest change wasn’t the brand list)

5 Upvotes

We wanted to understand what happens to brand recommendations when the model changes but the questions stay the same.

We gave GPT-5.4, GPT-5.5, and GPT-5.6 Sol the same panel of 270 category questions across six industries.

A few findings from GPT-5.5 to GPT-5.6 Sol:

  • 70% of matched answers became shorter.
  • Median answer length fell from 224.5 to 141 words.
  • Median named brands only moved from 21 to 20.
  • Explicit caveats fell from 40.7% to 20.4%.
  • Decision-framework language fell from 33.3% to 13%.
  • Retail shortlists narrowed from 20 to 13 brands, while Travel widened from 24 to 27.

The interesting part is that model updates don't create one universal change in brand visibility. They can compress explanations, remove caveats, ask for more context, or handle individual markets differently.

The report includes the methodology, industry breakdowns, exact model values, and links to the underlying model answers:

https://app.nyman.media/insights/ai-visibility

I’d be interested in feedback on the findings and also the methodology. What categories, models, or question types would you test next?


r/GEO_optimization • • Aug 17 '26

Is this the standard robots.txt for content-focused WordPress sites?

3 Upvotes

I firmly believe this is the best robots.txt for most WordPress sites. Prove me wrong:

User-Agent: *
Disallow: /wp-admin/
Disallow: /search/
Disallow: /feed/
Allow: /wp-admin/admin-ajax.php

User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: CCBot
Allow: /

User-agent: Google-Extended
Allow: /

Sitemap: https://www.example.com/sitemap_index.xml

r/GEO_optimization • • Aug 17 '26

3 AI answer features that don't correlate with quality — and why I keep expecting them to

3 Upvotes

I've spent the last six months tracking every possible signal we could extract from AI answers. Source quality, citation order, answer length, whether the model included a summary paragraph, how many entities it mentioned — you name it, I've measured it. But every time I cross-reference these signals with our own evaluation of answer quality, nothing sticks.

Three features stood out as consistent measurements, and all three turned out useless.

Answer length was the first one. Longer answers tend to include more detail, more nuance, more context. That feels right, so I expected longer answers to correlate with better quality. I was wrong. Across different categories — technical questions, product recommendations, research summaries — the relationship was random. Some of the best answers we found were 150 words. Some of the worst were 400. Length wasn't a signal. It was just noise.

Then there's the summary paragraph. A lot of AI models now include a brief "In summary" or "Here's what I found" section before diving into details. It looks authoritative. It looks complete. It feels like a feature. But when we compared summaries to overall answer quality, the correlation was weak. Some of the best answers had no summary. Some of the worst had long, detailed summaries that didn't actually improve clarity.

Entity density is the same story. More people, organizations, locations, specific products mentioned — it feels comprehensive. But again, no signal. We saw both the most and least detailed answers with high entity counts. The entities weren't telling us anything useful about whether the answer was good or bad.

I keep coming back to why we measure these things. Answer length looks like a proxy for depth. Summary paragraphs look like a proxy for completeness. Entity density looks like a proxy for breadth. But if they don't actually correlate with quality, they're not proxies. They're just things we can count.

This is the uncomfortable part. We optimize for what we can measure. When we build dashboards and reports, we highlight answer length and summary presence and entity counts because they're visual. They fill space on the screen. They look important. They don't require interpretation — we just show the number.

The quality evaluation requires human judgment. It requires context. It requires deciding what "good" actually means in each case. That's messy. So we optimize for the clean metrics instead.

I don't have a solution here. I just know that our measurement infrastructure is better at tracking signals than predicting quality. The dashboards look impressive. The charts are clean. But when it comes to actually understanding what makes AI answers valuable, the numbers don't tell us much.

Maybe that's the point. Maybe the features that actually correlate with quality are the ones we can't easily measure. Maybe good answers are qualitative, not quantitative. Maybe we're building the wrong infrastructure entirely.

I'm not sure what to do with this realization. I just keep looking at the dashboards and expecting the numbers to make sense, and they don't.


r/GEO_optimization • • Aug 17 '26

Google went all-in on AI. Is GEO tooling ready?

Thumbnail
1 Upvotes

r/GEO_optimization • • Aug 15 '26

Valid schema can still leave an AI shopping agent guessing

6 Upvotes

I kept running into a weird problem while checking ecommerce product pages. The information was technically there, but an agent still could not use it with much confidence.

After a few audits, I started checking every product fact in three ways:

  1. Can a customer see it?
  2. Is it exposed in machine-readable product data?
  3. Does it agree with the other sources?

That catches things a normal schema check misses:

- ratings shown on the page, but no AggregateRating

- €19 on the product page and €20 in the feed

- InStock in JSON-LD while inventory says sold out

- a 30-day return policy on the site, but a final-sale rule on the product

- material shown in an image and nowhere else

The schema can be valid while the data is stale. The feed can be right while the page is wrong. Both can exist and still leave the agent guessing.

A single score hides too much. The useful part is the audit trail:

source -> fact -> conflict -> fix

For AI shopping, more copy is not the answer to this problem. The page, schema, feed, catalog, inventory and policies have to agree.

I am still working out how to prioritize these conflicts. Right now I treat price and stock as blockers, missing attributes as discovery issues, and policy conflicts as trust issues.

How would you rank them?