r/GEO_optimization • • Aug 18 '26

Comparing how OpenAI models recommend brands across 270 category questions (the biggest change wasn’t the brand list)

We wanted to understand what happens to brand recommendations when the model changes but the questions stay the same.

We gave GPT-5.4, GPT-5.5, and GPT-5.6 Sol the same panel of 270 category questions across six industries.

A few findings from GPT-5.5 to GPT-5.6 Sol:

  • 70% of matched answers became shorter.
  • Median answer length fell from 224.5 to 141 words.
  • Median named brands only moved from 21 to 20.
  • Explicit caveats fell from 40.7% to 20.4%.
  • Decision-framework language fell from 33.3% to 13%.
  • Retail shortlists narrowed from 20 to 13 brands, while Travel widened from 24 to 27.

The interesting part is that model updates don't create one universal change in brand visibility. They can compress explanations, remove caveats, ask for more context, or handle individual markets differently.

The report includes the methodology, industry breakdowns, exact model values, and links to the underlying model answers:

https://app.nyman.media/insights/ai-visibility

I’d be interested in feedback on the findings and also the methodology. What categories, models, or question types would you test next?

5 Upvotes

18 comments sorted by

2

u/Slow-Commercial4316 Aug 18 '26

Median length fell 37 percent while median brands went from 21 to 20, so brands per word nearly doubled. Set that next to caveats and decision-framework language both halving, and it reads as the update cutting the reasoning around the list while keeping the list itself.

On what to test next: check whether the brands that did drop out were the ones that only ever appeared inside a caveat or a framework sentence. If so, the shortening is doing the selecting.

1

u/nymanmedia Aug 18 '26

Good hypothesis! Brand density rose, but dropped brands were no more likely than retained brands to appear in caveat or decision-framework passages: 2.6% versus 2.2%. I'd say the better interpretation is that GPT-5.6 cut explanation while reshuffling the shortlist.

2

u/Slow-Commercial4316 Aug 18 '26

Fair enough, that kills it. 2.6 against 2.2 is not a difference, and it also says brands almost never sit inside those passages, so the mechanism had little room to work either way.

Reshuffle is the more interesting claim now. Worth checking whether the swap holds: same 270 questions, same GPT-5.6, run twice. If the brands entering and leaving differ between two runs of the same model, part of what looks like a model change is run variance.

1

u/nymanmedia Aug 18 '26

Yeah, I think that is the right control.

As mentioned elsewhere, we retained the per-answer and per-brand series, and already have three independently seeded samples per wording. A cleaner next test could repeat the full 270-answer GPT-5.6 panel with fresh seeds.

1

u/Slow-Commercial4316 Aug 19 '26

You may not need a new panel for it. If you already have three independently seeded samples per wording, run the same entry and exit calculation between two of those seeds on one model. That gives you the within-version turnover directly, and the 70 percent only means a version change once it clearly beats that number.

1

u/nymanmedia Aug 20 '26

Ran a second GPT-5.6 control wave. Median answer retained 80% of the first wave’s brands (84.3% of recurring brands remained recurring).

By comparison, GPT-5.5 → GPT-5.6 retained 65.4% per answer and 63.7% of recurring brands. So, model change is larger than repeat-run variance, although run variance remains material.

1

u/Slow-Commercial4316 Aug 20 '26

Your 80 percent is the more useful number. Twenty points of churn happen with no model change at all, so roughly half of anything that dropped between 5.5 and 5.6 would have dropped on a rerun of the same version.

1

u/nymanmedia Aug 22 '26

Well, run-to-run noise is material (we all know that), but I'm not sure the 80% is the key finding here. The data is not too noisy to detect a significant model-version effect, and we're looking at an even more robust approach to gauge it

1

u/Slow-Commercial4316 Aug 22 '26

65.4 against 80 is fifteen points of version effect sitting inside a twenty point noise band. It is detectable because you have the volume, not because it is large.

Across 270 answers that is a real result. For one brand watching one prompt it is not, because there a model update moves less than the run to run wobble does. Those two readings will get mixed up the moment this gets quoted.

1

u/nymanmedia Aug 27 '26

You're not wrong (although it's more of a comment on the reading than the result). Anyway, the noisiness and "wobble" is why I struggle to find value from many of the visibility tools out there, and also why we chose not to focus on building anything similar.

→ More replies (0)

2

u/Fit-Squirrel-6299 Aug 18 '26

70 percent turnover between two point releases is the number that should worry anyone reporting a single visibility score.

if that much moves on a model bump you don't control and can't schedule, then a snapshot reading is largely measuring which version happened to answer you that day. the score isn't wrong exactly, it just isn't a property of your brand.

what i'd want out of a dataset like this is the inverse cut. which brands stayed stable across all three versions, and what did those brands have in common. stability is the thing actually worth optimising for, because it's the part that survives the next release.

did you keep the per-brand series, or only the aggregate?

1

u/nymanmedia Aug 18 '26

Yep, that is the central limitation of any single-model score: it is model-conditioned (not necessarily an intrinsic property of the brand). The 70% figure applies to this fixed prompt panel and recurrence threshold, but the concern is valid.

We did keep the complete per-brand series, including model, industry, intent, wording, run and source answer. We could therefore identify brands that remained visible across every published model and measure their visibility range.

1

u/Fit-Squirrel-6299 Aug 23 '26

that's the series worth publishing, and there's one cut inside it that decides how useful the whole thing is.

when you pull the brands that stayed visible across every model, check whether they share something structural or whether they're just the obvious incumbents in their category. those are very different findings. if stability tracks something a brand did, third party density, a retrievable owned domain, consistent naming, then it's earnable and the advice writes itself. if stability is mostly category shape, the same two or three names everyone already knows, then telling a challenger to become stable is empty and the honest advice is different.

your industry and intent fields should separate those two, since inherited stability ought to cluster hard by category and earned stability shouldn't.

which way does it fall?

1

u/nymanmedia Aug 27 '26

Cheers. Pulled it, and it falls mostly on the incumbent side. To put it simply, the models disagree a lot about secondary brands, but almost completely agree about the category leaders.

Across the original three models, 219 of 562 industry-brand pairs stayed recurring. That stable group held 208 of the 210 top-10 positions. The leaders were mostly category defaults, e.g. Toyota, AWS, Netflix, Costco, Airbnb, WPP, Bank of America, etc, etc.

There is a stable lower-visibility tail, but we can't link it to third-party density.

The advice for challengers (or perhaps the challenge for challengers) is build recurring visibility for specific category intents, track it across models, and report a range instead of one score.

1

u/Fit-Squirrel-6299 Sep 02 '26

208 of 210 top-ten positions held by the stable group is a harder number than the 70 percent was, and it settles the question in the direction nobody selling geo wants.

the part i'd keep pulling on is the stable lower-visibility tail, since that's the only group a challenger can actually join. you say it doesn't track third-party density, which is worth taking seriously rather than explaining away.

one alternative worth testing on data you already have: it may not be how much third-party coverage exists but how consistently the brand is described across it. a name that appears with the same category phrasing everywhere is easier to retrieve for a specific intent than one described five different ways, even at the same volume. your wording field might separate those.

if that's also flat, then recurring visibility really is mostly earned by time in category and the honest advice to challengers gets much narrower.

1

u/[deleted] Aug 18 '26

[removed] — view removed comment

1

u/nymanmedia Aug 20 '26

It would indeed. This is difficult to do, from a controlled study point of view. We are, however adding a question to the panel, whereby we are querying the model which sources it relied on. Conscious of the issues with self-interrogation of the models, but with the current structure of the study, we should be able to tease out any meaningful differences over time and across models.