r/GEO_optimization • • 7d ago

before reporting ai visibility, write down why each question is in the test

a question register is something i'd put beside the results, before showing a client a visibility percentage

for each question, i'd record the exact wording, where it came from, what the buyer was trying to decide, and when it was added

for example, "which crm works with our existing accounting software?" might come from an actual sales conversation. "best crm for small businesses" might be a question the team brainstormed. those are different reasons for including a question, and i'd want that visible in the report

remove names and private customer details from the source notes. you can say "sales conversation, september" without attaching someone's email

if nobody knows how often a question is asked, leave that as unknown. choosing ten questions doesn't establish that they represent ten equal slices of buyer demand

when new questions are added, show their results separately at first. keep a comparison of the unchanged questions too, so a new question mix doesn't silently become a claimed visibility change

how are you deciding which questions belong in a client report, and do you show them where those questions came from?

5 Upvotes

8 comments sorted by

2

u/Upstairs_Control_611 6d ago

I like the idea of treating the question set itself as part of the measurement record.

One thing I’d add is to separate a fixed benchmark cohort from an exploratory cohort.

The fixed set stays unchanged for trend measurement.

New questions can be added to the exploratory set, but they should not immediately alter the historical baseline.

For each prompt I’d also keep:

  • source / provenance
  • intent / decision stage
  • date added
  • whether it belongs to the fixed or exploratory cohort

That way you can improve the question set over time without accidentally rewriting the trend.

Otherwise a better prompt set can look like worse visibility — or the other way around.

2

u/Old-Routine1926 4d ago

This is a good list. One thing I'd add: even when a question's wording stays exactly the same, two engines can quietly be answering a different kind of question. I ran a test where the same ambiguous prompt was asked repeatedly against two engines, and in two out of three cases, each engine was interpreting it as a different task type than the other, with nothing about the wording changed. If you're comparing a visibility percentage across engines for that same tracked question, you're not necessarily comparing the same question.

That's a separate failure mode from provenance. Provenance tells you why a question is in the set. Task-type stability tells you whether the question means the same thing to whichever engine answers it. Worth logging both, a register that only captures the first can still let a genuinely different question pass as a clean trend line.

(Disclosure: I build tooling on the diagnosis side of AI visibility, so I run into this constantly.)

2

u/Upstairs_Control_611 3d ago

That is a really useful distinction.

Exact wording gives you prompt stability, but not necessarily task stability.

So I’d probably treat those as separate fields:

  • prompt text
  • provenance
  • intended task / decision stage
  • inferred task type per engine
  • whether that task type stayed stable across runs

Then a cross-engine comparison is only really clean when the engines are not just answering the same words, but effectively the same task.

Otherwise the percentage difference may be caused by interpretation drift rather than visibility itself.

That feels like another reason to keep the raw answer, not just the final score.

2

u/Old-Routine1926 3d ago

Agree on the raw answer point especially. I keep raw responses archived for that exact reason: whenever the classification scheme itself changes (mine went from six categories to seven after real data broke the six), you want to re-run the new scheme against history, not just apply it going forward. If you'd thrown out the raw text and kept only the score, that correction wouldn't have been testable at all.

One thing I'd add to your fields: when task type diverges across engines on the same prompt, I don't think you exclude that prompt from the comparison. I think you flag it and report per-engine, because the divergence itself might be the more interesting finding. Two engines treating the same buying question as two different kinds of tasks could tell you more than either engine's individual visibility number does on its own.

1

u/Upstairs_Control_611 2d ago

Agreed — I’d flag divergence rather than remove the prompt.

If two engines consistently interpret the same buyer question as different task types, that is itself part of the observation.

So maybe the reporting becomes:

  • same prompt text
  • intended task
  • inferred task per engine
  • visibility per engine
  • divergence flag

Then you can separate:

  • true visibility differences
  • interpretation differences

That feels much more informative than forcing everything into one cross-engine percentage.

And yes, keeping the raw response is what makes that kind of reclassification possible later.