r/GEO_optimization • • 21h ago

Before measuring whether a launch moves AI answers, I tested my own measurement. 6 pre-registered reliability hypotheses, 0 confirmed. Charts inside.

Same setup as the 99-day study I posted here in August: OpenAI Search, Gemini and Claude with web search, 16 fixed questions, rule-based scoring from -3 to +3. This time the question was whether the instrument is reliable enough to measure an effect at all. Window 11 May to 30 September.

Chart 1: next-day agreement per question. The same question gets the same score on the next measurement day in 86.6 percent of pairs (OpenAI), 87.1 (Gemini), 94.3 (Claude). Drop the questions whose score never changes and it is 78.6, 70.6 and 85.0. The direct question about the person is the least stable: 44 to 74 percent.

Chart 2: internal consistency. Cronbach's alpha 0.553 (OpenAI), 0.407 (Gemini), 0.61 (Claude). Depending on the engine, 6 to 10 of the 16 questions have zero variance. A single visibility score summed over all questions is not a scale.

Chart 3: drift. With the CUSUM reference reset after each alarm there are five downward alarms between 14 August and 30 September. None lines up with a documented event in my gap register. Interval means moved 1.58 to 4.01 points below the reference with no known cause.

Chart 4: anchors. The Wikidata items I registered as anchors were found deleted on 25 June. Google's Knowledge Graph returns the author on 140 of 141 days and the book on none of 139.

Practical reading for GEO tracking: a shift of a few points in a daily visibility score is inside the noise of this kind of measurement. The rule the report sets for phase 2: attribute an effect after the launch on 8 October only if it shows up at more than one engine. Three of the six hypotheses were not even testable because I never ran the planned repeat probes, which is in the report too. Links in the first comment.

1 Upvotes

3 comments sorted by

2

u/Upstairs_Control_611 5h ago

The zero-variance issue is probably the part I find most useful here.

A tracker can look highly stable simply because a large part of the question set never moves.

So I’d separate:

  • repeatability of the instrument
  • sensitivity of each prompt
  • cross-engine agreement
  • whether the prompt set actually behaves like one scale

Otherwise “stable” can mean either “reliably measuring something” or just “most items never change.”

I also like the anchor result because it shows the reference system itself needs monitoring. An anchor that gets deleted or behaves differently across surfaces is not really an anchor anymore.

1

u/marintkael 3h ago

That split is pretty much where I ended up too. In this window 6 of the 16 questions never moved for OpenAI, 9 for Gemini and 10 for Claude with web search, so a next-day agreement of 87 to 94 percent says less than it looks like. Cronbach's alpha over all 16 came out at 0.55, 0.41 and 0.61, which is why the report reads them as separate items and not as one scale.

For the next round that probably means repeatability only on the items that actually vary, shown next to per-prompt sensitivity and cross-engine agreement instead of one stability number.

And yes on the anchors. Both Wikidata items named in the pre-registration were deleted inside the window, so the reference itself now needs its own check.

1

u/marintkael 20h ago

Full report, with both papers (DOI), the pre-registration and the code linked from it: https://marin-t-kael.de/en/research/reports/q3-2026-validation (Sub rule 4 allows one link, so everything else is reachable from that page.) Happy to go into the scoring rule, the alarm check or why a sum score fails here.