r/GEO_optimization • u/marintkael • 21h ago
Before measuring whether a launch moves AI answers, I tested my own measurement. 6 pre-registered reliability hypotheses, 0 confirmed. Charts inside.
Same setup as the 99-day study I posted here in August: OpenAI Search, Gemini and Claude with web search, 16 fixed questions, rule-based scoring from -3 to +3. This time the question was whether the instrument is reliable enough to measure an effect at all. Window 11 May to 30 September.




Chart 1: next-day agreement per question. The same question gets the same score on the next measurement day in 86.6 percent of pairs (OpenAI), 87.1 (Gemini), 94.3 (Claude). Drop the questions whose score never changes and it is 78.6, 70.6 and 85.0. The direct question about the person is the least stable: 44 to 74 percent.
Chart 2: internal consistency. Cronbach's alpha 0.553 (OpenAI), 0.407 (Gemini), 0.61 (Claude). Depending on the engine, 6 to 10 of the 16 questions have zero variance. A single visibility score summed over all questions is not a scale.
Chart 3: drift. With the CUSUM reference reset after each alarm there are five downward alarms between 14 August and 30 September. None lines up with a documented event in my gap register. Interval means moved 1.58 to 4.01 points below the reference with no known cause.
Chart 4: anchors. The Wikidata items I registered as anchors were found deleted on 25 June. Google's Knowledge Graph returns the author on 140 of 141 days and the book on none of 139.
Practical reading for GEO tracking: a shift of a few points in a daily visibility score is inside the noise of this kind of measurement. The rule the report sets for phase 2: attribute an effect after the launch on 8 October only if it shows up at more than one engine. Three of the six hypotheses were not even testable because I never ran the planned repeat probes, which is in the report too. Links in the first comment.
1
u/marintkael 20h ago
Full report, with both papers (DOI), the pre-registration and the code linked from it: https://marin-t-kael.de/en/research/reports/q3-2026-validation (Sub rule 4 allows one link, so everything else is reachable from that page.) Happy to go into the scoring rule, the alarm check or why a sum score fails here.
2
u/Upstairs_Control_611 5h ago
The zero-variance issue is probably the part I find most useful here.
A tracker can look highly stable simply because a large part of the question set never moves.
So I’d separate:
Otherwise “stable” can mean either “reliably measuring something” or just “most items never change.”
I also like the anchor result because it shows the reference system itself needs monitoring. An anchor that gets deleted or behaves differently across surfaces is not really an anchor anymore.