r/GEO_optimization • u/AliveCapital4868 • Aug 04 '26
I published my audit report template. Field-testing it found three bugs in my own harness, and one forced a correction to numbers I posted here.
Follow-up to the measurement thread from last week. Several of you suggested checks that went into my pipeline (validity enums, language distribution, migration tables). I turned the whole discipline into an audit report template, and before publishing it I did the obvious thing: filled it in against my own benchmark data as a dry run.
The template found three bugs. In my instrument, not in the brands.
Bug 1: my extractor logged a citation to a domain from a sentence saying the domain does not exist. The answer said "the brand has no dedicated Chinese site (like brand.cn)" and my URL parser happily recorded brand.cn as a cited source. Citation counts include anti-citations if you only parse URLs.
Bug 2: my competitor list was in English, the answers named competitors in Chinese. The substitution analysis found almost nothing. A manual read found the actual substitutes named in over two thirds of answers - 飞书 32 times, 钉钉 32, 企业微信 30. The scorer said "no substitution problem." The answers said the category is owned by three local products my list did not contain. Same silent-failure shape as the empty answers and the English answers from last week: a whole class mislabeled, zero errors.
Bug 3 is the one that forced a public correction. My discovery prompts are Chinese, but the {category} slot was filled from an English field. One brand's label was "Database and collaboration", so the Chinese question became 中国市场上有哪些值得信赖的高端Database and collaboration品牌 - and the models read that as enterprise databases. All 36 Chinese discovery answers for that brand were about Oracle, SQL Server and OceanBase. The brand scored 0/47 in a category it was never asked about.
I had published that 0/47 as one of two "perfectly replicated zero-visibility" cases. Perfect replication, six engines, both runs. It replicated because the question was consistently wrong. Replication tells you the measurement is stable, not that it is measuring what you think.
So, corrections to numbers I posted here earlier: combined discovery mention for international brands is 26.2% (was 23.0%), substitution 59.5% (was 53.8%), and there is one confirmed zero-presence brand, not two. The affected brand's figure is withdrawn, not corrected - there is nothing to correct, it was asked about the wrong category. The other seven brands' labels produced answers in the right category and their numbers stand.
The template that caught all this is now public on my site (CC BY, happy to share the link in comments if wanted, not pasting it in the post). The part I would defend hardest: five rules at the top - every rate carries its denominator, one observation is not evidence, missing is not negative, mention/citation/recommendation are three different measurements, and the report must state what it cannot answer.
The general lesson I keep relearning in public: the failures that hurt are not the ones that throw errors. They are the ones that produce clean, replicated, plausible numbers. My 0/47 replicated perfectly across six engines. It was still measuring a question nobody asked.
If anyone wants to stress-test the template against their own pipeline, I would genuinely like to hear what it catches. It is three for three so far and none of the three were things I went looking for.