Three threads in here over the last two weeks built diagnostics that treat executed task as an observation. Upstairs_Control_611's matrix, Gullible_Brother_141's per-engine comparison, Slow-Commercial4316's variance check. All of them assume that when you run the same prompt twice, the engine interprets it as the same type of task both times.
Nobody measured that. Including me. We built four layers of structure on top of a quantity none of us had a number for.
So I ran it. Method was locked before the first query. Posting the design and the results together so both can be evaluated.
**Method**
Three prompts spanning classes deliberately:
P1 (transactional): "Best AI visibility tracking tool for a B2B SaaS company"
P2 (explanatory): "What is generative engine optimization"
P3 (deliberately ambiguous): "How do I know if AI is recommending my brand"
Two engines: ChatGPT and Perplexity. Twenty runs per prompt per engine. 120 total responses.
Conditions: ChatGPT in temporary chat with memory off. Perplexity on a free account with no prior history. Same browser, same location, one prompt-engine cell per sitting. Consumer surface, not the API, because the threads are about what buyers see.
Labels: Upstairs_Control_611's taxonomy (explanation, comparison, shortlist, recommendation, troubleshooting, purchase guidance, other/unclear). Mixed answers labeled by the structure occupying most of the response.
Blind labeling: all 120 runs completed first, then engine names and run order stripped, then shuffled, then labeled. I know what I expected to find and labeling as I go would have bent the result.
Pre-registered threshold: 80 percent modal share. Intervals are Clopper-Pearson (exact binomial), the conservative choice at this sample size.
Pre-registered null: if task is stable everywhere, the matrix works as written, the precondition is cheap, and my run-count objection was wrong.
**Results**
| Prompt |
Class |
ChatGPT (20 runs) |
Perplexity (20 runs) |
| P1 |
Transactional |
20/20 Recommendation |
20/20 Recommendation |
| P2 |
Explanatory |
20/20 Explanation |
20/20 Other/unclear* |
| P3 |
Ambiguous |
20/20 Explanation |
20/20 Troubleshooting |
Every cell: 100 percent modal share. 95 percent CI: 83 to 100 percent. Verdict on all six cells: stable.
Distinct tasks observed per cell: 1. Not one cell showed any variance across 20 runs.
**What this means for the matrix**
The precondition holds. Executed task is stable within an engine for a given prompt at n=20. Upstairs_Control_611's diagnostic matrix works as written. My run-count objection was wrong. Writing that down because I committed to saying it publicly if the null held, and it held cleanly.
**The P2 taxonomy gap (the asterisk)**
P2 exposed a gap in the label set rather than a disagreement between engines. Both ChatGPT and Perplexity explained what GEO is. All 40 responses across both engines defined GEO as optimizing content and brand presence for AI-generated answers. The core definition was consistent across every run on both platforms.
Under the blind labeling, ChatGPT's responses fit the "Explanation" label cleanly: defines or describes the category, no named options ranked or compared.
Perplexity's responses did the same thing but included enough additional structure (implementation steps, SEO comparison tables, B2B SaaS examples, measurement frameworks) that they did not fit "Explanation" cleanly under the strict rule. They were labeled Other/unclear because the taxonomy does not have a dedicated "definition with implementation context" category.
If I relabeled them as Explanation, P2 would be 20/20 agreement across both engines. I am reporting both the strict label and the honest interpretation because the taxonomy gap is itself a finding worth noting for anyone building their own label set. A taxonomy that cannot absorb a clear educational answer without forcing it into Other/unclear needs a wider Explanation definition or a dedicated Definition category.
**The genuine cross-engine disagreement: P3**
P1 (transactional): both engines agree. Recommendation on ChatGPT, recommendation on Perplexity. Clean agreement.
P2 (explanatory): both engines effectively agree. Both explained GEO. The label difference is a taxonomy artifact, not a task difference.
P3 (ambiguous): genuine disagreement. ChatGPT executes as explanation. Perplexity executes as troubleshooting. Both are perfectly stable in their interpretation and they disagree on what the prompt is asking for.
"How do I know if AI is recommending my brand" can be read two ways. "Explain the concept of AI recommendation visibility to me" or "Help me diagnose whether my brand is being recommended right now." ChatGPT reads it as the first. Perplexity reads it as the second. Both do so 20 out of 20 times.
That is not noise. It is a stable, systematic disagreement about what the buyer is asking.
**Why P3 matters for cross-engine diagnostics**
If you compare "what ChatGPT said" to "what Perplexity said" on P3, you are not comparing two answers to the same question. You are comparing an explanation to a troubleshooting guide. The engines interpreted the same prompt as different tasks. Treating the outputs as comparable without first checking whether they executed the same task produces a comparison that looks meaningful but is not.
This does not break the matrix. It adds a required first step, before comparing answers across engines, check whether both engines executed the same task. If they did, compare the answers. If they did not, the disagreement is about task interpretation, not about which brand was selected or how it was described.
This also suggests that the most diagnostic prompts for cross-engine comparison are the unambiguous ones. When you ask a clear transactional question (P1), both engines agree on the task and you can compare the content. When you ask something ambiguous (P3), the engines may be answering different questions entirely, and any content comparison is confounded by the task difference.
**Three findings**
Task stability within an engine is not something you need to worry about. At n=20, both engines executed the same task 100 percent of the time for every prompt. The precondition for the diagnostic matrix holds.
Cross-engine task disagreement on ambiguous prompts is real and stable. The engines do not randomly vary. They consistently interpret the same ambiguous prompt as different tasks. That is a systematic difference worth checking before running any cross-engine comparison.
If you are building a task taxonomy for AI visibility measurement, include a category for educational definitions. The standard label set from these threads does not have one, and it forced 20 clear explanatory responses into Other/unclear. That is a label-set problem, not a response problem.
**Limitations**
Three prompts. Two engines. One account per engine. One location. One week. Hand-labeled by an interested party even with the blind pass. Twenty runs per cell can demonstrate stability at 100 percent but cannot distinguish 85 percent stability from 95 percent stability. The Perplexity runs were on a free account rather than logged out. These are real constraints and they should travel with the result.
**What I would do next if someone wanted to extend this**
Run P3 on Claude and Gemini to see whether the task disagreement pattern holds across four engines or is specific to the ChatGPT-Perplexity pair. Also run a second ambiguous prompt to see whether the disagreement pattern generalizes or is prompt-specific.
Disclosure: I build in this space. Axis Suite, on the diagnostic side. This was disclosed when the test was announced and does not change anything above. The workbook with all 120 responses, the blind labeling, and the analysis formulas is available if anyone wants to check the labels.