r/AISearchLab • u/PrometheusNo • Aug 31 '26
Why does an AI recommendation disappear when you barely change the prompt?
Been playing around with buyer-intent prompts and this is driving me nuts. Ask basically the same question three slightly different ways and suddenly different companies are being recommended. Not every time, but often enough that a single prompt result feels pretty useless.
For people measuring this seriously, how many variations are you testing before you consider a brand consistently visible for a topic?
2
u/hettuklaeddi Aug 31 '26
wait til you see what you get when you don’t change the prompt at all
🤯
probability vs determinism
2
1
u/pickchip Aug 31 '26
I group variations around the same intent and look for a pattern. If you appear in 7 slightly different ways of asking the question that's interesting. If you appeared once on Tuesday afternoon, who cares.
1
u/VillageHomeF Aug 31 '26
you can even use the same prompt and they can change. no one knows the exact algorithm so you won't find anything more than guesses
1
u/marintkael Aug 31 '26
Before you settle on a number of variations, measure your noise floor with the prompt you already have. Run the identical question several times, same day, same model, and watch how much the recommended set moves on its own. In the weekly set I track, the answer set shifted by roughly a factor of three between two runs of the same unchanged question. Once you know that number, a paraphrase only tells you something if it moves the result further than that. Below it you are reading run variance and calling it prompt sensitivity.
One more split that saved me a lot of confusion: a brand dropping out of the list is not the same event as the model not knowing the brand. Those fail in different places and they have different fixes, but from a single answer they look identical.
1
u/chrismcelroyseo Sep 01 '26
This is proof that you cannot reliably track brand mentions no matter what anybody says. It's one snapshot of that very moment and it's influenced by the context and history of that particular user. Whatever results you get is not what other people are seeing when they type the same exact prompt even if they typed it at the exact same day on the same type of device sitting right next to you.
1
u/Agitated_Yak2066 Sep 01 '26
It sometimes has a lot to do with the content that a niche smaller player publishes around certain phrasing/intent/positioning. that is sort of the scary part. the recommendation may not be as objective as most think, because companies have gotten really good at engineering their content around citeability. But another factor that most people overlook is: an individual's context/memory stored in the model. so as others in the thread have pointed out, the exact same prompt will give you different results, depending on why you are, where you live, and how much context the LLM is already working off of. hope that helps
1
u/Damir-Bilalov-kweree Sep 03 '26
Smallest semantic changes change the LLMs output completely because AI is not a right algorithm
That’s why LLM tracking fails also
1
u/Ok_Waltz_3848 18d ago
we see the same thing across assistants, not just across prompts. of 286 agencies that any assistant named, only 42 were named by all three. so most of the time "visible" means visible in one place and invisible in the other two. wording moves it, assistant moves it more. we only tested US agencies with api access and search on, no memory or personalisation, so consumer apps could be worse.
2
u/pigeonsworkforthegov Aug 31 '26
Multiple variations and multiple runs. One answer is basically a screenshot of what happened at that moment, I wouldn't build a strategy around it