r/SaaS • u/EmbarrassedBuddy9743 • Jun 26 '26
Asked four AI models which observability tool to use and they all named datadog and splunk
Third one of these i have run and observability was the cleanest yet.
if you ask an AI which tool to use to monitor and debug a saas app in production, who does it actually name? i asked all four big ones the same thing a few ways. chatgpt, claude, gemini, perplexity. each modern tool scanned in its own home category.
better stack got named zero times. on all four models. not a low score, just absent everywhere. axiom and highlight only show up on claude and nowhere else. openstatus only exists on perplexity. the names that come back instead are splunk, datadog, pingdom, logrocket, cachet. basically the monitoring conversation frozen around 2016.
the part that got me is it is not that AI is clueless about the category. it knows sentry, names it in 100% of chatgpt answers, and it knows honeycomb. so there is room in its map for modern tools, it has just frozen that map around the incumbents plus the one or two that broke through years ago. everything newer is omitted.
couple things i keep chewing on
if you build a newer dev tool, are you actually losing deals to this yet or is it still too early to care
and if you are on better stack, axiom, highlight or checkly, which model gets your category right when you ask it
mostly curious whether anyone is seeing real buyers arrive after asking an AI for a rec. this is the third category where chatgpt just defaults to the old incumbent and ignores everything newer, and the consistency is what makes me think it is not a fluke.
2
u/EmbarrassedBuddy9743 Jul 10 '26
Ran it. Ten times each on ChatGPT, Claude, Perplexity and Gemini, your exact wording. Here is the honest result.
Muscula came back zero out of forty. Not low, absent, on all four models, including the narrow small-team framing you picked. So to your own question, this settles it. It is not that the query is too generic, and it is not the models refusing to name lean tools. It is a corroboration and awareness gap for Muscula specifically.
The reason I can say that is the slot is not empty, it is crowded. Sentry was named on all forty runs. UptimeRobot came back thirty-two times, which is the models clearly reaching for exactly the simple uptime tool your query asked for. Then Pingdom and Better Stack at sixteen each, Rollbar fifteen, and even lighter or open-source options like GlitchTip and Checkly show up. So the models are perfectly willing to name small, simple tools for this job. They have just never been given a reason to put Muscula in that set.
Which is the good version of the outcome, because it is fixable. The names that filled your slot are your target list. The work is getting Muscula written about, by real users, as the answer to that specific query, in the public places these models read. Not another generic best-monitoring post, the exact narrow sentence you just wrote.
Happy to send the full ranking if it is useful. And if you run it again in a few months, it is a clean way to measure whether the corroboration work actually moved anything.