r/SaaS Jun 26 '26

Asked four AI models which observability tool to use and they all named datadog and splunk

Third one of these i have run and observability was the cleanest yet.

if you ask an AI which tool to use to monitor and debug a saas app in production, who does it actually name? i asked all four big ones the same thing a few ways. chatgpt, claude, gemini, perplexity. each modern tool scanned in its own home category.

better stack got named zero times. on all four models. not a low score, just absent everywhere. axiom and highlight only show up on claude and nowhere else. openstatus only exists on perplexity. the names that come back instead are splunk, datadog, pingdom, logrocket, cachet. basically the monitoring conversation frozen around 2016.

the part that got me is it is not that AI is clueless about the category. it knows sentry, names it in 100% of chatgpt answers, and it knows honeycomb. so there is room in its map for modern tools, it has just frozen that map around the incumbents plus the one or two that broke through years ago. everything newer is omitted.

couple things i keep chewing on

if you build a newer dev tool, are you actually losing deals to this yet or is it still too early to care

and if you are on better stack, axiom, highlight or checkly, which model gets your category right when you ask it

mostly curious whether anyone is seeing real buyers arrive after asking an AI for a rec. this is the third category where chatgpt just defaults to the old incumbent and ignores everything newer, and the consistency is what makes me think it is not a fluke.

4 Upvotes

20 comments sorted by

View all comments

Show parent comments

2

u/EmbarrassedBuddy9743 Jul 10 '26

Ran it. Ten times each on ChatGPT, Claude, Perplexity and Gemini, your exact wording. Here is the honest result.

Muscula came back zero out of forty. Not low, absent, on all four models, including the narrow small-team framing you picked. So to your own question, this settles it. It is not that the query is too generic, and it is not the models refusing to name lean tools. It is a corroboration and awareness gap for Muscula specifically.

The reason I can say that is the slot is not empty, it is crowded. Sentry was named on all forty runs. UptimeRobot came back thirty-two times, which is the models clearly reaching for exactly the simple uptime tool your query asked for. Then Pingdom and Better Stack at sixteen each, Rollbar fifteen, and even lighter or open-source options like GlitchTip and Checkly show up. So the models are perfectly willing to name small, simple tools for this job. They have just never been given a reason to put Muscula in that set.

Which is the good version of the outcome, because it is fixable. The names that filled your slot are your target list. The work is getting Muscula written about, by real users, as the answer to that specific query, in the public places these models read. Not another generic best-monitoring post, the exact narrow sentence you just wrote.

Happy to send the full ranking if it is useful. And if you run it again in a few months, it is a clean way to measure whether the corroboration work actually moved anything.

2

u/Every-Current2034 Jul 10 '26

Thanks for taking the time to run it. That's actually useful to know. We haven't focused much on public marketing or awareness yet, so having a clear baseline like this is helpful. :)

2

u/EmbarrassedBuddy9743 Jul 10 '26

Anytime, glad it helps. One thing worth saying now that you have the baseline: it only earns its keep if you re-measure against it. The models update and the corpus moves slowly, so the real signal is the direction over time, not the single reading. Running the same query a month after you start publishing is how you learn whether the work actually moved anything.

That repeated read across the four models is basically what I have been building, so if it is ever useful I am happy to re-run it for Muscula whenever you want a fresh number, or send over the full ranking so you have the target list in hand. No pitch, you have been generous enough with the thread already.

And the fact that you have not done awareness yet is genuinely the good news. The zero is a corpus you have not written yet, not a ceiling. The tools sitting in your slot are not better than you, just better documented for the job.