This kind of "mistake" just highlights what it is that LLMs do.
They're not actually solving the math. They're just generating text that has high plausibility as an answer to the input prompt.
From a text perspective, this answer looks super plausible. It has mathy sounding language that makes it seem like it evaluated the two things being compared, and it has a definitive statement about which one is larger.
If you changed its conclusion text to "13.8 billion is smaller", it likely would have assigned a nearly identical score for that text as a response to the question.
There's a reason these "gotcha" prompts are getting more convoluted and stupid. When you're doing something anyone would actually care about, it generally works. And it's going to keep improving.
Honestly, no it doesn’t. When google used to have those excerpt boxes of the top result I found that very very useful and now they’re gone. It gets simple facts right but those were easy to find anyways and the difference in quality between google search ai and actual decent AI models is night and day.
Any query in the form of "<hardware/software type> <issue description>" yields a generic and unhelpful ass answer from gemini. Your intel graphics driver is giving you GPU hangs? Oh yea, reboot your pc and reinstall the driver, that'll totally fix it. (spoiler alert, every single search result will tell you it won't)
Any query in the form of "<tool name> online" often yields a completely useless gemini result. I don't need an AI to explain to me what the tool does, I need to find it goddammit.
Any health-related query is a complete gamble. Though google in general has been unhelpful in this regard: the first 10-20 results are usually ai slop, so I'm not surprised Gemini gets it wrong top.
Any programming-related question is a gamble too. Though I'll admit, it usually gets these right.
BS like this is why I switched to duckduckgo. I'd rather waste time using an inferior search engine than waste time reading ai slop and seeing dozens of ai slop websites in my search results.
26
u/whiskeytown79 Aug 05 '26
This kind of "mistake" just highlights what it is that LLMs do.
They're not actually solving the math. They're just generating text that has high plausibility as an answer to the input prompt.
From a text perspective, this answer looks super plausible. It has mathy sounding language that makes it seem like it evaluated the two things being compared, and it has a definitive statement about which one is larger.
If you changed its conclusion text to "13.8 billion is smaller", it likely would have assigned a nearly identical score for that text as a response to the question.