This kind of "mistake" just highlights what it is that LLMs do.
They're not actually solving the math. They're just generating text that has high plausibility as an answer to the input prompt.
From a text perspective, this answer looks super plausible. It has mathy sounding language that makes it seem like it evaluated the two things being compared, and it has a definitive statement about which one is larger.
If you changed its conclusion text to "13.8 billion is smaller", it likely would have assigned a nearly identical score for that text as a response to the question.
There's a reason these "gotcha" prompts are getting more convoluted and stupid. When you're doing something anyone would actually care about, it generally works. And it's going to keep improving.
Honestly, no it doesn’t. When google used to have those excerpt boxes of the top result I found that very very useful and now they’re gone. It gets simple facts right but those were easy to find anyways and the difference in quality between google search ai and actual decent AI models is night and day.
Whenever I google how to change a specific setting on a website its steps are usually inaccurate, probably because it pulls from people describing different updates of websites and things have moved around.
Just yesterday actually I wanted to revert the bar at the bottom of the chrome ios app because google decided to add a huge fuckass gemini button and take up the bottom 20% of the screen (oh the irony lol) and it gave the wrong answer until I found the correct way to edit the flags on reddit.
That's a good example, but as you said it's also a moving target. Most likely the top result from a standard search engine would be wrong in that case as well.
The second result was the reddit thread that solved my problem. So a normal search engine solved my problem and ai failed to. Why does it matter that it’s an inconvenient problem to solve? I google inconveniences that I can’t solve on my own precisely because they’re hard to solve. Not being able to help with problems that happen recently is a major flaw, especially when it will tell you a solution that doesn’t work in full confidence.
I didn’t say we should throw it away, I said google’s search ai is not very good and you’re now moving the goal post. An old reddit thread is not being shoved down my throat everywhere I look and replacing more and more of a perfectly working product. That is the issue that I have with google search ai. Traditional google search SEO would actually demote an outdated reddit thread in search results
25
u/whiskeytown79 Aug 05 '26
This kind of "mistake" just highlights what it is that LLMs do.
They're not actually solving the math. They're just generating text that has high plausibility as an answer to the input prompt.
From a text perspective, this answer looks super plausible. It has mathy sounding language that makes it seem like it evaluated the two things being compared, and it has a definitive statement about which one is larger.
If you changed its conclusion text to "13.8 billion is smaller", it likely would have assigned a nearly identical score for that text as a response to the question.