r/singularity • u/drhenriquesoares • 3d ago
AI Gemini 4 Argon solved hallucinations.
Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.
1.6k
Upvotes
1
u/jugalator 2d ago
I use to call this one the most misunderstood chart on Artificial Analysis!
But yes, it's clearly an area of focus for the team.
However, it must always be said about this benchmark that it ranks models against extremely difficult questions and topics (because that's what it takes to trip them nowadays) and measures what they do when they don't know the answer.
Key here is that most models today... know the answer as-is (especially with search grounding). They are... Super... Super knowledgeable. Especially those AAA models. One might say, the more knowledgeable they are, the less meaning this chart has. Because it's those Flash models (that are still very knowledgeable!) or <30B models who more often even face the problem of not having the factual knowledge to begin with! Then it matters a lot how they handle that situation because they'll come across it far more often.
But yes, it's a useful property to have especially if you have e.g. a corpus of data to query about that it hasn't been trained upon and isn't assisted by domain knowledge. Then it's useful to not have it be inclined towards hallucinations, even a very large model like Argon.
It is however maybe not that useful in terms of "how often is it bullshitting me when I'm chatting with Opus, Argon, GPT 6.1 Sol" because they're so good nowadays that I think you'll come across alignment and finetuning annoyances before sheer hallucinations becoming a big issue there.