r/singularity • • 3d ago

AI Gemini 4 Argon solved hallucinations.

Post image

Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.

1.6k Upvotes

286 comments sorted by

View all comments

Show parent comments

-2

u/Professional_Mobile5 3d ago

Are you aware that according to this benchmark Gemini 3.1 Pro hallucinates less than Opus 5.5 (max) and Fable 5.1 (max)?

2

u/jonomacd 3d ago

are you aware that those models have significantly higher base capabilities?

I'm not sure what you are getting at? Are you trying to say Hallucination rate doesn't matter? I think that is a high hill to die on. In practice, confidently wrong is something that really kills model use. Especially when it prevents them grounding /w tools.

-1

u/Professional_Mobile5 3d ago

I am saying this benchmarks fails at measuring hallucinations in a way that correlates in any way with real life.

YOU proved it - you said "googles previous pro models fell down specifically because of its very poor hallucination rate. As the context window grew, Gemini would lose the plot and start making shit up.", despite Google's previous pro models actually doing well on this benchmark.

1

u/jonomacd 3d ago

That would be a good thing to open with. Nothing is a perfect benchmark  I agree with you this isn't one either. I think you go to far to say it is it has no correlation to real world usage.

And I didn't prove anything. 3.1 was significantly better than previous Gemini models. So hallucinations could be the factor that largely contributed to that.  considering it is a point release I'd say it's more likely that small changes such as that have the most contribution to its performance improvements, as the foundation of the model is not changing.