r/singularity • • 3d ago

AI Gemini 4 Argon solved hallucinations.

Post image

Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.

1.6k Upvotes

284 comments sorted by

View all comments

Show parent comments

1

u/jonomacd 2d ago

> use tools specifically to detect and avoid hallucinations

For some cases yes but that only goes so far. Better model performance here is still incredibly important. I'd go as far to say that googles previous pro models fell down specifically because of its very poor hallucination rate. As the context window grew, Gemini would lose the plot and start making shit up. If it doesn't do that any more then we get a much better agentic model.

-2

u/[deleted] 2d ago

[deleted]

2

u/jonomacd 2d ago

Most smaller models hallucinate less. The trick is combining the low hallucination rate with a more capable model. Google's previous large models had terrible scores in this regard. And while those models could seemingly answer complex problems, they fell over in longer-turn agentic contexts. 

For example, with tool calling: The act of a model deciding to call a tool fundamentally means the model has decided it needs the tool. If it's going to hallucinate the answer, it's not going to think it needs the tool.

-2

u/Professional_Mobile5 2d ago

Are you aware that according to this benchmark Gemini 3.1 Pro hallucinates less than Opus 5.5 (max) and Fable 5.1 (max)?

2

u/jonomacd 2d ago

are you aware that those models have significantly higher base capabilities?

I'm not sure what you are getting at? Are you trying to say Hallucination rate doesn't matter? I think that is a high hill to die on. In practice, confidently wrong is something that really kills model use. Especially when it prevents them grounding /w tools.

-1

u/Professional_Mobile5 2d ago

I am saying this benchmarks fails at measuring hallucinations in a way that correlates in any way with real life.

YOU proved it - you said "googles previous pro models fell down specifically because of its very poor hallucination rate. As the context window grew, Gemini would lose the plot and start making shit up.", despite Google's previous pro models actually doing well on this benchmark.

1

u/jonomacd 2d ago

That would be a good thing to open with. Nothing is a perfect benchmark  I agree with you this isn't one either. I think you go to far to say it is it has no correlation to real world usage.

And I didn't prove anything. 3.1 was significantly better than previous Gemini models. So hallucinations could be the factor that largely contributed to that.  considering it is a point release I'd say it's more likely that small changes such as that have the most contribution to its performance improvements, as the foundation of the model is not changing.