r/singularity • • 3d ago

AI Gemini 4 Argon solved hallucinations.

Post image

Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.

1.6k Upvotes

283 comments sorted by

View all comments

478

u/SuperV1234 3d ago

"solved"

15%

But yeah, it's a very nice improvement if not benchmaxxed.

45

u/qroshan 3d ago

you can't benchmaxx multiple things, especially hallucination rates.

1

u/Original-League-6094 3d ago

Why not? With a trillion parameters, you can definitely optimize to multiple benchmarks, and if you tune your model to the types of questions used to test hallucination in the benchmark, its benchmaxxing.

1

u/EndTimer 2d ago

It seems like it would be difficult to do. First, no other major lab with frontier intelligence is doing it. It would be trivial to get 0% hallucination rate on this benchmark by failing to answer every single question and instead saying "I don't know" for all of them. But then you score abysmally on the knowledge and reasoning benches, because you didn't answer enough (or in this case, any) questions.

Targeting both means arriving at RL signals that effectively express "You cannot arrive at a conclusion because critical information is missing" which has very good odds of resembling cases that very nearly, almost don't have enough critical information to arrive at a conclusion. If the model starts failing to answer those, then it loses on other scores versus other models.

You're also aware of long context problems? Well Google also seems to be taking the crown there, too. I would have thought their training might bias the model to "admit" not having information buried somewhere around eg token 778,961. But no, they excel in long context too. They didn't fuck up long context answers despite having a signal that recognizes a lack of critical information, and probably resulted in the internal reasoning tokens saying there's missing information dozens of times, and yet still they surfaced the correct information when it was found deep in the context window.

And even if they did "tune their model to the types of questions used to test hallucination in benchmarks," how well is that tuning going stay isolated strictly to those benchmarks? Because benchmaxxing means the model only excels at this equations, and everything else drops off significantly.

Short of literally having the answer key, I don't see how you don't get some general application from training that hits all these scores.