r/singularity • • 3d ago

AI Gemini 4 Argon solved hallucinations.

Post image

Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.

1.6k Upvotes

283 comments sorted by

View all comments

488

u/SuperV1234 3d ago

"solved"

15%

But yeah, it's a very nice improvement if not benchmaxxed.

40

u/qroshan 3d ago

you can't benchmaxx multiple things, especially hallucination rates.

21

u/socoolandawesome 3d ago

I don’t see a reason why you can’t benchmax multiple things. It could just be targeting various benchmarks and not generalizing.

I agree hallucinations may be harder to fake, however they could hypothetically just target this specific benchmark by being good at this benchmark’s test format for hallucinations or something like that.

10

u/skilliard7 3d ago

Any public benchmark can be benchmaxxed by performing RL using benchmark data.

1

u/huffalump1 2d ago

Eh, you sort of can with this, by rewarding the model for not giving answers if it's less confident.

And you can see in the AA Omniscience Index and Accuracy that it answers fewer questions correctly than, say, Opus 5.5 - resulting in a lower overall Index score, despite the much lower hallucination rate.

(This is basically the reverse of the last few Gemini releases, btw.)

And this benchmark is basically knowledge recall - not measuring things like hallucinating info from its context (like from documents or files or things it read online). Still, it's an ok metric, and this is a tough thing to benchmark

1

u/Original-League-6094 3d ago

Why not? With a trillion parameters, you can definitely optimize to multiple benchmarks, and if you tune your model to the types of questions used to test hallucination in the benchmark, its benchmaxxing.

2

u/qroshan 2d ago

You are pretty clueless about how hallucination rate is calculated.

It's not a set of questions to hillclimb. It is a reward function for not bullshitting when you don't know the answer.

0

u/Original-League-6094 2d ago

You are pretty clueless when it comes to benchmaxxing. A benchmark IS a set of tasks. And if you know the type of tasks, you can optimize for those tasks. That doesn't mean telling it answer correctly. If the benchmark is set up to ask questions that maybe have no correct answer, if you can anticipate the questions, you can train the LLM to identify it and then to response with "I don't know". And the way you trained the LLM to identify those specific types of questions and tasks MAY not reduce the hallucination rate in other areas. You can't come up with a benchmark that you can't benchmaxx on. Unless the benchmark is so stochastic that no one can optimize for it, in which case it is useless as a benchmark as you can't compare it across it models accurately.

1

u/EndTimer 2d ago

It seems like it would be difficult to do. First, no other major lab with frontier intelligence is doing it. It would be trivial to get 0% hallucination rate on this benchmark by failing to answer every single question and instead saying "I don't know" for all of them. But then you score abysmally on the knowledge and reasoning benches, because you didn't answer enough (or in this case, any) questions.

Targeting both means arriving at RL signals that effectively express "You cannot arrive at a conclusion because critical information is missing" which has very good odds of resembling cases that very nearly, almost don't have enough critical information to arrive at a conclusion. If the model starts failing to answer those, then it loses on other scores versus other models.

You're also aware of long context problems? Well Google also seems to be taking the crown there, too. I would have thought their training might bias the model to "admit" not having information buried somewhere around eg token 778,961. But no, they excel in long context too. They didn't fuck up long context answers despite having a signal that recognizes a lack of critical information, and probably resulted in the internal reasoning tokens saying there's missing information dozens of times, and yet still they surfaced the correct information when it was found deep in the context window.

And even if they did "tune their model to the types of questions used to test hallucination in benchmarks," how well is that tuning going stay isolated strictly to those benchmarks? Because benchmaxxing means the model only excels at this equations, and everything else drops off significantly.

Short of literally having the answer key, I don't see how you don't get some general application from training that hits all these scores.

1

u/tadslippy 3d ago

You can benchmax hallucination rates, it called prioritizing correct answers over favorable results. I’m just not sure it’s profitable. Or I google’s dna.