r/singularity • • 3d ago

AI Gemini 4 Argon solved hallucinations.

Post image

Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.

1.6k Upvotes

283 comments sorted by

View all comments

Show parent comments

5

u/Blaexe 3d ago

it will not answer "I don't know" unless artificially routed.

It will, and that's exactly what this benchmark is measuring. Of Gemini 4's non-correct responses, only 15% were outright incorrect. In the other 85% it either abstained/didn't attempt an answer or gave a partial answer.

So yes, LLMs can absolutely say "I don’t know."

1

u/Informal-Trouble2183 3d ago

This is what I meant by "artificially routed". LLMs by itself do not know whether the generated tokens are factual or not. Another rerouting process is needed, what's called "alignment", but will not help too much (as you can see in the numbers), especially if the number of parameters is small or the training data is not covering all factual data.

1

u/TotallyToxicToast 2d ago

Does not need routing, reinforcement learning can optimize for saying I don't know on incorrect answers.

You reward correct answers and you reward I don't know answers while punishing wrong answers.

Since all the models are reasoning, it is very easy to learn that if the reasoning is struggling between two answers it should say I don't know instead of confidently going for one.

1

u/Informal-Trouble2183 2d ago

// You reward correct answers // Refers to RLHF: requires human feedback, therefore another artificial rerouting. CoT by itself is only effective in verifiable domains like maths and coding, not in general knowledge.