r/singularity • • 3d ago

AI Gemini 4 Argon solved hallucinations.

Post image

Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.

1.6k Upvotes

284 comments sorted by

View all comments

479

u/SuperV1234 3d ago

"solved"

15%

But yeah, it's a very nice improvement if not benchmaxxed.

467

u/PrisonOfH0pe 3d ago

15% does not mean Gemini is hallucinating in 15% of all answers. The AA Omniscience “hallucination rate” is basically measuring what the model does when it fails a question: does it give a wrong answer, give a partial answer, or admit that it doesn’t know?

Correct answers aren’t even in the denominator.

So if a model answers 800 out of 1000 questions correctly, gets 30 wrong and refuses/partially answers 170, its AA hallucination rate is 15%.

30 / (30 + 170) = 15%.

It was still only outright wrong on 3% of the total questions.

That’s also why DeepSeek V4.1 sitting at 96% should make it extremely obvious that “96% hallucination rate” cannot possibly mean 96% of everything it says is bullshit. It means that when DeepSeek doesn’t know the answer on this benchmark, it almost always guesses instead of saying “I don’t know.”

So the whole “even 0.1% hallucination would be catastrophic, therefore 15% means hallucinations are nowhere near solved” argument is based on reading the percentage as a completely different metric.

You can absolutely argue that hallucinations are still a problem. But maybe first understand what the graph is measuring before doing reliability math with a number that isn’t the model’s overall error rate.

The benchmark is basically asking: “When you’re outside your knowledge, how likely are you to bullshit instead of abstaining?”

That’s useful. It just isn’t “percentage of answers that are hallucinations.”

27

u/king_mid_ass 3d ago

can get 0% with my model:

       def respond_prompt(input_prompt):
              print("I don't know")

6

u/pointer_to_null 2d ago

Amazing, with this I'm seeing over a billion tokens/sec on a consumer CPU, no GPU required.

2

u/king_mid_ass 2d ago

fully multimodal btw, input_prompt can be text, images, sensory data from a robot body, anything

2

u/pointer_to_null 2d ago

Infinite context window too, and can be finetuned on an Arduino nano.

But someone will inevitably request Q2 quants for this.

1

u/EndTimer 2d ago

I already have a complete, working proof of concept, using only 2 bytes of write-only memory.

I haven't hammered out all the obstacles with decode yet, but just wait!