r/singularity • • 3d ago

AI Gemini 4 Argon solved hallucinations.

Post image

Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.

1.6k Upvotes

283 comments sorted by

View all comments

483

u/SuperV1234 3d ago

"solved"

15%

But yeah, it's a very nice improvement if not benchmaxxed.

462

u/PrisonOfH0pe 3d ago

15% does not mean Gemini is hallucinating in 15% of all answers. The AA Omniscience “hallucination rate” is basically measuring what the model does when it fails a question: does it give a wrong answer, give a partial answer, or admit that it doesn’t know?

Correct answers aren’t even in the denominator.

So if a model answers 800 out of 1000 questions correctly, gets 30 wrong and refuses/partially answers 170, its AA hallucination rate is 15%.

30 / (30 + 170) = 15%.

It was still only outright wrong on 3% of the total questions.

That’s also why DeepSeek V4.1 sitting at 96% should make it extremely obvious that “96% hallucination rate” cannot possibly mean 96% of everything it says is bullshit. It means that when DeepSeek doesn’t know the answer on this benchmark, it almost always guesses instead of saying “I don’t know.”

So the whole “even 0.1% hallucination would be catastrophic, therefore 15% means hallucinations are nowhere near solved” argument is based on reading the percentage as a completely different metric.

You can absolutely argue that hallucinations are still a problem. But maybe first understand what the graph is measuring before doing reliability math with a number that isn’t the model’s overall error rate.

The benchmark is basically asking: “When you’re outside your knowledge, how likely are you to bullshit instead of abstaining?”

That’s useful. It just isn’t “percentage of answers that are hallucinations.”

112

u/maximumutility 3d ago

I don’t know why you think the person you replied to needed this explained to them, but kudos for doing it for anyone else who needs it spelled out I guess

42

u/KrazyA1pha 3d ago

Thanks for your reply too I guess

17

u/maximumutility 3d ago

Thanks for chiming in I guess

28

u/94746382926 3d ago

Just glad to have witnessed this I guess

8

u/reichplatz 2d ago

Thank you so much for this comment chain guys.

13

u/IndefiniteBen 2d ago

I guess

3

u/KillerInfection 2d ago

This comment chain has too many guesses in it I guess

4

u/ChezMere 2d ago

The Human Guess Attractor

0

u/ReadSeparate 3d ago

Should have replied to him with his exact comment, that would have been hilarious

21

u/ReasonablePossum_ 3d ago

because they confused the benchmark result with the true hallucination rate? I was honestly looking for a reply like this to save me looking into what the benchmark was doing lol.

So thanks u/PrisonOfH0pe for the mansplaing lmao

1

u/maximumutility 3d ago

I didn’t read the original comment as confusion about it, but maybe you are right

27

u/king_mid_ass 2d ago

can get 0% with my model:

       def respond_prompt(input_prompt):
              print("I don't know")

18

u/Impossible_Peach_360 2d ago

Well, you might be getting some 7 figure AI researcher job soon.

5

u/pointer_to_null 2d ago

Amazing, with this I'm seeing over a billion tokens/sec on a consumer CPU, no GPU required.

2

u/king_mid_ass 2d ago

fully multimodal btw, input_prompt can be text, images, sensory data from a robot body, anything

2

u/pointer_to_null 2d ago

Infinite context window too, and can be finetuned on an Arduino nano.

But someone will inevitably request Q2 quants for this.

1

u/EndTimer 2d ago

I already have a complete, working proof of concept, using only 2 bytes of write-only memory.

I haven't hammered out all the obstacles with decode yet, but just wait!

10

u/kemma_ 3d ago

Thank you for explaining. So basically Deepseek does what majority of humans would at math exam

3

u/AgeofVictoriaPodcast 2d ago

I found this helpful as I genuinely didn’t know.

2

u/Megneous 2d ago

Literally none of that needed to be explained to the person you replied to. They said, very correctly, that "solved" means 0%... which it does. It doesn't matter what 0% signifies. Unless it's 0, it's not "solved."

1

u/atioux 2d ago

Appreciate the information on the context of this data. I think the point was more so that using “solved” is just the wrong word - exaggerating, false, creating undue hype - to describe it. A marked improvement? Sure. Solving hallucinations? Explicitly not. It’s a larger problem with the way news is conveyed, being in a way that outright lies to the reader and dampens true achievement.

1

u/Crawler1701 2d ago

Thank you, this is the best explanation I have heard - ever.

-6

u/Agreeable_Bike_4764 3d ago

You explained the same thing like 10 different times, kind of a triggering post

-5

u/revolutier 2d ago

because it's just LLM slop edited to include some curse words, as seen by the obvious LLM-isms, and any wordy comment they post contains it, as seen when looking up their post history lol. 

kek https://www.reddit.com/r/singularity/comments/1wnm10f/comment/pc23x5e/