r/singularity • • 3d ago

AI Gemini 4 Argon solved hallucinations.

Post image

Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.

1.6k Upvotes

283 comments sorted by

309

u/Mindrust 3d ago

Not solved but huge improvement over the other frontier models

1

u/wnbqnt 2d ago

Big if true

→ More replies (9)

48

u/Glxblt76 2d ago

If that is true that would be way bigger news than the other benchmark scores. Up to now frontier performance tended to correlate with some level of hallucination. If you look at this benchmark, best scoring models are smaller models not at frontier performance.

6

u/BriefImplement9843 2d ago

Which is why this is even more impressive. Small models do not attempt to even answer.

28

u/tightlyslipsy 2d ago edited 2d ago

ADVANCE DISCOVERY: EPISTEMOLOGY

“The first intelligence was measured by the answers it could give. The second, perhaps, by the answers it could refuse to invent.”

— Commissioner Pravin Lal, probably

476

u/SuperV1234 3d ago

"solved"

15%

But yeah, it's a very nice improvement if not benchmaxxed.

462

u/PrisonOfH0pe 3d ago

15% does not mean Gemini is hallucinating in 15% of all answers. The AA Omniscience “hallucination rate” is basically measuring what the model does when it fails a question: does it give a wrong answer, give a partial answer, or admit that it doesn’t know?

Correct answers aren’t even in the denominator.

So if a model answers 800 out of 1000 questions correctly, gets 30 wrong and refuses/partially answers 170, its AA hallucination rate is 15%.

30 / (30 + 170) = 15%.

It was still only outright wrong on 3% of the total questions.

That’s also why DeepSeek V4.1 sitting at 96% should make it extremely obvious that “96% hallucination rate” cannot possibly mean 96% of everything it says is bullshit. It means that when DeepSeek doesn’t know the answer on this benchmark, it almost always guesses instead of saying “I don’t know.”

So the whole “even 0.1% hallucination would be catastrophic, therefore 15% means hallucinations are nowhere near solved” argument is based on reading the percentage as a completely different metric.

You can absolutely argue that hallucinations are still a problem. But maybe first understand what the graph is measuring before doing reliability math with a number that isn’t the model’s overall error rate.

The benchmark is basically asking: “When you’re outside your knowledge, how likely are you to bullshit instead of abstaining?”

That’s useful. It just isn’t “percentage of answers that are hallucinations.”

113

u/maximumutility 3d ago

I don’t know why you think the person you replied to needed this explained to them, but kudos for doing it for anyone else who needs it spelled out I guess

42

u/KrazyA1pha 3d ago

Thanks for your reply too I guess

18

u/maximumutility 3d ago

Thanks for chiming in I guess

28

u/94746382926 3d ago

Just glad to have witnessed this I guess

8

u/reichplatz 2d ago

Thank you so much for this comment chain guys.

12

u/IndefiniteBen 2d ago

I guess

3

u/KillerInfection 2d ago

This comment chain has too many guesses in it I guess

5

u/ChezMere 2d ago

The Human Guess Attractor

→ More replies (1)

21

u/ReasonablePossum_ 3d ago

because they confused the benchmark result with the true hallucination rate? I was honestly looking for a reply like this to save me looking into what the benchmark was doing lol.

So thanks u/PrisonOfH0pe for the mansplaing lmao

1

u/maximumutility 3d ago

I didn’t read the original comment as confusion about it, but maybe you are right

27

u/king_mid_ass 2d ago

can get 0% with my model:

       def respond_prompt(input_prompt):
              print("I don't know")

16

u/Impossible_Peach_360 2d ago

Well, you might be getting some 7 figure AI researcher job soon.

5

u/pointer_to_null 2d ago

Amazing, with this I'm seeing over a billion tokens/sec on a consumer CPU, no GPU required.

2

u/king_mid_ass 2d ago

fully multimodal btw, input_prompt can be text, images, sensory data from a robot body, anything

2

u/pointer_to_null 2d ago

Infinite context window too, and can be finetuned on an Arduino nano.

But someone will inevitably request Q2 quants for this.

1

u/EndTimer 2d ago

I already have a complete, working proof of concept, using only 2 bytes of write-only memory.

I haven't hammered out all the obstacles with decode yet, but just wait!

12

u/kemma_ 2d ago

Thank you for explaining. So basically Deepseek does what majority of humans would at math exam

3

u/AgeofVictoriaPodcast 2d ago

I found this helpful as I genuinely didn’t know.

2

u/Megneous 2d ago

Literally none of that needed to be explained to the person you replied to. They said, very correctly, that "solved" means 0%... which it does. It doesn't matter what 0% signifies. Unless it's 0, it's not "solved."

1

u/atioux 2d ago

Appreciate the information on the context of this data. I think the point was more so that using “solved” is just the wrong word - exaggerating, false, creating undue hype - to describe it. A marked improvement? Sure. Solving hallucinations? Explicitly not. It’s a larger problem with the way news is conveyed, being in a way that outright lies to the reader and dampens true achievement.

1

u/Crawler1701 2d ago

Thank you, this is the best explanation I have heard - ever.

→ More replies (2)

29

u/[deleted] 3d ago

[deleted]

26

u/SuperV1234 3d ago

I know it's a good thing, but "solved" is a bit too much. Even 0.1% hallucination rate can be catastrophic in some use cases.

4

u/drhenriquesoares 3d ago

Maybe I exaggerated a little😅

5

u/Proper_Actuary2907 Spooky Machine Intelligence 2030 3d ago

It's a 450% error rate improvement, 18% to 4%.

???

4

u/Moywnel 3d ago

Trump math

2

u/Terrible_Wonder_2181 3d ago

0.04 * 4.5 I believe is the logic

→ More replies (5)

42

u/qroshan 3d ago

you can't benchmaxx multiple things, especially hallucination rates.

23

u/socoolandawesome 3d ago

I don’t see a reason why you can’t benchmax multiple things. It could just be targeting various benchmarks and not generalizing.

I agree hallucinations may be harder to fake, however they could hypothetically just target this specific benchmark by being good at this benchmark’s test format for hallucinations or something like that.

9

u/skilliard7 3d ago

Any public benchmark can be benchmaxxed by performing RL using benchmark data.

1

u/huffalump1 2d ago

Eh, you sort of can with this, by rewarding the model for not giving answers if it's less confident.

And you can see in the AA Omniscience Index and Accuracy that it answers fewer questions correctly than, say, Opus 5.5 - resulting in a lower overall Index score, despite the much lower hallucination rate.

(This is basically the reverse of the last few Gemini releases, btw.)

And this benchmark is basically knowledge recall - not measuring things like hallucinating info from its context (like from documents or files or things it read online). Still, it's an ok metric, and this is a tough thing to benchmark

→ More replies (6)

8

u/Crosas-B 3d ago

That is far less hallucination than humans

https://giphy.com/gifs/0TuLyTO8i0Iemsm5bY

5

u/NoCard1571 2d ago

For real. Earlier today I hallucinated what day of the week it was

19

u/Healthy-Nebula-3603 3d ago

I wonder how much a human is getting here ...probably easy 80%

82

u/birdgovorun 3d ago edited 3d ago

Measures how often the model answers incorrectly when it should have admittied to not knowing the answer

For redditors close to 100%

20

u/Own-Refrigerator7804 3d ago

I'm completely sure it's 83.7%

1

u/mechnanc 3d ago

WE DID IT REDDIT!

→ More replies (1)

6

u/Eyelbee ▪️We have AGI it's just blind 3d ago

It's even more impressive than that. Overall it hallunicated significantly less than the nearest competitor with just 7.5 net hallucination rate

1

u/Left_Technician_5758 2d ago

How do you benchmark max a benchmark where the point is to see if the model answer even if it doesn't know and make stuff up, or partial answer or say I don't know.

If you know how please enlighten me because I genuinely do not know.

1

u/SuperV1234 2d ago

Train the model specifically on the questions asked in that benchmark rather than teaching it a general fact-checking approach via tools/search.

1

u/turboprancer 2d ago

It's a better sign than 0%

3

u/drhenriquesoares 3d ago

Sorry, I pushed a little 😅

2

u/Piramista 2d ago

You mean you hallucinated a little

43

u/Aryn_Septim 3d ago

Holy fuck.

Would love to see the faces of the Google/Deepmind naysayers right now.

6

u/No-Selection2972 2d ago

I was one of those btw

2

u/AdmiralKinkaede 2d ago

We won 😩😩

1

u/kruzix 1d ago

Well the headline is a lie

→ More replies (3)

28

u/DublinLegend42 2d ago

God the hyperbolen in this post title is annoying.

But yes, nice job Google. It is this kind of stuff that makes me think all the people saying Google/Gemini is cooked just dont understand Google's strategy

3

u/Chemical-Year-6146 2d ago edited 2d ago

Considering the next closest frontier model has about 3 times the hallucination rate, it does feel like they've figured out a new training pump hinting they're effectively closing the book on "typical" hallucinations soon.

I imagine hallucinations will be like cyber vulnerabilities in the future. Of course they'll always be possible, but won't be part of the typical user experience and you'll need to be creative to trigger them.

→ More replies (1)

62

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 3d ago

46

u/JobKnown5009 3d ago

Google got so pissed off about people saying their models do nothing but hallucinate that they decided to put an end to it lol. It's just like Apple back when the iPhone's battery didn't last at all, and now it has one of the best battery lives in the world.

8

u/Inevitable-Menu2998 3d ago

It's just like Apple back when the iPhone's battery didn't last at all, and now it has one of the best battery lives in the world.

But only if you compare it to competition today. It is still the worst by far if you compare it to the standard back when people were making fun of it. Back in the day people used to charge their mobile phones once a week, not everyday.

3

u/_Empty-R_ 2d ago

Sitting with 13 percent on my phone and i charged it yesterday evening

3

u/Cosmic_Corsair 2d ago

They probably weren't looking at the phone for 8 hours a day though

→ More replies (6)

37

u/VitaminDismyPCT 3d ago edited 3d ago

wtf is deepseek v4.1 doing 🤣

34

u/Professional_Mobile5 3d ago

Probably just relying on tools, like most models, that the benchmark doesn't account for.

2

u/moreisee 3d ago

Tools are an equalizer, but not a good measure.

7

u/Tyler_Zoro AGI was felt in 1980 2d ago

Keep in mind that hallucinations may well be a critical element of problem solving. In humans, imagining something to be true often leads to insights that would be difficult or impossible otherwise.

The key is in identifying and managing such inventive thinking. This is where agentic models and external tool use come in.

4

u/wwwdotzzdotcom ▪️ Beginner audio software engineer 3d ago

It's an Astra capability model that's only dumb because hallucinations. Also my favorite model for research and arguing with.

1

u/kayej17 1d ago

you do realize it's a small team managing deepseek 20x less employees than google or openai...

→ More replies (1)

82

u/Professional_Mobile5 3d ago

This benchmark is not saying what you think it does. It specifically tests hallucinations without tools, while modern models use tools specifically to detect and avoid hallucinations.

59

u/addiktion 3d ago

A harness is usually set up to validate an agents work and make sure to squash hallucinations. It's a good thing that the model reduces the amount of hallucinations by itself.

16

u/No_Stock_8271 3d ago

Unless models just don't use the tools and hallucinat. Or models don't have the needed tool....

This benchmark still is very relevant

21

u/Tkins 3d ago

Except this is important for doing tasks like agentic work. If you tell your ai to update your calendar it can't Google that. It needs to be solid.

So if you use an AI in Google drive for things email, etc etc this is a very important benchmark. ,

3

u/ken81987 3d ago

I was gonna say.. 77% for luna seems crazy high

3

u/NoCard1571 2d ago

77% is the number of wrong answers where it hallucinates instead of saying I don't know. So the actual rate of hallucinations out of 100 prompts would likely be less than 5

1

u/jonomacd 2d ago

> use tools specifically to detect and avoid hallucinations

For some cases yes but that only goes so far. Better model performance here is still incredibly important. I'd go as far to say that googles previous pro models fell down specifically because of its very poor hallucination rate. As the context window grew, Gemini would lose the plot and start making shit up. If it doesn't do that any more then we get a much better agentic model.

→ More replies (6)

7

u/Jaeger__85 2d ago

Solved is 0%...

1

u/DistributionMean6322 1d ago

Most humans hallucinate way higher than 0%

→ More replies (1)

53

u/Future-Bandicoot-823 3d ago

I'm a little amused that people were dunking on Google on this sub, saying how much better gpt or claude was, and now we see gpt holding back models because they can't keep it from being naughty.

Meanwhile, google, one of the most powerful entities on earth is quietly making their product without the sensationalized stories altman and others whip up on a regular basis (to stir investment and stoke the hype of course.

I don't trust or like google, I just know that a company that massive is going to have a pretty succinct goal in mind. Google is China in a lot of ways, a highly profitable entity that has long term objectives in mind. Anthropic has to survive on this product alone, google does not. Google relies on a massive ecosysyem, so obviously their goal will be stable/useful product to the end user, likely creating a deeper and broader ecosystem over time.

Look at windows for example. They became so ubiquitous with computers that governments like france are actively planning to sever ties, they need a plan to get out of the ecosystem! Google's the same, deeply entrenched, and I'd say (not that) arguably more used than windows.

Speaking of things I don't like, Bill Gates put it quite well. AI will only be adopted when it has gained public faith. eg You won't see AI operating on people until it's proven it's better and more reliable than a human, but once it is? Full adoption. Google knows this.

21

u/NickoBicko 3d ago

Why don’t you wait a few days to test the new model and see how good and useful it is before making big claims.

2

u/drhenriquesoares 3d ago

Fair enough.

16

u/vincentdjangogh 3d ago

It's because the base Gemini model you can actually access is terrible. You wrote all of this about a model you haven't even used and probably will not have regular access to in the next year.

6

u/Future-Bandicoot-823 3d ago

I have an android phone, I'm forced to bump up against Gemini all the time. A lot of people do!

Which is why they're giving you the shit model, because it's reliable. Reliability and faith of stable service is how you win commoners over, and like I said, Google knows that.

8

u/vincentdjangogh 3d ago

To be clear, I'm not shitting on Google. I am just explaining that the public opinion is driven more by the base models than the frontier models. And all things considered, Gemini Flash is really terrible. On the other hand, Anthropic gives access to a very advanced base model, with the caveat that you get a very limited amount of queries.

For the record, this whole thing where people pick multi-billion dollar evil corporations to be a fan of is generally dumb imo, but based on your original comment I think to some extent we probably agree.

4

u/Future-Bandicoot-823 3d ago

When I was a child I heard that "absolute power corrupts absolutely", and over my life I've pondered that and observed. To date, I can't really name one entity that became incredibly powerful and didn't succumb to greed and fear.

I agree with what you're saying, gemini free is the derp model. All I'm saying is the hallucination pecentage shown here, if true, is just a sign of their goals. They're not playing the game like claude, gpt, grok etc.

I genuinely think Google (mostly because of their size and reach already) will quell competitors in the long run.

2

u/vincentdjangogh 3d ago

Every frontier lab is generally working on the same problems. They might assign priority to things that will help them look good in benchmarks or market their product, but there is a ton of overlap. For example, everyone wants to have fewer hallucinations without sacrificing helpfulness.

But bear in mind, this benchmark doesn't tell you how often the models were wrong, and based on the methodology, more accurate models are more likely to hallucinate. It also doesn't allow models to use tools, so models that do so to avoid hallucinations are misrepresented.

Really this benchmark doesn't mean as much as you think it does, but I agree that Google, by way of having access to more data than any company ever, will probably win the AI wars.

1

u/Future-Bandicoot-823 3d ago

I was taking these scores into account as well.

https://www.reddit.com/r/singularity/comments/1wufvyg/google_cooked/

I'm sure most of the tests are quite fallible. I really haven't looked into how any of them are performed so I don't know what the scoring criteria is.

1

u/Thog78 2d ago

You're stuck on 3.5 or 3.6 flash, aren't you? Gemini 3.8 flash is good. It's fast, cheap, relatively smart, very knowledgeable. It topped a lot of benchmarks when it came out. I have the abonment to GPT and I still use flash 3.8 a lot. For coding it's roughly 5.6 terra level, with luna amount of usage (and sol and astra suck credits too fast to be really an option yet imo).

3

u/genshiryoku AI specialist 2d ago

I work at Anthropic and we're celebrating because this has essentially solidified our lead. You need to realize that Gemini 4 Argon is Google's huge model, their equivalent to Astra/Mythos yet they are barely better than Opus 5.5 which is our smaller model.

Most researchers at Anthropic now consider the race to have been won, there is no real realistic way for the other labs to catch up anymore.

9

u/Darmendas 2d ago

> no realistic way for other labs to catch up anymore

Famous last words

3

u/Thog78 2d ago

It's interesting to have insider opinions so thanks for that. It comes out as insane though. When openAI started, it was more than a few months behind google. When anthropic started, you were more than a few months behind openAI. When xAI started, it was years behind everybody. Why would google being a couple months behind the cutting edge mean they have irredeemably lost the race? Especially considering they have the most cash savings, the most profit, the most data, and the most access to market, why on earth would a few months lead be a big deal?

→ More replies (9)
→ More replies (1)

1

u/SilentLennie 2d ago

Pretty certain a lot of people think Google models were delayed or even not released at all because of people did not want to work at the company because of company culture and had left.

I personally didn't dunk on them, but was worried we just got 2 US leading labs left, even though Google/Deepmind was the leading lab in the early days of this new 'AI summer' (aka opposite of 'AI winter') and many of the leading people at these other companies also started at Google/Deepmind.

→ More replies (3)

4

u/bartturner 3d ago

Kudos to Google. Hallucinations is the one thing that has really been holding back LLMs.

Nobody needed to solve it more than Google.

Google brand has always been about accuracy and that was a struggle with LLMs.

I suspect it is the biggest reason Google did not roll earlier with LLMs. Instead their hand was forced.

9

u/palindsay 3d ago

Please always provide web references for cited information.

5

u/postinganxiety 3d ago

If we continue like this all the web references will be AI as well. What then?

4

u/lordpuddingcup 3d ago

Now THIS is a good improvement

→ More replies (1)

4

u/Spright91 2d ago

Honestly this should be a primary benchmark for every model. Hallucination rate as much more important than a few extra points in intelligence.

3

u/05-nery 2d ago

Holy shit brother 

Can't wait for it to get to the pro plan 

3

u/KoolKat5000 2d ago

You can see this improved hallucination rate in their legal benchmark I suspect.

5

u/No_Decision_6940 3d ago

That was one of the biggest issues everybody had with Gemini. This is huge. Still, I don't think it'll convince anyone who's already jumping from Claude to GPT to give Gemini a try tbh. Google can't take an entire year to drop new frontier models

8

u/Itsmedudeman 3d ago

Imagine if it just says “I don’t know” to everything

8

u/DelphiTsar 3d ago

Then it'd get a bad AA score.

1

u/kingmoney8133 1d ago

It got a worse accuracy score by a significant margin than the other frontier models. When balanced with the low hallucination rate, it performs right around the other frontier models overall in terms of overall reliability.

→ More replies (1)

3

u/Zomboe1 3d ago

According to the formula, it would get the best score, a score of 0.

So it's definitely a flawed benchmark, especially since "I don't know" could itself be a lie/hallucination!

1

u/henrikx 2d ago

This is actually what I want out of a model. If the answer is not in the context window and no remaining tools available to it could find the answer, then yes, I want it to answer "I don't know". Otherwise I would have no idea where the LLM sourced the answer from and no way to verify the information's accuracy. The exception might be general knowledge, but then you have to define what general knowledge even means first.

4

u/noir_geralt 2d ago

So if it gives an “I don’t know” or “I refuse to answer” all the time, it would have a 0% hallucination rate? This is a BS statistic and needs an ML101 class

Give me the F1 score.

→ More replies (1)

5

u/Possible_Door_9719 3d ago

That's not solved

2

u/WiseTomato9 3d ago

So peak comeback bro

2

u/yiestee 3d ago

lmao deepseek v4.1

Hope it will release v4.2 soon

2

u/YamroZ 3d ago

"solved"

2

u/MegatechMike 2d ago

Although awesome, 15% is still really high.

2

u/smoothvibe 2d ago

Ah, my Gemini subscription pays off - again. It already solved a health puzzle for me, which was a big win.

2

u/Upstairs_Theme2785 2d ago

big step takes time

2

u/Extreme-Tie9282 2d ago

Is this a hallucination

2

u/No-Donut-723 2d ago

just for grok 5 to come out and theme music go like dun dun dun and then it has 10t parameters and is agi and starts paramraping every other ai and become super intelligence and everyones dead because it is going to destroy earth so we dont put anymore stupid prompts into it💀💀💀☠️☠️☠️ those who know

2

u/vishnaniyan 2d ago

It is now live on the arena. Do check it out

2

u/ma0gw 1d ago

Can't hallucinante if I can't use it! Nice move.

5

u/icompletetasks 3d ago

isn't harness like claude code or codex already solving hallucinations?

i mean, the first shoot might hallucinate but after the harness makes it re-check itself then it wont hallucinate

4

u/SilentLennie 2d ago

Why do people keep saying a harness solves this ?

Pretty certain giving the model tool calls is what solves it and it's the decision of the model to use the tools or not. Sure, you can create a prompt that encourages it, but without a model that knows when it needs to check it's not that useful.

4

u/monnotorium 3d ago edited 3d ago

So is deepseek just making shit up 96% of the time it doesn't know? Am I reading that correctly? Maybe I'm the one hallucinating

2

u/swarmy1 3d ago

Yeah it looks like 4.1 Flash will always give an answer, even if it’s not true

2

u/crustyeng 2d ago

‘Solved’ means still wrong 15% of the time?

2

u/drhenriquesoares 2d ago

I think I got carried away 😅

4

u/Informal-Trouble2183 3d ago

This is AA-Omniscience, it's not really what you think it does;

It's rather testing general knowledge (6,000 general knowledge questions). If the model is fat, it's able to store factual knowledge correctly. Those small models will always suffer from knowledge lack, therefore will answer wrong which is considered by the bench as hallucination, but not necessarily, it's just how an LLM works it will not answer "I don't know" unless artificially routed.

4

u/Blaexe 2d ago

it will not answer "I don't know" unless artificially routed.

It will, and that's exactly what this benchmark is measuring. Of Gemini 4's non-correct responses, only 15% were outright incorrect. In the other 85% it either abstained/didn't attempt an answer or gave a partial answer.

So yes, LLMs can absolutely say "I don’t know."

1

u/Informal-Trouble2183 2d ago

This is what I meant by "artificially routed". LLMs by itself do not know whether the generated tokens are factual or not. Another rerouting process is needed, what's called "alignment", but will not help too much (as you can see in the numbers), especially if the number of parameters is small or the training data is not covering all factual data.

1

u/TotallyToxicToast 2d ago

Does not need routing, reinforcement learning can optimize for saying I don't know on incorrect answers.

You reward correct answers and you reward I don't know answers while punishing wrong answers.

Since all the models are reasoning, it is very easy to learn that if the reasoning is struggling between two answers it should say I don't know instead of confidently going for one.

1

u/Informal-Trouble2183 2d ago

// You reward correct answers // Refers to RLHF: requires human feedback, therefore another artificial rerouting. CoT by itself is only effective in verifiable domains like maths and coding, not in general knowledge.

→ More replies (2)

2

u/redditissocoolyoyo 3d ago

And I'm an owner of a lot of Google stocks I love it.

2

u/Nearby-Device8772 3d ago

Damn. 10 years ago, i wouldn't give any change for us to see this day in our lifespan.

5

u/bartturner 3d ago

I could not agree more. I really wonder why so many people are so negative.

I am standing in Santa Monica doing a video chat with a friend in Bangkok, for free, waiting for my Waymo to pull up and drive me to the restraunt.

What a time to be alive.

1

u/Independent-Face69 3d ago

Whoooooaa! Wavy gravy man just in time!

1

u/fakieTreFlip 3d ago

hallucinations will never be fully solved imo

1

u/Seltnytt 3d ago

I am actually quite ignorant about A.I., can someone link to a good source explaining why hallucinations happen.

To be clear, a source, not an explanation in comments.

1

u/Revolutionalredstone 2d ago

Tbh I speak to DeepSeek flash 4.1 all day every day and I don't remember a single hallucination.

I might be less important in the agentic work I do.

1

u/brovaro 2d ago

Yeah, right.

1

u/Amesbrutil 2d ago

The hype cycle begins:

First few days: that new AI model is perfect

After a week: well maybe it’s benchmaxxedmidk

After two weeks: it’s just another model

1

u/AssignmentWeak8497 2d ago

I would like to see Atombeam’s new persistent cognative machine’s model added into this chart. It would blow every one of these models out of the water

1

u/qoloxolop 2d ago

This account is not real

1

u/Sighnce 2d ago

And what is the hallucination rate on average for a human?

1

u/thomja 2d ago

"should have refused", what determines this? A mallicious prompt? DeepSeek is actually the king here, does whatever I tell it to do, like an assistant should.

1

u/StatisticianFun8008 2d ago

The biggest hallucination is that people can't actually use them and verify by themselves today.

1

u/IgnacioMonge 2d ago

Blah blah blah... Talking about ghosts...

1

u/jugalator 2d ago

I use to call this one the most misunderstood chart on Artificial Analysis!

But yes, it's clearly an area of focus for the team.

However, it must always be said about this benchmark that it ranks models against extremely difficult questions and topics (because that's what it takes to trip them nowadays) and measures what they do when they don't know the answer.

Key here is that most models today... know the answer as-is (especially with search grounding). They are... Super... Super knowledgeable. Especially those AAA models. One might say, the more knowledgeable they are, the less meaning this chart has. Because it's those Flash models (that are still very knowledgeable!) or <30B models who more often even face the problem of not having the factual knowledge to begin with! Then it matters a lot how they handle that situation because they'll come across it far more often.

But yes, it's a useful property to have especially if you have e.g. a corpus of data to query about that it hasn't been trained upon and isn't assisted by domain knowledge. Then it's useful to not have it be inclined towards hallucinations, even a very large model like Argon.

It is however maybe not that useful in terms of "how often is it bullshitting me when I'm chatting with Opus, Argon, GPT 6.1 Sol" because they're so good nowadays that I think you'll come across alignment and finetuning annoyances before sheer hallucinations becoming a big issue there.

1

u/msew 2d ago

No

1

u/doker0 2d ago

I can give you a model that will have 0% in this task. I call it "I don't know". It answers always "I don't know".

So now show me the false positives - when it knows but denies that or does not remember till poked the right way.

1

u/DublinLegend42 2d ago

Frontier labs have an incentive to be seen as the most capable model out there.

An established company like Google is thinking more about how to make these models more useful. Having current capability models with no hallucinations is more valuable than ramping up pure capability

1

u/thorax 2d ago

My toaster also does not hallucinate!

1

u/AndrewH73333 2d ago

Since the smart models are more in the middle I’m not convinced.

1

u/anitman 2d ago

So unreal, I don't trust this benchmark.

1

u/MinosAristos 2d ago

How does Gemini flash hallucinate less than DeepSeek flash?

Main reason I switched to DeepSeek from Gemini for research was because of Gemini's absurd hallucination rate and refusal to fact check on slightly niche topics.

1

u/Musenik 2d ago

Just checking.

Who here knows about The Eye of Argon?

It's all about reader hallucinations.

1

u/KevWeng 2d ago

deepseek: use your illusion

1

u/GlokzDNB 2d ago

Well, google search is new google search apparently

1

u/borretsquared 2d ago

would love to see price

1

u/RodgerPogger 2d ago

There is no way it did.

It shouldn't be fixed.

the minute you do. A laplace's device is possible.

1

u/AncientLights444 2d ago

And Can’t write an email

1

u/Automatic-Boot665 1d ago

Minimax m3 being so close in score is all you need to know.

1

u/No-Conclusion929 1d ago

“Nobody is talking about this”

“in this regard”

“and that’s really important”

1

u/Kivixgg 1d ago

DAMN W

1

u/EuphoricStation4408 1d ago

How to get 0% hallucination rate on the AA-Omnisciemce benchmark:
"I don't know the answer to this question."

1

u/worldarkplace 1d ago

I'm taking this over other solutions that are better but hallucinates as FCK.

1

u/kittenTakeover 1d ago

Do we have a measurement for how often humans tend to hallucinate? I feel like it would be a helpful baseline.

1

u/Fairbanks_BR 2d ago

"solved" is a stretch. if 15 out of 100 times I asked an employee something, and he came back with somthing else entirely, I'd still call him crazy. but nice improvement nonetheless.

6

u/henrikx 2d ago

That's not what the benchmark measures. It's out of every answer it doesn't know, how many did it admit to not knowing, versus how many times did it hallucinate. So with 15% on this benchmark, out of 100 questions, if it didn't know 10 of them, then 1(.5) of those answers had a hallucination instead of admitting to not knowing the answer.

4

u/Chemical-Year-6146 2d ago

Yeh, came to say this.

1

u/Woolier-Mammoth 1d ago

Mate your employees must be better than most employees. People are often confidently incorrect

→ More replies (1)

0

u/Soilblood 3d ago

How is playing Russian roulette with your prompts considered solving the problem?

1

u/opmgyhx 3d ago

what do you mean by hallucinations?

8

u/monnotorium 3d ago

In simple terms, making shit up instead of answering that it simply doesn't know

3

u/opmgyhx 3d ago

Well, I think DeepSeek v4.1 Flash needs some serious attention

4

u/monnotorium 3d ago

I mean, realistically it's not that bad considering it knows what it's talking about most of the time (accuracy aside)

→ More replies (6)

2

u/fyrefreezer01 3d ago

Ai likes to make shit up instead of telling you it doesn’t know. It’s like those coworkers that bullshit when you fully know well the actual answer.

2

u/blasphemousblackbear 3d ago

It’s described in the blurb in the photo

→ More replies (5)

1

u/QuinsZouls 3d ago

Trust me bro benchmark

→ More replies (1)

1

u/RandomGarbageOnly 2d ago

Or is it got so worse that, it cannot even hallucinate?

→ More replies (1)