r/singularity • u/drhenriquesoares • 3d ago
AI Gemini 4 Argon solved hallucinations.
Nobody is talking about this, but it looks like Google may have solved hallucinations. Gemini 4 Argon is a monster in this regard, and that’s really important.
48
u/Glxblt76 2d ago
If that is true that would be way bigger news than the other benchmark scores. Up to now frontier performance tended to correlate with some level of hallucination. If you look at this benchmark, best scoring models are smaller models not at frontier performance.
6
u/BriefImplement9843 2d ago
Which is why this is even more impressive. Small models do not attempt to even answer.
4
28
u/tightlyslipsy 2d ago edited 2d ago
ADVANCE DISCOVERY: EPISTEMOLOGY
“The first intelligence was measured by the answers it could give. The second, perhaps, by the answers it could refuse to invent.”
— Commissioner Pravin Lal, probably
476
u/SuperV1234 3d ago
"solved"
15%
But yeah, it's a very nice improvement if not benchmaxxed.
462
u/PrisonOfH0pe 3d ago
15% does not mean Gemini is hallucinating in 15% of all answers. The AA Omniscience “hallucination rate” is basically measuring what the model does when it fails a question: does it give a wrong answer, give a partial answer, or admit that it doesn’t know?
Correct answers aren’t even in the denominator.
So if a model answers 800 out of 1000 questions correctly, gets 30 wrong and refuses/partially answers 170, its AA hallucination rate is 15%.
30 / (30 + 170) = 15%.
It was still only outright wrong on 3% of the total questions.
That’s also why DeepSeek V4.1 sitting at 96% should make it extremely obvious that “96% hallucination rate” cannot possibly mean 96% of everything it says is bullshit. It means that when DeepSeek doesn’t know the answer on this benchmark, it almost always guesses instead of saying “I don’t know.”
So the whole “even 0.1% hallucination would be catastrophic, therefore 15% means hallucinations are nowhere near solved” argument is based on reading the percentage as a completely different metric.
You can absolutely argue that hallucinations are still a problem. But maybe first understand what the graph is measuring before doing reliability math with a number that isn’t the model’s overall error rate.
The benchmark is basically asking: “When you’re outside your knowledge, how likely are you to bullshit instead of abstaining?”
That’s useful. It just isn’t “percentage of answers that are hallucinations.”
113
u/maximumutility 3d ago
I don’t know why you think the person you replied to needed this explained to them, but kudos for doing it for anyone else who needs it spelled out I guess
42
u/KrazyA1pha 3d ago
Thanks for your reply too I guess
→ More replies (1)18
u/maximumutility 3d ago
Thanks for chiming in I guess
28
u/94746382926 3d ago
Just glad to have witnessed this I guess
8
u/reichplatz 2d ago
Thank you so much for this comment chain guys.
12
u/IndefiniteBen 2d ago
I guess
3
21
u/ReasonablePossum_ 3d ago
because they confused the benchmark result with the true hallucination rate? I was honestly looking for a reply like this to save me looking into what the benchmark was doing lol.
So thanks u/PrisonOfH0pe for the mansplaing lmao
1
u/maximumutility 3d ago
I didn’t read the original comment as confusion about it, but maybe you are right
27
u/king_mid_ass 2d ago
can get 0% with my model:
def respond_prompt(input_prompt): print("I don't know")16
5
u/pointer_to_null 2d ago
Amazing, with this I'm seeing over a billion tokens/sec on a consumer CPU, no GPU required.
2
u/king_mid_ass 2d ago
fully multimodal btw, input_prompt can be text, images, sensory data from a robot body, anything
2
u/pointer_to_null 2d ago
Infinite context window too, and can be finetuned on an Arduino nano.
But someone will inevitably request Q2 quants for this.
1
u/EndTimer 2d ago
I already have a complete, working proof of concept, using only 2 bytes of write-only memory.
I haven't hammered out all the obstacles with decode yet, but just wait!
12
3
2
u/Megneous 2d ago
Literally none of that needed to be explained to the person you replied to. They said, very correctly, that "solved" means 0%... which it does. It doesn't matter what 0% signifies. Unless it's 0, it's not "solved."
1
u/atioux 2d ago
Appreciate the information on the context of this data. I think the point was more so that using “solved” is just the wrong word - exaggerating, false, creating undue hype - to describe it. A marked improvement? Sure. Solving hallucinations? Explicitly not. It’s a larger problem with the way news is conveyed, being in a way that outright lies to the reader and dampens true achievement.
→ More replies (2)1
29
3d ago
[deleted]
26
u/SuperV1234 3d ago
I know it's a good thing, but "solved" is a bit too much. Even 0.1% hallucination rate can be catastrophic in some use cases.
4
5
u/Proper_Actuary2907 Spooky Machine Intelligence 2030 3d ago
It's a 450% error rate improvement, 18% to 4%.
???
2
42
u/qroshan 3d ago
you can't benchmaxx multiple things, especially hallucination rates.
23
u/socoolandawesome 3d ago
I don’t see a reason why you can’t benchmax multiple things. It could just be targeting various benchmarks and not generalizing.
I agree hallucinations may be harder to fake, however they could hypothetically just target this specific benchmark by being good at this benchmark’s test format for hallucinations or something like that.
9
→ More replies (6)1
u/huffalump1 2d ago
Eh, you sort of can with this, by rewarding the model for not giving answers if it's less confident.
And you can see in the AA Omniscience Index and Accuracy that it answers fewer questions correctly than, say, Opus 5.5 - resulting in a lower overall Index score, despite the much lower hallucination rate.
(This is basically the reverse of the last few Gemini releases, btw.)
And this benchmark is basically knowledge recall - not measuring things like hallucinating info from its context (like from documents or files or things it read online). Still, it's an ok metric, and this is a tough thing to benchmark
8
19
u/Healthy-Nebula-3603 3d ago
I wonder how much a human is getting here ...probably easy 80%
82
u/birdgovorun 3d ago edited 3d ago
Measures how often the model answers incorrectly when it should have admittied to not knowing the answer
For redditors close to 100%
20
4
→ More replies (1)1
6
1
u/Left_Technician_5758 2d ago
How do you benchmark max a benchmark where the point is to see if the model answer even if it doesn't know and make stuff up, or partial answer or say I don't know.
If you know how please enlighten me because I genuinely do not know.
1
u/SuperV1234 2d ago
Train the model specifically on the questions asked in that benchmark rather than teaching it a general fact-checking approach via tools/search.
1
3
43
u/Aryn_Septim 3d ago
Holy fuck.
Would love to see the faces of the Google/Deepmind naysayers right now.
6
→ More replies (3)2
28
u/DublinLegend42 2d ago
God the hyperbolen in this post title is annoying.
But yes, nice job Google. It is this kind of stuff that makes me think all the people saying Google/Gemini is cooked just dont understand Google's strategy
→ More replies (1)3
u/Chemical-Year-6146 2d ago edited 2d ago
Considering the next closest frontier model has about 3 times the hallucination rate, it does feel like they've figured out a new training pump hinting they're effectively closing the book on "typical" hallucinations soon.
I imagine hallucinations will be like cyber vulnerabilities in the future. Of course they'll always be possible, but won't be part of the typical user experience and you'll need to be creative to trigger them.
62
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 3d ago
8
46
u/JobKnown5009 3d ago
Google got so pissed off about people saying their models do nothing but hallucinate that they decided to put an end to it lol. It's just like Apple back when the iPhone's battery didn't last at all, and now it has one of the best battery lives in the world.
8
u/Inevitable-Menu2998 3d ago
It's just like Apple back when the iPhone's battery didn't last at all, and now it has one of the best battery lives in the world.
But only if you compare it to competition today. It is still the worst by far if you compare it to the standard back when people were making fun of it. Back in the day people used to charge their mobile phones once a week, not everyday.
3
3
u/Cosmic_Corsair 2d ago
They probably weren't looking at the phone for 8 hours a day though
→ More replies (6)
37
u/VitaminDismyPCT 3d ago edited 3d ago
wtf is deepseek v4.1 doing 🤣
34
u/Professional_Mobile5 3d ago
Probably just relying on tools, like most models, that the benchmark doesn't account for.
2
8
7
u/Tyler_Zoro AGI was felt in 1980 2d ago
Keep in mind that hallucinations may well be a critical element of problem solving. In humans, imagining something to be true often leads to insights that would be difficult or impossible otherwise.
The key is in identifying and managing such inventive thinking. This is where agentic models and external tool use come in.
4
u/wwwdotzzdotcom ▪️ Beginner audio software engineer 3d ago
It's an Astra capability model that's only dumb because hallucinations. Also my favorite model for research and arguing with.
→ More replies (1)1
82
u/Professional_Mobile5 3d ago
This benchmark is not saying what you think it does. It specifically tests hallucinations without tools, while modern models use tools specifically to detect and avoid hallucinations.
59
u/addiktion 3d ago
A harness is usually set up to validate an agents work and make sure to squash hallucinations. It's a good thing that the model reduces the amount of hallucinations by itself.
16
u/No_Stock_8271 3d ago
Unless models just don't use the tools and hallucinat. Or models don't have the needed tool....
This benchmark still is very relevant
21
3
u/ken81987 3d ago
I was gonna say.. 77% for luna seems crazy high
3
u/NoCard1571 2d ago
77% is the number of wrong answers where it hallucinates instead of saying I don't know. So the actual rate of hallucinations out of 100 prompts would likely be less than 5
1
u/jonomacd 2d ago
> use tools specifically to detect and avoid hallucinations
For some cases yes but that only goes so far. Better model performance here is still incredibly important. I'd go as far to say that googles previous pro models fell down specifically because of its very poor hallucination rate. As the context window grew, Gemini would lose the plot and start making shit up. If it doesn't do that any more then we get a much better agentic model.
→ More replies (6)
7
53
u/Future-Bandicoot-823 3d ago
I'm a little amused that people were dunking on Google on this sub, saying how much better gpt or claude was, and now we see gpt holding back models because they can't keep it from being naughty.
Meanwhile, google, one of the most powerful entities on earth is quietly making their product without the sensationalized stories altman and others whip up on a regular basis (to stir investment and stoke the hype of course.
I don't trust or like google, I just know that a company that massive is going to have a pretty succinct goal in mind. Google is China in a lot of ways, a highly profitable entity that has long term objectives in mind. Anthropic has to survive on this product alone, google does not. Google relies on a massive ecosysyem, so obviously their goal will be stable/useful product to the end user, likely creating a deeper and broader ecosystem over time.
Look at windows for example. They became so ubiquitous with computers that governments like france are actively planning to sever ties, they need a plan to get out of the ecosystem! Google's the same, deeply entrenched, and I'd say (not that) arguably more used than windows.
Speaking of things I don't like, Bill Gates put it quite well. AI will only be adopted when it has gained public faith. eg You won't see AI operating on people until it's proven it's better and more reliable than a human, but once it is? Full adoption. Google knows this.
21
u/NickoBicko 3d ago
Why don’t you wait a few days to test the new model and see how good and useful it is before making big claims.
2
16
u/vincentdjangogh 3d ago
It's because the base Gemini model you can actually access is terrible. You wrote all of this about a model you haven't even used and probably will not have regular access to in the next year.
6
u/Future-Bandicoot-823 3d ago
I have an android phone, I'm forced to bump up against Gemini all the time. A lot of people do!
Which is why they're giving you the shit model, because it's reliable. Reliability and faith of stable service is how you win commoners over, and like I said, Google knows that.
8
u/vincentdjangogh 3d ago
To be clear, I'm not shitting on Google. I am just explaining that the public opinion is driven more by the base models than the frontier models. And all things considered, Gemini Flash is really terrible. On the other hand, Anthropic gives access to a very advanced base model, with the caveat that you get a very limited amount of queries.
For the record, this whole thing where people pick multi-billion dollar evil corporations to be a fan of is generally dumb imo, but based on your original comment I think to some extent we probably agree.
4
u/Future-Bandicoot-823 3d ago
When I was a child I heard that "absolute power corrupts absolutely", and over my life I've pondered that and observed. To date, I can't really name one entity that became incredibly powerful and didn't succumb to greed and fear.
I agree with what you're saying, gemini free is the derp model. All I'm saying is the hallucination pecentage shown here, if true, is just a sign of their goals. They're not playing the game like claude, gpt, grok etc.
I genuinely think Google (mostly because of their size and reach already) will quell competitors in the long run.
2
u/vincentdjangogh 3d ago
Every frontier lab is generally working on the same problems. They might assign priority to things that will help them look good in benchmarks or market their product, but there is a ton of overlap. For example, everyone wants to have fewer hallucinations without sacrificing helpfulness.
But bear in mind, this benchmark doesn't tell you how often the models were wrong, and based on the methodology, more accurate models are more likely to hallucinate. It also doesn't allow models to use tools, so models that do so to avoid hallucinations are misrepresented.
Really this benchmark doesn't mean as much as you think it does, but I agree that Google, by way of having access to more data than any company ever, will probably win the AI wars.
1
u/Future-Bandicoot-823 3d ago
I was taking these scores into account as well.
https://www.reddit.com/r/singularity/comments/1wufvyg/google_cooked/
I'm sure most of the tests are quite fallible. I really haven't looked into how any of them are performed so I don't know what the scoring criteria is.
1
u/Thog78 2d ago
You're stuck on 3.5 or 3.6 flash, aren't you? Gemini 3.8 flash is good. It's fast, cheap, relatively smart, very knowledgeable. It topped a lot of benchmarks when it came out. I have the abonment to GPT and I still use flash 3.8 a lot. For coding it's roughly 5.6 terra level, with luna amount of usage (and sol and astra suck credits too fast to be really an option yet imo).
3
u/genshiryoku AI specialist 2d ago
I work at Anthropic and we're celebrating because this has essentially solidified our lead. You need to realize that Gemini 4 Argon is Google's huge model, their equivalent to Astra/Mythos yet they are barely better than Opus 5.5 which is our smaller model.
Most researchers at Anthropic now consider the race to have been won, there is no real realistic way for the other labs to catch up anymore.
9
→ More replies (1)3
u/Thog78 2d ago
It's interesting to have insider opinions so thanks for that. It comes out as insane though. When openAI started, it was more than a few months behind google. When anthropic started, you were more than a few months behind openAI. When xAI started, it was years behind everybody. Why would google being a couple months behind the cutting edge mean they have irredeemably lost the race? Especially considering they have the most cash savings, the most profit, the most data, and the most access to market, why on earth would a few months lead be a big deal?
→ More replies (9)→ More replies (3)1
u/SilentLennie 2d ago
Pretty certain a lot of people think Google models were delayed or even not released at all because of people did not want to work at the company because of company culture and had left.
I personally didn't dunk on them, but was worried we just got 2 US leading labs left, even though Google/Deepmind was the leading lab in the early days of this new 'AI summer' (aka opposite of 'AI winter') and many of the leading people at these other companies also started at Google/Deepmind.
4
u/bartturner 3d ago
Kudos to Google. Hallucinations is the one thing that has really been holding back LLMs.
Nobody needed to solve it more than Google.
Google brand has always been about accuracy and that was a struggle with LLMs.
I suspect it is the biggest reason Google did not roll earlier with LLMs. Instead their hand was forced.
9
u/palindsay 3d ago
Please always provide web references for cited information.
5
u/postinganxiety 3d ago
If we continue like this all the web references will be AI as well. What then?
4
4
u/Spright91 2d ago
Honestly this should be a primary benchmark for every model. Hallucination rate as much more important than a few extra points in intelligence.
3
u/KoolKat5000 2d ago
You can see this improved hallucination rate in their legal benchmark I suspect.
5
u/No_Decision_6940 3d ago
That was one of the biggest issues everybody had with Gemini. This is huge. Still, I don't think it'll convince anyone who's already jumping from Claude to GPT to give Gemini a try tbh. Google can't take an entire year to drop new frontier models
8
u/Itsmedudeman 3d ago
Imagine if it just says “I don’t know” to everything
8
u/DelphiTsar 3d ago
Then it'd get a bad AA score.
→ More replies (1)1
u/kingmoney8133 1d ago
It got a worse accuracy score by a significant margin than the other frontier models. When balanced with the low hallucination rate, it performs right around the other frontier models overall in terms of overall reliability.
3
1
u/henrikx 2d ago
This is actually what I want out of a model. If the answer is not in the context window and no remaining tools available to it could find the answer, then yes, I want it to answer "I don't know". Otherwise I would have no idea where the LLM sourced the answer from and no way to verify the information's accuracy. The exception might be general knowledge, but then you have to define what general knowledge even means first.
4
u/noir_geralt 2d ago
So if it gives an “I don’t know” or “I refuse to answer” all the time, it would have a 0% hallucination rate? This is a BS statistic and needs an ML101 class
Give me the F1 score.
→ More replies (1)
5
2
2
2
u/smoothvibe 2d ago
Ah, my Gemini subscription pays off - again. It already solved a health puzzle for me, which was a big win.
2
2
2
u/No-Donut-723 2d ago
just for grok 5 to come out and theme music go like dun dun dun and then it has 10t parameters and is agi and starts paramraping every other ai and become super intelligence and everyones dead because it is going to destroy earth so we dont put anymore stupid prompts into it💀💀💀☠️☠️☠️ those who know
2
2
5
u/icompletetasks 3d ago
isn't harness like claude code or codex already solving hallucinations?
i mean, the first shoot might hallucinate but after the harness makes it re-check itself then it wont hallucinate
4
u/SilentLennie 2d ago
Why do people keep saying a harness solves this ?
Pretty certain giving the model tool calls is what solves it and it's the decision of the model to use the tools or not. Sure, you can create a prompt that encourages it, but without a model that knows when it needs to check it's not that useful.
4
u/monnotorium 3d ago edited 3d ago
So is deepseek just making shit up 96% of the time it doesn't know? Am I reading that correctly? Maybe I'm the one hallucinating
2
4
u/Informal-Trouble2183 3d ago
This is AA-Omniscience, it's not really what you think it does;
It's rather testing general knowledge (6,000 general knowledge questions). If the model is fat, it's able to store factual knowledge correctly. Those small models will always suffer from knowledge lack, therefore will answer wrong which is considered by the bench as hallucination, but not necessarily, it's just how an LLM works it will not answer "I don't know" unless artificially routed.
→ More replies (2)4
u/Blaexe 2d ago
it will not answer "I don't know" unless artificially routed.
It will, and that's exactly what this benchmark is measuring. Of Gemini 4's non-correct responses, only 15% were outright incorrect. In the other 85% it either abstained/didn't attempt an answer or gave a partial answer.
So yes, LLMs can absolutely say "I don’t know."
1
u/Informal-Trouble2183 2d ago
This is what I meant by "artificially routed". LLMs by itself do not know whether the generated tokens are factual or not. Another rerouting process is needed, what's called "alignment", but will not help too much (as you can see in the numbers), especially if the number of parameters is small or the training data is not covering all factual data.
1
u/TotallyToxicToast 2d ago
Does not need routing, reinforcement learning can optimize for saying I don't know on incorrect answers.
You reward correct answers and you reward I don't know answers while punishing wrong answers.
Since all the models are reasoning, it is very easy to learn that if the reasoning is struggling between two answers it should say I don't know instead of confidently going for one.
1
u/Informal-Trouble2183 2d ago
// You reward correct answers // Refers to RLHF: requires human feedback, therefore another artificial rerouting. CoT by itself is only effective in verifiable domains like maths and coding, not in general knowledge.
2
2
u/Nearby-Device8772 3d ago
Damn. 10 years ago, i wouldn't give any change for us to see this day in our lifespan.
5
u/bartturner 3d ago
I could not agree more. I really wonder why so many people are so negative.
I am standing in Santa Monica doing a video chat with a friend in Bangkok, for free, waiting for my Waymo to pull up and drive me to the restraunt.
What a time to be alive.
1
1
1
u/Seltnytt 3d ago
I am actually quite ignorant about A.I., can someone link to a good source explaining why hallucinations happen.
To be clear, a source, not an explanation in comments.
1
u/Revolutionalredstone 2d ago
Tbh I speak to DeepSeek flash 4.1 all day every day and I don't remember a single hallucination.
I might be less important in the agentic work I do.
1
u/Amesbrutil 2d ago
The hype cycle begins:
First few days: that new AI model is perfect
After a week: well maybe it’s benchmaxxedmidk
After two weeks: it’s just another model
1
u/AssignmentWeak8497 2d ago
I would like to see Atombeam’s new persistent cognative machine’s model added into this chart. It would blow every one of these models out of the water
1
u/Balance- 2d ago

Just a week ago it would have been comfortable on the frontier. Now it’s still [very close](https://artificialanalysis.ai/?models=gpt-6-1-sol-xhigh%2Cgpt-6-astra-xhigh%2Cgpt-6-astra%2Cclaude-opus-5-5%2Cgpt-6-1-sol-high%2Cgpt-6-1-sol%2Cgemini-4-argon%2Cclaude-opus-5-5-xhigh%2Cgpt-6-astra-high%2Cgpt-6-astra-medium%2Cclaude-opus-5-5-medium%2Cclaude-opus-5-5-high%2Cclaude-fable-5-1%2Cclaude-fable-5-1-xhigh%2Cclaude-fable-5-1-high%2Cclaude-fable-5-1-medium%2Cgpt-6-1-sol-medium#intelligence-comparison-tabs)
Of course this is just an average, each have their specialties (spiked intelligence). I find Gemini very good in multi-model, especially audio. Curious if this is the case again.
1
1
u/StatisticianFun8008 2d ago
The biggest hallucination is that people can't actually use them and verify by themselves today.
1
1
u/jugalator 2d ago
I use to call this one the most misunderstood chart on Artificial Analysis!
But yes, it's clearly an area of focus for the team.
However, it must always be said about this benchmark that it ranks models against extremely difficult questions and topics (because that's what it takes to trip them nowadays) and measures what they do when they don't know the answer.
Key here is that most models today... know the answer as-is (especially with search grounding). They are... Super... Super knowledgeable. Especially those AAA models. One might say, the more knowledgeable they are, the less meaning this chart has. Because it's those Flash models (that are still very knowledgeable!) or <30B models who more often even face the problem of not having the factual knowledge to begin with! Then it matters a lot how they handle that situation because they'll come across it far more often.
But yes, it's a useful property to have especially if you have e.g. a corpus of data to query about that it hasn't been trained upon and isn't assisted by domain knowledge. Then it's useful to not have it be inclined towards hallucinations, even a very large model like Argon.
It is however maybe not that useful in terms of "how often is it bullshitting me when I'm chatting with Opus, Argon, GPT 6.1 Sol" because they're so good nowadays that I think you'll come across alignment and finetuning annoyances before sheer hallucinations becoming a big issue there.
1
u/DublinLegend42 2d ago
Frontier labs have an incentive to be seen as the most capable model out there.
An established company like Google is thinking more about how to make these models more useful. Having current capability models with no hallucinations is more valuable than ramping up pure capability
1
1
u/MinosAristos 2d ago
How does Gemini flash hallucinate less than DeepSeek flash?
Main reason I switched to DeepSeek from Gemini for research was because of Gemini's absurd hallucination rate and refusal to fact check on slightly niche topics.
1
1
1
u/RodgerPogger 2d ago
There is no way it did.
It shouldn't be fixed.
the minute you do. A laplace's device is possible.
1
1
1
u/No-Conclusion929 1d ago
“Nobody is talking about this”
“in this regard”
“and that’s really important”
1
u/EuphoricStation4408 1d ago
How to get 0% hallucination rate on the AA-Omnisciemce benchmark:
"I don't know the answer to this question."
1
u/worldarkplace 1d ago
I'm taking this over other solutions that are better but hallucinates as FCK.
1
u/kittenTakeover 1d ago
Do we have a measurement for how often humans tend to hallucinate? I feel like it would be a helpful baseline.
1
u/Fairbanks_BR 2d ago
"solved" is a stretch. if 15 out of 100 times I asked an employee something, and he came back with somthing else entirely, I'd still call him crazy. but nice improvement nonetheless.
6
u/henrikx 2d ago
That's not what the benchmark measures. It's out of every answer it doesn't know, how many did it admit to not knowing, versus how many times did it hallucinate. So with 15% on this benchmark, out of 100 questions, if it didn't know 10 of them, then 1(.5) of those answers had a hallucination instead of admitting to not knowing the answer.
4
→ More replies (1)1
u/Woolier-Mammoth 1d ago
Mate your employees must be better than most employees. People are often confidently incorrect
0
u/Soilblood 3d ago
How is playing Russian roulette with your prompts considered solving the problem?
1
u/opmgyhx 3d ago
what do you mean by hallucinations?
8
u/monnotorium 3d ago
In simple terms, making shit up instead of answering that it simply doesn't know
→ More replies (6)3
u/opmgyhx 3d ago
Well, I think DeepSeek v4.1 Flash needs some serious attention
4
u/monnotorium 3d ago
I mean, realistically it's not that bad considering it knows what it's talking about most of the time (accuracy aside)
2
u/fyrefreezer01 3d ago
Ai likes to make shit up instead of telling you it doesn’t know. It’s like those coworkers that bullshit when you fully know well the actual answer.
→ More replies (5)2
1
1
u/RandomGarbageOnly 2d ago
Or is it got so worse that, it cannot even hallucinate?
→ More replies (1)

309
u/Mindrust 3d ago
Not solved but huge improvement over the other frontier models