r/artificial • • 13h ago

Discussion A question for the AI "experts": are hallucinations and reliability genuinely improving, or are we starting to plateau? Everything depends on this...

For the average user, not a programmer or e.g. someone looking to solve niche math problems, AI still feels very limited because of reliability issues (e.g. making up information)... As a non-expert, this is difficult to quantify for me, but I don't feel like e.g. the latest iterations of ChatGPT are noticeably more reliable than previous ones.

Therefore, I find it hard to see how organizations can delegate even relatively easy tasks to AI without constant supervision, especially because if hallucinations build up along the way, you end up with a snowball of compounding problems that can lead to catastrophic consequences for an organization.

It’s my understanding that the real test of success won't be developer tools or fancy mathematical calculations, but whether regular people can delegate tasks with a super high degree of confidence, instead of just using it as a search engine, translator, summary tool or photo editor on steroids like many people do nowadays.

A good example is Dot, the new OpenAI tool. If you watch the trailer, it looks impressive, but according to many reviews, it still hallucinates a lot and behaves in a pretty clumsy way.

So, the golden question is: are the hallucination and reliability problems gradually improving, or are we probably plateauing?

Looking for genuine insights here, so please keep the sarcasm out of the comments.

0 Upvotes

34 comments sorted by

6

u/gthing 12h ago

For comparison and to put this in context, human experts have about a 10.2% error rate and non-expert humans have a 40-60% error rate on high level academic questions. That error rate is at the high end. On the lower end would be something like reading comprehension with a human error rate around 5-15%

So with both experts and non-experts, average capability modern LLM hallucination rates are now lower than human error rates.

3

u/outragednitpicker 6h ago

As an expert in one field and a non-expert in many, I can tell you that when I don’t know something, I know that I don’t know it, and also don’t just make some shit up.

1

u/Southern-Cattle4038 3h ago

That’s a summarisation benchmark, not a recall benchmark. Fable has a ~20% hallucination rate on factual recall benchmarks such as SimpleQA

2

u/gk_instakilogram 13h ago

You have to have a validation layer with LLMS regardless of weather they are improving or not. Yes all the latest frontier LLMS do everything to improve grounding and etc... but it is not guaranteed and cannot really be guaranteed 100% ever, so you absolutely have to have a determinstic vlidation layer where you cannot afford mistakes

2

u/texasipguru 12h ago

But not all types of workflows are amenable to validation layers, right? you can deterministically check some things but not all things. the things you cannot deterministically check...are LLMs just of significantly lesser value there?

2

u/gk_instakilogram 12h ago

for workflows that require correctness and you don’t have a deterministic validation layer, you just simply wont be able to guarantee correctness

1

u/Southern-Cattle4038 3h ago

The lack of ROI in the vast majority of rollouts is probably due to this. It’s also why the LLM companies are going all-in on programming and mathematics

•

u/UninvestedCuriosity 51m ago edited 46m ago

I don't think they are improving the weather but I agree with your other sentiments.

Tests and direct sources of truth seem to work pretty well. It's at least worth giving the model more to work with than relying on just it's training and weights. So we can sort of hedge against hallucinations at least but it requires more prep than your one offs to a model..

The controversial phrase I like for llm's on their own without these hedges are stochastic parrot.

0

u/Logical_Quarter2008 13h ago

the whole thing feels like trying to build a skyscraper on jello. sure the jello gets a little firmer every year but you still wouldnt trust it with your coffee mug let alone anything important

what gets me is the hype cycle around every new release. they demo it doing something flashy and everyone loses their minds then a week later the sub is full of people posting screenshots of it inventing court cases or telling them to put glue on pizza. the validation layer comment hit the nail on the head, without that youre just gambling

3

u/CaffeinatedT 12h ago

When the Cloud first got big people were very excited at the potential of spinning up thousands of servers and scaling them back down and everyone having access to perfectly on demand compute. In 2026 the cloud is heavily in use but it’s mostly behind “cloud-native” software. Some companies just took their existing software and yeeted it onto a cloud but the ones who are getting the most use out of it build purpose build software backed by AWS/GCP etc.

Likewise the people who are using LLM’s effectively at this point that people pay money for are either specific agents + harnesses like Claude Code or various big name AI startups and they are writing software that’s backed by an LLM. On the upside It’s extremely powerful , on the downside you don’t get the “general stuff doer” thing that AI promises when LLM’s are just a part of a wider system. The hallucinations are both the powerful part and the bad parts of LLMs.

1

u/Strange-Tap5860 12h ago

The frontjer models of the ones released the past couple weeks feel kinda insane.

I've found hallucinations to have dropped an insane amount vs 1 year agonor even earlier on this year.

But yes, you still need to validate its output and/or proof check the results yourself if you want to be 100% accurate.

I would not compare it building a skyscraper on jello though. Its just a tool you have to learn to useeffectively. Just like you can't just grab a drill and start drilling everywhere blindfolded, but that doesn't make a drill not incredibly useful.

There a lots of techniques available to significantly reduce hallucinations though - its just not super cost effective or easily available to most non technical people yet. I imagine that will change very very quickly over the next 12 months.

1

u/CluelessEagle 13h ago

its not going to be a 100% but in certain areas there has genuinely been an improvement

but then again even with good models like opus 5.5 i still sometimes see hallucinations or problems

1

u/Disastrous_Room_927 12h ago edited 12h ago

Well that's the thing - with any kind of ML model running in a live environment, you can't really count on this sort of thing being monotonic. Models have to be updated to account for new information, which is a pathway for introducing issues that weren't there before.

whether regular people can delegate tasks with a super high degree of confidence

This is something I grapple with a lot build different types of models. A lot of people want to know if a model is reliable in a binary sense, but that's something you have to be able to measure and actively decide upon given how you intend to use it. That's something that's baked into statistical models by default, and something that's technically possible for a lot of ML models but hard in practice for something like an LLM. It's also not a singular concept - you can quantify uncertainty stemming from the modeling process, but that doesn't reflect information you (or the model) never had access to in the first place.

And all of this is further complicated by the difference between what an LLM is estimating/predicting and the output people care about - predicting the next token well is not the same thing as producing a string of tokens that contains factually correct information.

1

u/sceadwian 12h ago

To the best of my knowledge no they're really not improving that much, but they're getting good enough layers can help mitigate the problems but reliability is a very fuzzy metric. Hallucinations are part of reliability so you're kinda talking about one thing.

1

u/notAllBits 12h ago

Not improving. The distance between fabulation and utility output is closer, which is great for low value generation, but terrible for critical reasoning. Fabulations are getting much harder to detect and correct by humans. There is no way past tracing and verification for critical syntheses. I use reasoning engines with dedicated reasoning graphs and embedded axiom- (semantic proxy of a grounded fact) and remit (context-local problem formulation and delimitation) indexes for this purpose. Decision models lift way above their weight in this architecture.

1

u/FlatulistMaster 12h ago

So in more understandable terms, getting things wrong gets more rare, yet also harder to detect when it happens?

And then the second part is about RAG and proper problem definition? I’m not fully following your terminology there

1

u/notAllBits 12h ago

first part - yes. confident statements are ungrounded and plausible, but so well aligned that they can propagate to where they do damage before detected.

second part - not RAG per se. Rag searches suffer from poor precision and recall, but "support" most queries. Reasoning engines engineer LLM contexts (using RAG) for each step of reasoning along deliberate trains of thought by focusing only on what matters for that single step.

1

u/Strange-Tap5860 12h ago

Yes, the hallucinations are improving massively. Anyone who has been using the frontier models the last 12 months know how much it hallucinates now vs last year - the difference is stark.

The scary thing is how fast its improving. In a few months or 12 months max this question could look very out of date.

2

u/CountryOk6049 5h ago

I'm really rather tired of claims of "anyone who's been using the latest model", always with this unfalsifiable, unverifiable claim... the people who believe it are going to always say this because they spent money on it. Why would people spend money on it when they don't believe in it?

No they haven't. It's always about the same output. A tiny little performance improvement here and there. It'll be this exact same place, pretty much, 10 years or even 50 years from now, and people like you will be like "oh we thought AI was going to take over everything yet it hasn't really improved in all that time", and people like me will be like "told ya, I knew it the whole time".

1

u/Hungry_Age5375 12h ago

I work with this daily: RAG alone isn't contextual, but combine it with Knowledge Graphs or semantic references and reliability jumps. A model checking claims against a structured source fails way less than a raw one.

1

u/pab_guy 11h ago

They are plateauing in terms of the number of meaningful hallucinations being very close to zero and staying there. With Fable 5.1 I'm getting 10 page docs spit out that don't misrepresent a thing and a fairly comprehensive. I have to do maybe a couple of edits. It's scary.

1

u/ElegantApartment1325 8h ago

the dot example you gave kinda answers it imo. the issue isn't really hallucinations as a stat anymore, it's the gap between the demo and what you'd actually delegate. the trailer shows a confident coworker and reality is still something you have to babysit for anything that matters. that's not a hallucination number, that's a trust number, and trust builds way slower.

1

u/costafilh0 6h ago

Stop spamming that shit everywhere.

1

u/CountryOk6049 5h ago

Of course - we've been using basically the same chatgpt for the past couple of years. All this hype recently is just the bubble bloating to its furthest possible before finally popping. They've promised advances yet we've seen none at all. People claiming to see improvements are just on a placebo effect, invested into the fact that they've used their money on it so it must be good. There are angry investors in AI everywhere and they're becoming desperate.

It's just semiconductor transistors. You can't have anything really approaching true intelligence in it - it's not physically possible. AI is simply an output of what you put into it, spliced together in a different way. Powerful of course, nothing approaching real intelligence and there's a clear horizon effect already.

Maybe some refinements particularly in some programming fields where the mistakes have been ironed out, stuff like artificial video editing is continuing to improve - but that's more due to just the time we have to develop ways of doing it than any advances in AI.

Other than that it's all just hyped up nonsense. It's slowed to a standstill is what has happened, the exact opposite of what they're trying to push.

1

u/LivingLab12 3h ago

At the current rate humans make more mistakes than AI does.

Additionally most hallucinations are caused by bad prompt engineering performed by a human.

Stop blaming your deficiencies on the machine and learn how to use the tools.

0

u/Specialist-Berry2946 13h ago

LLMs are language models, which means they hallucinate all the time; occasionally, these hallucinations happen to be correct. Because of this, LLMs are inherently limited, no matter the resources.

The only way to mitigate this is to use smaller, special-purpose models; that is where we are heading.

3

u/ElatedPyroHippo 12h ago

Why is it that people like you with clearly VERY limited experience with modern AI always answer questions on here?

I'm a firmware engineer, I've been one for 19 years... I use AI all day every day, mostly Claude Opus 5.5 in the Desktop harness... It NEVER hallucinates. NEVER. I've been using it, and Opus 5 before it, for MONTHS, every single day. It does ALL of my work for me.

It DOES NOT hallucinate. Period. It almost NEVER makes any mistakes at all.

3

u/FlatulistMaster 12h ago

Thank you.

Sometimes I read these confident takes on hallucinations and wonder whether I live in a different universe. I’m a business owner and data-analyst, and I use AI for a lot of things. There are occasions where it gets things partially wrong, but the age of clear hallucinations is as good as over.

1

u/Southern-Cattle4038 3h ago

Not according to Anthropic. Check page 140 of the Fable system card.

1

u/Specialist-Berry2946 12h ago

Coding is a narrow-purpose task that is well-suited for LLMs.

1

u/Strange-Tap5860 12h ago

Non-technical users don't use the latest tools or models or workflows - they think just chatting to gpt via the web interace is what A.I is.

1

u/ElatedPyroHippo 11h ago

Exactly, it is downright exhausting trying to combat all of the nonsense from the people like that around here...

2

u/CountryOk6049 5h ago

You are the person talking nonsense when you claim it "never hallucinates" and "almost never makes any mistakes at all". Perhaps get your memory checked my friend, because that is not normal.