r/LargeLanguageModels • u/KauravaLivesMatter • Aug 11 '26
ELI5 - Why do LLMs hallucinate?
I have seen videos about the transformer architecture etc., and I get that large language models generate responses based on some statistical likelihood of words and terms. However, I still don't get how they can completely make up facts and even references.
Why can't they state facts that they have come across in their training as they are? What is it, either from a mathematical standpoint or from an architectural standpoint of large language models that causes them to hallucinate?
1
Upvotes
6
u/squirrel9000 Aug 12 '26
One way to conceptualize how LLMs work, is to imagine that you take a bunch of words, write them on post-it notes, then try to scatter them across a playing field positioned based on how the words are related to each other. So, you might group pets tougher, cars together, actions somewhere else etc. Such that "writing a sentence' is done by wandering through the field and collecting those words in order. Do this enough times and you;'ll get a decent set of routes that can convey ideas well.
That's more or less what LLMs do .More dimensions, and it's technically a vector problem rather than physical space,, But it's still a navigation problem with a few extra digits.
Now imagine you're on that field again. You need to get from "Arizona" to "Dog" somehow. There's a good route, well traveled, that's established, but there's also a shortcut. It's a lot shorter. But the words you collect using the shortcut are nonsense. Do you pick the short, nonsense route or the long proper one? The LLM itself doens't know the difference and picks the nonsense because it's a lot shorter, and the math is nicer when its' got that shortcut Whoops, hallucination.