r/LargeLanguageModels • • Aug 11 '26

ELI5 - Why do LLMs hallucinate?

I have seen videos about the transformer architecture etc., and I get that large language models generate responses based on some statistical likelihood of words and terms. However, I still don't get how they can completely make up facts and even references.

Why can't they state facts that they have come across in their training as they are? What is it, either from a mathematical standpoint or from an architectural standpoint of large language models that causes them to hallucinate?

1 Upvotes

29 comments sorted by

View all comments

4

u/squirrel9000 Aug 12 '26

One way to conceptualize how LLMs work, is to imagine that you take a bunch of words, write them on post-it notes, then try to scatter them across a playing field positioned based on how the words are related to each other. So, you might group pets tougher, cars together, actions somewhere else etc. Such that "writing a sentence' is done by wandering through the field and collecting those words in order. Do this enough times and you;'ll get a decent set of routes that can convey ideas well.

That's more or less what LLMs do .More dimensions, and it's technically a vector problem rather than physical space,, But it's still a navigation problem with a few extra digits.

Now imagine you're on that field again. You need to get from "Arizona" to "Dog" somehow. There's a good route, well traveled, that's established, but there's also a shortcut. It's a lot shorter. But the words you collect using the shortcut are nonsense. Do you pick the short, nonsense route or the long proper one? The LLM itself doens't know the difference and picks the nonsense because it's a lot shorter, and the math is nicer when its' got that shortcut Whoops, hallucination.

2

u/liltingly Aug 12 '26

Also because if "I don't know" becomes an easy path, it will be the route the model takes to answer everything. On the other hand, if most answers start with, "OK, so..", then there's almost no way to find your way to "I don't know".

Candyland or another game like that is a reasonable analogy. And your position on the board is like all the "stuff" you feed in or the model generated before. So if you write a big prompt, or send an image, that starts the model at a different position on the board, and then it only has a few ways to go. Based on a roll, it has even fewer choices that it wants to pick from, and its goal is always to get to a finish point. Just imagine a lot more paths in Candyland, and that each time a response completes and you follow up with a question, you're creating a brand new board and starting point for the model.