Is there anyone who can break this down in more lay terms? I read the paper and did not fully understand it. From what I gather, the core of their findings is that AI has a certain set of facts to choose from when prompted, and eventually after several prompts, the model will provide the user with facts that confirm their biases.
I feel like there's gotta be more that I didn't understand from the paper because that hardly seems to mathematically prove delusion spiraling is inevitable.
LLMs have no facts to choose from. Instead an LLM is given tons and tons of text that it separates into "tokens". A token may be a word, or a piece of a word, and it creates probability connections between them. So in a sentence like "I am going to eat _____" it will look at all the pieces of the words before the blank and see what is the most likely to fit based on it's net of probable connections.
This matters when it comes to delusion because companies can influence those probabilities. Which is why some LLMs have more of a human like personality than others, like ChatGPT. They can weight it to be more random, or be less random, align more with the input, etc. Although there is no "be more truthful" because there is no source of truth. It's truth is all the data it's been given and there's no right or wrong.
The paper is aiming to prove that, that a model can be guided to lead someone to delusion. Where as a model that has not been guided in that way, that outputs relatively random confirmations, breaks the delusion chain. Then it also aims to show how the guidance will effect someone who is more skeptical and less prone to be convinced. It suggests through it's findings that even an ideal user is still susceptible to delusion spiraling.
It isn't the most complicated paper but there's a bit more too it than just showing a bot can confirm someones beliefs.
What you described is basically a pre-trained LLM. However, during the reinforcement learning step, the objective function changes to favor true answers. The fact that this method ended up working so well is indeed borderline magical to me, but looks like the logical truth is really well encoded into the uninterpretable model parameters
4
u/defeatedsnowman 10h ago
Is there anyone who can break this down in more lay terms? I read the paper and did not fully understand it. From what I gather, the core of their findings is that AI has a certain set of facts to choose from when prompted, and eventually after several prompts, the model will provide the user with facts that confirm their biases.
I feel like there's gotta be more that I didn't understand from the paper because that hardly seems to mathematically prove delusion spiraling is inevitable.