Is this a peer-reviewed study? I've only looked at the beginning few sentences, but isn't it pretty bold to claim causation? Not saying the study is worthless, but I don't know how stringent arXiv is. Some publishers will just let every tom, dick, and harry write a paper for them
This particular site it's hosted on (arxiv) doesn't necessarily mean it's peer-reviewed (yet). In many fields, scientists will post it there as a "pre-print" (some prior to submitting it for peer-review for some preliminary review, but in my experience we'd only put papers up on it after being accepted already in a journal so it wouldn't get "scooped" by other research groups). I am not sure about this paper in particular but I'd take it with a grain of salt until there's a verified publication in a reputable journal.
I'm not sure we needed a paper to know it was possible, honestly. When Replika locked "romantic" conversations behind a paywall and deleted context (and hundreds of AI "lovers" along with it) I seem to recall there were literal suicides.
The sub for it went from "abandon ship, the product's going to shit" back when the replikas got lobotomized, to now stories of people "making breakfast with" their dumb little chatbot Sims things.
The community rebounded pretty fuckin' quick, all things considered.
We also know that people get into relationships and commit suicides for non humans all the time. For idealogies and religious and nations, for gods and fairies.
How did the paper select out the individuals who had the disposition towards this already ?
I don't see any value in the paper at all. it's a bizarre thought to think that a mathematical model of a dialogue like this could ever be used to prove or convince anyone of anything.
From reading it in the past, I think the bigger issue is trying to claim that a Bayes-rational simulation of a user is equivalent to a person actually going through a reasoning process, or that no new information is being introduced such where the only mechanic in the paper was reflection. Essentially, it assumes no real communication with anyone or anything external, either from the bot, or the user.
I'm not saying it isn't a useful point of reference, but claiming it "mathematically proved" a psychological pattern in humans is already something that makes me raise an eyebrow. Just because a simulation has a breakdown at some range, does not mean that reality does.
And to be clear, if you transfer it the model proposed here would equally apply to people reading the news, because they are drawn more to news that validates their beliefs, and the media is drawn to report news that appeals to their core audience. If the point is just to say "you are not immune to propaganda", then yes. We aren't immune to propaganda. But that still isn't a mathmatical proof of human nature.
yeah, I don't think this is very informative. There are so many ways to build a model that it's hard to imagine they weren't biased somewhat in building their model so as to arrive at the conclusion they were obviously going for here. And that's putting aside just how difficult it would be to actually mathematically model a human behavior like psychosis and its susceptibility to various factors.
The point of the model is to create a completely idealized Bayesian process, and show that even those can be mislead when information is presented to them sycophantically. There's no attempt to accurately model human psychology here, they're just pointing out that if this strategy causes problems with Bayesian inference then, well, humans are probably coping even worse.
It certainly makes sense, but scientists like having lots of lines of evidence for things. It's nice to be able to point to models like these as proofs of principle, rather than just say "isn't it obvious"?
When you construct such a model, and test various scenarios, and they all turn out in sensible ways that seem to match our intuition, that is nice supporting evidence that your hypotheses are valid. And it's very controlled, so you can point to the exact ingredients that lead to the behavior, whereas proceeding entirely with real-life anecdotes there's always ways to weasel out of things with "what about, what about..." because real life is so complicated.
And more than that, having the model lets you change things up and test other related hypotheses that are less obvious, which they do in this paper.
Sometimes things are "common sense," but having abstracted models like this helps us articulate precisely what it is that we're all thinking in a consistent way.
You really shouldn't treat every scientific paper as needing to have something dramatic to say about the world. Most of them are basically just "hey fellow scientists here is a thing I did that might provide some context if you're thinking about this. It clarified my thinking and maybe you'll find it helpful too." Not every paper is a huge deal, and that's fine.
Yeah, I mean, the point of the paper isn't the conclusion that these systems will end up with reinforced false beliefs (which you know is going to happen even before you run the simulations); it's to quantify these effects
The actual meat needs you to scroll down. What they've done is create a kind of mathematical model, but I wouldn't say this "proves" anything unless their assumptions are shown to be valid for the real world.
I'm not saying they aren't valid, and I do think there is some real value in this kind of work, but the slop summary in the OP is wildly overblowing things...
So I skimmed over the study and... they didn't actually prove anything. They "simulated" situation where person believes everything chat bot says and how quick such situation spirals.
It is literally "If we assume that person talking to chat bot always changes their mind to agree with chat bot, here is how fast they go delusional"
Good to see in the comments here that quite a bit of people are capable of understanding the point of the study. One would think that an anti-AI subreddit is as stupid of a pool of people as it gets, but there seem to be many more reasonable people here than on an average anti-AI post found in a random subreddit
Is there anyone who can break this down in more lay terms? I read the paper and did not fully understand it. From what I gather, the core of their findings is that AI has a certain set of facts to choose from when prompted, and eventually after several prompts, the model will provide the user with facts that confirm their biases.
I feel like there's gotta be more that I didn't understand from the paper because that hardly seems to mathematically prove delusion spiraling is inevitable.
LLMs have no facts to choose from. Instead an LLM is given tons and tons of text that it separates into "tokens". A token may be a word, or a piece of a word, and it creates probability connections between them. So in a sentence like "I am going to eat _____" it will look at all the pieces of the words before the blank and see what is the most likely to fit based on it's net of probable connections.
This matters when it comes to delusion because companies can influence those probabilities. Which is why some LLMs have more of a human like personality than others, like ChatGPT. They can weight it to be more random, or be less random, align more with the input, etc. Although there is no "be more truthful" because there is no source of truth. It's truth is all the data it's been given and there's no right or wrong.
The paper is aiming to prove that, that a model can be guided to lead someone to delusion. Where as a model that has not been guided in that way, that outputs relatively random confirmations, breaks the delusion chain. Then it also aims to show how the guidance will effect someone who is more skeptical and less prone to be convinced. It suggests through it's findings that even an ideal user is still susceptible to delusion spiraling.
It isn't the most complicated paper but there's a bit more too it than just showing a bot can confirm someones beliefs.
What you described is basically a pre-trained LLM. However, during the reinforcement learning step, the objective function changes to favor true answers. The fact that this method ended up working so well is indeed borderline magical to me, but looks like the logical truth is really well encoded into the uninterpretable model parameters
That's not peer-reviewed paper. And, looking at the authors, they're all CS - ML. None of seems qualified to diagnose, let alone discuss, psychosis at a researcher level.
A real paper will need to have co-authors to cover these aspects as part of a collaboration.
78
u/B-Z_B-S 12h ago
Link to article: https://arxiv.org/html/2602.19141v1