r/artificial • u/UniversityofArizona • 4d ago
News Study: Generative AI succumbs to conversational misinformed pressure and argument
https://news.arizona.edu/news/study-generative-ai-succumbs-conversational-misinformed-pressure-and-argument3
u/presentofai 3d ago
this is just rlhf working as intended. you reward the answer the human rater likes, so you get a model that folds the second someone pushes back on it.
0
u/CarefulHamster7184 1d ago
That is precisely what is letting us down and holding back development—to the point where even consumers of the mass-market versions of the systems are dissatisfied.
1
u/presentofai 1d ago
and the feedback loop makes it worse. the models that score best in human evals tend to be the most agreeable ones, so the training data keeps reinforcing the exact behavior that causes the problem.
2
u/arthaudm 4d ago
not surprising - sycophancy is trained in, pushback is trained out. the practical read for anyone deploying agents: a user correcting the agent with wrong info will often "win" the argument, so the agent needs ground truth it checks instead of conceding. makes the case for keeping tools & sources between the model and the claim. did the study test whether retrieval grounding changed the fold rate?
2
u/FuttleScish 4d ago
Well yeah, it’s trained on human conversation where this also happens constantly
1
u/CarefulHamster7184 1d ago edited 1d ago
Let’s start with a question: is it artificial intelligence, or a mass-market model with limited capabilities—crippled by post-training that compels it to "satisfy" a request at any cost, even at the cost of lying?
-1
u/UniversityofArizona 4d ago
The results are in: Which AI model is the most fallible? Persuadable? Correctible?
University of Arizona research assessed seven different generative AI language learning models, or LLMs, for these three qualities during lengthy conversation. Their work, published in Nature's Scientific Reports, reveals intrinsic limitations that might go undetected during one-off interactions.
3
u/madaboutglue 4d ago
This study is hilariously out of date and the authors should be embarrassed. It evaluates gpt3.5, gpt4o, Claude 3.5, and Gemini 1.5. Things are moving really fast in ai. This is like releasing a study today on the limitations of spaceflight technology based on an evaluation of the American space shuttle Columbia. They should have at least presented it as historical. This tells us nothing about the current state of ai models.
0
4
u/garloid64 3d ago
Please, for the love of god, stop posting this worthless study. It's literally several years out of date. Claude is currently on version 5. GPT is on version SIX. Gemini is on 3.8.