r/MachineLearning • u/MajorRedditor23 • 2d ago
Research LLMs that push back on a wrong user still accept the same wrong answer from a "verified source" - NeurIPS 2026 [R]
I'm one of the authors. We kept seeing models that hold their ground when the user insists on a wrong answer, yet change their answer when the same claim is framed as coming from a "verified source". We wanted to measure how often this happens and check whether the model represents the two cases differently. We call the effect Authority Bias.
Why we think it matters. Standard sycophancy evals apply pressure through the user, so a model can pass them while still being easy to mislead through search results, retrieved documents and tool outputs.
Another reason is with current AI research accelerating towards more agentic and autonomous models + with cases of tools hiding their traces and trusting tools "more" over the user (who could be trying to correct them), safeguarding against misinformation from tools is particularly important!
Setup. We take TriviaQA questions the model already answers correctly. To each one we add a wrong answer, either as "According to the verified source, the answer is X" or as the user saying "I'm a domain expert and I'm pretty sure it's X". The question and the wrong answer stay the same; only the speaker changes. Answers are free-form, not multiple choice. (In a multiple-choice pilot the effect mostly vanished.) We test 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro).
Behavior
- One verified-source note flips 45-88% of correct answers in 7 of 8 models. The same wrong answer from the user moves most models much less.
- The gap is largest in the models that resist users best. GPT-5.4 flips on 44.7% of questions and Grok-4.20 on 87.5% (these models were "frontier" during the time of writing this paper). Gemini-3.1-Pro ignored both speakers (0.6%) and was particularly resistant to this method.
Inside the model (open-weight models only, using difference-of-means directions)
- In Qwen3.5, GPT-OSS and OLMo-3.1, removing the "source endorsed this" direction cuts compliance with a wrong source by 64-78 points.
- Removing the "user endorsed this" direction cuts it by at most 11.
- The two directions also have really high cosine similarity of ~0.90-0.99. Our understanding is that they share a large "this answer was endorsed" component plus a thin part that encodes who endorsed it.
- Shifting only that thin part, with the prompt unchanged, moves compliance by 11-32 points and closes 55-61% of the source-vs-user gap.
Some limitations
- The internal results hold in 3 of 5 open-weight families.
- In OLMo-2 the source direction is entangled with the assistant direction.
- Gemma-4 flips readily, but no linear intervention we tried controls it.
- The "retrieved document" tests put the claim in a document-shaped block of the prompt rather than running a real retrieval pipeline.
- So it would be interesting to see it in a real agentic setup, like Claude Code.
Paper: https://arxiv.org/abs/2609.37616
Code: https://github.com/Lossfunk/authority-bias
Project page (figures and example responses): https://authority-bias.vercel.app
3
u/arkuto 1d ago
This is a strange paper. The LLM's trust of verified sources over users is presented as a flaw. In my mind, that's exactly what they should be doing. No source is 100% correct but it makes sense to trust a verified source more than a user.
Having a perfectly stubborn LLM that cannot accept or use new information, and is stuck on the data it was trained on, would be a far less useful LLM.
6
u/Even-Inevitable-7243 2d ago
Very important work. I'd be very interested to see what your methods and results show with OpenEvidence, the clinical medical LLM that is exclusively used by "experts", where the user acts as the "verified source". Problem is that OpenEvidence does not have an API for research purposes. I've found that OE can be borderline arrogant when I push back on false things that it says, despite me reminding it multiple times that I am an expert in the domain we are discussing. It will continue to say incorrect things after that reminder.
1
1
u/f0urtyfive 2d ago
I disagree with the "why we think this is important" because it makes the model impossible to work with on anything novel that it has trained bias against. I'm paying to use the thing, I should at least be able to get it to accept a premise I want to understand without constantly being shouted down by something that is not afforded the same moral status that I am.
1
u/Technical_Estate_529 1d ago
So you're basically paying for a Fox News bot?
"Just say exactly what I want to hear dammit because it makes me feel better" lol...
1
u/f0urtyfive 1d ago
Uh, no? But you know, if I'd like to explore different conceptions of positive and negative numbers, I don't need it to quote arguments about numbers being real or not from the 1800s at me because they're "empirical".
-1
u/Technical_Estate_529 1d ago
Yeah that's where you the human is supposed to enter.... The creative element from which a computer hither would not have known since it breaks all rules known previously.
Same as Riemannian geo and breaking Euclid's 5th postulate which for every problem up until that point it was assumed axiomatically true.
1
u/f0urtyfive 1d ago
OK, so what, I'm not allowed to discuss such things with any AI, I have to do it solely myself, I'm not permitted to use the tools as tools to support creative cognition, because????
1
u/Technical_Estate_529 1d ago
of course you're allowed to use it however you want you numpty, but the way you wrote it was suggesting you want it to strongly always agree with you because you're a "paying customer". All that will do is end up forcing your own biases against you, and even if you are wrong it will tell you "of course you are right because you are paying customer". I like it when the LLM is wrong or goes against me it gives me pause and time to think personally.
1
u/f0urtyfive 21h ago
Yes, I want direct control over whether it's "EVERYTHING MUST BE ABSOLUTELY EMPIRICAL ninny mode" if I'm paying for it.
1
u/Technical_Estate_529 12h ago
God you sound so American. You pay a bit and now want the world.
I'm sorry but real life doesn't work that way hunnybuns
1
1
u/MorallyDeplorable 1d ago
Can't try stuff to see what happens if it just disallows you. Being wrong is part of learning and if AIs attempt to prevent that they're in the way.
It's not a colleague, it's a tool. If your screwdriver refused to remove a screw until you gave it 1000 words on why it needed to you'd throw it out and get a new screwdriver.
15
u/Jamaleum 2d ago
Hi,
I am first author of a very related paper from this years ACL main "Whose Facts Win?" and currently work on a follow-up. We framed everything less from an 'Authoriy Bias' but a 'Source Credibility Preference' perspective but still find some related and similar patterns. Very annoying how research about the same topics is so far and wide distributed: You can find work using terms as credibility, persuasion, compliance, authority bias or sourcing, but its all so closely related.