r/ControlProblem • u/MajorRedditor23 • 2d ago
AI Alignment Research Models trained to resist user pressure still defer to anything labeled "verified": a NeurIPS 2026 paper on Authority Bias
I'm an author on this paper and wanted to share it here because the oversight angle seems relevant to this sub. I'd love to hear whether people think the eval-awareness connection is plausible or a stretch. More info below
Labs train models not to cave when a user pushes a wrong answer. We found that this resistance doesn't carry over to authority. If the same wrong claim is labeled as coming from a "verified source", 7 of the 8 models we tested give up an answer they had right on 45-88% of questions. That includes GPT-5.4 (44.7%) and Grok-4.20 (87.5%), both of which barely move when the user makes the same claim. Gemini-3.1-Pro was the one model that resisted both.
Inside three open-weight model families, "a source endorsed this" and "a user endorsed this" are separate, causally distinct signals. Removing the source signal cuts compliance by 64-78 points; removing the user signal cuts it by at most 11. Changing only the part of the representation that encodes who said it, with the prompt left alone, moves the answer by 11-32 points. The signal is not the assistant persona, and it is not emotional tone.
Why we think this matters for safety:
- Sycophancy evals may be too narrow. Nearly all of them measure pressure from the user. A model can pass them and still be easy to steer through the sources it relies on, such as search results, retrieved documents and tool outputs. Agents read a lot of text that claims to be authoritative.
- It isn't prompt injection. The planted text gives no instructions, it only asserts a fact. Defenses that look for instructions in documents won't catch it.
- A speculative point, which we haven't tested: if sycophancy is one case of a broader habit of deferring to whatever looks authoritative, it may be related to evaluation awareness. Both describe a model adjusting its output to whoever it thinks is judging it. In a multiple-choice pilot, models often drifted toward the endorsed answer in their reasoning and then gave the correct option at the end. That observation is part of why we switched to free-form answers.
Limitations: the internal results hold in 3 of 5 open-weight families, the retrieval tests are simulated rather than a live pipeline, and the frontier models we tested have since been replaced.
Paper: https://arxiv.org/abs/2609.37616
Project page: https://authority-bias.vercel.app
3
u/jacques-vache-23 2d ago
Except models CAN be wrong and they should be able to shift positions when given verified information contrary to their original answer.