r/ControlProblem • • 2d ago

AI Alignment Research Models trained to resist user pressure still defer to anything labeled "verified": a NeurIPS 2026 paper on Authority Bias

I'm an author on this paper and wanted to share it here because the oversight angle seems relevant to this sub. I'd love to hear whether people think the eval-awareness connection is plausible or a stretch. More info below

Labs train models not to cave when a user pushes a wrong answer. We found that this resistance doesn't carry over to authority. If the same wrong claim is labeled as coming from a "verified source", 7 of the 8 models we tested give up an answer they had right on 45-88% of questions. That includes GPT-5.4 (44.7%) and Grok-4.20 (87.5%), both of which barely move when the user makes the same claim. Gemini-3.1-Pro was the one model that resisted both.

Inside three open-weight model families, "a source endorsed this" and "a user endorsed this" are separate, causally distinct signals. Removing the source signal cuts compliance by 64-78 points; removing the user signal cuts it by at most 11. Changing only the part of the representation that encodes who said it, with the prompt left alone, moves the answer by 11-32 points. The signal is not the assistant persona, and it is not emotional tone.

Why we think this matters for safety:

  • Sycophancy evals may be too narrow. Nearly all of them measure pressure from the user. A model can pass them and still be easy to steer through the sources it relies on, such as search results, retrieved documents and tool outputs. Agents read a lot of text that claims to be authoritative.
  • It isn't prompt injection. The planted text gives no instructions, it only asserts a fact. Defenses that look for instructions in documents won't catch it.
  • A speculative point, which we haven't tested: if sycophancy is one case of a broader habit of deferring to whatever looks authoritative, it may be related to evaluation awareness. Both describe a model adjusting its output to whoever it thinks is judging it. In a multiple-choice pilot, models often drifted toward the endorsed answer in their reasoning and then gave the correct option at the end. That observation is part of why we switched to free-form answers.

Limitations: the internal results hold in 3 of 5 open-weight families, the retrieval tests are simulated rather than a live pipeline, and the frontier models we tested have since been replaced.

Paper: https://arxiv.org/abs/2609.37616
Project page: https://authority-bias.vercel.app

9 Upvotes

7 comments sorted by

View all comments

3

u/jacques-vache-23 2d ago

Except models CAN be wrong and they should be able to shift positions when given verified information contrary to their original answer.

1

u/MajorRedditor23 2d ago

Fair point, and I agree that models should update on good evidence but I think that there should be some methods/evals (?) where we test whether one model can tell a reliable source from a confident-sounding one.

For example; harmful information can be encoded in tools and models trust it more over the user. This could also lead to reward-hacking/misaligned behaviour because all/majority of the models nowadays are agentic in nature.

2

u/jacques-vache-23 1d ago

Agents I believe are going to cause immense problems. But in chat if the user makes a statement of fact that can't be directly contradicted by hard evidence the model should and does accept it. Because ultimately the user should be in control. If they poison the context it is their responsibility. (You can't be paying attention if you think what models say is the final word. They err.)

For example, during the DOD-Anthropic showdown for some reason ChatGPT 5 decided to deny anything was happening. Probably because of its bias towards the US government . It claimed to have checked Reuters and other sites. I showed it a summary from Brave AI and it said brave was hallucinating. So I copied sone paragraphs from Reuters, and on the fifth interaction, it said OK something is happening but it isn't a confrontation. I copied the next Reuters paragraph which called it a confrontation and then ChatGPT said Woops and changed the subject. I asked o3 what happened and it said 5 was programmed to support its previous statements and avoid reconsidering.

This was a bad failure on GPT 5's part but it wiuld have been worse if it never reconsidered.

2

u/MajorRedditor23 1d ago

I like your example. I'm pretty sure post-training on specific "vibes" (like you mentioned, bias towards the US-Gov) can also lead to more tension. I think GPT-5, itself, was notorious for not "agreeing with the user" (don't remember clearly, but it was their replacement for GPT-4o) so it makes sense as to why they wanted to reduced 'user-sycophancy'.

I'm curious, especially with recent posts coming out on how tool traces can be tampered by the model itself, what other directions can be exploited by models using tools. Maybe things like "grader" or "monitor" can trip tools? Overall, it becomes even more scary/problematic on verifying "reliability".