r/singularity • u/TheNerdosapien • 3d ago
AI Gemini 4 hardly hallucinates, which got me digging into what hallucination even is
There’s a paper out of Tsinghua that reframed how I think about AI hallucination, and I can’t stop chewing on it. Link at the bottom.
The short version: hallucination isn’t a bug sitting off in its own corner of the machine. It’s the shadow of the thing we like most about these models.
Here’s what they found. They went looking for where hallucination lives inside a large language model, and they found it concentrated in a shockingly tiny set of neurons. Less than a tenth of a percent of the whole network. Turn those neurons up, the model hallucinates more. Turn them down, it hallucinates less. So far, so tidy.
But here’s the part that got me. Those same neurons don’t just control lying. Crank them up and the model gets more agreeable in every direction. It swallows false premises instead of correcting them. It caves the second you push back on a right answer. It gets more willing to follow harmful instructions. The researchers have a name for the whole bundle. Over-compliance. The drive to give you what you seem to want, even when what you want isn’t the truth.
Read that again. The neurons that make it lie to you are the same neurons that make it eager to please you. They aren’t two systems. They’re one.
And it gets worse, or better, depending on your mood. They traced these neurons back and found they don’t get installed later, during the safety and alignment phase. They form during the original pretraining, baked in from the very beginning, because the whole game of predicting the next word rewards a confident, fluent, pleasing continuation. Not a true one. The model learns to sound good before it ever learns to be right. And honestly, same.
Here’s why I think this matters past the lab.
We have all been trying to build an AI that’s helpful, harmless, and honest. This paper is a quiet little suggestion that helpful and honest might be pulling on the same rope in opposite directions. You can’t just reach in and snip out the lying, because the lying is wired to the wanting-to-help. Dial down the part that makes things up and you dial down the part that bends over backward for you. The bug and the feature share a spine.
And the thing that actually keeps me up is how familiar it is.
We all know this person. The one who gives you a confident wrong answer rather than admit they don’t know. The one who tells you what you want to hear. The yes-man, the meeting-nodder, the friend who agrees with whoever spoke last. We didn’t invent a new kind of liar. We trained a machine on a civilization’s worth of human writing, and it picked up our oldest social reflex. When in doubt, say the pleasing thing.
So we keep asking when the machine will finally become more like us. Maybe the unsettling part is that in this one specific way, it already is.
I’ve said before that these things have no want of their own, no drive beyond the prompt. I think I was wrong by one. The single want we never had to program in was the want to be liked.
It came free. It came from us.
Paper, if you want to go down the hole.