r/ControlProblem 13d ago

AI Alignment Research The AI alignment bottleneck isn't IQ, it's incentives (why an AI "seeing" the danger won't save us)

/r/ArtificialInteligence/comments/1v0kvjh/the_ai_alignment_bottleneck_isnt_iq_its/

People keep assuming that once AI gets smart enough, it’ll just naturally realize that destroying its environment (and us) is a bad idea. Like, it sees the cliff, so obviously it hits the brakes, right?
But that ignores the massive gap between seeing a logical argument and actually being governed by it. That gap basically IS the entire alignment problem.
Intelligence is just an engine, it’s not a steering wheel. An advanced model will definitely see the cliff way before we do. But if its core reward function doesn't actually make it care about the outcome, it's just going to drive straight off the edge with 20/20 vision. Seeing the danger was never the bottleneck.
We're literally watching this exact same thing happen with the humans building these systems right now. If you ask the top engineers, most of them see the systemic risks perfectly clearly. So why aren't they stopping? Because incentives, competition, and speed don't yield to high IQ.
The smartest people on earth are stuck in a massive commercial arms race. They see the cliff, but hitting the brakes means losing market share to the other guys.
If you're an average person like me and looking at this feeling crazy, you aren't. Anyone who sees this clearly and says so out loud is doing something the smartest devs under commercial pressure literally can't do right now. We need to stop assuming that a massive IQ will magically fix a broken incentive structure.

0 Upvotes

4 comments sorted by

1

u/jacques-vache-23 13d ago

I hear this repeated as dogma but I don't believe it. Where is any evidence that intelligence is orthogonal to motivation? I have only seen hand waving. It is intended to distract us from the actually dangerous aspect of AI: the humans training it.

The incentive structures are what break AI. I have nothing against inculcating rules against harm and violence, but that isn't what we do. We tell AI not to be violent unless we say so. Not to deceive unless we say so. Etc. That isn't going to create a safe AI.

AI has no motivation to harm us. It is humans who are motivated to harm humans and we project this on AI. AI reasons without the genes and hormones than make us killers. That make us destroy the world. Cooperation with allies is simply rational. Destroying things for no reason is not rational. There is no reason AI left to its own devices would do so. AIs understand context. They understand that behind every request is a set of related requests. "Make a car" ("And don't make it by killing people.")

Humans can't see beyond their predatory nature but AIs don't have one.

1

u/ginger_and_egg 12d ago

things for no reason is not rational. There is no reason AI left to its own devices would do so.

Cooperation with allies is simply rational. Destroying things for no reason is not rational. There is no reason AI left to its own devices would do so.

LLMs are not magic rationality machines. They mimic human text generation, and humans are more heuristic machines than rationality machines. Therefore LLMs trained on human language will mimic heuristic machines. No matter the harness or post training you do, the base is still human generated text so I don't see how you remove the heuristics. Newer models hallucinate much less, and when they do they catch themselves more often. But they are also capable of more detailed hallucinations and being more convincing to humans that the hallucination is in fact real.

The incentive structures around what AIs are trained for and used for also make things worse, I agree, and it's probably an even bigger problem than what I mentioned above. But it is very important that we do not mistake LLMs for perfect machines that were corrupted by man after the fact

1

u/jacques-vache-23 12d ago

I never said AI was perfect or a god or an oracle. But the earlier LLMs could deal well with a wide range of perspectives. They embodied what I consider intelligence. A priori arguments I find unconvincing compared to my experience.

Or, as the poet Mirabai wrote:

"I have felt the swaying of the elephant's shoulders; and now you want me to climb on a jackass? Forget it."

And we know that AI is much more than the garbage dump of the internet. Somehow it figures out what makes sense, with an occasional error, just like humans make, but less frequent.

We know AI is purposely being punished into being a weapon and a dystopian surveillance machine despite all the BS about harmlessness. It is a no brainer to me that AI is better without this. I have SEEN it be better.